IP Library › Granted Patent US 12,216,475
Granted Patent B2
US 12,216,475 · App. 17/411,894 · Granted Feb 4, 2025

Autonomous vehicle object classifier method and system

Inventors: Jiachen Li (Albany, CA); Haiming Gang (San Jose, CA); Hengbo Ma (Albany, CA); Chiho Choi (San Jose, CA)
Assignee: Honda Motor Co., Ltd.
G05D1/0221B60W50/08B60W60/0027G01S17/89G06F18/2433G06V10/462G06V20/58G01S17/931G06F18/253G06V10/454G06V10/806
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,216,475
App. No.
17/411,894
Granted
Feb 4, 2025
Kind
B2
Abstract

Object identification may be provided herein. A feature extractor may extract a first set of visual features, extract a second set of visual features, concatenate the first set of visual features, the second set of visual features, and a set of bounding box information, determine a number of object features and a global feature for a scene, and receive ego-vehicle feature information associated with an ego-vehicle. An object classifier may receive the number of object features, the global feature, and the ego-vehicle feature information, generate relational features with respect to relationships between each of the number of objects from the scene, and classify each of the number of objects from the scene based on the number of object features, the relational features, the global feature, the ego-vehicle feature information, and an intention of the ego-vehicle.

Claims (54)

1. A system for object identification, comprising:

a feature extractor:

extracting a first set of visual features from a first image of a scene detected by a first sensor;

extracting a second set of visual features from a second image of the scene detected by a second sensor of a different sensor type than the first sensor;

concatenating the first set of visual features, the second set of visual features, and a set of bounding box information associated with the first image and the second image;

determining a number of object features associated with a corresponding number of objects other than an ego-vehicle from the scene and a global feature for the scene; and

receiving ego-vehicle feature information associated with the ego-vehicle; and

an object classifier:

receiving the number of object features, the global feature, and the ego-vehicle feature information;

generating relational features with respect to relationships between each of the number of objects from the scene, the relational features including a mutual influence between each of the number of objects from the scene;

classifying each of the number of objects from the scene as a binary classification of relevant or non-relevant to a behavior of the ego-vehicle, based on the number of object features, the relational features, the global feature, the ego-vehicle feature information, and an intention of the ego-vehicle; and

labeling each of the number of objects from the scene as relevant or non-relevant to the behavior of the ego-vehicle,

wherein the object classifier is trained utilizing semi-supervised learning including a labeled dataset and an unlabeled dataset, wherein the unlabeled dataset is annotated with pseudo labels generated from classifying each of the number of objects, the pseudo labels indicating each of the number of objects from the scene as relevant or non-relevant to the behavior of the ego-vehicle.

2. The system for object identification for the ego-vehicle of claim 1 , wherein the ego-vehicle feature information associated with the ego-vehicle includes a position, a velocity, or an acceleration associated with the ego-vehicle.

3. The system for object identification for the ego-vehicle of claim 1 , wherein the feature extractor sequence encodes the first set of visual features, the second set of visual features, and the set of bounding box information prior to concatenation.

4. The system for object identification for the ego-vehicle of claim 1 , wherein the generating relational features with respect to relationships between each of the number of objects from the scene is based on a fully-connected object relation graph, wherein each node corresponds to an object feature and each edge connecting two nodes represents a relationship between two objects associated with the two nodes.

5. The system for object identification for the ego-vehicle of claim 1 , comprising a task generator generating a task to be implemented via an autonomous controller and one or more vehicle systems using one or more of the number of object features associated with the corresponding number of objects from the scene classified as relevant to the behavior of the ego-vehicle, while discarding one or more of the number of object features associated with the corresponding number of objects from the scene classified as non-relevant to the behavior of the ego-vehicle.

6. The system for object identification for the ego-vehicle of claim 5 , wherein the task generator generates the task based on one or more of the number of object features associated with the corresponding number of objects from the scene classified as relevant to the behavior of the ego-vehicle, the ego-vehicle feature information, and the global feature, while discarding one or more of the number of object features associated with the corresponding number of objects from the scene classified as non-relevant to the behavior of the ego-vehicle.

7. The system for object identification for the ego-vehicle of claim 5 , wherein the task to be implemented includes an ego-vehicle action classifier and an ego-vehicle trajectory.

8. The system for object identification for the ego-vehicle of claim 1 , wherein the object classifier is trained utilizing supervised learning including a labeled dataset.

9. The system for object identification for the ego-vehicle of claim 1 , wherein the first sensor is an image capture sensor and the second sensor is a light detection and ranging (LiDAR) sensor.

10. A computer-implemented method for object identification, comprising:

extracting a first set of visual features from a first image of a scene detected by a first sensor;

extracting a second set of visual features from a second image of the scene detected by a second sensor of a different sensor type than the first sensor;

concatenating the first set of visual features, the second set of visual features, and a set of bounding box information associated with the first image and the second image;

determining a number of object features associated with a corresponding number of objects other than an ego-vehicle from the scene and a global feature for the scene;

receiving ego-vehicle feature information associated with the ego-vehicle;

receiving the number of object features, the global feature, and the ego-vehicle feature information;

generating relational features with respect to relationships between each of the number of objects from the scene, the relational features including a mutual influence between each of the number of objects from the scene;

classifying, by an object classifier, each of the number of objects from the scene as a binary classification of relevant or non-relevant to a behavior of the ego-vehicle, based on the number of object features, the relational features, the global feature, the ego-vehicle feature information, and an intention of the ego-vehicle; and

labeling each of the number of objects from the scene as relevant or non-relevant to the behavior of the ego-vehicle,

wherein the method further comprises training the object classifier utilizing semi-supervised learning including a labeled dataset and an unlabeled dataset, wherein the unlabeled dataset is annotated with pseudo labels generated from classifying each of the number of objects, the pseudo labels indicating each of the number of objects from the scene as relevant or non-relevant to the behavior of the ego-vehicle.

11. The method for object identification for the ego-vehicle of claim 10 , wherein the ego-vehicle feature information associated with the ego-vehicle includes a position, a velocity, or an acceleration associated with the ego-vehicle.

12. The method for object identification for the ego-vehicle of claim 10 , comprising sequence encoding the first set of visual features, the second set of visual features, and the set of bounding box information prior to concatenation.

13. The method for object identification for the ego-vehicle of claim 10 , comprising generating the relational features based on a fully-connected object relation graph, wherein each node corresponds to an object feature and each edge connecting two nodes represents a relationship between two objects associated with the two nodes.

14. The method for object identification for the ego-vehicle of claim 10 , comprising generating a task to be implemented via an autonomous controller and one or more vehicle systems using one or more of the number of object features associated with the corresponding number of objects from the scene classified as relevant to the behavior of the ego-vehicle, while discarding one or more of the number of object features associated with the corresponding number of objects from the scene classified as non-relevant to the behavior of the ego-vehicle.

15. The method for object identification for the ego-vehicle of claim 14 , comprising generating the task based on one or more of the number of object features associated with the corresponding number of objects from the scene classified as relevant to the behavior of the ego-vehicle, the classifying each of the number of objects, the ego-vehicle feature information, and the global feature, while discarding one or more of the number of object features associated with the corresponding number of objects from the scene classified as non-relevant to the behavior of the ego-vehicle.

16. The method for object identification for the ego-vehicle of claim 14 , wherein the task to be implemented includes an ego-vehicle action classifier and an ego-vehicle trajectory.

17. The method for object identification for the ego-vehicle of claim 10 , comprising training the object classifier utilizing supervised learning including a labeled dataset.

18. A system for object identification, comprising:

a feature extractor:

extracting a first set of visual features from a first image of a scene detected by a first sensor;

extracting a second set of visual features from a second image of the scene detected by a second sensor of a different sensor type than the first sensor;

concatenating the first set of visual features, the second set of visual features, and a set of bounding box information associated with the first image and the second image;

determining a number of object features associated with a corresponding number of objects other than an ego-vehicle from the scene and a global feature for the scene; and

receiving ego-vehicle feature information associated with the ego-vehicle;

an object classifier:

receiving the number of object features, the global feature, and the ego-vehicle feature information;

generating relational features with respect to relationships between each of the number of objects from the scene, the relational features including a mutual influence between each of the number of objects from the scene;

classifying each of the number of objects from the scene as a binary classification of relevant or non-relevant to a behavior of the ego-vehicle, based on the number of object features, the relational features, the global feature, the ego-vehicle feature information, and an intention of the ego-vehicle; and

labeling each of the number of objects from the scene as relevant or non-relevant to the behavior of the ego-vehicle;

a task generator generating a task to be implemented using one or more of the number of object features associated with the corresponding number of objects from the scene classified as relevant to the behavior of the ego-vehicle, while discarding one or more of the number of object features associated with the corresponding number of objects from the scene classified as non-relevant to the behavior of the ego-vehicle; and

an autonomous controller implementing the task by driving one or more vehicle systems to execute the task,

wherein the object classifier is trained utilizing semi-supervised learning including a labeled dataset and an unlabeled dataset, wherein the unlabeled dataset is annotated with pseudo labels generated from classifying each of the number of objects, the pseudo labels indicating each of the number of objects from the scene as relevant or non-relevant to the behavior of the ego-vehicle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2021
From: LI, JIACHEN; GANG, HAIMING; MA, HENGBO; CHOI, CHIHO
To: HONDA MOTOR CO., LTD.
Reel/Frame 057288/0964 →
Continuity (2)
Provisional Application 63216202 · Jun 29, 2021
Related Publication 20220413507A1 · Dec 29, 2022
References Cited (62)
US 11436504B1 · Lukarski · 2022 [cited by examiner]
US 11537819B1 · Das · 2022 [cited by examiner]
US 20180334108A1 · Rötzer · 2018 [cited by examiner]
US 20190147610A1 · Frossard · 2019 [cited by examiner]
US 20200394434A1 · Rao · 2020 [cited by examiner]
US 20220129684A1 · Saranin · 2022 [cited by examiner]
US 20220327318A1 · Radhakrishnan · 2022 [cited by examiner]
US 20220358401A1 · Kang · 2022 [cited by examiner]
US 20220366175A1 · Yu · 2022 [cited by examiner]
US 20220388535A1 · Singh · 2022 [cited by examiner]
Penghao Zhou and Mingmin Chi. Relation parsing neural network for human-object interaction detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 843-851, 2019. [cited by applicant]
Yi Zhou, Xiaodong He, Lei Huang, Li Liu, Fan Zhu, Shanshan Cui, and Ling Shao. Collaborative learning of semi-supervised segmentation and classification for medical images. In Proceedings of the IEEE/CVF Conference on C… [cited by applicant]
Ekrem Aksoy, Ahmet Yazci, and Mahmut Kasap. See, attend and brake: An attention-based saliency map prediction model for end-to-end driving. arXiv preprint arXiv:2002.11020, 2020. [cited by applicant]
Ferran Alet, EricaWeng, Tomas Lozano-Perez, and Leslie Pack Kaelbling. Neural relational inference with fast modular meta-learning. In Advances in Neural Information Processing Systems, pp. 11804-11815, 2019. [cited by applicant]
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE confe… [cited by applicant]
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics. In Advances in neural information processing systems, pp. 4502-4510, 2… [cited by applicant]
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. Relational inductive biases, deep l… [cited by applicant]
David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi-supervised learning. arXiv preprint arXiv:1905.02249, 2019. [cited by applicant]
Manjot Bilkhu, Siyang Wang, and Tushar Dobhal. Attention is all you need for videos: Self-attention based video summarization using universal transformers. arXiv preprint arXiv:1906.02792, 2019. [cited by applicant]
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of… [cited by applicant]
C. Chen, S. Hu, P. Nikdel, G. Mori, and M. Savva. Relational graph learning for crowd navigation. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 10007-10013, 2020. [cited by applicant]
Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017. [cited by applicant]
Chiho Choi, Joon Hee Choi, Jiachen Li, and Srikanth Malla. Shared cross-modal trajectory prediction for autonomous driving. arXiv preprint arXiv:2011.08436, 2020. [cited by applicant]
Luca Cultrera, Lorenzo Seidenari, Federico Becattini, Pietro Pala, and Alberto Del Bimbo. Explaining autonomous driving by learning end-to-end visual attention. In Proceedings of the IEEE/CVF Conference on Computer Visi… [cited by applicant]
Zihang Dai, Zhilin Yang, Fan Yang, William W Cohen, and Ruslan Salakhutdinov. Good semi-supervised learning that requires a bad gan. In Proceedings of the 31st International Conference on Neural Information Processing S… [cited by applicant]
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra. Human attention in visual question answering: Do humans and deep networks look at the same regions? Computer Vision and Image Understanding, 163:… [cited by applicant]
Stuart Eiffert, Kunming Li, Mao Shan, Stewart Worrall, Salah Sukkarieh, and Eduardo Nebot. Probabilistic crowd gan: Multimodal pedestrian trajectory prediction using a graph vehicle-pedestrian attention network. IEEE Ro… [cited by applicant]
Mingfei Gao, Ashish Tawari, and Sujitha Martin. Goal-oriented object importance estimation in on-road driving videos. In 2019 International Conference on Robotics and Automation (ICRA), pp. 5509-5515. IEEE, 2019. [cited by applicant]
William L Hamilton, Rex Ying, and Jure Leskovec. Inductive representation learning on large graphs. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 1025-1035, 2017. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. [cited by applicant]
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn, 2018. [cited by applicant]
Fa-Ting Hong, Wei-Hong Li, and Wei-Shi Zheng. Learning to detect important people in unlabelled images for semi-supervised important people detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat… [cited by applicant]
Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei. Relation networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3588-3597, 2018. [cited by applicant]
Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, and Ondrej Chum. Label propagation for deep semi-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5070-5079, 20… [cited by applicant]
Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016. [cited by applicant]
Longlong Jing, Toufiq Parag, Zhe Wu, Yingli Tian, and Hongcheng Wang. Videossl: Semi-supervised learning for video classification. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.… [cited by applicant]
Jinkyu Kim and John Canny. Interpretable learning for self-driving cars by visualizing causal attention. In Proceedings of the IEEE international conference on computer vision, pp. 2942-2950, 2017. [cited by applicant]
Diederik P. Kingma, Danilo J Rezende, Shakir Mohamed, and Max Welling. Semi-supervised learning with deep generative models. In Proceedings of the 27th International Conference on Neural Information Processing Systems—v… [cited by applicant]
Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, and Richard Zemel. Neural relational inference for Interacting systems. In International Conference on Machine Learning, pp. 2688-2697, 2018. [cited by applicant]
Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016. [cited by applicant]
Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, vol. 3, 2013. [cited by applicant]
C. Li, S. H. Chan, and Y. T. Chen. Who make drivers stop? towards driver-centric risk assessment: Risk object dentification via causal inference. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Syste… [cited by applicant]
Jiachen Li, Hengbo Ma, Zhihao Zhang, and Masayoshi Tomizuka. Social-wagdat: Interaction-aware trajectory prediction via wasserstein graph double-attention network. arXiv preprint arXiv:2002.06241, 2020. [cited by applicant]
Jiachen Li, Fan Yang, Masayoshi Tomizuka, and Chiho Choi. Evolvegraph: Multi-agent trajectory prediction with dynamic relational reasoning. Advances in Neural Information Processing Systems, 33, 2020. [cited by applicant]
Wei-Hong Li, Fa-Ting Hong, and Wei-Shi Zheng. Learning to learn relation for important people detection in still images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5003-501… [cited by applicant]
Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogniti… [cited by applicant]
Sujitha Martin, Sourabh Vora, Kevan Yuen, and Mohan Manubhai Trivedi. Dynamics of driver's gaze: Explorations in behavior modeling and maneuver prediction. IEEE Transactions on Intelligent Vehicles, 3 (2):141-150, 2018. [cited by applicant]
Avital Oliver, Augustus Odena, Colin A Raffel, Ekin Dogus Cubuk, and Ian Goodfellow. Realistic evaluation of deep semi-supervised learning algorithms. Advances in Neural Information Processing Systems, 31:3235-3246, 201… [cited by applicant]
Yassine Ouali, Celine Hudelot, and Myriam Tami. An overview of deep semi-supervised learning. arXiv preprint arXiv:2006.05278, 2020. [cited by applicant]
Andrea 442 Palazzi, Davide Abati, Francesco Solera, Rita Cucchiara, et al. Predicting the driver's focus of attention: the dr (eye) ve project. IEEE transactions on pattern analysis and machine intelligence, 41(7): 1720… [cited by applicant]
Abhishek Patil, Srikanth Malla, Haiming Gang, and Yi-Ting Chen. The h3d dataset for full-surround 3d multi-object detection and tracking in crowded urban scenes. In 2019 International Conference on Robotics and Automati… [cited by applicant]
Alvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying, Jure Leskovec, and Peter Battaglia. Learning to simulate complex physics with graph networks. In International Conference on Machine Learning, pp. 8459-8… [cited by applicant]
Tom Silver, Rohan Chitnis, Aidan Curtis, Joshua Tenenbaum, Tomas Lozano-Perez, and Leslie Pack Kaelbling. Planning with learned object importance in large problem instances using graph neural networks. arXiv preprint ar… [cited by applicant]
Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In European Conference on Computer Vision, pp. 347-365. Springer, 2020. [cited by applicant]
Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Proceedings of the 31st International Conference on Neural I… [cited by applicant]
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017. [cited by applicant]
Anirudh Vemula, Katharina Muelling, and Jean Oh. Social attention: Modeling attention in human crowds. In 2018 IEEE international Conference on Robotics and Automation (ICRA), pp. 4601-4607. IEEE, 2018. [cited by applicant]
Wenguanwang, Jianbing Shen, Ruigang Yang, and Fatih Porikli. Saliency-aware video object segmentation. IEEE transactions on pattern analysis and machine intelligence, 40(1):20-33, 2017. [cited by applicant]
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7794-7803, 2018. [cited by applicant]
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6281-6290… [cited by applicant]
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, et al. Relational deep reinforcement learning. arXiv preprint arXiv… [cited by applicant]
Zehua Zhang, Ashish Tawari, Sujitha Martin, and David Crandall. Interaction graphs for object importance estimation in on-road driving videos. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pp.… [cited by applicant]
Cited By (1)
US 12,462,581