IP Library › Granted Patent US 12,361,732
Granted Patent B2
US 12,361,732 · App. 17/554,671 · Granted Jul 15, 2025

System and method for efficient visual navigation

Inventors: Han-Pang Chiu (West Windsor, NJ); Zachary Seymour (Pennington, NJ); Niluthpol C. Mithun (Lawrenceville, NJ); Supun Samarasekera (Skillman, NJ); Rakesh Kumar (West Windsor, NJ); Kowshik Thopalli (Tempe, AZ); Muhammad Zubair Irshad (Atlanta, GA)
Assignee: SRI International
G06V20/70G01C21/3635G06N3/02G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,732
App. No.
17/554,671
Filed
Dec 17, 2021
Granted
Jul 15, 2025
Kind
B2
Art Unit
2666
USPC
382/104
Abstract

A method, apparatus and system for efficient navigation in a navigation space includes determining semantic features and respective 3D positional information of the semantic features for scenes of captured image content and depth-related content in the navigation space, combining information of the determined semantic features of the scene with respective 3D positional information using neural networks to determine an intermediate representation of the scene which provides information regarding positions of the semantic features in the scene and spatial relationships among the semantic features, and using the information regarding the positions of the semantic features and the spatial relationships among the semantic features in a machine learning process to provide at least one of a navigation path in the navigation space, a model of the navigation space, and an explanation of a navigation action by the single, mobile agent in the navigation space.

Claims (47)

1. A method for efficient navigation of a single mobile agent in a navigation space, comprising:

determining semantic features and respective 3D positional information of the semantic features for scenes of the navigation space using captured image content of the scenes and captured depth-related content of the scenes;

combining information of the determined semantic features of at least one of the scenes with respective 3D positional information using neural networks to determine an intermediate representation of the at least one of the scenes which provides information regarding positions of the semantic features in the at least one of the scenes and spatial relationships among the semantic features;

using the information regarding the positions of the semantic features and the spatial relationships among the semantic features from the intermediate representations determined for the at least one of the scenes of the navigation space in a machine learning process to provide at least one of a navigation path in the navigation space, a model of the navigation space, and an explanation of a navigation action by the single, mobile agent in the navigation space; and

utilizing at least one region of interest in at least one scene of the captured content to identify which semantic features of the at least one scene of the captured content to include in the intermediate representation.

2. The method of claim 1 further comprising;

using the information regarding the positions of the semantic features and the spatial relationships among the semantic features from the intermediate representations to train the single, mobile agent to at least one of learn and navigate the navigation space, wherein the use of the intermediate representations to train the single, mobile agent reduces an amount of training data required to successfully train the single, mobile agent to navigate the navigation space compared to training instances in which the intermediate representations are not used.

3. The method of claim 1 , further comprising:

applying an attention mechanism to at least one of the intermediate representation of the at least one of the scenes or the at least one of the scenes of the captured content to assist in providing the at least one of a navigation path in the navigation space, the model of the navigation space, and the explanation of a navigation action by the single, mobile agent in the navigation space;

wherein a focus of the attention mechanism is based on a spatial relationship between at least one of the semantic features of the at least one of the scenes of the captured content or the semantic features in the intermediate representation and the single, mobile agent.

4. The method of claim 1 , further comprising:

training a navigation model for the single, mobile agent by:

causing sensors associated with the single, mobile agent to capture scenes while traversing multiple environments;

determining respective intermediate representations for the captured scenes; and

inputting information from the respective intermediate representations into a machine learning process for determining the navigation model.

5. The method of claim 4 , wherein the navigation model and an inference of a machine learning process are implemented to assist the single, mobile agent to at least one of learn or navigate an unknown environment.

6. The method of claim 1 , wherein the intermediate representation comprises a scene graph comprising at least two nodes representative of respective, at least two semantic features of at least one captured scene of the navigation space and at least one edge representative of a spatial relationship between the at least two semantic features of at least one captured scene of the navigation space.

7. The method of claim 1 , wherein the intermediate representation comprises a semantic map representation.

8. The method of claim 7 , wherein the semantic map representation depicts a spatial relationship between semantic features of captured scenes with respect to a position of the single mobile agent in the navigation space.

9. The method of claim 1 , further comprising identifying a spatial relationship between at least one of semantic features of a scene and positional information of the semantic features of the scene and at least a portion of language instructions provided for directing the single, mobile agent through the navigation space.

10. The method of claim 1 , further comprising separating semantic features of a scene by class and combining respective classes of semantic features with respective 3D positional information using neural networks to determine an intermediate representation of the scene categorized according to semantic class.

11. The method of claim 1 , further comprising predicting a next action for the agent in the navigation space based on a previous action taken by the agent and information from at least one determined intermediate representation.

12. A non-transitory machine-readable medium having stored thereon at least one program, the at least one program including instructions which, when executed by a processor, cause the processor to perform a method in a processor-based system for efficient navigation in a navigation space, comprising:

determining semantic features and respective 3D positional information of the semantic features for scenes of the navigation space using captured image content of the scenes and captured depth-related content of the scenes;

combining information of the determined semantic features of at least one of the scenes with respective 3D positional information using neural networks to determine an intermediate representation of the at least one of the scenes which provides information regarding positions of the semantic features in the at least one of the scenes and spatial relationships among the semantic features;

using the information regarding the positions of the semantic features and the spatial relationships among the semantic features from the intermediate representations determined for the at least one of the scenes of the navigation space in a machine learning process to provide at least one of a navigation path in the navigation space, a model of the navigation space, and an explanation of a navigation action by the single, mobile agent in the navigation space; and

utilizing at least one region of interest in at least one scene of the captured content to identify which semantic features of the at least one scene of the captured content to include in the intermediate representation.

13. The non-transitory machine-readable medium of claim 12 , further comprising:

using the information regarding the positions of the semantic features and the spatial relationships among the semantic features from the intermediate representations in a deep reinforced learning process to train the single, mobile agent to at least one of learn and navigate the navigation space, wherein the use of the intermediate representations to train the single, mobile agent reduces an amount of training data required to successfully train the single, mobile agent to navigate the navigation space compared to training instances in which the intermediate representations are not used.

14. The non-transitory machine-readable medium of claim 12 , further comprising:

training a navigation model for the agent by:

causing sensors associated with the single, mobile agent to capture scenes while traversing multiple environments;

determining respective intermediate representations for the captured scenes; and

inputting information from the respective intermediate representations into a machine learning process for determining the navigation model, wherein, the navigation model and an inference of a machine learning process are implemented to assist the single, mobile agent to at least one of learn or navigate an unknown environment.

15. The non-transitory machine-readable medium of claim 12 , further comprising identifying a spatial relationship between at least one of semantic features of a scene and positional information of the semantic features of the scene and at least a portion of language instructions provided for directing the single, mobile agent through the navigation space.

16. A system for efficient navigation in a navigation space, comprising:

a processor; and

a memory coupled to the processor, the memory having stored therein at least one of programs or instructions executable by the processor to configure the system to:

determine semantic features and respective 3D positional information of the semantic features for scenes of the navigation space using captured image content of the scenes and captured depth-related content of the scenes;

combine information of the determined semantic features of at least one of the scenes with respective 3D positional information using neural networks to determine an intermediate representation of the at least one of the scenes which provides information regarding positions of the semantic features in the at least one of the scenes and spatial relationships among the semantic features;

use the information regarding the positions of the semantic features and the spatial relationships among the semantic features from the intermediate representations determined for the at least one of the scenes of the navigation space in a machine learning process to provide at least one of a navigation path in the navigation space, a model of the navigation space, and an explanation of a navigation action by the single, mobile agent in the navigation space; and

utilize at least one region of interest in at least one scene of the captured content to identify which semantic features of the at least one scene of the captured content to include in the intermediate representation.

17. The system of claim 16 , wherein the system is further configured to use the information regarding the positions of the semantic features and the spatial relationships among the semantic features from the intermediate representations to train the single, mobile agent to at least one of learn and navigate the navigation space, wherein the use of the intermediate representations to train the single, mobile agent reduces an amount of training data required to successfully train the single, mobile agent to navigate the navigation space compared to training instances in which the intermediate representations are not used.

18. The system of claim 16 , wherein the intermediate representation comprises at least one of a scene graph comprising at least two nodes representative of respective, at least two semantic features of at least one captured scene of the navigation space and at least one edge representative of a spatial relationship between the at least two semantic features of at least one captured scene of the navigation space and a semantic map representation which depicts a spatial relationship between semantic features of captured scenes with respect to a position of the single mobile agent in the navigation space.

19. The system of claim 16 , wherein the system is further configured to identify a spatial relationship between at least one of semantic features of a scene and positional information of the semantic features of the scene and at least a portion of language instructions provided for directing the single, mobile agent through the navigation space.

20. The system of claim 16 , wherein the system is further configured to separate semantic features of a scene by class and combining respective classes of semantic features with respective 3D positional information using neural networks to determine an intermediate representation of the scene categorized according to semantic class.

21. The system of claim 16 , wherein the system is further configured to train a navigation model for the agent by causing sensors associated with the single, mobile agent to capture scenes while traversing multiple environments, determining respective intermediate representations for the captured scenes, and inputting the information from the respective intermediate representations into a machine learning process for determining the navigation model, wherein, the navigation model and an inference of a machine learning process are implemented to assist the single, mobile agent to at least one of learn and navigate an unknown environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: CHIU, HAN-PANG; SEYMOUR, ZACHARY; MITHUN, NILUTHPOL C.; SAMARASEKERA, SUPUN; KUMAR, RAKESH; THOPALLI, KOWSHIK; IRSHAD, MUHAMMAD ZUBAIR
To: SRI INTERNATIONAL
Reel/Frame 058489/0487 →
Continuity (2)
Provisional Application 63126981 · Dec 17, 2020
Related Publication 20220198813A1 · Jun 23, 2022
References Cited (28)
US 11830253B2 · Tang · 2023 [cited by examiner]
US 20180161986A1 · Kee · 2018 [cited by examiner]
US 20210141383A1 · Silander · 2021 [cited by examiner]
US 20210248375A1 · Geng · 2021 [cited by examiner]
US 20210276595A1 · Casas · 2021 [cited by examiner]
US 20220111869A1 · Liu · 2022 [cited by examiner]
WO WO2021058090A1 · 2021 [cited by examiner]
Lydia E Kavraki, Petr Svestka, J-C Latombe, and Mark H Overmars. Probabilistic roadmaps for path planning in high-dimensional configuration spaces. IEEE transactions on Robotics and Automation, 12(4):566-580, 1996. [cited by applicant]
Andrew J Davison and David W Murray. Mobile robot localisation using active vision. In ECCV, pp. 809-825, 1998. [cited by applicant]
Sebastian Thrun, Maren Bennewitz, Wolfram Burgard, Armin B Cremers, Frank Dellaert, Dieter Fox, Dirk Hahnel, Charles Rosenberg, Nicholas Roy, Jamieson Schulte, et al. Minerva: A second-generation museum tour-guide robot… [cited by applicant]
Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. In CVPR, pp. 1534-1543, 2016. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pp. 770-778, 2016. [cited by applicant]
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu. Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397, 2016. [cited by applicant]
Manolis Savva, Angel X. Chang, Alexey Dosovitskiy, Thomas A. Funkhouser, and Vladlen Koltun. Minos: Multimodal indoor simulator for navigation in complex environments. arXiv preprint arXiv:1712.03931, 2017. [cited by applicant]
Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik. Cognitive mapping and planning for visual navigation. In CVPR, pp. 7272-7281, 2017. [cited by applicant]
Saurabh Gupta, David Fouhey, Sergey Levine, and Jitendra Malik. Unifying map and landmark based representations for visual navigation. arXiv preprint arXiv:1712.08125, 2017. [cited by applicant]
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In CVPR, pp. 1746-1754, 2017. [cited by applicant]
Claudia Yan, Dipendra Misra, Andrew Bennnett, Aaron Walsman, Yonatan Bisk, and Yoav Artzi. Chalet: Cornell house agent learning environment. arXiv preprint arXiv:1801.07357, 2018. [cited by applicant]
Fei Xia, Amir Roshan Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson Env: Real-world perception for embodied agents. In CVPR, pp. 9068-9079, 2018. [cited by applicant]
Peter Anderson, Angel X. Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, and Amir Roshan Zamir. On evaluation of embodied … [cited by applicant]
Tianyi Wu, Sheng Tang, Rui Zhang, and Yongdong Zhang. Cgnet: A light-weight context guided network for semantic segmentation. arXiv preprint arXiv: 1811.08201, 2018. [cited by applicant]
Dmytro Mishkin, Alexey Dosovitskiy, and Vladlen Koltun. Benchmarking classic and learned navigation in complex 3d environments. arXiv preprint arXiv:1901.10915, 2019. [cited by applicant]
Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, and Dhruv Batra. Decentralized distributed ppo: Solving pointgoal navigation. arXiv preprint arXiv:1911.00357, 2019. [cited by applicant]
Federico Landi, Lorenzo Baraldi, Massimiliano Corsini, and Rita Cucchiara. Embodied vision-and-language navigation with dynamic convolutional filters. In BMVC, 2019. [cited by applicant]
William B. Shen, Danfei Xu, Yuke Zhu, Leonidas J. Guibas, Fei-Fei Li, and Silvio Savarese. Situational fusion of visual representation for visual navigation. In ICCV, 2019. [cited by applicant]
Wei Yang, XiaolongWang, Ali Farhadi, Abhinav Gupta, and Roozbeh Mottaghi. Visual semantic navigation using scene priors. arXiv preprint arXiv:1810.06543, 2018. [cited by applicant]
Han Zhang, Ian J. Goodfellow, Dimitris N. Metaxas, and Augustus Odena. Self-attention generative adversarial networks. In ICML, vol. 97 of Proceedings of Machine Learning Research, pp. 7354-7363. PMLR, 2019. [cited by applicant]
Guilherme N DeSouza and Avinash C Kak. Vision for mobile robot navigation: A survey. IEEE transactions on pattern analysis and machine intelligence, 24(2):237-267, 2002. [cited by applicant]