IP Library Granted Patent US 12,315,187
Granted Patent B2
US 12,315,187 · App. 17/485,054 · Granted May 27, 2025

Heterogeneous multi-threaded visual odometry in autonomous driving

Inventors: Weizhao Shao (Santa Clara, CA); Ankit Kumar Jain (Mountain View, CA); Xxx Xinjilefu (Pittsburgh, PA); Gang Pan (Fremont, CA); Brendan Christopher Byrne (Detroit, MI)
Assignee: VOLKSWAGEN GROUP OF AMERICA INVESTMENTS, LLC
G06T7/73B60W40/02B60W60/00G06T1/20G06T3/40G06T7/11G06T7/55G06T9/00B60W2420/403G01S19/485G06T2207/20016G06T2207/30244G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,187
App. No.
17/485,054
Granted
May 27, 2025
Kind
B2
Abstract

A system and method for performing visual odometry is disclosed. In aspects, the system implements methods to generate an image pyramid based on an input image received. A refined pose prior information representing a location and orientation of the autonomous vehicle can be generated based on one or more images of the image pyramid. One or more seed points can be selected from the one or more images of the image pyramid. One or more refined seed representing the one or more seed points with added depth values can be generated. One or more scene points can be generated based on the one or more refined seed points. A point cloud can be generated based on the one or more scene points.

Claims (65)

1. A computer implemented method for performing visual-odometry, the method comprising:

identifying, by an on-board computing device of an autonomous vehicle (AV), a failure of the AV's connection with a guidance system;

based on the identification of the failure of the AV's connection with the guidance system:

generating, by a graphics processing unit (GPU), an image pyramid based on an input image received, wherein the input image received represents an image of an environment in which an autonomous vehicle is being operated;

generating, by the GPU, a refined pose prior information representing a location and orientation of the autonomous vehicle based on one or more images of the image pyramid;

selecting, by the GPU, one or more seed points from the one or more images of the image pyramid, the one or more seed points representing pixel locations within one or more images of the image pyramid representing estimations of where an object is likely located;

generating, by the GPU, one or more refined seed points representing the one or more seed points with added depth values;

generating, by a central processing unit (CPU), one or more scene points based on the one or more refined seed points;

generating, by the CPU, a point cloud based on the one or more scene points; and

navigating, by the on-board computing device and using the point cloud, the AV to an area of safety for it to be shut down.

2. The method of claim 1 , further comprising generating, by the GPU or the CPU, one or more first scene points and first keyframes as initial values for performing the visual-odometry, based on performing an initialization process utilizing the image pyramid.

3. The method of claim 2 , wherein the initialization process comprises:

selecting one or more first seed points from one or more first images of a first image pyramid;

generating one or more first refined seed points representing the one or more first seed points with added depth values;

generating one or more first scene points based on the one or more refined first seed points; and

generating the first keyframes based on the one or more first scene points.

4. The method of claim 1 , further comprising generating, by the CPU, a keyframe based on the input image received.

5. The method of claim 1 , further comprising decoding, by the GPU, the input image received to generate a decoded image.

6. The method of claim 5 , wherein generating the image pyramid comprises resizing the decoded image to generate a set of images, wherein the set of images comprise the images of the image pyramid.

7. The method of claim 1 , wherein selecting the one or more seed points from the one or more images of the image pyramid further comprises:

dividing the one or more images of the image pyramid into block regions; and

selecting a pixel in each block region of the block regions with the largest gradient as a seed point for each block region.

8. A non-transitory computer readable medium including instructions for causing one or more processors to perform operations for performing visual-odometry, the operations comprising:

identifying a failure of an autonomous vehicle's(AV) connection with a guidance system;

based on the identification of the failure of the AV's connection with the guidance system:

generating, by a graphics processing unit (GPU), an image pyramid based on an input image received, wherein the input image received represents an image of an environment in which an autonomous vehicle is being operated;

generating, by the GPU, a refined pose prior information representing a location and orientation of the autonomous vehicle based on one or more images of the image pyramid;

selecting, by the GPU, one or more seed points from the one or more images of the image pyramid, the one or more seed points representing pixel locations within one or more images of the image pyramid representing estimations of where an object is likely located;

generating, by the GPU, one or more refined seed points representing the one or more seed points with added depth values;

generating, by a central processing unit (CPU), one or more scene points based on the one or more refined seed points;

generating, by the CPU, a point cloud based on the one or more first scene points; and

navigating, using the point cloud, the AV to an area of safety for it to be shut down.

9. The non-transitory computer readable medium of claim 8 , wherein the operations further comprise generating, by the GPU or the CPU, one or more first scene points and first keyframes as initial values for performing the visual-odometry, based on performing an initialization process utilizing the image pyramid.

10. The non-transitory computer readable medium of claim 9 , wherein the initialization process comprises:

selecting one or more first seed points from one or more first images of a first image pyramid;

generating one or more first refined seed points representing the one or more first seed points with added depth values;

generating one or more first scene points based on the one or more refined first seed points; and

generating the first keyframes based on the one or more first scene points.

11. The non-transitory computer readable medium of claim 8 , the operations further comprising generating, by the CPU, a keyframe based on the input image received.

12. The non-transitory computer readable medium of claim 8 , the operations further comprising decoding, by the GPU, the input image received to generate a decoded image.

13. The non-transitory computer readable medium of claim 12 , wherein generating the image pyramid comprises resizing the decoded image to generate a set of images, wherein the set of images comprise the images of the image pyramid.

14. The non-transitory computer readable medium of claim 8 , wherein selecting the one or more seed points from the one or more images of the image pyramid further comprises:

dividing the one or more images of the image pyramid into block regions; and

selecting a pixel in each block region of the block regions with the largest gradient as a seed point for each block region.

15. A computing system of an autonomous vehicle (AV) for performing visual-odometry comprising:

a storage unit to store instructions;

an on-board computing device coupled to the storage unit configured to:

identifying a failure of the AV's connection with a guidance system; and

based on identifying the failure of the AV's connection with the guidance system:

generate, with a graphics processing unit (GPU), an image pyramid based on an input image received, wherein the input image received represents an image of an environment in which an autonomous vehicle is being operated,

generate, with the GPU, a refined pose prior information representing a location and orientation of the autonomous vehicle based on one or more images of the image pyramid,

select, with the GPU, one or more seed points from the one or more images of the image pyramid, the one or more seed points representing pixel locations within one or more images of the image pyramid representing estimations of where an object is likely located,

generate, with the GPU, one or more refined seed points representing the one or more seed points with added depth values, and

generate, with a central processing unit (CPU), one or more scene points based on the one or more refined seed points;

generate, with the CPU, a point cloud based on the one or more scene points; and

navigate, using the point cloud, the AV to an area of safety for it to be shut down.

16. The computing system of claim 15 , wherein the GPU or the CPU is further configured to generate one or more first scene points and first keyframes as initial values for performing the visual-odometry, based on performing an initialization process utilizing the image pyramid.

17. The computing system of claim 16 , wherein the GPU or the CPU is further configured to perform the initialization process by:

selecting one or more first seed points from one or more first images of a first image pyramid;

generating one or more first refined seed points representing the one or more first seed points with added depth values;

generating one or more first scene points based on the one or more refined first seed points; and

generating the first keyframes based on the one or more first scene points.

18. The computing system of claim 15 , wherein the CPU is further configured to generate a keyframe based on the input image received.

19. The computing system of claim 15 , wherein the GPU is further configured to decode the input image received to generate a decoded image.

20. The computing system of claim 19 , wherein the GPU is further configured to generate the image pyramid by resizing the decoded image to generate a set of images, wherein the set of images comprise the images of the image pyramid.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2024
From: ARGO AI, LLC
To: VOLKSWAGEN GROUP OF AMERICA INVESTMENTS, LLC
Reel/Frame 069113/0265 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2021
From: SHAO, WEIZHAO; JAIN, ANKIT KUMAR; XINJILEFU, XXX; PAN, GANG; BYRNE, BRENDAN CHRISTOPHER
To: ARGO AI, LLC
Reel/Frame 057607/0321 →
Continuity (1)
Related Publication 20230162386A1 · May 25, 2023
References Cited (24)
US 8542745B2 · Zhang et al. · 2013 [cited by applicant]
US 9665403B2 · Siepmann et al. · 2017 [cited by applicant]
US 9870624B1 · Narang et al. · 2018 [cited by applicant]
US 10338223B1 · Englard et al. · 2019 [cited by applicant]
US 10401866B2 · Rust · 2019 [cited by applicant]
US 10614541B2 · Storey et al. · 2020 [cited by applicant]
US 20070031064A1 · Zhao et al. · 2007 [cited by applicant]
US 20140240501A1 · Newman · 2014 [cited by examiner]
US 20180025235A1 · Fridman · 2018 [cited by examiner]
US 20180052463A1 · Mays · 2018 [cited by examiner]
US 20180341019A1 · Sakai et al. · 2018 [cited by applicant]
US 20180373564A1 · Hushchyn · 2018 [cited by applicant]
US 20190025853A1 · Julian · 2019 [cited by examiner]
US 20190061771A1 · Bier et al. · 2019 [cited by applicant]
US 20190080166A1 · Zhu et al. · 2019 [cited by applicant]
US 20190080167A1 · Zhu et al. · 2019 [cited by applicant]
US 20190080470A1 · Zhu et al. · 2019 [cited by applicant]
US 20190128677A1 · Naman · 2019 [cited by examiner]
US 20200098161A1 · Smith et al. · 2020 [cited by applicant]
US 20210011161A1 · Chen et al. · 2021 [cited by applicant]
US 20210033706A1 · Funaya · 2021 [cited by applicant]
Zhao X., et al: “A Robust Stereo Semi-direct SLAM System Based on Hybrid Pyramid”, 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, Nov. 3, 2019, pp. 5376-5382, XP033695584, DOI: 10… [cited by applicant]
Georges Y., et al: “Keyframe-based monocular SLAM: design, survey, and future directions”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jul. 2, 2016, XP081351338, DOI: 10.… [cited by applicant]
Engel et al., “Direct Sparse Odometry,” arXiv:1607.02565v2 [cs.CV] Oct. 7, 2016. [cited by applicant]