IP Library Granted Patent US 12,450,774
Granted Patent B2
US 12,450,774 · App. 18/324,398 · Granted Oct 21, 2025

Visual odometry for operating a movable device

Inventors: Shenbagaraj Kannapiran (Tempe, AZ); Nalin Bendapudi (Ann Arbor, MI); Devarth Parikh (Ann Arbor, MI); Ankit Girish Vora (Northville, MI)
Assignee: Ford Global Technologies, LLC
G06T7/74G05D1/0251G06T7/174G06T2207/20084G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,774
App. No.
18/324,398
Granted
Oct 21, 2025
Kind
B2
Abstract

A computer that includes a processor and a memory, the memory including instructions executable by the processor to select first side stereo images of first and second pairs of stereo images acquired at first and second time steps, respectively, and mask first and second first side stereo images, determine point features in masked first and second first side stereo images and determine line features in masked first and second first side stereo images. Matching point features in the masked first and second first side stereo images can be determined using a first attention/graph neural network (attn/GNN) based on keypoints determined based on the line features. Matching line features in the masked first and second first side stereo images can be determined using a second attn/GNN based on keypoints determined based on the line features. Three-dimensional (3D) locations in a scene can be determined by determining stereo disparity based on the matched point features included in the first first side stereo image and point features determined in a first second side stereo image of first and second stereo pairs of images and a 3D stereo camera pose can be determined by determining a perspective-n-point and line algorithm on the 3D locations.

Claims (37)

1. A system, comprising:

a computer that includes a processor and a memory, the memory including instructions executable by the processor to:

select first side first and second stereo images of first and second pairs of stereo images acquired at first and second time steps by a first side camera of a stereo camera, respectively;

mask first and second first side stereo images;

determine point features in the masked first and second first side stereo images;

determine line features in the masked first and second first side stereo images;

determine matching point features in the masked first and second first side stereo images using a first attention/graph neural network (attn/GNN) based on keypoints determined based on the line features;

determine matching line features in the masked first and second first side stereo images using a second attn/GNN based on keypoints determined based on the line features;

determine three-dimensional (3D) locations in a scene by determining stereo disparity based on the matched point features included in the first first side stereo image and point features determined in a first second side stereo image of first and second pairs of stereo images; and

determine a 3D stereo camera pose by determining a perspective-n-points-and-line algorithm on the 3D locations.

2. The system of claim 1 , the instructions including further instructions to mask the first and second first side stereo images by determining first and second segmented images for the first and second first side stereo images by an image segmentor and masking the first and second first side stereo images based on the first and second segmented images, respectively.

3. The system of claim 1 , wherein masking the first and second first side stereo images removes dynamic objects and sky.

4. The system of claim 1 , the instructions including further instructions to determine the point features in the masked first and second first side stereo images by self-supervised machine learning using the first attn/GNN followed by a differentiable optimal transport processor.

5. The system of claim 1 , the instructions including further instructions to determine the line features in the masked first and second first side stereo images by self-supervised machine learning using the second attn/GNN followed by a differentiable optimal transport processor.

6. The system of claim 1 , the instructions including further instructions to determine the point features based on a Superpoint algorithm.

7. The system of claim 1 , the instructions including further instructions to determine the line features based on a SOLD2 algorithm.

8. The system of claim 1 , the instructions including further instructions to determine the keypoints by sampling the line features determined in the masked first and second first side stereo images.

9. The system of claim 1 , the instructions including further instructions to determine point features in the first second side stereo image of first and second pairs of stereo images based on correlating pixel neighborhoods surrounding point features in the first left stereo image.

10. The system of claim 1 , the instructions including further instructions to determine stereo disparity by determining a distance between point features in the first second side stereo image and the masked first first side stereo image and based on one or more of a baseline between stereo cameras, a focal distance, and a pixel scale.

11. The system of claim 1 , the instructions including further instructions to combine the 3D stereo camera pose with one or more of global navigation satellite system (GNSS) data, or inertial measurement unit (IMU) to determine a 3D vehicle pose based on an extended Kalman filter.

12. The system of claim 11 , the instructions including further instructions to determine a path polynomial based on the 3D vehicle pose, and operate a vehicle based on the path polynomial.

13. A method, comprising:

selecting first side first and second stereo images of first and second pairs of stereo images acquired at first and second time steps, respectively;

masking first and second first side stereo images;

determining point features in masked first and second first side stereo images;

determining line features in masked first and second first side stereo images;

determining matching point features in the masked first and second first side stereo images using a first attention/graph neural network (attn/GNN) based on keypoints determined based on the line features;

determining matching line features in the masked first and second first side stereo images using a second attn/GNN based on keypoints determined based on the line features;

determining three-dimensional (3D) locations in a scene by determining stereo disparity based on the matched point features included in the first first side stereo image and point features determined in a first second side stereo image of first and second pairs of stereo images; and

determining a 3D stereo camera pose by determining a perspective-n-points-and-line algorithm on the 3D locations.

14. The method of claim 13 , further comprising masking the first and second first side stereo images by determining first and second segmented images for the first and second first side stereo images by an image segmentor and masking the first and second first side stereo images based on the first and second segmented images, respectively.

15. The method of claim 13 , wherein masking the first and second first side stereo images removes dynamic objects and sky.

16. The method of claim 13 , further comprising determining the point features in the masked first and second first side stereo images by self-supervised machine learning using the first attn/GNN followed by a differentiable optimal transport processor.

17. The method of claim 13 , further comprising determining the line features in the masked first and second first side stereo images by self-supervised machine learning using the second attn/GNN followed by a differentiable optimal transport processor.

18. The method of claim 13 , further comprising determining the point features based on a Superpoint algorithm.

19. The method of claim 13 , further comprising determining the line features based on a SOLD2 algorithm.

20. The method of claim 13 , further comprising determining the keypoints by sampling the line features determined in the masked first and second first side stereo images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: KANNAPIRAN, SHENBAGARAJ; BENDAPUDI, NALIN; PARIKH, DEVARTH; VORA, ANKIT GIRISH
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 063773/0370 →
Continuity (1)
Related Publication 20240394915A1 · Nov 28, 2024
References Cited (23)
US 9329598B2 · Pack et al. · 2016 [cited by applicant]
US 10096129B2 · Narang et al. · 2018 [cited by applicant]
US 11270148B2 · Li et al. · 2022 [cited by applicant]
US 20170359561A1 · Vallespi-Gonzalez · 2017 [cited by examiner]
US 20230316571A1 · Kadambi · 2023 [cited by examiner]
CN 108682027A · 2018 [cited by applicant]
CN 110375732A · 2019 [cited by applicant]
CN 112649016A · 2021 [cited by applicant]
Zhao, Wanqing et al. “Learning Deep Network for Detecting 3D Object Keypoints and 6D Poses.” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2020. 14122-14130. Web. (Year: 2020). [cited by examiner]
Zhang et al., “DynPL-SVO: A New Method Using Point and Line Features for Stereo Visual Odometry in Dynamic Scenes”, arXiv: 2205.08207v2 [cs. CV] Sep. 29, 2022. (Year: 2022). [cited by examiner]
Sarlin, Paul-Edouard et al. “SuperGlue: Learning Feature Matching With Graph Neural Networks.” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2020. 4937-4946. Web. (Year: 2020). [cited by examiner]
Agostinho, Sérgio, João Gomes, and Alessio Del Bue. “CvxPnPL: A Unified Convex Solution to the Absolute Pose Estimation Problem from Point and Line Correspondences.” Journal of mathematical imaging and vision 65.3 (2023… [cited by examiner]
Agostinho et al., CvxPnPL: A Unified Convex Solution to the Absolute Pose Estimation Problem from Point and Line Correspondences, arXiv:1907.10545v2 [cs.CV] Aug. 9, 2019. [cited by applicant]
Bay et al., “SURF: Speeded Up Robust Features”, https://link.springer.com/chapter/10.1007/11744023_32. [cited by applicant]
DeTone et al., “SuperPoint: Self-Supervised Interest Point Detection and Description”, This CVPR workshop paper is the Open Access version, provided by the Computer Vision Foundation. It is identical to the version avai… [cited by applicant]
Dosovitskiy et al., “CARLA: An Open Urban Driving Simulator”, arXiv:1711.03938v1 [cs.LG] Nov. 10, 2017. [cited by applicant]
Lowe, “Distinctive Image Features from Scale-Invariant Keypoints”, Accepted for publication in the International Journal of Computer Vision, 2004. [cited by applicant]
Luo et al., “Accurate Line Reconstruction for Point and Line-Based Stereo Visual Odometry”, Received Dec. 4, 2019, accepted Dec. 16, 2019, date of publication Dec. 19, 2019, date of current version Dec. 31, 2019. Digita… [cited by applicant]
Pautrat et al., “SOLD2: Self-supervised Occlusion-aware Line Description and Detection”, This CVPR 2021 paper is the Open Access version, provided by the Computer Vision Foundation. It is identical to the accepted versi… [cited by applicant]
Rublee et al., “ORB: an efficient alternative to SIFT or SURF”, https://ieeexplore.ieee.org/document/6126544. [cited by applicant]
Sarlin et al., “SuperGlue: Learning Feature Matching with Graph Neural Networks”, This CVPR 2020 paper is the Open Access version, provided by the Computer Vision Foundation. It is identical to the accepted version; the… [cited by applicant]
Yi et al., “LIFT: Learned Invariant Feature Transform”, https://www.researchgate.net/publication/308277668, conference Paper ⋅Oct. 2016 DOI: 10.1007/978-3-319-46466-4_28. [cited by applicant]
Zhang et al., “DynPL-SVO: A New Method Using Point and Line Features for Stereo Visual Odometry in Dynamic Scenes”, arXiv:2205.08207v2 [cs.CV] Sep. 29, 2022. [cited by applicant]
Cited By (1)
US 12,694,566