IP Library Granted Patent US 12,470,683
Granted Patent B1
US 12,470,683 · App. 18/817,540 · Granted Nov 11, 2025

Flow-guided online stereo rectification

Inventors: Felix Heide (New York City, NY); Anush Kumar (Austin, TX); Shile Li (Munich, DE); Omid Hosseini Jafari (Munich, DE); Fahim Mannan (Toronto, CA)
Assignee: Torc Robotics, Inc.
H04N13/246G06V10/44G06V10/761G06V10/771H04N13/239
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,470,683
App. No.
18/817,540
Granted
Nov 11, 2025
Kind
B1
Abstract

An autonomy computing system and a method of an autonomous vehicle for rectifying stereo images includes a memory storing computer executable instructions and a processor coupled to the memory, the processor, upon execution of the computer executable instructions, configured to: receive an image pair captured using respective cameras in the stereo camera pair; predict a rotation matrix between the first image and the second image by: extracting a first feature map and a second feature map; applying positional feature enhancement on the feature maps to derive a pair of enhanced feature maps; computing a correlation volume across the enhanced feature maps; determining a set of likely matches between the enhanced feature maps; computing a predicted relative pose; and computing the rotation matrix. The system and method further include calibrating the stereo camera pair to rectify the first image and the second image based on the rotation matrix.

Claims (72)

1 . An autonomous vehicle comprising:

a stereo camera pair disposed on the autonomous vehicle, the stereo camera pair comprising a first camera and a second camera separated by a baseline distance, the first camera and the second camera configured to capture a first image and a second image, respectively;

at least one memory device storing computer executable instructions; and

at least one processor coupled to the at least one memory device and the stereo camera pair, the least one processor, upon execution of the computer executable instructions, configured to:

receive the first image and the second image captured using respective cameras in the stereo camera pair;

predict, using a neural network model, a rotation matrix between the first image and the second image by:

extracting a first feature map and a second feature map based on the first image and the second image;

applying positional feature enhancement on the first feature map and the second feature map to derive a first enhanced feature map and a second enhanced feature map;

computing a correlation volume across the first enhanced feature map and the second enhanced feature map;

determining a set of likely matches between the first enhanced feature map and the second enhanced feature map based on the correlation volume;

computing a predicted relative pose based on the set of likely matches; and

computing the rotation matrix based on the predicted relative pose; and

calibrate the stereo camera pair by:

employing differentiable rectification to rectify the first image and the second image based on the rotation matrix.

2 . The autonomous vehicle of claim 1 wherein the first camera and the second camera are positioned laterally at least 60 cm apart.

3 . The autonomous vehicle of claim 1 , wherein the at least one processor is further configured to:

train the neural network model by a self-supervised learning.

4 . The autonomous vehicle of claim 1 , wherein the at least one processor is further configured to:

train the neural network model by optimizing a loss function including a vertical optical flow of the first image and the second image.

5 . The autonomous vehicle of claim 4 , wherein the at least one processor is further configured to train the neural network model by:

employing the differentiable rectification to the first image and the second image; and

adjusting the neural network model to minimize the vertical optical flow of the first image and the second image.

6 . The autonomous vehicle of claim 1 , wherein the at least one processor is further configured to:

calibrate the stereo camera pair by calibrating the camera pair while an autonomous vehicle equipped with the camera pair is operating.

7 . The autonomous vehicle of claim 1 , wherein the at least one processor is further configured to:

extract the first feature map and the second feature map further by:

employing the differentiable rectification to rectify the first feature map and the second feature map; and

apply the positional feature enhancement by:

applying the positional feature enhancement to the first rectified feature map and the second rectified feature map.

8 . The autonomous vehicle of claim 1 , wherein the at least one processor is further configured to:

compute the correlation volume by:

flattening the first feature map along height by width of the first feature map to produce a two-dimensional first feature map; and

flattening the second feature map along height by width of the second feature map to produce a two-dimensional second feature map.

9 . A computer-implemented method of calibrating a stereo camera pair, the method comprising:

capturing a first image and a second image using respective cameras in a stereo camera pair;

predicting a rotation matrix between the first image and the second image by:

extracting a first feature map and a second feature map based on the first image and the second image;

applying positional feature enhancement on the first feature map and the second feature map to derive a first enhanced feature map and a second enhanced feature map;

computing a correlation volume across the first enhanced feature map and the second enhanced feature map;

determining a set of likely matches between the first enhanced feature map and the second enhanced feature map based on the correlation volume;

computing a predicted relative pose based on the set of likely matches; and

computing the rotation matrix based on the predicted relative pose; and

calibrating the stereo camera pair by:

employing differentiable rectification to rectify the first image and the second image based on the rotation matrix.

10 . The method of claim 9 further comprising:

training the neural network model by a self-supervised learning.

11 . The method of claim 9 , further comprising:

training the neural network model by:

optimizing a loss function including a vertical optical flow of the first image and the second image.

12 . The method of claim 11 , wherein the training the neural network model further comprises:

employing the differentiable rectification to the first image and the second image; and

adjusting the neural network model to minimize the vertical optical flow of the first image and the second image.

13 . The method of claim 9 , wherein the calibrating the stereo camera pair further comprises:

calibrating the camera while a machine equipped with the camera pair is operating.

14 . The method of claim 9 , wherein:

the extracting the first feature map and the second feature map further comprises:

employing the differentiable rectification to rectify the first feature map and the second feature map; and

the applying the positional feature enhancement further comprises:

applying the positional feature enhancement to the first rectified feature map and the second rectified feature map.

15 . The method of claim 9 , wherein the computing the correlation volume further comprises:

flattening the first feature map along height by width of the first feature map to produce a two-dimensional first feature map; and

flattening the second feature map along height by width of the second feature map to produce a two-dimensional second feature map.

16 . The method of claim 15 , wherein the computing the correlation volume further comprises:

applying a soft-max along the last two dimensions of the correlation volume to convert the correlation volume into a likelihood of the set of likely matches.

17 . The method of claim 9 , wherein the computing the predicted relative pose further comprises:

flattening the set of likely matches to a one-dimensional list.

18 . The method of claim 9 , wherein the computing the predicted relative pose further comprises:

orthogonalizing the predicted relative pose using Gram-Schmidt orthogonalization.

19 . The method of claim 9 , wherein the determining the set of likely matches comprises:

determining a set of likely matches using a decoder layer of the neural network model.

20 . The method of claim 9 , further comprising:

training the neural network model according to a loss function: L=λ 1 L rot +λ 2 L flow , where λ 1 , λ 2 are scalar weights, L rot is a pose loss supervised on ground truth calibration data, and L flow is a self-supervised vertical-flow loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2024
From: HEIDE, FELIX; KUMAR, ANUSH; LI, SHILE; HOSSEINI JAFARI, OMID; MANNAN, FAHIM
To: TORC ROBOTICS, INC.
Reel/Frame 068423/0851 →
Continuity (1)
Provisional Application 63667026 · Jul 2, 2024
References Cited (37)
US 10832062B1 · Evans · 2020 [cited by examiner]
US 20190004533A1 · Huang · 2019 [cited by examiner]
US 20200025935A1 · Liang · 2020 [cited by examiner]
US 20200103909A1 · Feinson · 2020 [cited by examiner]
US 20210379992A1 · Domeyer · 2021 [cited by examiner]
US 20210380137A1 · Domeyer · 2021 [cited by examiner]
US 20220214457A1 · Liang · 2022 [cited by examiner]
US 20230306718A1 · Revaud · 2023 [cited by examiner]
US 20240282105A1 · Bharathwaj · 2024 [cited by examiner]
Arnold et al., “Map-free Visual Relocalization: Metric Pose Relative to a Single Image”, European Conference on Computer Vision, 2022, pp. 1-18. [cited by applicant]
Ayache et al., “Rectification of images for binocular and trinocular stereovision”, 9th International Conference on Pattern Recognition, 1988, 1, pp. 1-32. [cited by applicant]
Brachmann et al., “DSAC-Differentiable RANSAC for Camera Localization”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 6684-6692. [cited by applicant]
Chen et al., “Wide-Baseline Relative Camera Pose Estimation with Directional Learning”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 3258-3268. [cited by applicant]
Dang et al., “Continuous Stereo Self-Calibration by Camera Parameter Tracking”, IEEE Transactions on Image Processing, 2009, 18(7), pp. 1536-1550. [cited by applicant]
En et al., “RPNet: An End-to-End Network for Relative Camera Pose Estimation”, ECCV Workshops, 2018, pp. 1-8. [cited by applicant]
Fusiello et al., “Quasi-Euclidean uncalibrated epipolar rectification”, 19th International Conference on Pattern Recognition, 2008, pp. 1-4. [cited by applicant]
Georgiev et al., “A fast and accurate re-calibration technique for misaligned stereo cameras”, IEEE International Conference on Image Processing, 2013, pp. 1-5. [cited by applicant]
Gluckman et al., “Rectifying Transformations That Minimize Resampling Effects”, Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 2001, pp. 1111-1117. [cited by applicant]
Hansen et al., “Online Continuous Stereo Extrinsic Parameter Estimation”, 2012 IEEE Conference on Computer Vision and Pattern Recognition, 2012, pp. 1-8. [cited by applicant]
Hartley, “Theory and Practice of Projective Rectification”, International Journal of Computer Vision, 1999, 35, pp. 1-19. [cited by applicant]
Ji et al., “Deep View Morphing”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1-28. [cited by applicant]
Kendall et al., “PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization”, 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 2938-2946. [cited by applicant]
Kendall et al., “Geometric loss functions for camera pose regression with deep learning”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5974-5983. [cited by applicant]
Laskar et al., “Camera Relocalization by Computing Pairwise Relative Poses Using Convolutional Neural Network”, 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), 2017, pp. 929-938. [cited by applicant]
Li et al., “Practical Stereo Matching via Cascaded Recurrent Network with Adaptive Correlation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 16263-16272. [cited by applicant]
Ling et al., “High-Precision Online Markerless Stereo Extrinsic Calibration”, 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2016, pp. 1771-1778. [cited by applicant]
Loop et al., “Computing Rectifying Homographies for Stereo Vision”, 1999 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 1999, 1, pp. 125-131. [cited by applicant]
Luo et al., “Unsupervised Learning of Depth Estimation from Imperfect Rectified Stereo Laparoscopic Images”, Computers in biology and medicine, 2021, 140, pp. 1-20. [cited by applicant]
Mallon et al., “Projective Rectification from the Fundamental Matrix”, Image Vis Comput, 2005, 23, pp. 1-16. [cited by applicant]
Melekhov et al., “Relative Camera Pose Estimation Using Convolutional Neural Networks”, ArXiv, 2017, pp. 1-12. [cited by applicant]
Parameshwara et al., “DiffPoseNet: Direct Differentiable Camera Pose Estimation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 6845-6854. [cited by applicant]
Sarlin et al., “SuperGlue: Learning Feature Matching with Graph Neural Networks”, CoRR, 2019, pp. 1-17. [cited by applicant]
Wang et al., “A Practical Stereo Depth System for Smart Glasses”, ArXiv, 2022, pp. 1-11. [cited by applicant]
Wang et al., “Stereo Rectification Based on Epipolar Constrained Neural Network”, 2021 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2021, pp. 2105-2109. [cited by applicant]
Xiao et al., “DSR: Direct Self-rectification for Uncalibrated Dual-lens Cameras”, 2018 International Conference on 3D Vision (3DV), 2018, pp. 1-9. [cited by applicant]
Zhang et al., “End-to-end learning of self-rectification and self-supervised disparity prediction for stereo vision”, Neurocomputing, 2022, 494, pp. 308-319. [cited by applicant]
Zilly et al., “Joint Estimation of Epipolar Geometry and Rectification Parameters using Point Correspondences for Stereoscopic TV Sequences”, Proceedings of 3DPVT, 2010, pp. 1-7. [cited by applicant]