IP Library Granted Patent US 12,205,324
Granted Patent B2
US 12,205,324 · App. 17/519,894 · Granted Jan 21, 2025

Learning to fuse geometrical and CNN relative camera pose via uncertainty

Inventors: Bingbing Zhuang (San Jose, CA); Manmohan Chandraker (Santa Clara, CA)
Assignee: NEC Corporation
G06T7/75G06T1/0014G06T7/77G06T2207/20076G06T2207/20081G06T2207/20084G06T2207/20221G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,324
App. No.
17/519,894
Granted
Jan 21, 2025
Kind
B2
Abstract

A computer-implemented method for fusing geometrical and Convolutional Neural Network (CNN) relative camera pose is provided. The method includes receiving two images having different camera poses. The method further includes inputting the two images into a geometric solver branch to return, as a first solution, an estimated camera pose and an associated pose uncertainty value determined from a Jacobian of a reproduction error function. The method also includes inputting the two images into a CNN branch to return, as a second solution, a predicted camera pose and an associated pose uncertainty value. The method additionally includes fusing, by a processor device, the first solution and the second solution in a probabilistic manner using Bayes' rule to obtain a fused pose.

Claims (34)

1. A computer-implemented method for fusing geometrical and Convolutional Neural Network (CNN) relative camera pose, comprising:

receiving two images having different camera poses;

inputting the two images into a geometric solver branch to return, as a first solution, an estimated camera pose and an associated pose uncertainty value determined from a Jacobian of a reproduction error function;

inputting the two images into a CNN branch, having a pose branch multi-layer perceptron (MLP) and an uncertainty branch MLP, to return, as a second solution, a predicted camera pose by the pose branch MLP and an associated pose uncertainty value predicted by the uncertainty branch MLP by extracting appearance features with a feature extractor and geometric features with an attentional graph neural network and a geometric feature MLP; and

fusing, by a processor device, the first solution and the second solution in a probabilistic manner based on the uncertainty predictions using Bayes' rule to obtain a fused pose.

2. The computer-implemented method of claim 1 , wherein the geometric solver branch comprises a correspondence point minimal solver.

3. The computer-implemented method of claim 1 , wherein the geometric solver branch performs a Bundle Adjustment process to nonlinearly minimize a reprojection error.

4. The computer-implemented method of claim 1 , wherein the CNN branch parameterizes a translation direction of the predicted camera pose by an azimuth angle and an elevation angle.

5. The computer-implemented method of claim 1 , wherein the predicted camera pose and the pose uncertainty value are interpreted as a mean and a variance, respectively, of an underlying Gaussian distribution.

6. The computer-implemented method of claim 5 , further comprising concatenating the appearance and the geometric features extracted from the two images prior to predicting the mean and inverse variance of the underlying Gaussian distribution.

7. The computer-implemented method of claim 1 , wherein said fusing step uses circular fusion based on parameters of a circle.

8. The computer-implemented method of claim 1 , wherein said fusing step comprising performing a weighted averaging of the first solution and the second solution based on parameters of a circle.

9. The computer-implemented method of claim 1 , further comprising performing robotic movement control using the fused pose.

10. A computer program product for fusing geometrical and Convolutional Neural Network (CNN) relative camera pose, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

receiving, by a processor device of the computer, two images having different camera poses;

inputting, by the processor device, the two images into a geometric solver branch to return, as a first solution, an estimated camera pose and an associated pose uncertainty value determined from a Jacobian of a reproduction error function;

inputting, by the processor device, the two images into a CNN branch, having a pose branch multi-layer perceptron (MLP) and an uncertainty branch MLP, to return, as a second solution, a predicted camera pose by the pose branch MLP and an associated pose uncertainty value predicted by the uncertainty branch MLP by extracting appearance features with a feature extractor and geometric features with an attentional graph neural network and a geometric feature MLP; and

fusing, by the processor device, the first solution and the second solution in a probabilistic manner based on the uncertainty predictions using Bayes' rule to obtain a fused pose.

11. The computer program product of claim 10 , wherein the geometric solver branch comprises a correspondence point minimal solver.

12. The computer program product of claim 10 , wherein the geometric solver branch performs a Bundle Adjustment process to nonlinearly minimize a reprojection error.

13. The computer program product of claim 10 , wherein the CNN branch parameterizes a translation direction of the predicted camera pose by an azimuth angle and an elevation angle.

14. The computer program product of claim 13 , further comprising concatenating the appearance and the geometric features extracted from the two images prior to predicting the mean and inverse variance of the underlying Gaussian distribution.

15. The computer program product of claim 10 , wherein the predicted camera pose and the pose uncertainty value are interpreted as a mean and a variance, respectively, of an underlying Gaussian distribution.

16. The computer program product of claim 10 , wherein said fusing step uses circular fusion based on parameters of a circle.

17. The computer program product of claim 10 , wherein said fusing step comprising performing a weighted averaging of the first solution and the second solution based on parameters of a circle.

18. The computer program product of claim 10 , wherein the method further comprises performing robotic movement control using the fused pose.

19. A computer processing system for fusing geometrical and Convolutional Neural Network (CNN) relative camera pose, comprising:

a memory device for storing program code; and

a processor device operatively coupled to the memory device for running the program code to:

receive two images having different camera poses;

input the two images into a geometric solver branch to return, as a first solution, an estimated camera pose and an associated pose uncertainty value determined from a Jacobian of a reproduction error function;

input the two images into a CNN branch, having a pose branch multi-layer perceptron (MLP) and an uncertainty branch MLP, to return, as a second solution, a predicted camera pose by the pose branch MLP and an associated pose uncertainty value predicted by the uncertainty branch MLP by extracting appearance features with a feature extractor and geometric features with an attentional graph neural network and a geometric feature MLP; and

fuse the first solution and the second solution in a probabilistic manner based on the uncertainty predictions using Bayes' rule to obtain a fused pose.

20. The computer processing system of claim 19 , wherein the geometric solver branch comprises a correspondence point minimal solver.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069540/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2021
From: ZHUANG, BINGBING; CHANDRAKER, MANMOHAN
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 058030/0906 →
Continuity (3)
Provisional Application 63113961 · Nov 15, 2020
Provisional Application 63111274 · Nov 9, 2020
Related Publication 20220148220A1 · May 12, 2022
References Cited (14)
US 20190279366A1 · Sick · 2019 [cited by examiner]
US 20190361460A1 · Medeiros · 2019 [cited by examiner]
US 20200013188A1 · Nakashima · 2020 [cited by examiner]
US 20200106942A1 · Jiang · 2020 [cited by examiner]
US 20200273190A1 · Ye · 2020 [cited by examiner]
Laidlow et al., Towards the Probabilistic Fusion of Learned Priors into Standard Pipelines for 3D Reconstruction, 2020 (Year: 2020). [cited by examiner]
Hedborg et al., Towards the Probabilistic Fusion of Learned Priors into Standard Pipelines for 3D Reconstruction, 2012 (Year: 2012). [cited by examiner]
Kendall et al., Modelling Uncertainty in Deep Learning for Camera Relocalization, 2016 (Year: 2016). [cited by examiner]
Klodt, Maria, et al. “Supervising the new with the old: learning SFM from SFM”, InProceedings of the European Conference on Computer Vision (ECCV). Sep. 2018, pp. 698-713. [cited by applicant]
Nister, David, et al. “An efficient solution to the five-point relative pose problem”, IEEE transactions on pattern analysis and machine intelligence, vol. 26, No. 6. Jun. 2004, pp. 756-770. [cited by applicant]
Y. Nakajima et al., “Fast and Accurate Semantic Mapping Through Geometric-based Incremental Segmentation”, 2018 IEEE/RSJ International Confernce on Intelligent Robots and System (IROS), Jan. 7, 2019, pp. 385-388. [cited by applicant]
T.D. Barfoot et al., “Associating Uncertainty With Three-Dimensional Poses for Use in Estimation Problems”, IEEE Transactions on Robotics, vol. 30, issue 3, Jan. 28, 2014, pp. 682, 686, 688, 690. [cited by applicant]
A. Vakhitov et al., “Stereo Relative Pose From Line and Point Feature Triplets”, arXiv:1907.00276v1, Jun. 2019, pp. 1, 10. [cited by applicant]
G. Stienne et al., “A multi-temporal multi-sensor circular fusion filter”, Information Fusion 18, pp. 86-100, Jun. 2013, pp. 93, 97-98. [cited by applicant]