IP Library Granted Patent US 12,650,598
Granted Patent B2
US 12,650,598 · App. 18/417,523 · Granted Jun 9, 2026

Systems and methods for performing self-improving visual odometry

Inventors: Daniel DeTone (San Francisco, CA); Tomasz Jan Malisiewicz (Mountain View, CA); Andrew Rabinovich (San Francisco, CA)
Assignee: Magic Leap, Inc.
G02B27/0172G06N3/08G06T7/33G06T7/74G06V10/40G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,650,598
App. No.
18/417,523
Granted
Jun 9, 2026
Kind
B2
Abstract

In an example method of training a neural network for performing visual odometry, the neural network receives a plurality of images of an environment, determines, for each image, a respective set of interest points and a respective descriptor, and determines a correspondence between the plurality of images. Determining the correspondence includes determining one or point correspondences between the sets of interest points, and determining a set of candidate interest points based on the one or more point correspondences, each candidate interest point indicating a respective feature in the environment in three-dimensional space). The neural network determines, for each candidate interest point, a respective stability metric and a respective stability metric. The neural network is modified based on the one or more candidate interest points.

Claims (82)

1 . A method of training a neural network for performing visual odometry, the method comprising:

accessing, by the neural network implemented using one or more computer systems, a plurality of images of an environment;

determining, by the neural network based on the plurality of images, a plurality of three-dimensional points in the environment;

tracking, by the neural network, the plurality of three-dimensional points across the plurality of images;

selecting a subset of the three-dimensional points, wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:

determining a stability metric of the three-dimensional point based on (i) a number of images of the plurality of images that depict the three-dimensional point, and (ii) a re-projection error associated with the three-dimensional point, and

determining whether to select the three-dimensional point based on the stability metric; and

modifying the neural network based on the subset of the three-dimensional points,

wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:

classifying, based on the stability metric, the three-dimensional point as one of:

a first classification representing stable points,

a second classification representing unstable points, or

a third classification other than the first classification and the second classification,

wherein classifying the three-dimensional point as the first classification comprises:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to a threshold number, and

determining that the re-projection error associated with the three-dimensional point is less than or equal to a first threshold error level, and

wherein classifying the three-dimensional point as the second classification comprises:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to the threshold number, and

determining that the re-projection error associated with the three-dimensional point is greater than or equal to a second threshold error level, wherein the second threshold error level is different from the first threshold error level.

2 . The method of claim 1 , wherein selecting the subset of the three-dimensional points comprises:

selecting the three-dimensional points having the first classification or the second classification.

3 . The method of claim 2 , wherein selecting the subset of the three-dimensional points comprises:

refraining from selecting the three-dimensional points of the plurality of three-dimensional points having the third classification.

4 . The method of claim 1 , wherein classifying the three-dimensional point as the third classification comprises at least one of:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is less than the threshold number, or

determining that the re-projection error associated with three-dimensional point is between the first threshold error level and the second threshold error level.

5 . The method of claim 1 , wherein the plurality of images comprise two-dimensional images extracted from a video sequence.

6 . The method of claim 5 , wherein the plurality of images correspond to non-contiguous frames of the video sequence.

7 . The method of claim 1 , further comprising:

subsequent to modifying the neural network, receiving, by the neural network, a second plurality of images of a second environment from a head-mounted display device; and

determining, by the neural network based on the second plurality of images, a second plurality of three-dimensional points in the second environment.

8 . The method of claim 7 , wherein performing visual odometry with respect to the second environment comprises determining a position and orientation of the head-mounted display device using the second plurality of three-dimensional points as landmarks.

9 . A system comprising:

one or more processors;

one or more non-transitory computer-readable media including one or more sequences of instructions which, when executed by the one or more processors, causes the one or more processors to perform operations comprising:

accessing, by a neural network, a plurality of images of an environment;

determining, by the neural network based on the plurality of images, a plurality of three-dimensional points in the environment;

tracking, by the neural network, the plurality of three-dimensional points across the plurality of images;

selecting a subset of the three-dimensional points, wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:

determining a stability metric of the three-dimensional point based on (i) a number of images of the plurality of images that depict the three-dimensional point, and (ii) a re-projection error associated with the three-dimensional point, and

determining whether to select the three-dimensional point based on the stability metric; and

modifying the neural network based on the subset of the three-dimensional points,

wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:

classifying, based on the stability metric, the three-dimensional point as one of:

a first classification representing stable points,

a second classification representing unstable points, or

a third classification other than the first classification and the second classification,

wherein classifying the three-dimensional point as the first classification comprises:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to a threshold number, and

determining that the re-projection error associated with the three-dimensional point is less than or equal to a first threshold error level,

wherein classifying the three-dimensional point as the second classification comprises:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to the threshold number, and

determining that the re-projection error associated with the three-dimensional point is greater than or equal to a second threshold error level, wherein the second threshold error level is different from the first threshold error level.

10 . The system of claim 9 , wherein selecting the subset of the three-dimensional points comprises:

selecting the three-dimensional points having the first classification or the second classification.

11 . The system of claim 10 , wherein selecting the subset of the three-dimensional points comprises:

refraining from selecting the three-dimensional points of the plurality of three-dimensional points having the third classification.

12 . The system of claim 9 , wherein classifying the three-dimensional point as the third classification comprises at least one of:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is less than the threshold number, or

determining that the re-projection error associated with three-dimensional point is between the first threshold error level and the second threshold error level.

13 . The system of claim 9 , further comprising:

subsequent to modifying the neural network, receiving, by the neural network, a second plurality of images of a second environment from a head-mounted display device; and

determining, by the neural network based on the second plurality of images, a second plurality of three-dimensional points in the second environment.

14 . One or more non-transitory computer-readable media including one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform operations comprising:

accessing, by a neural network, a plurality of images of an environment;

determining, by the neural network based on the plurality of images, a plurality of three-dimensional points in the environment;

tracking, by the neural network, the plurality of three-dimensional points across the plurality of images;

selecting a subset of the three-dimensional points, wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:

determining a stability metric of the three-dimensional point based on (i) a number of images of the plurality of images that depict the three-dimensional point, and (ii) a re-projection error associated with the three-dimensional point, and

determining whether to select the three-dimensional point based on the stability metric; and

modifying the neural network based on the subset of the three-dimensional points,

wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:

classifying, based on the stability metric, the three-dimensional point as one of:

a first classification representing stable points,

a second classification representing unstable points, or

a third classification other than the first classification and the second classification,

wherein classifying the three-dimensional point as the first classification comprises:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to a threshold number, and

determining that the re-projection error associated with the three-dimensional point is less than or equal to a first threshold error level, and

wherein classifying the three-dimensional point as the second classification comprises:

determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to the threshold number, and

determining that the re-projection error associated with the three-dimensional point is greater than or equal to a second threshold error level, wherein the second threshold error level is different from the first threshold error level.

Assignments (3)
SECURITY INTEREST Recorded Oct 28, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073388/0027 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2024
From: DETONE, DANIEL; RABINOVICH, ANDREW
To: MAGIC LEAP, INC.
Reel/Frame 066262/0954 →
EMPLOYMENT AGREEMENT Recorded Jan 26, 2024
From: MALISIEWICZ, TOMASZ
To: MAGIC LEAP, INC.
Reel/Frame 066374/0163 →
Continuity (4)
Continuation 17293772
Provisional Application 62913378 · Oct 10, 2019
Provisional Application 62767887 · Nov 15, 2018
Related Publication 20240231102A1 · Jul 11, 2024
References Cited (35)
US 20160012311A1 · Romanik et al. · 2016 [cited by applicant]
US 20160180510A1 · Grau · 2016 [cited by applicant]
US 20160196659A1 · Vrcelj et al. · 2016 [cited by applicant]
US 20180012411A1 · Richey et al. · 2018 [cited by applicant]
US 20180039745A1 · Chevalier et al. · 2018 [cited by applicant]
US 20180053056A1 · Rabinovich et al. · 2018 [cited by applicant]
US 20180129907A1 · Saklatvala · 2018 [cited by applicant]
US 20180211401A1 · Lee et al. · 2018 [cited by applicant]
US 20180268256A1 · Di Febbo et al. · 2018 [cited by applicant]
US 20180293738A1 · Yang et al. · 2018 [cited by applicant]
US 20190026956A1 · Gausebeck · 2019 [cited by examiner]
US 20210174539A1 · Duong · 2021 [cited by examiner]
US 20220028110A1 · Detone · 2022 [cited by examiner]
CN 106708037 · 2017 [cited by applicant]
CN 107924579 · 2018 [cited by applicant]
CN 107958460A · 2018 [cited by applicant]
WO WO2017143239 · 2017 [cited by applicant]
WO WO2017168899 · 2017 [cited by applicant]
WO WO2018125812 · 2018 [cited by applicant]
WO WO2018138782 · 2018 [cited by applicant]
DeTone et al., “Self-improving visual odometry,” arXiv preprint, Dec. 8, 2018, arXiv: 1812.03245v1, 9 pages. [cited by applicant]
Office Action in European Appln. No. 19885433.3, dated Sep. 17, 2024, 7 pages. [cited by applicant]
Cieslewski et al., “SIPS: Unsupervised Succinct Interest Points,” cs.CV, Computer Vision and Pattern Recognition, Submitted on May 3, 2018, arXiv:1805.01358v1, 23 pages. [cited by applicant]
Cieslewski et al., “SIPS: unsupervised succinct interest points,” Paper, Presented at Proceedings of the 2019 IEEE International Conference on 3D Vision (3DV), Quebec City, Canada, Sep. 16-19, 2019, 23 pages. [cited by applicant]
DeTone et al., “SuperPoint: Self-Supervised Interest Point Detection and Description,” Paper, Presented at Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt L… [cited by applicant]
DeTone et al., “Toward Geometric Deep SLAM,” cs.CV, Computer Vision and Pattern Recognition, Submitted Jul. 24, 2017, arXiv:1707.07410v1, 14 pages. [cited by applicant]
Extended European Search Report in European Appln. No. 19885433.3, dated Jul. 29, 2022, 12 pages. [cited by applicant]
Han et al., “A Cnn based framework for stable image feature selection,” Paper, Presented at Proceedings of the 2017 IEEE Global Conference on Signal and Information Processing, Montreal, QC, Canada, Nov. 14-16, 2017, 6 … [cited by applicant]
Office Action in Chinese Appln. No. 201980087289.6, dated Dec. 12, 2023, 18 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2021-526271, dated Aug. 16, 2023, 6 pages (with English translation). [cited by applicant]
PCT International Search Report and Written Opinion in International Appln. No. PCT/US2019/061272, dated Feb. 6, 2020, 12 pages. [cited by applicant]
Sarlin et al., “SuperGlue: Learning Feature Matching with Graph Neural Networks,” cs.CV, Submitted on Nov. 26, 2019, arXiv:1911.11763v1, 17 pages. [cited by applicant]
Weerasekera et al., “Learning deeply supervised good features to match for dense monocular reconstruction,” Paper, Presented at Proceedings of the 14th Asian Conference on Computer Vision, Perth, Australia, Dec. 2-6, 20… [cited by applicant]
Zhou et al., “Paper: Unsupervised learning of depth and ego-motion from video,” Paper, Presented at Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, Jul. 21-26, 2017, 10 pages. [cited by applicant]
Notice of Allowance in Chinese Appln. No. 201980087289.6, dated Apr. 21, 2024, 6 pages (with English translation). [cited by applicant]