IP Library Granted Patent US 12,631,469
Granted Patent B2
US 12,631,469 · App. 18/493,984 · Granted May 19, 2026

Machine localization

Inventors: Yu Sheng (San Deigo, CA); Amir Akbarzadeh (Alamo, CA); Vishisht Gupta (Santa Clara, CA); Jordan Marr (Campbell, CA); Shaun Liu (San Jose, CA)
Assignee: NVIDIA CORPORATION
G01C21/3848G01C21/3453G06T7/536G06T7/75G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,631,469
App. No.
18/493,984
Granted
May 19, 2026
Kind
B2
Abstract

Embodiments of the present disclosure relate to a system and method used to localize one or more systems using 2D map data. The method may include determining an image location of a representation of a portion of an object in an image corresponding to an environment. In some embodiments, the method may additionally include determining one or more predicted image locations corresponding to the image location of the representation of the portion of the object. The method may additionally include comparing one or more ground plane locations of the portion of the object with the one or more predicted image locations, and determining a cost based at least on the comparison between the one or more ground plane locations and the one or more predicted image locations. Further, the method may include localizing a system to the 2D map data based on the determined cost.

Claims (50)

1 . A method comprising:

determining an image location of a representation of a portion of an object in an image corresponding to an environment, the image location being defined in two-dimensional (2D) image space based at least on image data corresponding to the image;

determining one or more predicted image locations corresponding to the image location of the representation of the portion of the object

comparing one or more ground plane locations corresponding to the portion of the object with the one or more predicted image locations, the one or more ground plane locations determined using 2D map data corresponding to a 2D map of the environment projected into the 2D image space;

updating the one or more ground plane locations until a cost determined based at least on the comparing satisfies a cost threshold;

localizing a machine to the 2D map data based at least on the cost; and

performing, by the machine, one or more autonomous operations based at least on the localizing to the 2D map data.

2 . The method of claim 1 , wherein the one or more predicted image locations are determined based at least on a vanishing point projected in the image space and the vanishing point is determined based at least on an orientation of a camera, used to generate the image data, relative to a gravitational vector, the gravitational vector indicating a direction of gravity projected into the 2D image space.

3 . The method of claim 1 , wherein the one or more predicted image locations are determined based at least on a vanishing point projected in the image space and the vanishing point is further determined using one or more characteristics of a camera used to generate the image data.

4 . The method of claim 1 , wherein a representation of the one or more predicted image locations includes a line segment including at least the image location of the representation of the portion of the object in the image corresponding to the environment.

5 . The method of claim 1 , wherein the image data is generated using a single, monocular camera.

6 . The method of claim 1 , wherein the one or more ground plane locations determined using the 2D map of the environment projected into the 2D image space are projected using a pinhole projection model.

7 . A method comprising:

determining one or more predicted image locations corresponding to an image location, in a 2D image space associated with an image, of a portion of an object;

updating one or more ground plane locations corresponding to the portion of the object until a difference between the one or more predicted image locations and the one or more ground plane locations satisfies an error threshold, the one or more ground plane locations corresponding to one or more positions of the portion of the object in a 2D map of an environment projected into the 2D image space;

localizing a machine to the 2D map based at least on the difference satisfying the error threshold; and

performing, by the machine, one or more autonomous operations based at least on the localizing to the 2D map data.

8 . The method of claim 7 , wherein the one or more predicted image locations is determined based at least on:

a representation of the ground in the image space;

a vanishing point projected into the 2D image space; and

the image location of the representation of the portion of the object.

9 . The method of claim 8 , wherein the vanishing point is determined based at least on an orientation of a camera, used to capture the image, relative to a gravitational vector, the gravitational vector indicating a direction of gravity projected into the 2D image space.

10 . The method of claim 9 , wherein the vanishing point is further determined using one or more characteristics of the camera.

11 . The method of claim 8 , wherein a representation of the one or more predicted image locations includes a line segment including at least the image location of the representation of the portion of the object in the image corresponding to the environment and the vanishing point projected in the image space.

12 . The method of claim 7 , wherein the image is captured using a single, monocular camera.

13 . The method of claim 7 , wherein the one or more ground plane locations determined using the 2D map of the environment projected into the 2D image space are projected using a pinhole projection model.

14 . One or more processors comprising processing circuitry to perform operations comprising:

determining one or more predicted image locations corresponding to an image location, in a 2D image space associated with an image, of a portion of an object, the one or more predicted image locations determined based at least on a vanishing point projected into the 2D image space and determined based at least on one or more characteristics of a camera used to capture the image;

comparing one or more ground plane locations corresponding to the portion of the object with the one or more predicted image locations, the one or more ground plane locations determined using a 2D map of an environment projected into the 2D image space;

localizing a machine to the 2D map based at least on the comparing; and

causing the machine to perform one or more autonomous operations based at least on the localizing to the 2D map data.

15 . The one or more processors of claim 14 , wherein the one or more predicted image locations are determined further based at least on:

a representation of the ground in a real-world environment;

and

the image location of the representation of the portion of the object.

16 . The one or more processors of claim 15 , wherein the one or more characteristics of the camera include an orientation of the camera relative to a gravitational vector, the gravitational vector indicating a direction of gravity projected into the 2D image space.

17 . The one or more processors of claim 14 , wherein the processor is included in a system, the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations;

a system for presenting at least one of augmented reality content, virtual reality content, or mixed reality content;

a system for hosting one or more real-time streaming applications; a system implemented using an edge device;

a system implemented using a robot;

a system for performing conversational AI operations;

a system implementing one or more large language models (LLMs); a system for performing generative AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2023
From: SHENG, YU; AKBARZADEH, AMIR; GUPTA, VISHISHT; MARR, JORDAN; LIU, SHAUN
To: NVIDIA CORPORATION
Reel/Frame 065342/0042 →
Continuity (1)
Related Publication 20250137813A1 · May 1, 2025
References Cited (12)
US 11462023B2 · Kehl · 2022 [cited by examiner]
US 11472442B2 · Duan · 2022 [cited by examiner]
US 11494937B2 · Urtasun · 2022 [cited by examiner]
US 11713978B2 · Akbarzadeh · 2023 [cited by examiner]
US 12014520B2 · Liu · 2024 [cited by examiner]
US 12223677B1 · Foucard · 2025 [cited by examiner]
US 12293543B1 · Wang · 2025 [cited by examiner]
US 20210150230A1 · Smolyanskiy · 2021 [cited by examiner]
US 20230176216A1 · Desai · 2023 [cited by examiner]
US 20230245469A1 · Holicki · 2023 [cited by examiner]
US 20240230335A1 · Roumeliotis · 2024 [cited by examiner]
US 20250014200A1 · Ahmed · 2025 [cited by examiner]