IP Library › Granted Patent US 12,260,574
Granted Patent B2
US 12,260,574 · App. 17/808,045 · Granted Mar 25, 2025

Image-based keypoint generation

Inventors: Ronghua Zhang (Campbell, CA); Derik Schroeter (Fremont, CA); Mengxi Wu (Mountain View, CA); Di Zeng (Sunnyvale, CA)
Assignee: NVIDIA CORPORATION
G06T7/521G01C21/32G01S17/89G06T7/73
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,574
App. No.
17/808,045
Granted
Mar 25, 2025
Kind
B2
Abstract

Operations may comprise obtaining a plurality of light detection and ranging (LIDAR) scans of a region. The operations may also comprise identifying a plurality of LIDAR poses that correspond to the plurality of LIDAR scans. In addition, the operations may comprise identifying, as a plurality of keyframes, a plurality of images of the region that are captured during capturing of the plurality of LIDAR scans. The operations may also comprise determining, based on the plurality of LIDAR poses, a plurality of camera poses that correspond to the keyframes. Further, the operations may comprise identifying a plurality of two-dimensional (2D) keypoints in the keyframes. The operations also may comprise generating one or more three-dimensional (3D) keypoints based on the plurality of 2D keypoints and the respective camera poses of the plurality of keyframes.

Claims (59)

1. A method, comprising:

selecting a first keypoint included in a first image, the first keypoint corresponding to a feature of an area represented by the first image, the selecting of the first keypoint being based at least on:

determining that a second image includes a second keypoint that also corresponds to the feature; and

determining that one or more points of a light detection and ranging (LIDAR) scan also correspond to the feature; and

annotating, based at least on the selecting of the first keypoint, map data with keypoint data that corresponds to the first keypoint, the keypoint data describing the feature corresponding to the first keypoint, wherein one or more pose parameters corresponding to a machine are determined based at least on the map data being annotated, based at least on the selecting of the first keypoint, with keypoint data that corresponds to the first keypoint.

2. The method of claim 1 , wherein the keypoint data used to annotate the map data correspond to a two-dimensional representation of the first keypoint.

3. The method of claim 1 , wherein the keypoint data used to annotate the map data corresponds to a three-dimensional representation of the first keypoint.

4. The method of claim 1 , wherein the keypoint data used to annotate the map data has been converted from corresponding to a two-dimensional representation of the first keypoint to corresponding to a three-dimensional representation of the first keypoint.

5. The method of claim 1 , wherein the first keypoint is identified based at least on one or more of:

a pixel intensity;

a color;

a size;

a shape;

a pixel transition;

a feature transition;

a texture transition;

a pattern;

a texture; or

a semantic label.

6. The method of claim 1 , wherein the first keypoint is further selected based at least on a distance between a first point of the LIDAR scan that is identified as corresponding to the first keypoint and a second point of the LIDAR scan that is identified as corresponding to the second keypoint.

7. A processor comprising processing circuitry to cause performance of operations, the operations comprising:

determining one or more pose parameters of a machine based at least on a comparison of sensor data with first keypoint data included in map data corresponding to a map of an area, the first keypoint data corresponding to a first keypoint included in a first image, the first keypoint corresponding to a feature of the area and the first keypoint data describing the feature, the first keypoint data being included in the map data based at least on the first keypoint being selected from the first image, the selecting of the first keypoint being based at least on one or more of:

determining that a second image includes a second keypoint that also corresponds to the feature; or

determining that one or more points of a light detection and ranging (LIDAR) scan also correspond to the feature.

8. The processor of claim 7 , wherein the first keypoint data included in the map data has been converted from corresponding to a two-dimensional representation of the first keypoint to corresponding to a three-dimensional representation of the first keypoint.

9. The processor of claim 7 , wherein the first keypoint is identified based at least on one or more of:

a pixel intensity;

a color;

a size;

a shape;

a pixel transition;

a feature transition;

a pattern;

a texture; or

a semantic label.

10. The processor of claim 7 , wherein the determining of the one or more pose parameters is further based at least on a comparison between the first keypoint and a point included in a third image.

11. The processor of claim 7 , wherein the first keypoint is selected further based at least on a distance between a first point of the LIDAR san that is identified as corresponding to the first keypoint and a second point of the LIDAR scan that is identified as corresponding to the second keypoint.

12. A system comprising one or more processing units to cause performance of operations, the operations comprising:

identifying a plurality of keypoints included in a first image, the plurality of keypoints respectively corresponding to one or more features of an area corresponding to the first image;

filtering the plurality of keypoints to identify a subset of keypoints, the filtering being based at least on one or more of:

a second image corresponding to the area; or

a light detection and ranging (LIDAR) scan corresponding to the area; and

annotating map data with keypoint data corresponding to subset of keypoints, wherein one or more pose parameters corresponding to a machine are determined based at least on the map data being annotated, based at least on the subset of keypoints, with keypoint data that corresponds to the subset of keypoints.

13. The system of claim 12 , wherein the filtering of the plurality of keypoints based at least on the second image includes selecting for inclusion in the subset of keypoints, one or more keypoints whose corresponding feature is included in the second image.

14. The system of claim 12 , wherein the filtering of the plurality of keypoints based at least on the LIDAR scan includes selecting for inclusion in the subset of keypoints, one or more keypoints whose corresponding feature also corresponds to one or more points of the LIDAR scan.

15. The system of claim 12 , wherein the keypoint data used to annotate the map data has been converted from corresponding to a two-dimensional representation of the subset of keypoints to corresponding to a three-dimensional representation of the subset of keypoints.

16. The system of claim 12 , wherein one or more of the plurality of keypoints are identified based at least on one or more of:

a pixel intensity;

a color;

a size;

a shape;

a pixel transition;

a feature transition;

a texture transition;

a pattern;

a texture; or

a semantic label.

17. The system of claim 12 , further wherein one or more pose parameters of a machine are determined based at least on the map data as annotated with the keypoint data.

18. The system of claim 12 wherein the filtering of the plurality of keypoints is further based at least on a distance between a first point of the LIDAR scan that is identified as corresponding to a first keypoint of the plurality of keypoints and a second point of the LIDAR scan that is identified as corresponding to a second keypoint corresponding to the second image.

Continuity (3)
Continuation 16912549 · Jun 25, 2020
Provisional Application 62866362 · Jun 25, 2019
Related Publication 20230018923A1 · Jan 19, 2023
References Cited (14)
US 7187809B2 · Zhao et al. · 2007 [cited by applicant]
US 11222442B2 · Bao · 2022 [cited by examiner]
US 20170045362A1 · Song · 2017 [cited by examiner]
US 20180188026A1 · Zhang · 2018 [cited by examiner]
US 20180357773A1 · Wang · 2018 [cited by examiner]
US 20190107839A1 · Parashar · 2019 [cited by examiner]
US 20200394824A1 · Kanzawa · 2020 [cited by examiner]
US 20200410701A1 · Chen · 2020 [cited by examiner]
US 20210227139A1 · Wang · 2021 [cited by examiner]
US 20210381845A1 · Gokhale · 2021 [cited by examiner]
WO 2019024793A1 · 2019 [cited by applicant]
Feng et al, (2D3D-MatchNet: Learning to Match Keypoints Across 2D Image and 3D Point Cloud, IEEE, pp. 4790-4796, May 2019 (Year: 2019). [cited by examiner]
Feng et al., “2D3D-MatchNet: Learning to Match Keypoints Across 2D Image and 3D Point Cloud.” 2019 International Conference on Robotics and Automation, Apr. 22, 2019. [cited by applicant]
PCT International Search Report and Written Opinion issued in corresponding application No. PCT/US2020/039703, dated Sep. 30, 2020. [cited by applicant]