IP Library Granted Patent US 12,373,960
Granted Patent B2
US 12,373,960 · App. 17/672,402 · Granted Jul 29, 2025

Dynamic object detection using LiDAR data for autonomous machine systems and applications

Inventors: Jens Christian Bo Joergensen (Flushing, NY); Ollin Boer Bohan (Redmond, WA); Joachim Pehserl (Lynnwood, WA); Nikolai Smolyanskiy (Seattle, WA)
Assignee: NVIDIA Corporation
G06T7/254G06V10/454G06T2207/10028G06T2207/20084G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,960
App. No.
17/672,402
Granted
Jul 29, 2025
Kind
B2
Abstract

In various examples, systems and methods of the present disclosure detect and/or track objects in an environment using projection images generated from LiDAR. For example, a machine learning model—such as a deep neural network (DNN)—may be used to compute a motion mask indicative of motion corresponding to points representing objects in an environment. Various input channels may be provided as input to the machine learning model to compute a motion mask. One or more comparison images may be generated based on comparing depth values projected from a current range image to a coordinate space of a previous range image to depth values of the previous range image. The machine learning model may use the one or more projection images, the one or more comparison images, and/or the one or more range images to compute a motion mask and/or a motion vector output representation.

Claims (60)

1. At least one processor comprising:

one or more circuits to:

generate, using sensor data obtained using one or more sensors of an ego-machine, a first range image corresponding to a first time and a second range image corresponding to a second time subsequent to the first time;

generate a first projected image based at least on projecting first depth values from the first range image to a first coordinate space of the second range image and a second projected image based at least on projecting second depth values from the second range image to a second coordinate space of the first range image;

compute, from the first projected image and the second projected image, input data to a deep neural network (DNN), the input data indicating differences in depth between the first range image and the second range image and differences in spatial correspondences of scene points across the first coordinate space and the second coordinate space; and

compute, based at least on the DNN processing the input data, output data indicative of motion at one or more pixels of the second range image.

2. The at least one processor of claim 1 , further comprising one or more circuits to:

generate a comparison image based at least on part on comparing the second depth values from the second range image in the second coordinate space of the first range image to depth values of the first range image; and

wherein the input data is computed from the comparison image.

3. The at least one processor of claim 1 , the input data includes a first representation of one or more regions of the first projected image and a second representation of one or more regions of the second projected image.

4. The at least one processor of claim 1 , wherein the first projected image and the second projected image are generated using tracked ego-motion of the ego-machine between the first time and the second time.

5. The at least one processor of claim 1 , wherein the input data includes a first channel corresponding to a first distance between a first location for the ego-machine at the first time and a first (three-dimensional) 3D point that projects to a same pixel as a second 3D point when viewed from a second location for the ego-machine at the second time, and a second channel corresponding to a second distance between the second location and the second 3D point.

6. The at least one processor of claim 1 , wherein the output data computed using the DNN is representative of a motion mask, the motion mask having one or more first values being indicative of one or more first objects depicted using the second range image being in motion at the second time and having one or more second values being indicative of one or more second objects depicted using the second range image being static at the second time.

7. The at least one processor of claim 6 , wherein the motion mask comprises, for each of the one or more pixels of the second range image, a confidence value between 0 and 1, the confidence value being indicative of whether a feature or object represented using a respective pixel of the one or more pixels corresponds to an object in motion at the second time.

8. The at least one processor of claim 1 , wherein the one or more circuits are further to perform, using the output data indicative of motion, one or more control operations corresponding to the ego-machine based at least on the motion at the one or more pixels of the second range image.

9. The at least one processor of claim 1 , wherein the DNN is trained to predict a likelihood that one or more pixels correspond to a static or a dynamic object.

10. The at least one processor of claim 1 , wherein the at least one processor is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

11. A method comprising:

generating, using a sensor of an ego-machine, a first range image at a first time and a second range image at a second time subsequent the first time;

generating a first projected image based at least on projecting one or more first depth values from the first range image to a first coordinate space of the second range image;

generating a second projected image based at least on projecting one or more second depth values from the second range image to a second coordinate space of the first range image;

computing, from the first projected image and the second projected image, input data to a deep neural network (DNN), the input data indicating differences in depth between the first range image and the second range image and differences in spatial correspondences of scene points across the first coordinate space and the second coordinate space; and

computing, based at least on a deep neural network (DNN) the DNN processing the input data, output data indicative of motion at one or more pixels of the second range image.

12. The method of claim 11 , further comprising:

generating a comparison image based at least in part on comparing the second depth values from the second range image in the second coordinate space of the first range image to depth values of the first range image,

wherein the input data is computed from the comparison image.

13. The method of claim 11 , wherein the input data includes a first representation of one or more regions of the first projected image and a second representation of one or more regions of the second projected image.

14. The method of claim 11 , wherein the generating the first projected image and the second projected image are based at least on tracked ego-motion of the ego-machine between the first time and the second time.

15. A system comprising:

one or more processors to perform operations including:

generating, using a sensor of an ego-machine, a first range image at a first time and a second range image at a second time subsequent the first time;

generating a first projected image based at least on projecting one or more first depth values from the first range image to a first coordinate space of the second range image;

generating a second projected image based at least on projecting one or more second depth values from the second range image to a second coordinate space of the first range image;

computing, from the first projected image and the second projected image, input data to a deep neural network (DNN), the input data indicating differences in depth between the first range image and the second range image and differences in spatial correspondences of scene points across the first coordinate space and the second coordinate space; and

computing, using the DNN having the input data, output data indicative of motion at one or more pixels of the second range image.

16. The system of claim 15 , wherein the operations further include:

generating a comparison image based at least on part on comparing the second depth values from the second range image in the second coordinate space of the first range image to depth values of the first range image,

wherein the computing the input data is computed from the comparison image.

17. The system of claim 15 , wherein, for each range image of n number of range images prior to the second range image, a respective first projected image, second projected image, and comparison image are generated using the second range image.

18. The system of claim 15 , wherein the generating the first projected image and the second projected image are based at least on tracked ego-motion of the ego-machine between the first time and the second time.

19. The system of claim 15 , wherein the operations further include computing, based at least on the output data computed using the DNN, one or more object detections.

20. The system of claim 15 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2022
From: JOERGENSEN, JENS CHRISTIAN BO; BOER BOHAN, OLLIN; PEHSERL, JOACHIM; SMOLYANSKIY, NIKOLAI
To: NVIDIA CORPORATION
Reel/Frame 059090/0147 →
Continuity (1)
Related Publication 20230260136A1 · Aug 17, 2023
References Cited (30)
US 10885698B2 · Muthler et al. · 2021 [cited by applicant]
US 11062454B1 · Cohen · 2021 [cited by examiner]
US 11741631B2 · Serackis · 2023 [cited by examiner]
US 20180322640A1 · Kim · 2018 [cited by examiner]
US 20180364717A1 · Douillard · 2018 [cited by examiner]
US 20190332939A1 · Alletto · 2019 [cited by examiner]
US 20200081448A1 · Creusot · 2020 [cited by examiner]
US 20200084427A1 · Sun · 2020 [cited by examiner]
US 20200226769A1 · Das · 2020 [cited by examiner]
US 20200258249A1 · Angelova · 2020 [cited by examiner]
US 20200377105A1 · Murashkin · 2020 [cited by examiner]
US 20210150230A1 · Smolyanskiy · 2021 [cited by examiner]
US 20210150736A1 · Lv · 2021 [cited by examiner]
US 20210319578A1 · Casser · 2021 [cited by examiner]
US 20210342608A1 · Smolyanskiy · 2021 [cited by examiner]
US 20210358137A1 · Lee · 2021 [cited by examiner]
US 20210383553A1 · Guizilini · 2021 [cited by examiner]
US 20220284221A1 · Slutsky · 2022 [cited by examiner]
EP 2525000A1 · 2019 [cited by applicant]
EP 3832532A2 · 2021 [cited by applicant]
WO 2023158556A1 · 2023 [cited by applicant]
Behl, A., Paschalidou, D., Donné, S., & Geiger, A. (2018). PointFlowNet: Learning Representations for Rigid Motion Estimation From Point Clouds. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)… [cited by examiner]
S. A. Baur, F. Moosmann, S. Wirges and C. B. Rist, “Real-time 3D LiDAR Flow for Autonomous Vehicles,” 2019 IEEE Intelligent Vehicles Symposium (IV), Paris, France, 2019, pp. 1288-1295, doi: 10.1109/IVS.2019.8814094. (Ye… [cited by examiner]
X. Qi, Z. Liu, Q. Chen and J. Jia, “3D Motion Decomposition for RGBD Future Dynamic Scene Synthesis,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 7665-7674,… [cited by examiner]
Joegensen, et al.; International Search Report and Written Opinion for International Application No. PCT/US2023/012097, filed Feb. 1, 2023, mailed Apr. 25, 2023, 13 pgs. [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, National Highway Traffic Safety Administration (NHTSA), A Division of the US Department of Transportation, and the S… [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, National Highway Traffic Safety Administration (NHTSA), A Division of the US Department of Transportation, and the S… [cited by applicant]
ISO 26262, “Road vehicle—Functional safety,” International standard for functional safety of electronic system, Retrieved from Internet URL: https://en.wikipedia.org/wiki/ISO_26262, accessed on Sep. 13, 2021, 8 pages. [cited by applicant]
IEC 61508, “Functional Safety of Electrical/Electronic/Programmable Electronic Safety-related Systems,” Retrieved from Internet URL: https://en.wikipedia.org/wiki/IEC_61508, accessed on Apr. 1, 2022, 7 pages. [cited by applicant]
Joergensen, Jens Christian Bo; International Preliminary Report on Patentability for PCT Application No. PCT/US2023/012097, filed Feb. 1, 2023, mailed Aug. 29, 2024, 10 pgs. [cited by applicant]