IP Library Granted Patent US 12,259,694
Granted Patent B2
US 12,259,694 · App. 18/656,210 · Granted Mar 25, 2025

Systems and methods for sensor data processing and object detection and motion prediction for robotic platforms

Inventors: Abhishek Mohta (San Mateo, CA); Fang-Chieh Chou (Redwood City, CA); Carlos Vallespi-Gonzalez (Wexford, PA); Brian C. Becker (Pittsburgh, PA); Nemanja Djuric (Pittsburgh, PA)
Assignee: AURORA OPERATIONS, INC.
G05B13/0265B60W60/001B60W50/00B60W2050/0052B60W2050/0083B60W2420/403B60W2420/408B60W2556/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,259,694
App. No.
18/656,210
Granted
Mar 25, 2025
Kind
B2
Abstract

Systems and methods are disclosed for detecting and predicting the motion of objects within the surrounding environment of a system such as an autonomous vehicle. For example, an autonomous vehicle can obtain sensor data from a plurality of sensors comprising at least two different sensor modalities (e.g., RADAR, LIDAR, camera) and fused together to create a fused sensor sample. The fused sensor sample can then be provided as input to a machine learning model (e.g., a machine learning model for object detection and/or motion prediction). The machine learning model can have been trained by independently applying sensor dropout to the at least two different sensor modalities. Outputs received from the machine learning model in response to receipt of the fused sensor samples are characterized by improved generalization performance over multiple sensor modalities, thus yielding improved performance in detecting objects and predicting their future locations, as well as improved navigation performance.

Claims (47)

1. A computer-implemented method comprising:

(a) obtaining sensor data from a plurality of sensors comprising at least two different sensor modalities;

(b) applying independent sensor dropout to the at least two different sensor modalities;

(c) fusing the sensor data from the at least two different sensor modalities with sensor dropout independently applied thereto to generate a fused sensor sample;

(d) generating a training data set comprising the fused sensor sample; and

(e) training a machine-learned model for object detection using the training data set, wherein the trained machine-learned model is employed by a robotic platform operating within an environment.

2. The computer-implemented method of claim 1 , wherein the robotic platform comprises an autonomous vehicle.

3. The computer-implemented method of claim 1 , wherein the environment comprises a real-world environment or a simulated environment.

4. The computer-implemented method of claim 1 , wherein the trained machine-learned model comprises an end-to-end model that is configured to jointly perform object detection and motion prediction.

5. The computer-implemented method of claim 1 , wherein (b) comprises independently applying sensor dropout to each of the at least two different sensor modalities at a fixed probability associated with the sensor modality.

6. The computer-implemented method of claim 1 , wherein the plurality of sensors comprise a RADAR system, a LIDAR system, and a camera.

7. The computer-implemented method of claim 6 , wherein:

the at least two different sensor modalities comprise at least one of the RADAR system or the camera; and

(b) comprises zeroing out a final feature vector for a portion of the sensor data obtained from the at least one of the RADAR system or the camera.

8. The computer-implemented method of claim 6 , wherein:

the at least two different sensor modalities comprise the LIDAR system; and

(b) comprises replacing a LIDAR intensity value with a sentinel value for a portion of the sensor data obtained from the LIDAR system.

9. The computer-implemented method of claim 1 , wherein (e) comprises:

inputting the fused sensor sample into the machine-learned model;

generating a loss metric for the machine-learned model based on output of at least a portion of the machine-learned model in response to the fused sensor sample as input; and

modifying at least a portion of the machine-learned model based on the loss metric.

10. The computer-implemented method of claim 9 , wherein the loss metric comprises at least one of a regression loss, a classification loss, an adversarial loss, a multi-task loss, or a perceptual loss.

11. The computer-implemented method of claim 9 , wherein the loss metric is associated with a plurality of loss terms, the plurality of loss terms comprising at least a first loss term associated with a determination or generation of bounding shapes.

12. The computer-implemented method of claim 11 , the plurality of loss terms further comprising at least a second loss term associated with a classification of features.

13. A training computing system, comprising:

one or more processors; and

one or more non-transitory computer-readable medium storing instructions that when executed by the one or more processors cause the training computing system to perform operations, the operations comprising:

(a) obtaining sensor data from a plurality of sensors comprising at least two different sensor modalities;

(b) applying independent sensor dropout to the at least two sensor modalities;

(c) fusing the sensor data from the at least two different sensor modalities with sensor dropout independently applied thereto to generate a fused sensor sample;

(d) generating a training data set comprising the fused sensor sample; and

(e) training a machine-learned model for object detection using the training data set, wherein the trained machine-learned model is employed by a robotic platform operating within an environment.

14. The training computing system of claim 13 , wherein the robotic platform comprises an autonomous vehicle.

15. The training computing system of claim 13 , wherein the environment comprises a real-world environment or a simulated environment.

16. The training computing system of claim 13 , wherein the trained machine-learned model comprises an end-to-end model that is configured to jointly perform object detection and motion prediction.

17. The training computing system of claim 13 , wherein (b) comprises independently applying sensor dropout to each of the at least two different sensor modalities at a fixed probability associated with the sensor modality.

18. The training computing system of claim 13 , wherein (e) comprises:

inputting the fused sensor sample into the machine-learned model;

generating a loss metric for the machine-learned model based on output of at least a portion of the machine-learned model in response to the fused sensor sample as input; and

modifying at least a portion of the machine-learned model based on the loss metric.

19. The training computing system of claim 18 , wherein the loss metric comprises one or more of a regression loss and a classification loss.

20. One or more non-transitory computer-readable medium storing instructions that when executed by one or more processors cause the one or more processors to perform operations, the operations comprising:

(a) obtaining sensor data from a plurality of sensors comprising at least two different sensor modalities;

(b) applying independent sensor dropout to the at least two sensor modalities;

(c) fusing the sensor data from the at least two different sensor modalities with sensor dropout independently applied thereto to generate a fused sensor sample;

(d) generating a training data set comprising the fused sensor sample; and

(e) training a machine-learned model for object detection using the training data set, wherein the trained machine-learned model is employed by a robotic platform operating within an environment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2024
From: MOHTA, ABHISHEK; BECKER, BRIAN C.; VALLESPI-GONZALEZ, CARLOS; CHOU, FANG-CHIEH; DJURIC, NEMANJA
To: UATC, LLC
Reel/Frame 067390/0381 →
Continuity (3)
Continuation 17501614 · Oct 14, 2021
Provisional Application 63091401 · Oct 14, 2020
Related Publication 20240369977A1 · Nov 7, 2024
References Cited (20)
US 11422546B2 · Giering · 2022 [cited by examiner]
US 20200174490A1 · Ogale · 2020 [cited by examiner]
US 20200225321A1 · Kruglick · 2020 [cited by examiner]
US 20210122364A1 · Lee · 2021 [cited by applicant]
US 20210304046A1 · Miki · 2021 [cited by examiner]
US 20210398338A1 · Philion · 2021 [cited by examiner]
US 20220026568A1 · Meuter et al. · 2022 [cited by applicant]
Casas et al., “IntentNet: Learning to Predict Intention from Raw Sensor Data”, 2 [cited by applicant]
Djuric et al., “MultiXNet: Multiclass Multistage Multimodal Motion Prediction”, arXiv:2006.02000v3, 8 pages. [cited by applicant]
Fadadu et al., “Multi-View Fusion of Sensor Data for Improved Perception and Prediction in Autonomous Driving”, arXiv:2008.11901v1, 10 pages. [cited by applicant]
Liang et al., “Deep Continuous Fusion for Multi-Sensor 3D Object Detection”, European Conference on Computer Vision (ECCV), 2018, 16 pages. [cited by applicant]
Liu et al., “Learning End-to-End Multimodal Sensor Policies for Autonomous Navigation”, arXiv:1705.10422v2, 13 pages. [cited by applicant]
Luo et al., “Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net”, IEEE CVPR, 2018, pp. 3569-3577. [cited by applicant]
Manivasagam et al., “LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World”, IEEE CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11167-11176. [cited by applicant]
Meyer et al., “LaserFlow: Efficient and Probabilistic Object Detection and Motion Forecasting”, arXiv:2003.05982v4, 8 pages. [cited by applicant]
Meyer et al., “Sensor Fusion for Joint 3D Object Detection and Semantic Segmentation”, IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2019, 8 pages. [cited by applicant]
Shah et al., “LiRaNet: End-to-End Trajectory Prediction using Spatio-Temporal Radar Fusion”, arXiv:2010.00731v3, 17 pages. [cited by applicant]
Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, Journal of Machine Learning Research, vol. 15, 2014, pp. 1929-1958. [cited by applicant]
Urmson et al., “Self-Driving Cars and the Urban Challenge”, IEEE Intelligent Transportation Systems, Mar./Apr. 2008, pp. 66-68. [cited by applicant]
Yang et al., “RadarNet: Exploiting Radar for Robust Perception of Dynamic Objects”, arXiv:2007.14366v1, 16 pages. [cited by applicant]