IP Library Granted Patent US 12,319,319
Granted Patent B2
US 12,319,319 · App. 18/465,128 · Granted Jun 3, 2025

Multi-task machine-learned models for object intention determination in autonomous driving

Inventors: Sergio Casas (Toronto, CA); Raquel Urtasun (Toronto, CA); Wenjie Luo (Toronto, CA)
Assignee: AURORA OPERATIONS, INC.
B60W60/0027B60W30/0956G06N20/00G06V10/764G06V10/82G06V20/58G06V40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,319,319
App. No.
18/465,128
Granted
Jun 3, 2025
Kind
B2
Abstract

Generally, the disclosed systems and methods utilize multi-task machine-learned models for object intention determination in autonomous driving applications. For example, a computing system can receive sensor data obtained relative to an autonomous vehicle and map data associated with a surrounding geographic environment of the autonomous vehicle. The sensor data and map data can be provided as input to a machine-learned intent model. The computing system can receive a jointly determined prediction from the machine-learned intent model for multiple outputs including at least one detection output indicative of one or more objects detected within the surrounding environment of the autonomous vehicle, a first corresponding forecasting output descriptive of a trajectory indicative of an expected path of the one or more objects towards a goal location, and/or a second corresponding forecasting output descriptive of a discrete behavior intention determined from a predefined group of possible behavior intentions.

Claims (50)

1. A method for detecting and forecasting an actor in an environment of an autonomous vehicle, comprising:

obtaining sensor data descriptive of the environment of the autonomous vehicle, the environment containing the actor;

obtaining map data associated with the environment; and

processing the sensor data and the map data with a machine-learned model comprising (i) one or more shared layers that are trained to generate intermediate features and (ii) one or more detection layers and one or more forecasting layers, wherein:

the one or more detection layers process the intermediate features to generate a detection output descriptive of the actor in the environment, and

the one or more forecasting layers process the intermediate features to generate:

a first forecasting output descriptive of a trajectory indicative of an expected path of the actor towards a goal location, and

a second forecasting output descriptive of a discrete behavior intention determined from a predefined group of possible behavior intentions for the actor.

2. The method of claim 1 , wherein the one or more detection layers and one or more forecasting layers are positioned structurally after the one or more shared layers.

3. The method of claim 1 , wherein:

the machine-learned model comprises a convolutional neural network; and

at least one of the (i) one or more shared layers that are trained to generate intermediate features and (ii) one or more detection layers and one or more forecasting layers comprises a convolutional layer of the convolutional neural network.

4. The method of claim 1 , wherein the detection output comprises a detection score indicative of a likelihood of the actor being in one of a plurality of predetermined classes.

5. The method of claim 1 , wherein the detection output comprises a bounding shape associated with the actor.

6. The method of claim 1 , wherein the first forecasting output is represented by trajectory data comprising a sequence of bounding shapes at a plurality of time stamps.

7. The method of claim 1 , wherein the predefined group of possible behavior intentions for the actor comprises at least one of keep lane, turn left, turn right, left change lane, right change lane, stopped, parked, or reverse driving.

8. The method of claim 1 , wherein processing the sensor data and the map data with the machine-learned model comprises processing a fused representation of a given view of the sensor data obtained relative to the autonomous vehicle and the map data in the given view.

9. The method of claim 8 , wherein the given view of the sensor data and the map data comprises a birds-eye view.

10. The method of claim 8 , wherein the given view of the sensor data is represented as a multi-dimensional tensor having at least one of a height dimension or a time dimension stacked into a channel dimension with the multi-dimensional tensor.

11. An autonomous vehicle (AV) control system, comprising:

one or more processors; and

one or more non-transitory computer-readable media that store instructions for execution by the one or more processors that cause the AV control system to perform operations, the operations comprising:

obtaining sensor data descriptive of an environment of the autonomous vehicle, the environment containing an actor;

obtaining map data associated with the environment; and

processing the sensor data and the map data with a machine-learned model comprising (i) one or more shared layers that are trained to generate intermediate features and (ii) one or more detection layers and one or more forecasting layers, wherein:

the one or more detection layers process the intermediate features to generate a detection output descriptive of the actor in the environment, and

the one or more forecasting layers process the intermediate features to generate:

a first forecasting output descriptive of a trajectory indicative of an expected path of the actor towards a goal location, and

a second forecasting output descriptive of a discrete behavior intention determined from a predefined group of possible behavior intentions for the actor.

12. The AV control system of claim 11 , wherein the one or more detection layers and one or more forecasting layers are positioned structurally after the one or more shared layers.

13. The AV control system of claim 11 , wherein:

the machine-learned model comprises a convolutional neural network; and

at least one of the (i) one or more shared layers that are trained to generate intermediate features and (ii) one or more detection layers and one or more forecasting layers comprises a convolutional layer of the convolutional neural network.

14. The AV control system of claim 11 , wherein processing the sensor data and the map data with the machine-learned model comprises processing a fused representation of a given view of the sensor data obtained relative to the autonomous vehicle and the map data in the given view.

15. The AV control system of claim 11 , wherein the detection output comprises a detection score indicative of a likelihood of the actor being in one of a plurality of predetermined classes.

16. The AV control system of claim 11 , wherein the first forecasting output is represented by trajectory data comprising a sequence of bounding shapes at a plurality of time stamps.

17. The AV control system of claim 11 , wherein the predefined group of possible behavior intentions for the actor comprises at least one of keep lane, turn left, turn right, left change lane, right change lane, stopped, parked, or reverse driving.

18. An autonomous vehicle, comprising:

one or more sensors that generate sensor data descriptive of an environment of the autonomous vehicle;

one or more processors; and

one or more non-transitory computer-readable media that store instructions for execution by the one or more processors that cause the one or more processors to perform operations, the operations comprising:

processing the sensor data and map data associated with the environment with a machine-learned model comprising (i) one or more shared layers that are trained to generate intermediate features and (ii) one or more detection layers and one or more forecasting layers, wherein:

the one or more detection layers process the intermediate features to generate a detection output descriptive of an actor in the environment, and

the one or more forecasting layers process the intermediate features to generate:

a first forecasting output descriptive of a trajectory indicative of an expected path of the actor towards a goal location, and

a second forecasting output descriptive of a discrete behavior intention determined from a predefined group of possible behavior intentions for the actor.

19. The autonomous vehicle of claim 18 , wherein the one or more detection layers and one or more forecasting layers are positioned structurally after the one or more shared layers.

20. The autonomous vehicle of claim 18 , wherein:

the machine-learned model comprises a convolutional neural network; and

at least one of the (i) one or more shared layers that are trained to generate intermediate features and (ii) one or more detection layers and one or more forecasting layers comprises a convolutional layer of the convolutional neural network.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2024
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 067073/0861 →
EMPLOYEE AGREEMENT Recorded Mar 28, 2024
From: LUO, WENJIE
To: UBER TECHNOLOGIES, INC.
Reel/Frame 066928/0337 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2024
From: CASAS, SERGIO
To: UBER TECHNOLOGIES, INC.
Reel/Frame 066928/0792 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2024
From: URTASUN, RAQUEL
To: UATC, LLC
Reel/Frame 066928/0795 →
Continuity (5)
Continuation 17749841 · May 20, 2022
Continuation 16420686 · May 23, 2019
Provisional Application 62748057 · Oct 19, 2018
Provisional Application 62685708 · Jun 15, 2018
Related Publication 20230415788A1 · Dec 28, 2023
References Cited (41)
US 10310087B2 · Laddha et al. · 2019 [cited by applicant]
US 10809361B2 · Vallespi-Gonzalez et al. · 2020 [cited by applicant]
US 20190025841A1 · Haynes et al. · 2019 [cited by applicant]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190302767A1 · Sapp et al. · 2019 [cited by applicant]
US 20190367020A1 · Yan et al. · 2019 [cited by applicant]
Ballan et al., “Knowledge transfer for scene-specific motion prediction”, European Conference on Computer Vision, 2016, pp. 697-713. [cited by applicant]
Chen et al., “Multi-view 3d object detection network for autonomous driving”, 2017 IEEE Conference on Computer Vision and Pattern Recognition, 2017, p. 3. [cited by applicant]
Dai et al., “R-fcn: Object detection via region-based fully convolutional networks”, Advances in Neutral Information Processing Systems, 2016, pp. 379-387. [cited by applicant]
Engelcke et al., “Vote3deep: Fast object detection in 3d point clouds using efficient convolutional neural networks”, 2017 IEEE International Conference on Robotics and Automation (IRCA), 2017, pp. 1355-1361. [cited by applicant]
Fathi et al., “Learning to recognize daily actions using gaze”, European Conference on Computer Vision, 2012, pp. 314-327. [cited by applicant]
Geiger et al., “Vision meets robotics: The KITTI dataset”, The International Journal of Robotics Research, 2013, pp. 1231-1237. [cited by applicant]
Girshick et al., “Rich feature hierarchies for accurate object detection and semantic segmentation”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 580-587. [cited by applicant]
He et al., “Deep residual learning for image recognition”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Hoermann et al., “Dynamic occupancy grid prediction for urban autonomous driving: A deep learning approach with fully automatic labeling”, 2018 IEEE International Conference on Robotics and Automation (IRCA), 2017, pp. … [cited by applicant]
Howard et al., “Mobilenets: Efficient convolutional neural networks for mobile vision applications”, arXiv preprint arXiv: 1704.04861, 2017. [cited by applicant]
Hu et al., “Probabilistic prediction of vehicle semantic intention and motion”, 2018 IEEE Intelligent Vehicles Symposium (IV), 2018, pp. 307-313. [cited by applicant]
Iandola et al., “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <0.5 mb model size”, arXiv preprint arXiv: 1602.07360, 2017. [cited by applicant]
Jain et al., “Car that knows before you do: Anticipating maneuvers via learning temporal driving models”, Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 3182-3190. [cited by applicant]
Kim et al., “Prediction of drivers intention of lane change by augmenting sensor information using machine learning techniques”, Sensors, 2017. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization”, Machine Learning, 2014. [cited by applicant]
Lee et al., “Desire: Distant future prediction in dynamic scenes with interacting agents”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2165-2174. [cited by applicant]
Lewis et al., “Sensor fusion weighting measures in audio-visual speech recognition”, Proceedings of the 27 [cited by applicant]
Li et al., “Vehicle detection from 3d lidar using fully convolutional network”, arXiv prepring arXiv: 1608.07916, 2016. [cited by applicant]
Lin et al., “Focal loss for dense object detection”, 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2999-3007. [cited by applicant]
Liu et al., “Ssd: Single shot multibox detector”, European Conference on Computer Vision, 2016, pp. 21-37. [cited by applicant]
Luo et al., “Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net”, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 3569-3… [cited by applicant]
Luo et al., “Understanding the effective receptive field in deep convolutional neural networks”, Advances in Neural Information Processing Systems, 2016, pp. 4898-4906. [cited by applicant]
Ma et al., “Forecasting interactive dynamics of pedestrians with fictitious play”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4643-4644. [cited by applicant]
Park et al., “Egocentric future localization”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4697-4705. [cited by applicant]
Phillips et al., “Generalizable intention prediction of human drivers at intersections”, Intelligent Vehicles Symposium (IV), 2017, pp. 1665-1670. [cited by applicant]
Qi et al., “Pointnet: Deep learning on point sets for 3d classification and segmentation”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 652-660. [cited by applicant]
Redmon et al., “You only look once: Unified, real-time object detection”, Proceedings of the IEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 779-788. [cited by applicant]
Ren et al., “Faster r-cnn: Towards real-time object detection with region proposal networks”, Advances in Neural Information Processing Systems, 2015, pp. 91-99. [cited by applicant]
Simon et al., “Complex-yolo: Real-time 3d object detection on point clouds”, European Conference on Computer Vision, 2018, pp. 197-209. [cited by applicant]
Snoek et al., “Early versus late fusion in semantic video analysis”, Proceedings of the 13 [cited by applicant]
Streubel et al., “Prediction of driver intended path at intersections”, Intelligent Vehicles Symposium (IV), 2014, pp. 134-139. [cited by applicant]
Sutton et al., [cited by applicant]
Tran et al., “Learning spatiotemporal features with 3d convolutional networks”, 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 4489-4497. [cited by applicant]
Yang et al., “Pixor: Real-time 3d object detection from point clouds”, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7652-7660. [cited by applicant]
Zhang et al., “Sensor fusion for semantic segmentation of urban scenes”, 2015 IEEE International Conference on Robotics and Automation (ICRA), 2015, pp. 1850-1857. [cited by applicant]