IP Library Granted Patent US 12,430,534
Granted Patent B2
US 12,430,534 · App. 18/656,150 · Granted Sep 30, 2025

Systems and methods for generating motion forecast data for actors with respect to an autonomous vehicle and training a machine learned model for the same

Inventors: Raquel Urtasun (Toronto, CA); Renjie Liao (Toronto, CA); Sergio Casas (Toronto, CA); Cole Christian Gulino (Pittsburgh, PA)
Assignee: AURORA OPERATIONS, INC.
G06N3/04G01C21/3626G06N3/045G06N3/06G06N3/08G06N3/084G06N7/01G06N7/046G08G1/0133G08G1/0141G08G1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,534
App. No.
18/656,150
Granted
Sep 30, 2025
Kind
B2
Abstract

Systems and methods for generating motion forecast data for actors with respect to an autonomous vehicle and training a machine learned model for the same are disclosed. The computing system can include an object detection model and a graph neural network including a plurality of nodes and a plurality of edges. The computing system can be configured to input sensor data into the object detection model; receive object detection data describing the location of the plurality of the actors relative to the autonomous vehicle as an output of the object detection model; input the object detection data into the graph neural network; iteratively update a plurality of node states respectively associated with the plurality of nodes; and receive, as an output of the graph neural network, the motion forecast data with respect to the plurality of actors.

Claims (43)

1. A computer-implemented method comprising:

generating, using an object detection model, object detection data describing a location of an actor within an environment of an autonomous vehicle;

generating, using a graph neural network, motion forecast data associated with the actor, wherein:

the graph neural network comprises a plurality of nodes, wherein a first node of the plurality of nodes corresponds to the actor, and

the graph neural network is configured to model anticipated interactions between the actor and other actors in the environment by passing one or more messages among the plurality of nodes to update at least one node state of the first node, wherein the anticipated interactions comprise an interaction between the actor and at least one other actor which alters a trajectory of the actor or the at least one other actor in the environment;

determining a motion plan for the autonomous vehicle based on the motion forecast data; and

controlling the autonomous vehicle based on the motion plan.

2. The computer-implemented method of claim 1 , further comprising:

inputting at least one of: (i) sensor data or (ii) map data into the object detection model.

3. The computer-implemented method of claim 2 , wherein the sensor data is captured by a remote computing system.

4. The computer-implemented method of claim 2 , wherein the sensor data and the map data are concatenated prior to being input into the object detection model.

5. The computer-implemented method of claim 1 , wherein the object detection data comprises a region of interest within the environment, the region of interest associated with the actor.

6. The computer-implemented method of claim 1 , wherein at least a portion of the plurality of nodes correspond to the other actors in the environment.

7. The computer-implemented method of claim 1 , wherein the anticipated interactions comprise an interaction within a threshold distance between the actor and at least one other actor in the environment.

8. A computing system comprising:

one or more processors, and

one or more memory resources storing instructions executable by the one or more processors to cause the one or more processors to:

generate, using an object detection model, object detection data describing a location of an actor within an environment of an autonomous vehicle;

generate, using a graph neural network, motion forecast data associated with the actor, wherein:

the graph neural network comprises a plurality of nodes, wherein a first node of the plurality of nodes corresponds to the actor, and

the graph neural network is configured to model anticipated interactions between the actor and other actors in the environment by passing one or more messages among the plurality of nodes to update at least one node state of the first node, wherein the anticipated interactions comprise an interaction between the actor and at least one other actor which alters a trajectory of the actor or the at least one other actor in the environment;

determine a motion plan for the autonomous vehicle based on the motion forecast data; and

control the autonomous vehicle based on the motion plan.

9. The computing system of claim 8 , wherein the one or more processors:

input at least one of: (i) sensor data or (ii) map data into the object detection model.

10. The computing system of claim 9 , wherein the sensor data is captured by a remote computing system.

11. The computing system of claim 9 , wherein the sensor data and the map data are concatenated prior to being input into the object detection model.

12. The computing system of claim 8 , wherein the object detection data comprises a region of interest within the environment, the region of interest associated with the actor.

13. The computing system of claim 8 , wherein at least a portion of the plurality of nodes correspond to the other actors in the environment.

14. The computing system of claim 8 , wherein the anticipated interactions comprise an interaction within a threshold distance between the actor and at least one other actor in the environment.

15. A non-transitory computer-readable media storing instructions executable by one or more processors of an autonomous vehicle computing system to cause the one or more processors to:

generate, using an object detection model, object detection data describing a location of an actor within an environment of an autonomous vehicle;

generate, using a graph neural network, motion forecast data associated with the actor, wherein:

the graph neural network comprises a plurality of nodes, wherein a first node of the plurality of nodes corresponds to the actor, and

the graph neural network is configured to model anticipated interactions between the actor and other actors in the environment by passing one or more messages among the plurality of nodes to update at least one node state of the first node, wherein the anticipated interactions comprise an interaction between the actor and at least one other actor which alters a trajectory of the actor or the at least one other actor in the environment;

determine a motion plan for the autonomous vehicle based on the motion forecast data; and

control the autonomous vehicle based on the motion plan.

16. The non-transitory computer-readable media of claim 15 , wherein the one or more processors:

input at least one of: (i) sensor data or (ii) map data into the object detection model.

17. The non-transitory computer-readable media of claim 16 , wherein the sensor data is captured by a remote computing system.

18. The non-transitory computer-readable media of claim 16 , wherein the sensor data and the map data are concatenated prior to being input into the object detection model.

19. The non-transitory computer-readable media of claim 15 , wherein the object detection data comprises a region of interest within the environment, the region of interest associated with the actor.

20. The non-transitory computer-readable media of claim 15 , wherein at least a portion of the plurality of nodes correspond to the other actors in the environment.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: LIAO, RENJIE; CASAS, SERGIO; GULINO, COLE CHRISTIAN
To: UATC, LLC
Reel/Frame 067364/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: URTASUN, RAQUEL
To: UATC, LLC
Reel/Frame 067364/0868 →
Continuity (5)
Continuation 18186718 · Mar 20, 2023
Continuation 16816671 · Mar 12, 2020
Provisional Application 62926826 · Oct 28, 2019
Provisional Application 62871452 · Jul 8, 2019
Related Publication 20240320466A1 · Sep 26, 2024
References Cited (94)
US 6333703B1 · Alewine · 2001 [cited by examiner]
US 6366701B1 · Chalom · 2002 [cited by examiner]
US 8254633B1 · Moon · 2012 [cited by examiner]
US 9248834B1 · Ferguson · 2016 [cited by examiner]
US 10127465B2 · Cohen · 2018 [cited by examiner]
US 10719712B2 · Gupta · 2020 [cited by examiner]
US 11112796B2 · Djuric · 2021 [cited by examiner]
US 20170131719A1 · Micks · 2017 [cited by examiner]
US 20170263130A1 · Sane · 2017 [cited by examiner]
US 20180074505A1 · Lv · 2018 [cited by examiner]
US 20190012574A1 · Anthony · 2019 [cited by examiner]
US 20190049970A1 · Djuric · 2019 [cited by examiner]
US 20190096086A1 · Xu · 2019 [cited by examiner]
US 20190147610A1 · Frossard · 2019 [cited by examiner]
US 20190152490A1 · Lan · 2019 [cited by examiner]
US 20190248411A1 · Peng · 2019 [cited by examiner]
US 20190339082A1 · Doig · 2019 [cited by examiner]
US 20200139960A1 · Newman · 2020 [cited by examiner]
US 20200159225A1 · Zeng · 2020 [cited by examiner]
US 20210090447A1 · Gnoth · 2021 [cited by examiner]
US 20210146963A1 · Li · 2021 [cited by examiner]
US 20210182604A1 · Anthony · 2021 [cited by examiner]
US 20210182605A1 · Anthony · 2021 [cited by examiner]
US 20210332766A1 · Hareyama · 2021 [cited by examiner]
US 20220230216A1 · Buibas · 2022 [cited by examiner]
EP 3016069A1 · 2016 [cited by examiner]
EP 3646254B1 · 2024 [cited by examiner]
Alahi et al., “Social LSTM: Human Trajectory Prediction in Crowded Spaces”, IEEE/Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2016), Jun. 26-Jul. 1, 2016, Las Vegas, NV, 11 pages. [cited by applicant]
Bansal et al., “ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst”, arXiv: 1812.03079v1, Dec. 7, 2018, 20 Pages. [cited by applicant]
Bickson, “Gaussian Belief Propagation: Theory and Application”, arXiv:0811.2518v1, Nov. 15, 2008, 86 pages. [cited by applicant]
Brown, “Iterative Solutions Of Games By Fictitious Play”, Activity Analysis of Production and Allocation, vol. 13, No. 1, 1951, pp. 374-376. [cited by applicant]
Casas et al., “IntentNet: Learning to Predict Intention from Raw Sensor Data”, Conference on Robot Learning (CoRL 2018), Oct. 29-31, 2018, Zurich, Switzerland, 10 pages. [cited by applicant]
Chen et al., “Multi-View 3D Object Detection Network for Autonomous Driving”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, Honolulu, HI, pp. 1907-1915. [cited by applicant]
Chen et al., “3D Object Proposals for Accurate Object Class Detection”, NIPS 2015, Dec. 7-12, 2015, Montreal, Canada, 9 pages. [cited by applicant]
Chen et al., “Monocular 3D Object Detection for Autonomous Driving”, IEEE/Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2016), Jun. 26-Jul. 1, 2016, Las Vegas, NV, pp. 2147-2156. [cited by applicant]
Chou et al., “Predicting Motion of Vulnerable Road Users using High-Definition Maps and Efficient ConvNets”, arXiv:1906.08469v1, Jun. 20, 2019, 9 pages. [cited by applicant]
Connolly et al., “Some Contagion Models of Speeding”, Accident Analysis & Prevention, vol. 25, Issue 1, Feb. 1993, pp. 57-66. [cited by applicant]
Cosgun et al., “Towards Full Automated Drive in Urban Environments: A Demonstration in GoMentum Station, California”, arXiv:1705.01187v1, May 2, 2017, 8 pages. [cited by applicant]
Cui et al., “Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks”, arXiv:1809.10732v2, Mar. 1, 2019, 7 pages. [cited by applicant]
Cui et al., “Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks”, arXiv:1809.10732v1, Sep. 18, 2018, 7 pages. [cited by applicant]
Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database”, IEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), Jun. 20-25, 2009, Miami, FL, pp. 1-8. [cited by applicant]
Deo et al., “Convolutional Social Pooling for Vehicle Trajectory Prediction”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 18-22, 2018, Salt Lake City, Utah, 9 pages. [cited by applicant]
Djuric et al., “Motion Prediction of Traffic Actors for Autonomous Driving using Deep Convolutional Networks”, arXiv:1808.05819v2, Sep. 16, 2018, 7 pages. [cited by applicant]
Engelcke et al., “Vote3Deep: Fast Object Detection in 3D Point Clouds Using Efficient Convolutional Neural Networks”, arXiv:1609.06666v2, Mar. 5, 2017, 7 pages. [cited by applicant]
Girshick, “Fast R-CNN”, 2015 IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, Santiago, Chile, pp. 1440-1448. [cited by applicant]
Gupta et al., “Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 18-22, 2018, Salt Lake City, Utah, pp. … [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, IEEE/Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2016), Jun. 26-Jul. 1, 2016, Las Vegas, NV, pp. 770-778. [cited by applicant]
He et al., “Mask R-CNN”, 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy, pp. 2961-2969. [cited by applicant]
Hu et al., “Probabilistic Prediction of Vehicle Semantic Intention and Motion”, arXiv:1804.03629v1, Apr. 10, 2018, 7 pages. [cited by applicant]
Kim et al., “Probabilistic Vehicle Trajectory Prediction Over Occupancy Grid Map via Recurrent Neural Network”. arXiv:1704.02049v2, Sep. 1, 2017, 6 pages. [cited by applicant]
Kingma et al., “Adam: A Method for a Stochastic Optimization”, arXiv:1412.6980v8, Jul. 23, 2015, 15 pages. [cited by applicant]
Kipf et al., “Neural Relational Inference for Interacting Systems”, arXiv:1802.04687v2, Jun. 6, 2018, 17 pages. [cited by applicant]
Ku et al., “Joint 3D Proposal Generation and Object Detection from View Aggregation”, arXiv:1712.02294v4, Jul. 12, 2018, 8 pages. [cited by applicant]
Lee et al., “DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, Honolulu, HI, pp. 336… [cited by applicant]
Li et al., “Situation Recognition with Graph Neural Networks”, 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy, pp. 4173-4182. [cited by applicant]
Li et al., “Stereo R-CNN based 3D Object Detection for Autonomous Driving”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019, Long Beach, CA, pp. 7644-7652. [cited by applicant]
Li et al., “Vehicle Detection from 3D Lidar Using Fully Convolutional Network”, arXiv:1608.07916v1, Aug. 29, 2016, 8 pages. [cited by applicant]
Li, “3D Fully Convolutional Network for Vehicle Detection in Point Cloud”, arXiv:1611.08069v2, Jan. 16, 2017, 5 pages. [cited by applicant]
Liang et al., “Deep Continuous Fusion for Multi-Sensor 3D Object Detection”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 16 pages. [cited by applicant]
Liang et al., “Multi-Task Multi-Sensor Fusion for 3D Object Detection”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019 Long Beach, CA, pp. 7345-7353. [cited by applicant]
Lin et al., “Focal Loss for Dense Object Detection”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, Honolulu, HI, pp. 2980-2988. [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, Honolulu, HI, pp. 2117-2125. [cited by applicant]
Luo et al., “Fast and Furious: Real Time End-to-End 3d Detection, Tracking and Motion Forecasting With a Single Convolutional Net”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 18-22… [cited by applicant]
Ma et al., “Arbitrary-Oriented Scene Text Detection via Rotation Proposals”, arXvi:1703.01086v3, Mar. 15, 2008, 11 pages. [cited by applicant]
Ma et al., “Forecasting Interactive Dynamics of Pedestrians With Fictitious Play”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu, HI, 9 pages. [cited by applicant]
Ma et al., “TrafficPredict: Trajectory Prediction for Heterogeneous Traffic-Agents”, The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), Jan. 27-Feb. 1, 2019, Honolulu, HI, pp. 6120-6127. [cited by applicant]
McNabb et al., “I'll Show You the Way: Risky Diver Behavior When “Following a Friend”.” Frontiers in Psychology, vol. 8, Article 705, May 9, 2017, 6 pages. [cited by applicant]
Meyer et al., “LaserNet: An Efficient Probabilistic 3D Object Detector for Autonomous Driving”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019 Long Beach, CA, pp. 12677-1268… [cited by applicant]
Meyer et al., “Sensor Fusion for Joint 3D Object Detection and Semantic Segmentation”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019 Long Beach, CA, pp. 41-48. [cited by applicant]
Qi et al., “3D Graph Neural Networks for RGBD Semantic Segmentation”, 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy, pp. 5199-5208. [cited by applicant]
Qi et al., “Frustum PointNets for 3D Object Detection from RGB-D Data”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 18-22, 2018, Salt Lake City, Utah, pp. 918-927. [cited by applicant]
Qi et al., “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, Honolulu, HI, pp. 652… [cited by applicant]
Qi et al., “PointNet++: Deep Hierarchical Feature Learning on Point Sets in Metric Space”, 31 [cited by applicant]
Sadeghian et al., “SoPhie: An Attentive GAN for Predicting Paths Compliant to Social and Physical Constraints”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019, Long Beach, C… [cited by applicant]
Santoro et al., “A simple neural network module for relational reasoning”, 31 [cited by applicant]
Scarselli et al., “The graph neural network”, IEEE Transactions on Neural Networks, vol. 20, No. 1, Jan. 2009, pp. 59-80. [cited by applicant]
Schlichtkrull et al., “Modeling Relational Data with Graph Convolutional Networks”, arXiv:1703.06103v4, Oct. 26, 2017, 9 pages. [cited by applicant]
Shi et al., “PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019 Long Beach, CA, pp. 770-779. [cited by applicant]
Simon et al., “Complex-YOLO: An Euler-Region-Proposal for Real-time 3D Object Detection on Point Clouds”, arXiv:1803.06199v2, Sep. 24, 2018, 14 pages. [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, arXiv:1409.1556v5, Dec. 23, 2014, 13 pages. [cited by applicant]
Soo Park et al., “Egocentric Future Localization” IEEE/Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2016), Jun. 26-Jul. 1, 2016, Las Vegas, NV, pp. 4697-4705. [cited by applicant]
Sun et al., “Actor-Centric Relation Network”, European Conference on Computer Vision, Sep. 8- 14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Sun et al., “Courteous Autonomous Cars”, arXiv:1808.02633, Aug. 16, 2018, 9 pages. [cited by applicant]
Sun et al., “Relational Action Forecasting”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019, Long Beach, CA, pp. 273-283 pages. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, 31 [cited by applicant]
Wang et al., “Deep Parametric Continuous Convolutional Neural Networks”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 18-22, 2018, Salt Lake City, Utah, pp. 2589-2597. [cited by applicant]
Weiss et al., “Correctness of belief propagation in Gaussian graphical models of arbitrary topology”, Neural Information Processing Systems 2000, Nov. 27-Dec. 2, 2000, Denver, CO, pp. 673-679. [cited by applicant]
Wilde, “Social Interaction Patterns in Driver Behavior: An Introductory Review” Human Factors: The Journal of the Human Factors and Ergonomics Society, vol. 11, No. 5, Oct. 1976, pp. 477-492. [cited by applicant]
Yang et al., “HDNET: Exploiting HD Maps for 3D Object Detection”, 2 [cited by applicant]
Yang et al., “PIXOR: Real-Time 3D Object Detection From Point Clouds”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 18-22, 2018 Salt Lake City, Utah, pp. 7652-7660. [cited by applicant]
Yi et al., “Pedestrian Behavior Understanding and Prediction with Deep Neural Networks”, European Conference on Computer Vision, Oct. 8-16, 2016, Amsterdam, The Netherlands, pp. 263-279. [cited by applicant]
Zeng et al., “End-to-end Interpretable Neural Motion Planner”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Jun. 16-20, 2019, Long Beach, CA, 10 pages. [cited by applicant]
Zhou et al., “VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 18-22, 2018, Salt Lake City, Utah, 10 pages. [cited by applicant]
Ziegler et al., “Making Bertha Drive—An Autonomous Journey on a Historic Route”, IEEE Intelligent Transportation Systems Magazine, vol. 6, No. 2, 2014, pp. 8-20. [cited by applicant]