IP Library Granted Patent US 12,570,329
Granted Patent B2
US 12,570,329 · App. 18/346,518 · Granted Mar 10, 2026

Systems and methods for actor motion forecasting within a surrounding environment of an autonomous vehicle

Inventors: Wenyuan Zeng (Toronto, CA); Renjie Liao (Toronto, CA); Raquel Urtasun (Toronto, CA); Ming Liang (Toronto, CA)
Assignee: AURORA OPERATIONS, INC.
B60W60/00276B60W40/072G06N20/00B60W2554/4041B60W2554/4044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,570,329
App. No.
18/346,518
Granted
Mar 10, 2026
Kind
B2
Abstract

Systems and methods are provided for forecasting the motion of actors within a surrounding environment of an autonomous platform. For example, a computing system of an autonomous platform can use machine-learned model(s) to generate actor-specific graphs with past motions of actors and the local map topology. The computing system can project the actor-specific graphs of all actors to a global graph. The global graph can allow the computing system to determine which actors may interact with one another by propagating information over the global graph. The computing system can distribute the interactions determined using the global graph to the individual actor-specific graphs. The computing system can then predict a motion trajectory for an actor based on the associated actor-specific graph, which captures the actor-to-actor interactions and actor-to-map relations.

Claims (73)

1 . A computer-implemented method comprising:

(a) obtaining data associated with a plurality of actors within an environment of an autonomous vehicle;

(b) obtaining map data indicating a plurality of lanes of the environment;

(c) for a respective actor of the plurality of actors:

determining one or more relevant lanes from the plurality of lanes; and

associating a plurality of local node embeddings with a plurality of nodes representing lane segments of the one or more relevant lanes and one or more edges representing relationships between at least a portion of the lane segments, wherein the plurality of local node embeddings encode information related to the respective actor;

(d) generating, based on the local node embeddings for the plurality of actors, a plurality of global node embeddings for the plurality of nodes;

(e) determining, based on the plurality of global node embeddings, a plurality of motion trajectories for the plurality of actors; and

(f) controlling the autonomous vehicle based on a vehicle motion trajectory selected for the autonomous vehicle based on the plurality of motion trajectories for the plurality of actors.

2 . The computer-implemented method of claim 1 , wherein the one or more edges indicate one or more of:

that a particular lane segment is left of another lane segment,

that a particular lane segment is right of another lane segment,

that a particular lane segment is a predecessor of another lane segment, or

that a particular lane segment is a successor of another lane segment.

3 . The computer-implemented method of claim 1 , wherein the information related to the respective actor comprises data describing a past motion of the respective actor and map features.

4 . The computer-implemented method of claim 1 , wherein at least one local node embedding of the plurality of local node embeddings encodes map information comprising one or more of: a geometric lane feature or a semantic lane feature.

5 . The computer-implemented method of claim 4 , wherein the geometric lane feature comprises feature data indicating one or more of:

a center location of a particular lane segment,

an orientation of a particular lane segment, or

a curvature of a particular lane segment.

6 . The computer-implemented method of claim 4 , wherein the semantic lane feature comprises feature data indicating a nature and intended purpose of an associated lane.

7 . The computer-implemented method of claim 6 , wherein the semantic lane feature comprises feature data indicating one or more of:

a type of a particular lane segment, or

an association of a particular lane segment with a traffic sign, a traffic light, or another type of traffic element.

8 . The computer-implemented method of claim 1 , wherein (d) comprises, for a respective global node and a corresponding global node embedding of the plurality of global node embeddings:

pooling information from a respective plurality of the local node embeddings for a respective plurality of actors that could interact with the respective global node; and

adding the pooled information to the respective global node.

9 . The computer-implemented method of claim 1 , comprising:

determining interactions between the plurality of actors by propagating information over the global node embeddings.

10 . An autonomous vehicle control system for controlling an autonomous vehicle, the autonomous vehicle control system comprising:

one or more processors; and

one or more non-transitory, computer-readable media storing instructions that are executable by the one or more processors to cause the autonomous vehicle control system to perform operations, the operations comprising:

(a) obtaining data associated with a plurality of actors within an environment of an autonomous vehicle;

(b) obtaining map data indicating a plurality of lanes of the environment;

(c) for a respective actor of the plurality of actors:

determining one or more relevant lanes from the plurality of lanes; and

associating a plurality of local node embeddings with a plurality of nodes representing lane segments of the one or more relevant lanes and one or more edges representing relationships between at least a portion of the lane segments, wherein the plurality of local node embeddings encode information related to the respective actor;

(d) generating, based on the local node embeddings for the plurality of actors, a plurality of global node embeddings for the plurality of nodes;

(e) determining, based on the plurality of global node embeddings, a plurality of motion trajectories for the plurality of actors; and

(f) controlling the autonomous vehicle based on a vehicle motion trajectory selected for the autonomous vehicle based on the plurality of motion trajectories for the plurality of actors.

11 . The autonomous vehicle control system of claim 10 , wherein the one or more edges indicate one or more of:

that a particular lane segment is left of another lane segment,

that a particular lane segment is right of another lane segment,

that a particular lane segment is a predecessor of another lane segment, or

that a particular lane segment is a successor of another lane segment.

12 . The autonomous vehicle control system of claim 10 , wherein at least one local node embedding of the plurality of local node embeddings encodes map information comprising one or more of: a geometric lane feature or a semantic lane feature.

13 . The autonomous vehicle control system of claim 12 , wherein the geometric lane feature comprises feature data indicating one or more of:

a center location of a particular lane segment,

an orientation of a particular lane segment, or

a curvature of a particular lane segment.

14 . The autonomous vehicle control system of claim 12 , wherein the semantic lane feature comprises feature data indicating a nature and intended purpose of an associated lane.

15 . The autonomous vehicle control system of claim 14 , wherein the semantic lane feature comprises feature data indicating one or more of:

a type of a particular lane segment, or

an association of a particular lane segment with a traffic sign, a traffic light, or another type of traffic element.

16 . The autonomous vehicle control system of claim 10 , wherein (d) comprises, for a respective global node and a corresponding global node embedding of the plurality of global node embeddings:

pooling information from a respective plurality of the local node embeddings for a respective plurality of actors that could interact with the respective global node; and

adding the pooled information to the respective global node.

17 . The autonomous vehicle control system of claim 10 , comprising:

determining interactions between the plurality of actors by propagating information over the global node embeddings.

18 . One or more non-transitory, computer-readable media storing instructions that are executable by one or more processors to cause an autonomous vehicle control system to perform operations, the operations comprising:

(a) obtaining data associated with a plurality of actors within an environment of an autonomous vehicle;

(b) obtaining map data indicating a plurality of lanes of the environment;

(c) for a respective actor of the plurality of actors:

determining one or more relevant lanes from the plurality of lanes; and

associating a plurality of local node embeddings with a plurality of nodes representing lane segments of the one or more relevant lanes and one or more edges representing relationships between at least a portion of the lane segments, wherein the plurality of local node embeddings encode information related to the respective actor,

(d) generating, based on the local node embeddings for the plurality of actors, a plurality of global node embeddings for the plurality of nodes;

(e) determining, based on the plurality of global node embeddings, a plurality of motion trajectories for the plurality of actors; and

(f) controlling the autonomous vehicle based on a vehicle motion trajectory selected for the autonomous vehicle based on the plurality of motion trajectories for the plurality of actors.

19 . The one or more non-transitory, computer-readable media of claim 18 , wherein the one or more edges indicate one or more of:

that a particular lane segment is left of another lane segment,

that a particular lane segment is right of another lane segment,

that a particular lane segment is a predecessor of another lane segment, or

that a particular lane segment is a successor of another lane segment.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 067069/0257 →
EMPLOYEE AGREEMENT Recorded Mar 29, 2024
From: LIANG, MING
To: UBER TECHNOLOGIES, INC.
Reel/Frame 066952/0892 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2024
From: ZENG, WENYUAN; LIAO, RENJIE; URTASUN, RAQUEL
To: UATC, LLC
Reel/Frame 066954/0078 →
Continuity (3)
Continuation 17528661 · Nov 17, 2021
Provisional Application 63114855 · Nov 17, 2020
Related Publication 20230347941A1 · Nov 2, 2023
References Cited (72)
US 20170263130A1 · Sane et al. · 2017 [cited by applicant]
US 20190049970A1 · Djuric et al. · 2019 [cited by applicant]
US 20190152490A1 · Lan et al. · 2019 [cited by applicant]
US 20190318041A1 · Bai et al. · 2019 [cited by applicant]
US 20200207339A1 · Neil et al. · 2020 [cited by applicant]
US 20200307563A1 · Ghafarianzadeh et al. · 2020 [cited by applicant]
US 20210009163A1 · Urtasun · 2021 [cited by examiner]
US 20210078603A1 · Nakhaei Sarvedani et al. · 2021 [cited by applicant]
US 20210139026A1 · Phan et al. · 2021 [cited by applicant]
US 20210174668A1 · Sun et al. · 2021 [cited by applicant]
US 20210295171A1 · Kamenev et al. · 2021 [cited by applicant]
US 20210312811A1 · Ohlarik et al. · 2021 [cited by applicant]
Argoverse Motion Forecasting Competition. Dec. 1, 2019, https://eval.ai/challenge/454/evaluation, retrieved on Jun. 8, 2022, 14 pages. [cited by applicant]
Alahi et al, “Social LSTM: Human Trajectory Prediction in Crowded Spaces”, Conference in Computer Vision and Pattern Recogition, Jun. 26-Jul. 1, 2016, Las Vegas, Nevada, United States, 11 pages. [cited by applicant]
Ba et al, “Layer Normalization”, arXiv:1607.06450v1, Jul. 21, 2016, 14 pages. [cited by applicant]
Bansal et al, “ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst”, arXiv:1812.03079v1, Dec. 7, 2018, 20 pages. [cited by applicant]
Bruna et al, “Spectral Networks and Locally Connected Networks on Graphs”, arXiv:1312.6203v3, May 21, 2014, 14 pages. [cited by applicant]
Casas et al, “SpAGNN: Spatially-Aware Graph Neural Networks for Relational Behavior Forecasting from Sensor Data”, arXiv:1910.08233V1, Oct. 18, 2019, 11 pages. [cited by applicant]
Casas et al, “Implicit Latent Variable Model for Scene-Consistent Motion Forecasting”, arXiv:2007.12036v1, Jul. 23, 2020, 44 pages. [cited by applicant]
Casas et al, “IntentNet: Learning to Predict Intention from Raw Sensor Data”, arXiv:2101.07907v1, Jan. 20, 2021, 10 pages. [cited by applicant]
Chai et al, “MultiPath: Multiple Probabilistic Anchor Trajectory Hypotheses for Behavior Prediction”, arXiv:1910.05449v1, Oct. 12, 2019, 14 pages. [cited by applicant]
Chang et al, “Argoverse: 3D Tracking and Forecasting with Rich Maps”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, United States, pp. 8748-8757. [cited by applicant]
Choi et al, “A Unified Framework for Multi-Target Tracking and Collective Activity Recognition”, European Conference on Computer Vision, Oct. 7-13, 2012, Firenze, Italy, 14 pages. [cited by applicant]
Choi et al, “Understanding Collective Activities of People from Videos”, Transactions on Pattern Analysis and Machine Intelligence, vol. 36, No. 6, Jun. 2014, pp. 1242-1257. [cited by applicant]
Cui et al, “Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks”, arXiv:1809.10732v2, Mar. 1, 2019, 7 pages. [cited by applicant]
Deo et al, “How Would Surround Vehicles Move? A Unified Framework for Maneuver Classification and Motion Prediction”, arXiv:1801.06523v1, Jan. 19, 2018, 12 pages. [cited by applicant]
Deo et al, “Convolutional Social Pooling for Vehicle Trajectory Prediction”, arXiv:1805.06771v1, May 15, 2018, 9 pages. [cited by applicant]
Gao et al, “VectorNet: Encoding HD Maps and Agent Dynamics for Vectorized Representation”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 11525-11533. [cited by applicant]
Garcia et al, “Few-Shot Learning with Graph Neural Networks”, arXiv:1711.04043v3, Feb. 20, 2018, 13 pages. [cited by applicant]
Gupta et al, “Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks”, arXiv:1803.10892v1, Mar. 29, 2018, 10 pages. [cited by applicant]
Hamilton et al, Inductive Representation Learning on Large Graphs, Conference on Neural Information Processing Systems, Dec. 4-9, 2017, Long Beach, California, United States, 11 pages. [cited by applicant]
He et al, “Mask R-CNN”, arXiv:1703.06870v3, Jan. 24, 2018, 12 pages. [cited by applicant]
Helbing et al, “Social Force Model for Pedestrian Dynamics”, Physical Review E, vol. 51, No. 5, May 1995, pp. 4282-4286. [cited by applicant]
Hu; “Collaborative Motion Prediction via Neural Motion Message Passing,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Oct. 2020, pp. 6318-6327, doi: 10.1109/CVPR42600.2020.00635; https://i… [cited by applicant]
Jain et al, “Discrete Residual Flow for Probabilistic Pedestrian Behavior Prediction”, arXiv:1910.08041v1, Oct. 17, 2019, 13 pages. [cited by applicant]
Khandelwal; “What-If Motion Prediction for Autonomous Driving”; Aug. 2020; Co RR, 2008.10587; https://arxiv.org/abs/2008.10587 ( Year: 2020). [cited by applicant]
Kingma et al, “Adam: A Method for Stochastic Optimization”, arXiv:1412.6980v9, Jan. 30, 2017, 15 pages. [cited by applicant]
Kipf et al, “Semi-Supervised Classification with Graph Convolutional Networks”, arXiv:1609.02907v4, Feb. 22, 2017, 14 pages. [cited by applicant]
Lee et al, “Desire: Distant Future Prediction in Dynamic Scenes with Interacting Agents”, arXiv:1704.04394v1, Apr. 14, 2017, 10 pages. [cited by applicant]
Li et al, “End-to-End Contextual Perception and Prediction with Interaction Transformer”, arXiv:2008.05927v1, Aug. 13, 2020, 8 pages. [cited by applicant]
Li et al, “Situation Recognition with Graph Neural Networks”, arXiv:1708.04320v1, Aug. 14, 2017, 22 pages. [cited by applicant]
Li et al, “Gated Graph Sequence Neural Networks”, arXiv:1511.05493v4, Sep. 22, 2017, 20 pages. [cited by applicant]
Liang et al, “Learning Lane Graph Representations for Motion Forecasting”, arXiv:2007.13732v1, Jul. 27, 2020, 18 pages. [cited by applicant]
Liao et al, “LanczosNet: Multi-Scale Deep Graph Convolutional Networks”, arXiv:1901.01484v2, Oct. 23, 2019, 18 pages. [cited by applicant]
Ma et al, “Forecasting Interactive Dynamics of Pedestrians with Fictitious Play”, arXiv: 1604.01431v3, Mar. 28, 2017, 9 pages. [cited by applicant]
Mehran et al, “Abnormal Crowd Behavior Detection using Social Force Model”, Conference on Computer Vision and Pattern Recognition, Jun. 22-24, 2009, Miami, Florida, United States, 8 pages. [cited by applicant]
Mercat et al, “Multi-Head Attention for Multi-Modal Joint Vehicle Motion Forecasting”, arXiv:1910.03650v3, Dec. 20, 2019, 7 pages. [cited by applicant]
Monti et al, “Geometric Deep Learning on Graphs and Manifolds using Mixture Model CNNs”, arXiv:1611.08402v3, Dec. 6, 2016, 13 pages. [cited by applicant]
Nair et al, “Rectified Linear Units Improve Restricted Boltzmann Machines”, International Conference on Machine Learning, Jun. 21-24, 2010, Haifa, Israel, 8 pages. [cited by applicant]
Phan-Minh et al., “CoverNet: Multimodal Behavior Prediction using Trajectory Sets”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 14074-14083. [cited by applicant]
Qi et al, 3D Graph Neural Networks for RGBD Semantic Segmentation, International Conference on Computer Vision, Oct. 22-29, 2017, Venice, Italy, pp. 5199-5208. [cited by applicant]
Ren et al, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, arXiv:1506.01497v3, Jan. 6, 2016, 14 pages. [cited by applicant]
Rhinehart et al, “R2P2: A ReparameteRized Pushforward Policy for Diverse, Precise Generative Path Forecasting”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Rhinehart et al, PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent Settings, arXiv:1905.01296v3, Sep. 30, 2019, 24 pages. [cited by applicant]
Sadat et al, “Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable Semantic Representations”, arXiv:2008.05930v1, Aug. 13, 2020, 28 pages. [cited by applicant]
Sadeghian et al, “SoPhie: An Attentive GAN for Predicting Paths Compliant to Social and Physical Constraints”, arXiv:1806.01482v2, Sep. 20, 2018, 9 pages. [cited by applicant]
Sadeghian et al, “CAR-Net: Clairvoyant Attentive Recurrent Network”, arXiv:1711.10661v3, Jul. 31, 2018, 17 pages. [cited by applicant]
Scarselli et al, “The Graph Neural Network Model”, Transactions on Neural Networks, vol. 20, No. 1, Jan. 2009, pp. 61-80. [cited by applicant]
Shrivastava et al, “Training Region-Based Object Detectors with Online Hard Example Mining”, arXiv:1604.03540v1, Apr. 12, 2016, 9 pages. [cited by applicant]
Song et al, “PiP: Planning Informed Trajectory Prediction for Autonomous Driving”, arXiv:2003.11476v2, Jan. 18, 2021, 16 pages. [cited by applicant]
Sun et al, “Relational Action Forecasting”, arXiv:1904.04231v1, Apr. 8, 2019, 11 pages. [cited by applicant]
Tang et al, “Multiple Futures Prediction”, arXiv:1911.00997v2, Dec. 6, 2019, 17 pages. [cited by applicant]
Teney et al, “Graph-Structured Representations for Visual Question Answering”, arXiv:1609.05600v2, Mar. 30, 2017, 17 pages. [cited by applicant]
Vaswani et al, “Attention Is All You Need”, arXiv:1706.03762v5, Dec. 6, 2017, 15 pages. [cited by applicant]
Vemula et al, “Social Attention: Modeling Attention in Human Crowds”, arXiv:1710.04689v2, Oct. 29, 2018, 7 pages. [cited by applicant]
Wang et al, “Deep Parametric Continuous Convolutional Neural Networks”, arXiv:2101.06742v1, Jan. 17, 2021, 18 pages. [cited by applicant]
Yamaguchi et al, “Who Are You With and Where Are You Going?”, Conference on Computer Vision and Pattern Recognition, Jun. 21-23, 2011, Colorado Springs, Colorado, United States, pp. 1345-1352. [cited by applicant]
Yu et al, “Multi-Scale Context Aggregation by Dilated Convolutions”, arXiv:1511.07122v3, Apr. 20, 2016, 13 pages. [cited by applicant]
Zeng et al, “End-to-End Interpretable Neural Motion Planner”, arXiv:2101.06679v1, Jan. 17, 2021, 10 pages. [cited by applicant]
Zeng et al., “DSDNet: Deep Structured Self-Driving Network”, arXiv:2008.06041v1, Aug. 13, 2020, 24 pages. [cited by applicant]
Zhao et al., “TNT: Target-DriveN Trajectory Prediction”, arXiv:2008.08294v2, Aug. 21, 2020, 12 pages. [cited by applicant]
Zhao et al., “Multi-Agent Tensor Fusion for Contextual Trajectory Prediction”, arXiv:1904.04776v2, Jul. 28, 2019, 9 pages. [cited by applicant]