IP Library › Granted Patent US 12,662,165
Granted Patent B2
US 12,662,165 · App. 18/600,159 · Granted Jun 23, 2026

Trajectory prediction by sampling sequences of discrete motion tokens

Inventors: Ari Seff (Mountain View, CA); Rami Al-Rfou (Menlo Park, CA); Angelo Brian Cera (Mountain View, CA); Nigamaa Nayakanti (San Jose, CA); Aurick Qikun Zhou (San Francisco, CA); Mason Ng (Palo Alto, CA); Benjamin Sapp (Marina del Rey, CA); Dian Chen (Auxtin, TX); Khaled Refaat (Mountain View, CA)
Assignee: Waymo LLC
B60W60/0027B60W50/0097B60W2420/00B60W2554/4044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,662,165
App. No.
18/600,159
Filed
Mar 8, 2024
Granted
Jun 23, 2026
Kind
B2
Art Unit
3665
USPC
701/27
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating trajectory predictions for one or more agents in an environment. In one aspect, a method comprises: obtaining scene context data characterizing a scene in an environment at a current time point and generating a respective predicted future trajectory for each of a plurality of agents in the scene at the current time point by sampling a sequence of discrete motion tokens that defines a joint future trajectory for the plurality of agents using a trajectory prediction neural network that is conditioned on the scene context data.

Claims (58)

1 . A method performed by one or more computers, the method comprising:

obtaining scene context data characterizing a scene in an environment at a current time point; and

generating a respective predicted future trajectory for each of a plurality of agents in the scene in the environment at the current time point,

the generating comprising:

sampling a sequence of discrete motion tokens that comprises a respective discrete motion token at each of a plurality of time steps and defines a joint future trajectory for the plurality of agents using a trajectory prediction neural network that is conditioned on the scene context data, the sampling comprising:

for a particular time step of the plurality of time steps, sampling the respective discrete motion token at the particular time step from a vocabulary of discrete motion tokens based on processing the discrete motion tokens corresponding to future time points that precede a particular future time point that corresponds to the particular time step using a trajectory decoder neural network conditioned on the scene context data.

2 . The method of claim 1 , wherein:

the scene context data comprises data generated from data captured by one or more sensors of an autonomous vehicle, and

the plurality of agents are agents in a vicinity of the autonomous vehicle in the environment.

3 . The method of claim 2 , further comprising:

controlling the autonomous vehicle based on the respective predicted future trajectories of the plurality of agents.

4 . The method of claim 1 , wherein:

each respective future trajectory comprises a respective predicted agent state for the agent at each of a plurality of future time points,

the sequence of discrete motion tokens comprises a respective discrete motion token at each of a plurality of time steps, and

each time step corresponds to a respective one of the plurality of agents and a respective future time point and the discrete motion token at the time step defines the predicted agent state for the corresponding agent at the corresponding future time point.

5 . The method of claim 4 , wherein each discrete motion token is selected from a vocabulary of motion tokens that each correspond to a different delta to be applied to a preceding agent state, and wherein the discrete motion token at the time step specifies a delta to be applied to a preceding agent state of the corresponding agent at a preceding time point that immediately precedes the corresponding future time point to generate the predicted agent state for the corresponding agent at the corresponding future time point.

6 . The method of claim 4 , wherein each respective predicted agent state comprises a predicted two-dimensional waypoint location of the corresponding agent at the corresponding future time point, and wherein each discrete motion token specifies a respective delta value for each of the two dimensions.

7 . The method of claim 4 , wherein the trajectory prediction neural network comprises (i) a scene encoder neural network.

8 . The method of claim 7 , wherein sampling a sequence of discrete motion tokens further comprises:

processing the scene context data using the scene encoder neural network to generate a respective scene encoding for each of the plurality of agents; and

wherein, for the particular time step of the plurality of time steps, sampling the respective discrete motion token at the particular time step from a vocabulary of discrete motion tokens based on processing the discrete motion tokens corresponding to future time points that precede a particular future time point that corresponds to the particular time step using a trajectory decoder neural network conditioned on the scene context data comprises:

processing an input sequence comprising the discrete motion tokens corresponding to future time points that precede a particular future time point that corresponds to the particular time step using the trajectory decoder neural network while conditioned on the respective scene encoding for the agent corresponding the particular time step to generate a score distribution over a vocabulary of discrete motion tokens; and

sampling the discrete motion token for the particular time step from the score distribution over the vocabulary of discrete motion tokens.

9 . The method of claim 8 , wherein processing the scene context data using the scene encoder neural network to generate a respective scene encoding for each of the plurality of agents comprises:

for each agent, extracting features from the scene context with respect to a frame of reference of the agent to generate agent-specific scene context data and processing the agent-specific scene context data using the scene encoder neural network to generate the respective scene encoding for the agent.

10 . The method of claim 8 , wherein the trajectory decoder neural network comprises one or more self-attention layers that perform self-attention over the input sequence and one or more cross-attention layers that perform cross-attention into the respective scene encoding for the agent corresponding to the particular time step.

11 . The method of claim 8 , wherein the trajectory decoder neural network implements temporally causal conditioning such that the discrete motion tokens at each particular time step are sampled conditioned only on discrete motion tokens corresponding to future time points that precede the future time point that corresponds to the particular time step and not any discrete motion tokens corresponding to future time points that are after the future time point that corresponds to the particular time step.

12 . The method of claim 1 , further comprising:

sampling one or more additional sequences of discrete motion tokens that each define a respective additional joint future trajectory for the plurality of agents using the trajectory prediction neural network that is conditioned on the scene context data; and

aggregating a plurality of joint future trajectories comprising the joint future trajectory and the additional joint future trajectories to generate (i) a plurality of predicted trajectory modes and (ii) a respective probability for each predicted trajectory mode.

13 . The method of claim 12 , wherein the plurality of joint future trajectories comprise a plurality of further joint future trajectories that are each defined by a respective further sequence of discrete motion tokens generated by a respective replica of the trajectory prediction neural network that is conditioned on the scene context data.

14 . The method of claim 1 , wherein the trajectory prediction neural network has been trained through imitation learning.

15 . The method of claim 1 , further comprising:

receiving a planned future trajectory for a particular one of the plurality of agents; and

determining a set of discrete motion tokens that represent the planned future trajectory, wherein sampling a sequence of discrete motion tokens that defines a joint future trajectory for the plurality of agents using a trajectory prediction neural network that is conditioned on the scene context data comprises fixing each discrete motion token in the sequence that correspond to the particular agent to be equal to a corresponding discrete motion token from the set of discrete motion tokens that represent the planned future trajectory.

16 . A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining scene context data characterizing a scene in an environment at a current time point; and

generating a respective predicted future trajectory for each of a plurality of agents in the scene in the environment at the current time point,

the generating comprising:

sampling a sequence of discrete motion tokens that comprises a respective discrete motion token at each of a plurality of time steps and defines a joint future trajectory for the plurality of agents using a trajectory prediction neural network that is conditioned on the scene context data, the sampling comprising:

for a particular time step of the plurality of time steps, sampling the respective discrete motion token at the particular time step from a vocabulary of discrete motion tokens based on processing the discrete motion tokens corresponding to future time points that precede a particular future time point that corresponds to the particular time step using a trajectory decoder neural network conditioned on the scene context data.

17 . The system of claim 16 , wherein:

the scene context data comprises data generated from data captured by one or more sensors of an autonomous vehicle, and

the plurality of agents are agents in a vicinity of the autonomous vehicle in the environment.

18 . The system of claim 17 , the operations further comprising:

controlling the autonomous vehicle based on the respective predicted future trajectories of the plurality of agents.

19 . The system of claim 16 , wherein:

each respective future trajectory comprises a respective predicted agent state for the agent at each of a plurality of future time points,

the sequence of discrete motion tokens comprises a respective discrete motion token at each of a plurality of time steps, and

each time step corresponds to a respective one of the plurality of agents and a respective future time point and the discrete motion token at the time step defines the predicted agent state for the corresponding agent at the corresponding future time point.

20 . One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining scene context data characterizing a scene in an environment at a current time point; and

generating a respective predicted future trajectory for each of a plurality of agents in the scene in the environment at the current time point,

the generating comprising:

sampling a sequence of discrete motion tokens that comprises a respective discrete motion token at each of a plurality of time steps and defines a joint future trajectory for the plurality of agents using a trajectory prediction neural network that is conditioned on the scene context data, the sampling comprising:

for a particular time step of the plurality of time steps, sampling the respective discrete motion token at the particular time step from a vocabulary of discrete motion tokens based on processing the discrete motion tokens corresponding to future time points that precede a particular future time point that corresponds to the particular time step using a trajectory decoder neural network conditioned on the scene context data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2024
From: SEFF, ARI; AL-RFOU, RAMI; CERA, ANGELO BRIAN; NAYAKANTI, NIGAMAA; ZHOU, AURICK QIKUN; NG, MASON; SAPP, BENJAMIN; CHEN, DIAN; REFAAT, KHALED
To: WAYMO LLC
Reel/Frame 068079/0319 →
Continuity (2)
Provisional Application 63450953 · Mar 8, 2023
Related Publication 20240300542A1 · Sep 12, 2024
References Cited (51)
US 12296857B2 · Afshar · 2025 [cited by examiner]
US 12311972B2 · Pronovost · 2025 [cited by examiner]
US 20200159232A1 · Refaat · 2020 [cited by examiner]
US 20220355825A1 · Deo · 2022 [cited by examiner]
US 20230406360A1 · Al-Rfou · 2023 [cited by examiner]
US 20240101150A1 · Pronovost · 2024 [cited by examiner]
Hussein, Ahmed et al.. “Imitation Learning: A Survey of Learning Methods.” Archived Mar. 6, 2020. URL: https://web.archive.org/web/20200306111131/https://www.open-access.bcu.ac.uk/5045/1/lmitation%20Learning%20A%20Surve… [cited by examiner]
Amirloo et al., “LatentFormer: Multi-agent-transformer based interaction modeling and trajectory prediction,” CoRR, submitted on Mar. 3, 2022, arXiv:2203.01880v1, 10 pages. [cited by applicant]
Caesar et al., “nuscenes: A multimodal dataset for autonomous driving,” In CVPR, 2020, p. 11621-11631. [cited by applicant]
Casas et al., “Implicit latent variable model for scene-consistent motion forecasting,” CoRR, submitted on Jul. 23, 2020, arXiv:2007.12036v1, 44 pages. [cited by applicant]
Casas et al., “Intentnet: Learning to predict intention from raw sensor data,” In Conference on Robot Learning, Oct. 23, 2018, p. 947-956. [cited by applicant]
Casas et al., “Spagnn: Spatially-aware graph neural networks for relational behavior forecasting from sensor data,” CoRR, submitted on Oct. 18, 2019, arXiv:1910.08233v1, 11 pages. [cited by applicant]
Chai et al., “Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction,” CoRR, submitted on Oct. 12, 2019, arXiv:1910.05449v1, 14 pages. [cited by applicant]
Chang et al., “Argoverse: 3d tracking and forecasting with rich maps,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, p. 8748-8757. [cited by applicant]
Cui et al., “Lookout: Diverse multi-future prediction and planning for self-driving,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, p. 16107-16116. [cited by applicant]
Cui et al., “Multimodal trajectory predictions for autonomous driving using deep convolutional networks,” CoRR, submitted on Sep. 18, 2018, arXiv:1809.10732v2, 7 pages. [cited by applicant]
Engstrom et al., “Modeling road user response timing in naturalistic settings: a surprise-based framework,” CoRR, submitted on Aug. 18, 2022, arXiv:2208.08651v2, 19 pages. [cited by applicant]
Ettinger et al., “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” CoRR, submitted on Apr. 20, 2021, arXiv:2104.10133v1, 15 pages. [cited by applicant]
Girgis et al., “Latent variable sequential set transformers for joint multi-agent motion prediction,” CoRR, submitted on Feb. 19, 2021, arXiv:2104.00563v3, 26 pages. [cited by applicant]
Holtzman et al., “The curious case of neural text degeneration,” CoRR, submitted on Apr. 22, 2019, arXiv:1904.09751v2, 16 pages. [cited by applicant]
Hong et al., “Rules of the road: Predicting driving behavior with a convolutional model of semantic interactions,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, p. 8454-8462. [cited by applicant]
Jaegle et al., “Perceiver: General perception with iterative attention,” In ICML, Jul. 1, 2021, p. 4651-4664. [cited by applicant]
Jia et al., “Hdgt: Heterogeneous driving graph transformer for multi-agent trajectory prediction via scene encoding,” CoRR, submitted on Apr. 30, 2022, arXiv:2205.09753v2, 15 pages. [cited by applicant]
Khandelwal et al., “What-if motion prediction for autonomous driving,” CoRR, submitted on Aug. 24, 2020, arXiv:2008.10587v1, 16 pages. [cited by applicant]
Konev, “Mpa: Multipath++ based architecture for motion prediction,” CoRR, submitted on Jun. 20, 2022, arXiv:2206.10041v1, 3 pages. [cited by applicant]
Lee et al., “Desire: Distant future prediction in dynamic scenes with interacting agents,” In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, p. 336-345. [cited by applicant]
Lee et al., “Set transformer: A framework for attention-based permutation-invariant neural networks,” In ICML, May 24, 2019, p. 3744-3753. [cited by applicant]
Liang et al., “Learning lane graph representations for motion forecasting,” In European Conference on Computer Vision, 2020, p. 541-556. [cited by applicant]
Liu et al., “Deep structured reactive planning,” CoRR, submitted on Jan. 18, 2021, arXiv:2101.06832v2, 15 pages. [cited by applicant]
Lu et al., “Kemp: Keyframe-based hierarchical end-to-end deep model for long-term trajectory prediction,” CoRR, submitted on May 10, 2022, arXiv:2205.04624v1, 7 pages. [cited by applicant]
Luo et al., “Densetnt: efficient vehicle type classification neural network using satellite imagery,” CoRR, submitted on Sep. 27, 2022, arXiv:2209.13500v1, 10 pages. [cited by applicant]
Luo et al., “JFP: Structured multi-agent interactive trajectories forecasting for autonomous driving,” CoRR, submitted on Dec. 16, 2022, arXiv:2212.08710v1, 13pages. [cited by applicant]
Nash et al., “PolyGen: An autoregressive generative model of 3D meshes,” International Conference on Machine Learning, Nov. 21, 2020, p. 7220-7229. [cited by applicant]
Nayakanti et al., “Wayformer: Motion forecasting via simple & efficient attention networks,” CoRR, submitted on Jul. 12, 2022, arXiv:2207.05844v1, 20 pages. [cited by applicant]
Ngiam et al., “Scene transformer: A unified architecture for predicting future trajectories of multiple agents,” CoRR, submitted on Jun. 15, 2021, arXiv:2106.08417v3, 25 pages. [cited by applicant]
Oord et al., “WaveNet: A generative model for raw audio,” CoRR, submitted on Sep. 12, 2016, arXiv:1609.03499v2, 15 pages. [cited by applicant]
Ortega et al., “Shaking the foundations: delusions in sequence models for interaction and control,” CoRR, submitted on Oct. 20, 2021, arXiv:2110.10819v1, 16 pages. [cited by applicant]
Rhinehart et al., “PRECOG: Prediction conditioned on goals in visual multi-agent settings,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, p. 2821-2830. [cited by applicant]
Robicquet et al., “Learning social etiquette: Human trajectory prediction in crowded scenes,” In European Conference on Computer Vision (ECCV), 2020, 17 pages. [cited by applicant]
Ross et al., “A reduction of imitation learning and structured prediction to noregret online learning,” In AISTATS, Jun. 14, 2011, p. 627-635. [cited by applicant]
Salzmann et al., “Trajectron++: Multi-agent generative trajectory forecasting with heterogeneous data for control,” CoRR, submitted on Jan. 9, 2020, arXiv:2001.03093v1, 13 pages. [cited by applicant]
Shi et al., “Motion transformer with global intention localization and local movement refinement,” Advances in Neural Information Processing Systems, Dec. 6, 2022, 35:6531-43. [cited by applicant]
Song et al., “PiP: Planning-informed trajectory prediction for autonomous driving,” CoRR, submitted on Mar. 25, 2020, arXiv:2003.11476v2, 16 pages. [cited by applicant]
Sun et al., “M2I: From factored marginal trajectory prediction to interactive prediction,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, p. 6543-6552. [cited by applicant]
Tang et al., “Multiple futures prediction,” In Advances in neural information processing systems, 2019, 11 pages. [cited by applicant]
Tolstaya et al., “Identifying driver interactions via conditional behavior prediction,” CoRR, submitted on Apr. 20, 2021, arXiv:2104.09959v2, 7 pages. [cited by applicant]
Varadarajan et al., “Multipath++: Efficient information fusion and trajectory aggregation for behavior prediction,” CoRR, submitted on Nov. 29, 2021, arXiv:2111.14973v3, 22 pages. [cited by applicant]
Vaswani et al., “Attention is all you need,” In Advances in Neural Information Processing Systems, 2017, 11 pages. [cited by applicant]
Yuan et al., “Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, p. 9813-9823. [cited by applicant]
Zhan, et al., “Interaction dataset: An international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps,” CoRR, submitted on Sep. 30, 2019, arXiv:1910.03088v1, 13 pages. [cited by applicant]
Zhao et al., “Tnt: Target-driven trajectory prediction,” In Conference on Robot Learning, Oct. 4, 2021, p. 895-904. [cited by applicant]