IP Library › Granted Patent US 12,314,853
Granted Patent B2
US 12,314,853 · App. 17/947,052 · Granted May 27, 2025

Training agent trajectory prediction neural networks using distillation

Inventors: Bertrand Robert Douillard (San Francisco, CA); Dijia Su (Jersey City, NJ)
Assignee: Waymo LLC
G06N3/08G06T7/20G06T2207/20081G06T2207/20084G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,853
App. No.
17/947,052
Granted
May 27, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training trajectory prediction neural networks using distillation.

Claims (63)

1. A method performed by one or more computers and for training a scene-centric trajectory prediction neural network that has a plurality of scene-centric parameters and that is configured to receive a scene input comprising features of a scene in an environment that includes a plurality of agents and to process the features in accordance with the scene-centric parameters to generate as output a respective trajectory prediction for each of the plurality of agents, the method comprising:

obtaining a batch of one or more training examples, each training example comprising data characterizing a respective scene in the environment that includes at least a respective plurality of agents;

for each training example in the batch:

generating, from the training example, a respective agent-centric input for each of the respective plurality of agents in the respective scene characterized by the training example, the respective agent-centric input for each agent comprising features that characterize the respective scene and that are represented in an agent-specific coordinate frame that is specific to the agent;

for each agent in the respective scene, processing the agent-centric input using a trained agent-centric trajectory prediction neural network, wherein the trained agent-centric trajectory prediction neural network has been trained to receive the agent-centric input that comprises features that are represented in the agent-specific coordinate frame for the agent to generate a trajectory prediction for the agent;

generating, from the training example, a scene input characterizing the respective scene characterized by the training example in a shared coordinate frame that is common to all of the agents in the respective scene; and

processing the scene input using the scene-centric trajectory prediction neural network and in accordance with current values of the scene-centric parameters to generate as output a respective trajectory prediction for each of the respective plurality of agents in the respective scene;

determining a gradient with respect to the scene-centric parameters of a loss function that includes one or more terms that measure, for each training example and for each of the respective plurality of agents in the respective scene characterized by the training example, a difference between (i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network; and

updating, using the gradient, the current values of the scene-centric parameters of the scene-centric trajectory prediction neural network.

2. The method of claim 1 , wherein each training example further comprises a respective ground truth observed trajectory for one or more of the respective plurality of agents in the respective scene characterized by the training example, and wherein the loss function further comprises an additional term that measures, for each training example and for each of the one or more agents for which a respective ground truth observed trajectory is included in the training example, a difference between (i) the respective ground truth observed trajectory for the agent and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network.

3. The method of claim 2 , wherein the additional term is only included in the loss function after the scene-centric trajectory prediction neural network has already been trained on a threshold number of batches of training examples.

4. The method of claim 1 , wherein, for each of the respective plurality of agents in each of the training examples:

(i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network defines a second likelihood distribution over possible future trajectories for the agent and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network defines a first likelihood distribution over possible future trajectories, and

wherein the one or more terms measure a difference between the first likelihood distribution and the second likelihood distribution.

5. The method of claim 4 , wherein, for each of the respective plurality of agents in each of the training examples:

(i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network includes data defining a plurality of second possible future trajectories and a respective second score for each of the plurality of second possible future trajectories,

and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network includes data defining a plurality of first possible future trajectories and a respective first score for each of the plurality of first possible future trajectories.

6. The method of claim 5 , wherein the one or more terms include, for each of the first possible future trajectories:

a first term that measures a difference between the first score for the first possible future trajectory and a second score for a corresponding second possible future trajectory.

7. The method of claim 5 , wherein:

for each first possible future trajectory, the data defining the first possible future trajectory are parameters of a probability distribution around the first possible future trajectory, and

for each second possible future trajectory, the data defining the second possible future trajectory are parameters of a probability distribution around the second possible future trajectory.

8. The method of claim 7 , wherein the one or more terms include, for each of the first possible future trajectories:

a third term that measures a likelihood assigned to a corresponding second possible future trajectory by the probability distribution around the first possible future trajectory.

9. The method of claim 7 , wherein the one or more terms include, for each of the first possible future trajectories:

a fourth term that measures a divergence between the probability distribution around the first possible future trajectory and the probability distribution around a corresponding second possible future trajectory.

10. The method of claim 7 , wherein the one or more terms include:

(i) a fifth term that measures a first score assigned to a first possible future trajectory that is closest to a possible future trajectory that has been sampled using the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network.

11. The method of claim 10 , wherein the one or more terms include

(ii) a sixth term that measures a likelihood assigned to the sampled possible future trajectory by the probability distribution around the second possible future trajectory that is closest to the sampled possible future trajectory.

12. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for training a scene-centric trajectory prediction neural network that has a plurality of scene-centric parameters and that is configured to receive a scene input comprising features of a scene in an environment that includes a plurality of agents and to process the features in accordance with the scene-centric parameters to generate as output a respective trajectory prediction for each of the plurality of agents, the operations comprising:

obtaining a batch of one or more training examples, each training example comprising data characterizing a respective scene in the environment that includes at least a respective plurality of agents;

for each training example in the batch:

generating, from the training example, a respective agent-centric input for each of the respective plurality of agents in the respective scene characterized by the training example, the respective agent-centric input for each agent comprising features that characterize the respective scene and that are represented in an agent-specific coordinate frame that is specific to the agent;

for each agent in the respective scene, processing the agent-centric input using a trained agent-centric trajectory prediction neural network, wherein the trained agent-centric trajectory prediction neural network has been trained to receive the agent-centric input that comprises features that are represented in the agent-specific coordinate frame for the agent to generate a trajectory prediction for the agent;

generating, from the training example, a scene input characterizing the respective scene characterized by the training example in a shared coordinate frame that is common to all of the agents in the respective scene; and

processing the scene input using the scene-centric trajectory prediction neural network and in accordance with current values of the scene-centric parameters to generate as output a respective trajectory prediction for each of the respective plurality of agents in the respective scene;

determining a gradient with respect to the scene-centric parameters of a loss function that includes one or more terms that measure, for each training example and for each of the respective plurality of agents in the respective scene characterized by the training example, a difference between (i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network; and

updating, using the gradient, the current values of the scene-centric parameters of the scene-centric trajectory prediction neural network.

13. The system of claim 12 , wherein each training example further comprises a respective ground truth observed trajectory for one or more of the respective plurality of agents in the respective scene characterized by the training example, and wherein the loss function further comprises an additional term that measures, for each training example and for each of the one or more agents for which a respective ground truth observed trajectory is included in the training example, a difference between (i) the respective ground truth observed trajectory for the agent and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network.

14. The system of claim 13 , wherein the additional term is only included in the loss function after the scene-centric trajectory prediction neural network has already been trained on a threshold number of batches of training examples.

15. The system of claim 12 , wherein, for each of the respective plurality of agents in each of the training examples:

(i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network defines a second likelihood distribution over possible future trajectories for the agent and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network defines a first likelihood distribution over possible future trajectories, and

wherein the one or more terms measure a difference between the first likelihood distribution and the second likelihood distribution.

16. The system of claim 15 , wherein, for each of the respective plurality of agents in each of the training examples:

(i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network includes data defining a plurality of second possible future trajectories and a respective second score for each of the plurality of second possible future trajectories,

and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network includes data defining a plurality of first possible future trajectories and a respective first score for each of the plurality of first possible future trajectories.

17. The system of claim 16 , wherein the one or more terms include, for each of the first possible future trajectories:

a first term that measures a difference between the first score for the first possible future trajectory and a second score for a corresponding second possible future trajectory.

18. The system of claim 16 , wherein:

for each first possible future trajectory, the data defining the first possible future trajectory are parameters of a probability distribution around the first possible future trajectory, and

for each second possible future trajectory, the data defining the second possible future trajectory are parameters of a probability distribution around the second possible future trajectory.

19. The system of claim 18 , wherein the one or more terms include, for each of the first possible future trajectories:

a third term that measures a likelihood assigned to a corresponding second possible future trajectory by the probability distribution around the first possible future trajectory.

20. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a scene-centric trajectory prediction neural network that has a plurality of scene-centric parameters and that is configured to receive a scene input comprising features of a scene in an environment that includes a plurality of agents and to process the features in accordance with the scene-centric parameters to generate as output a respective trajectory prediction for each of the plurality of agents, the operations comprising:

obtaining a batch of one or more training examples, each training example comprising data characterizing a respective scene in the environment that includes at least a respective plurality of agents;

for each training example in the batch:

generating, from the training example, a respective agent-centric input for each of the respective plurality of agents in the respective scene characterized by the training example, the respective agent-centric input for each agent comprising features that characterize the respective scene and that are represented in an agent-specific coordinate frame that is specific to the agent;

for each agent in the respective scene, processing the agent-centric input using a trained agent-centric trajectory prediction neural network, wherein the trained agent-centric trajectory prediction neural network has been trained to receive the agent-centric input that comprises features that are represented in the agent-specific coordinate frame for the agent to generate a trajectory prediction for the agent;

generating, from the training example, a scene input characterizing the respective scene characterized by the training example in a shared coordinate frame that is common to all of the agents in the respective scene; and

processing the scene input using the scene-centric trajectory prediction neural network and in accordance with current values of the scene-centric parameters to generate as output a respective trajectory prediction for each of the respective plurality of agents in the respective scene;

determining a gradient with respect to the scene-centric parameters of a loss function that includes one or more terms that measure, for each training example and for each of the respective plurality of agents in the respective scene characterized by the training example, a difference between (i) the trajectory prediction generated for the agent by the trained agent-centric trajectory prediction neural network and (ii) the trajectory prediction generated for the agent by the scene-centric trajectory prediction neural network; and

updating, using the gradient, the current values of the scene-centric parameters of the scene-centric trajectory prediction neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2022
From: DOUILLARD, BERTRAND ROBERT; SU, DIJIA
To: WAYMO LLC
Reel/Frame 061477/0262 →
Continuity (3)
Provisional Application 63248950 · Sep 27, 2021
Provisional Application 63245173 · Sep 16, 2021
Related Publication 20230082079A1 · Mar 16, 2023
References Cited (44)
US 20190033085A1 · Ogale · 2019 [cited by examiner]
US 20190034794A1 · Ogale · 2019 [cited by examiner]
US 20200174490A1 · Ogale · 2020 [cited by examiner]
US 20210001897A1 · Chai · 2021 [cited by examiner]
[No Author Listed]., “Narrowing the coordinate-frame gap in behavior prediction models: Distillation for efficient and accurate scene-centric motion forecasting,” Presented at Proceedings of ICRA 2022, Philadelphia, PA,… [cited by applicant]
Alahi et al., “Social LSTM: Human Trajectory Prediction in Crowded Spaces,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, 961-971. [cited by applicant]
Bansal et al., “Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst,” csRO, Submitted on Dec. 7, 2018, arXiv:1812.03079, 20 pages. [cited by applicant]
Buhet et al. “Plop: Probabilistic polynomial objects trajectory planning for autonomous driving,” csCV, Submitted on Mar. 9, 2020, arXiv.2003.08744, 20 pages. [cited by applicant]
Caesar et al., “nuscenes: A multimodal dataset for autonomous driving,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 11621-11631. [cited by applicant]
Casas et al., “Intentnet: Learning to predict intention from raw sensor data,” Proceedings of the 2nd Conference on Robot Learning, PMLR, Oct. 2018, 87:947-956. [cited by applicant]
Casas et al., “Spagnn: Spatially-aware graph neural networks for relational behavior forecasting from sensor data,” csCV, Submitted on Oct. 18, 2019, arXiv:1910.08233, 11 pages. [cited by applicant]
Chai et al, “Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction,” csLG, Submitted on Oct. 12, 2019, arXiv:1910.05449, 1-14. [cited by applicant]
Chang et al., “Argoverse: 3d tracking and forecasting with rich maps,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 8748-8757. [cited by applicant]
Cheng et al, “Explaining knowledge distillation by quantifying the knowledge,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 12925-12935. [cited by applicant]
Ettinger et al., “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, 9… [cited by applicant]
Gao et al, “Vectornet: Encoding hd maps and agent dynamics from vectorized representation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 11525-11533. [cited by applicant]
Gupta et al, “Social GAN: Socially acceptable trajectories with generative adversarial networks,” Proceeding of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018, 2255-2264. [cited by applicant]
Hinton et al., “Distilling the knowledge in a neural network,” statML, Submitted on Mar. 9, 2015, arXiv:1503.02531, 9 pages. [cited by applicant]
Hong et al, “Rules of the road: Predicting driving behavior with a convolutional model of semantic interactions,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 8454… [cited by applicant]
Khandelwal et al, “What-if motion prediction for autonomous driving,” csLG, Submitted on Aug. 24, 2020, arXiv:2008.10587, 16 pages. [cited by applicant]
Kim et al., “Sequence-level knowledge distillation,” csCL, Submitted on Jun. 25, 2016, arXiv:1606.07947, 2016, 11 pages. [cited by applicant]
Lang et al, “Pointpillars: Fast encoders for object detection from point clouds,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 12697-12705. [cited by applicant]
Lee et al., “Desire: Distant future prediction in dynamic scenes with interacting agents,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, 336-345. [cited by applicant]
Liang et al., “Learning lane graph representations for motion forecasting,” csCV Submitted on Jul. 27, 2020, arXiv:2007.13732, 18 pages. [cited by applicant]
Liu et al., “Multimodal motion prediction with stacked transformers,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, 7577-7586. [cited by applicant]
Mercat et al., “Multi-head attention for joint multi-modal vehicle motion forecasting,” 2020 IEEE International Conference on Robotics and Automation (ICRA), HAL open science, May 2020, 8 pages. [cited by applicant]
Mercat et al., “Multi-head attention for multi-modal joint vehicle motion forecasting,” csLG, Submitted on Oct. 8, 2019, arXiv:1910.03650, 7 pages. [cited by applicant]
Ngiam et al., “Scene transformer: A unified multi-task model for behavior prediction and planning,” csCV, Submitte Jun. 15, 2021, arXiv:2106.08417, 21 pages. [cited by applicant]
Phan-Minh et al, “Covernet: Multimodal behavior prediction using trajectory sets,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 14074-14083. [cited by applicant]
Phuong et al., “Towards understanding knowledge distillation,” Proceedings of the 36th International Conference on Machine Learning, PMLR, Jun. 2019, 97:5142-5151. [cited by applicant]
Rhinehart et al., “PRECOG: Prediction conditioned on goals in visual multi-agent settings,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Jun. 2019, 2821-2830. [cited by applicant]
Rhinehart et al., “R2p2: A reparameterized pushforward policy for diverse, precise generative path forecasting,” Proceedings of the European Conference on Computer Vision (ECCV), Jun. 2018, 772-788. [cited by applicant]
Salzmann et al., “Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data,” Proceeding of the European Conference on Computer vision (ECCV), Dec. 4, 2020, 683-700. [cited by applicant]
Salzmann et al., “Trajectron++: Multi-agent generative trajectory forecasting with heterogeneous data for control,” csRO, Submitted on Jan. 9, 2020, arXiv:2001.03093, 13 pages. [cited by applicant]
Su et al., “Narrowing the coordinate-frame gap in behavior prediction models: Distillation for efficient and accurate scene-centric motion forecasting,” 2022 International Conference on Robotics and Automation (ICRA), M… [cited by applicant]
Tan et al, “Efficientnet: Rethinking model scaling for convolutional neural networks,” Proceedings of the 36th International Conference on Machine Learning. PMLR, Jun. 2019, 97:6105-6114. [cited by applicant]
Tang et al, “Multiple futures prediction,” Advances in Neural Information Processing Systems, Dec. 8-14, 2019, vol. 32, 11 pages. [cited by applicant]
Ye et al., “Tpcn: Temporal point cloud networks for motion forecasting,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, 11318-11327. [cited by applicant]
Yuan et al., “Diverse trajectory forecasting with determinantal point processes,” csCV, Submitted on Jul. 11, 2019, arXiv:1907.04967, 15 pages. [cited by applicant]
Zeng et al., “Lanercnn: Distributed representations for graph-centric motion forecasting,” csCV, Submitted on Jan. 17, 2021, arXiv2101.06653, 14 pages. [cited by applicant]
Zhan et al., “Interaction dataset: An international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps,” csRO, Submitted on Sep. 30, 2019, arXiv:1910.03088, 13 pages. [cited by applicant]
Zhao et al, “Multi-agent tensor fusion for contextual trajectory prediction,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 12126-12134. [cited by applicant]
Zhao et al., “Tnt: Target-driven trajectory prediction,” csCV, Submitted on Aug. 19, 2020, arXiv:2008.08294, 12 pages. [cited by applicant]
Zhou et al, “Understanding knowledge distillation in non-autoregressive machine translation,” csCL, Submitted on Nov. 7, 2019, arXiv:1911.02727, 14 pages. [cited by applicant]
Cited By (1)
US 12,565,234