IP Library Granted Patent US 12,434,739
Granted Patent B2
US 12,434,739 · App. 18/087,540 · Granted Oct 7, 2025

Latent variable determination by a diffusion model

Inventor: Ethan Miller Pronovost (Redwood City, CA)
Assignee: Zoox, Inc.
B60W60/0027B60W40/04G06N3/04B60W2554/404B60W2554/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,434,739
App. No.
18/087,540
Granted
Oct 7, 2025
Kind
B2
Abstract

A computing device can implement techniques for predicting an object trajectory or scene information. For example, the techniques may include inputting latent variable data into a machine learned model. The machine learned model may output an object trajectory (e.g., position data, velocity data, acceleration data, etc.) for one or more objects in the environment based on the latent variable data. The object trajectory can be sent to a vehicle computing device for consideration during vehicle planning, which may include simulation.

Claims (70)

1. A method comprising:

receiving, by a diffusion model, map data representing an environment, wherein the diffusion model comprises one or more self-attention layers;

receiving, by the diffusion model, condition data representing a state or an action of an object in the environment;

determining, by the diffusion model and based at least in part on the map data, the condition data, and output data associated with the one or more self-attention layers, latent variable data associated with the object;

inputting the latent variable data into a machine learned model;

determining, by the machine learned model and based at least in part on the latent variable data, an object trajectory for the object to follow in the environment; and

causing control of a vehicle in the environment by executing one or more vehicle maneuvers based at least in part on the object trajectory to avoid potential interaction between the vehicle and the object.

2. The method of claim 1 , wherein:

the machine learned model comprises at least one of a Generative Adversarial Network (GAN), a Graph Neural Network (GNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a transformer model.

3. The method of claim 1 , wherein:

the diffusion model is trained based at least in part on a set of conditions, at least one condition of the set of conditions comprising a previous action, a previous position, a previous trajectory, or a previous acceleration of the object.

4. The method of claim 1 , wherein:

the condition data comprises a token from a codebook associated with a transformer model, and

the token representing a behavior of the object.

5. The method of claim 1 , wherein:

the condition data comprises a node from a Graph Neural Network, and

the node representing a behavior of the object.

6. The method of claim 1 , wherein:

the condition data represents one of: a yield action, a drive straight action, a left turn action, a right turn action, a brake action, an acceleration action, a steering action, a lane change action, a position, a heading, or an acceleration of the object.

7. The method of claim 1 , wherein:

the diffusion model is trained by incrementally denoising data to generate an output based on a conditional input.

8. The method of claim 1 , further comprising:

mapping the latent variable data to continuous variables associated with feature vectors,

wherein inputting the latent variable data into the machine learned model comprises inputting the continuous variables associated with the feature vectors, and

determining, by the machine learned model, the object trajectory is based at least in part on the continuous variables associated with the feature vectors.

9. The method of claim 1 , wherein:

the object is a first object,

the condition data represents a potential action of a second object in the environment, and

the latent variable data represents a potential interaction between the second object and the first object.

10. The method of claim 1 , wherein:

the environment is a simulated environment,

the condition data represents a feature of the simulated environment, and further comprising:

determining, by the machine learned model and based at least in part on the feature of the simulated environment, scene data for testing or verifying a scenario in the simulated environment.

11. One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform actions comprising:

receiving, by a diffusion model, map data representing an environment, wherein the diffusion model comprises one or more self-attention layers;

receiving, by the diffusion model, condition data representing a state or an action of an object in the environment;

determining, by the diffusion model and based at least in part on the map data, the condition data, and output data associated with the one or more self-attention layers, latent variable data associated with the object;

inputting the latent variable data into a machine learned model;

determining, by the machine learned model and based at least in part on the latent variable data, an object trajectory for the object to follow in the environment; and

causing control of a vehicle in the environment by executing one or more vehicle maneuvers based at least in part on the object trajectory to avoid potential interaction between the vehicle and the object.

12. The one or more non-transitory computer-readable media of claim 11 , wherein:

the machine learned model comprises at least one of a Generative Adversarial Network (GAN), a Graph Neural Network (GNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a transformer model.

13. The one or more non-transitory computer-readable media of claim 11 , wherein:

the diffusion model is trained based at least in part on a set of conditions, at least one condition of the set of conditions comprising a previous action, a previous position, a previous trajectory, or a previous acceleration of the object.

14. The one or more non-transitory computer-readable media of claim 11 ,

wherein:

the object is a first object,

the condition data represents a potential action of a second object in the environment, and

the latent variable data represents a potential interaction between the second object and the first object.

15. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the one or more processors to perform actions comprising:

receiving, by a diffusion model, map data representing an environment, wherein the diffusion model comprises one or more self-attention layers;

receiving, by the diffusion model, condition data representing a state or an action of an object in the environment;

determining, by the diffusion model and based at least in part on the map data, the condition data, and output data associated with the one or more self-attention layers, latent variable data associated with the object;

inputting the latent variable data into a machine learned model;

determining, by the machine learned model and based at least in part on the latent variable data, an object trajectory for the object to follow in the environment; and

causing control of a vehicle in the environment by executing one or more vehicle maneuvers based at least in part on the object trajectory to avoid potential interaction between the vehicle and the object.

16. The system of claim 15 , wherein:

the machine learned model comprises at least one of a Generative Adversarial Network (GAN), a Graph Neural Network (GNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a transformer model.

17. The system of claim 15 , wherein:

the diffusion model is trained based at least in part on a set of conditions, at least one condition of the set of conditions comprising a previous action, a previous position, a previous trajectory, or a previous acceleration of the object.

18. The system of claim 15 , wherein:

the condition data comprises a token from a codebook associated with a transformer model, and

the token representing a behavior of the object.

19. The system of claim 15 , wherein:

the condition data comprises a node from a Graph Neural Network, and

the node representing a behavior of the object.

20. The system of claim 15 , wherein:

the condition data represents one of: a yield action, a drive straight action, a left turn action, a right turn action, a brake action, an acceleration action, a steering action, a lane change action, a position, a heading, or an acceleration of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2022
From: PRONOVOST, ETHAN MILLER
To: ZOOX, INC.
Reel/Frame 062190/0119 →
Continuity (3)
Continuation In Part 17885671 · Jun 30, 2022
Continuation In Part 17855696 · Jun 30, 2022
Related Publication 20240101157A1 · Mar 28, 2024
References Cited (98)
US 10019011B1 · Green et al. · 2018 [cited by applicant]
US 10086782B1 · Konrardy · 2018 [cited by examiner]
US 10421453B1 · Ferguson et al. · 2019 [cited by applicant]
US 10459444B1 · Kentley-Klay · 2019 [cited by applicant]
US 10671076B1 · Kobilarov et al. · 2020 [cited by applicant]
US 10678244B2 · Iandola · 2020 [cited by examiner]
US 10717004B2 · Buttner · 2020 [cited by examiner]
US 11200679B1 · Li et al. · 2021 [cited by applicant]
US 11370424B1 · Cohen et al. · 2022 [cited by applicant]
US 11667301B2 · Misra · 2023 [cited by examiner]
US 11731652B2 · Dolben · 2023 [cited by examiner]
US 11731663B2 · Zeng · 2023 [cited by examiner]
US 11772663B2 · Anthony · 2023 [cited by examiner]
US 11774978B2 · Song · 2023 [cited by examiner]
US 11960292B2 · Nayhouse et al. · 2024 [cited by applicant]
US 11975726B1 · Gu · 2024 [cited by examiner]
US 12361689B1 · Yazdani · 2025 [cited by examiner]
US 20170132334A1 · Levinson et al. · 2017 [cited by applicant]
US 20190072966A1 · Zhang et al. · 2019 [cited by applicant]
US 20190129436A1 · Sun · 2019 [cited by examiner]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190152490A1 · Lan et al. · 2019 [cited by applicant]
US 20190164007A1 · Liu et al. · 2019 [cited by applicant]
US 20190303759A1 · Farabet et al. · 2019 [cited by applicant]
US 20190332875A1 · Vallespi-Gonzalez et al. · 2019 [cited by applicant]
US 20200148201A1 · King et al. · 2020 [cited by applicant]
US 20200174481A1 · Van Heukelom et al. · 2020 [cited by applicant]
US 20200180647A1 · Anthony · 2020 [cited by applicant]
US 20200216085A1 · Bobier-Tiu et al. · 2020 [cited by applicant]
US 20200225669A1 · Silva et al. · 2020 [cited by applicant]
US 20200283016A1 · Blaiotta · 2020 [cited by applicant]
US 20200324795A1 · Bojarski et al. · 2020 [cited by applicant]
US 20200380085A1 · Behrendt · 2020 [cited by applicant]
US 20200409368A1 · Caldwell et al. · 2020 [cited by applicant]
US 20200409378A1 · Benisch et al. · 2020 [cited by applicant]
US 20210114617A1 · Phillips et al. · 2021 [cited by applicant]
US 20210220739A1 · Zinno · 2021 [cited by examiner]
US 20210286923A1 · Kristensen · 2021 [cited by examiner]
US 20210286924A1 · Wyrwas et al. · 2021 [cited by applicant]
US 20210294944A1 · Nassar · 2021 [cited by examiner]
US 20210341927A1 · Refaat et al. · 2021 [cited by applicant]
US 20210347382A1 · Huang et al. · 2021 [cited by applicant]
US 20220153309A1 · Cui · 2022 [cited by examiner]
US 20220153314A1 · Suo · 2022 [cited by examiner]
US 20220161811A1 · Lu et al. · 2022 [cited by applicant]
US 20220315049A1 · Stenson · 2022 [cited by examiner]
US 20230121388A1 · Taslim et al. · 2023 [cited by applicant]
US 20230150529A1 · Stenson · 2023 [cited by examiner]
US 20230177819A1 · Iandola · 2023 [cited by examiner]
US 20230202511A1 · Atsmon · 2023 [cited by examiner]
US 20230213945A1 · Sajjan · 2023 [cited by examiner]
US 20230286539A1 · Malloch · 2023 [cited by examiner]
US 20240101150A1 · Pronovost · 2024 [cited by applicant]
US 20240101157A1 · Pronovost · 2024 [cited by applicant]
US 20240104934A1 · Pronovost · 2024 [cited by applicant]
US 20240160888A1 · Rempe · 2024 [cited by examiner]
US 20240199071A1 · Atsmon · 2024 [cited by examiner]
US 20240210942A1 · Pronovost · 2024 [cited by applicant]
US 20240211731A1 · Pronovost · 2024 [cited by applicant]
US 20240211797A1 · Pronovost · 2024 [cited by applicant]
US 20240212360A1 · Pronovost · 2024 [cited by applicant]
US 20240217530A1 · Martin Bragado · 2024 [cited by examiner]
US 20240221178A1 · Hughes · 2024 [cited by examiner]
US 20240273261A1 · Brehmer · 2024 [cited by examiner]
US 20240296919A1 · Alesiani · 2024 [cited by examiner]
US 20240394944A1 · Liu · 2024 [cited by examiner]
US 20250021761A1 · Santhanam · 2025 [cited by examiner]
US 20250037298A1 · Liang · 2025 [cited by examiner]
US 20250058802A1 · Chen · 2025 [cited by examiner]
US 20250103779A1 · Zhang · 2025 [cited by examiner]
US 20250111552A1 · Yu · 2025 [cited by examiner]
US 20250162150A1 · Chen · 2025 [cited by examiner]
US 20250171017A1 · Chen · 2025 [cited by examiner]
US 20250218139A1 · Kharbanda · 2025 [cited by examiner]
US 20250225659A1 · Choi · 2025 [cited by examiner]
US 20250232471A1 · Zhou · 2025 [cited by examiner]
CN 113936243 · 2022 [cited by applicant]
CN 113936243A · 2022 [cited by applicant]
Ethan Pronovost et al., Scenario Diffusion: Controllable Driving Scenario Generation With Diffusion, Nov. 16, 2023, arXiv, pp. 1-22 (pdf). [cited by examiner]
Office Action for U.S. Appl. No. 18/087,609, mailed on Oct. 30, 2024, Pronovost, “Generating a Scenario Using a Variable Autoencoder Conditioned With a Diffusion Model”, 22 Pages. [cited by applicant]
ICLR 2021 Conference Paper 2345 Authors, Official Comment on Latent Optimization Variational Autoencoder for Conditional Molecular Generation [online], Nov. 13, 2020 (last modification date)Y [retrieved on Mar. 11, 2024… [cited by applicant]
Office Action for U.S. Appl. No. 17/855,671, Dated Jun. 21, 2024, 13 pages. [cited by applicant]
The PCT Search Report and Written Opinion mailed Apr. 29, 2024 for PCT Application No. PCT/US2023/084627 from PCT Summary, 11 pages. [cited by applicant]
The PCT Search Report and Written Opinion mailed Apr. 18, 2024 for PCT Application No. PCT/US2023/084618 from PCT Summary, 13 pages. [cited by applicant]
Bruno Sauvalle et al., Autoencoder-based background reconstruction and foreground segmentation with background noise estimation [online], Jun. 27, 2022 \Y [retrieved on Mar. 11, 2024]. Retrievedfrom<URL:https://www.rese… [cited by applicant]
Arsal Syed, Forecasting Pedestrian Trajectory Using Deep Learning, In: Unlv Theses, Dissertations, Professional Papers, and Capstones [online], 2021Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://digitalsch… [cited by applicant]
Hao Xue, Deep Learning Based Pedestrian Trajectory Prediction, In: Thesis—Doctor of Philosophy (research output) [online], 2020Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://research-repository.uwa.edu.au/… [cited by applicant]
Arsal Syed, Forecasting Pedestrian Trajectory Using Deep Learning, In: Unlv Theses, Dissertations, Professional Papers, and Capstones [online], 2021Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://digitalsch… [cited by applicant]
Balakrishnan, et al. “MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction” Submitted to Cornell University on Nov. 29, 2021; 22 pages. [cited by applicant]
Esser, “Taming Transformers for High-Resolution Image Synthesis” Submitted to Cornell University on Dec. 17, 2020; 52 pages. [cited by applicant]
Gilled, et al. “GOHOME: Graph-Oriented Heatmap Output for future Motion Estimation” submitted to Corness University on Sep. 4, 2021; 8 pages. [cited by applicant]
Giris, et al. “Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction” ICLR 2022 Spotlight; Sep. 29, 2021; 26 pages. [cited by applicant]
Janjos, et al; “StarNet: Joint Action-Space Prediction with Star Graphs and Implicit Global Frame Self-Attention” Submitted to Corness University on Nov. 26, 2021; 7 pages. [cited by applicant]
Nglam, et al. “Scene Transformer: A unified architecture for predicting multiple agent trajectories” Published as a conference paper at ICLR 2022; Mar. 4, 2022, 25 pages. [cited by applicant]
Rhinehart, et al. “PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings” Submitted to Cornell University on May 3, 2019; 24 pages. [cited by applicant]
Salzmann, et al. “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data” Submitted to Cornell University on Jan. 9, 2020; 23 pages. [cited by applicant]
Tang, et al. “Multiple Futures Prediction” Submitted to Cornell University on Nov. 4, 2019; 17 pages. [cited by applicant]
Yang, et al; “TPPO A Novel Trajectory Predictor with Pseudo Oracle” IEEE Dec. 29, 2021; 14 pages. [cited by applicant]