IP Library Granted Patent US 12,555,043
Granted Patent B2
US 12,555,043 · App. 18/087,598 · Granted Feb 17, 2026

Training a variable autoencoder using a diffusion model

Inventor: Ethan Miller Pronovost (Redwood City, CA)
Assignee: Zoox, Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,043
App. No.
18/087,598
Filed
Dec 22, 2022
Granted
Feb 17, 2026
Kind
B2
Art Unit
2682
USPC
382/104
Abstract

Techniques for training a variable autoencoder to output data associated with one or more objects in an environment are described herein. For example, the techniques can include training an encoder and a decoder of the variable autoencoder to improve predictions over time. The variable autoencoder can be trained to output object occupancy information (e.g., a bounding box, heatmap, feature vector) and/or object attribute information (e.g., an object state, an object type, etc.). A vehicle computing device can use an output from a trained autoencoder during vehicle planning, which may include simulation.

Claims (64)

1 . A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations comprising:

inputting, as first input data into an encoder of a variable autoencoder, map data representing an environment and occupancy data associated with an object in the environment;

receiving, from the encoder, first output data representing a compressed representation of the first input data;

receiving, from a diffusion model, discrete latent variable data associated with the object, the diffusion model configured to implement a diffusion process to add or remove noise from condition data received as input;

inputting, as second input data into a decoder of the variable autoencoder, the first output data from the encoder and the discrete latent variable data from the diffusion model;

receiving, from the decoder, second output data representing an occupancy representation for the object in the environment and object state data associated with the object; and

training the encoder or the decoder based at least in part on the second output data.

2 . The system of claim 1 , wherein the discrete latent variable data associated with the object indicates an acceleration action, a braking action, or a steering action of the object.

3 . The system of claim 1 , wherein:

the map data is associated with a simulated environment,

the compressed representation of the first input data represents a latent embedding of the first input data, and

the object state data indicates an acceleration, a velocity, an orientation, or a position of the object.

4 . The system of claim 1 , the operations further comprising:

performing a simulation to verify a response by a vehicle relative to the object; and

training, based at least in part on the response, the decoder to output bounding box data for the object.

5 . The system of claim 1 , the operations further comprising:

transmitting the second output data to a vehicle computing device to control a vehicle in the environment.

6 . One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform operations comprising:

inputting, into an encoder, map data representing an environment and occupancy data associated with an object in the environment;

receiving, from the encoder, first output data representing a compressed representation of the map data and the occupancy data;

receiving, from a diffusion model, discrete latent variable data associated with the object, the diffusion model configured to implement a diffusion process to add or remove noise from condition data received as input;

inputting, into a decoder, the first output data from the encoder and the discrete latent variable data from the diffusion model;

receiving, from the decoder, second output data representing an occupancy representation for the object in the environment and object state data associated with the object; and

training the encoder or the decoder based at least in part on the second output data.

7 . The one or more non-transitory computer-readable media of claim 6 , wherein the second output data comprises a trajectory, a velocity, an acceleration, or an orientation associated with the object.

8 . The one or more non-transitory computer-readable media of claim 6 , wherein:

the map data is associated with a simulated environment,

the compressed representation represents a latent embedding of data input to the encoder, and

the object state data indicates at least one of: a trajectory, an acceleration, a velocity, an orientation, width or length of the object, or a position of the object.

9 . The one or more non-transitory computer-readable media of claim 6 , the operations further comprising:

performing a simulation to verify a response by a vehicle relative to the object; and

training, based at least in part on the response, the decoder to output bounding box data for the object.

10 . The one or more non-transitory computer-readable media of claim 6 , the operations further comprising:

transmitting the second output data to a vehicle computing device to control a vehicle in the environment.

11 . The one or more non-transitory computer-readable media of claim 6 , wherein the encoder and the decoder are components of a variable autoencoder.

12 . The one or more non-transitory computer-readable media of claim 6 , wherein training the encoder or the decoder comprises:

comparing, as a comparison, the first output data or the second output data to ground truth; and

training the encoder or the decoder based at least in part on the comparison.

13 . The one or more non-transitory computer-readable media of claim 6 , wherein the occupancy representation comprises a bounding box to represent the object in the environment.

14 . The one or more non-transitory computer-readable media of claim 6 , wherein the occupancy representation comprises a feature vector to represent the object in the environment.

15 . The one or more non-transitory computer-readable media of claim 6 , wherein:

the object is a first object,

the second output data comprises:

a first bounding box associated with the first object,

a second bounding box associated with a second object, and

an orientation or a position of the first object.

16 . The one or more non-transitory computer-readable media of claim 6 , wherein the occupancy representation comprises a heatmap to represent the object in the environment.

17 . A method comprising:

inputting, into an encoder, map data representing an environment and occupancy data associated with an object in the environment;

receiving, from the encoder, first output data representing a compressed representation of the map data and the occupancy data;

receiving, from a diffusion model, discrete latent variable data associated with the object, the diffusion model configured to implement a diffusion process to add or remove noise from condition data received as input;

inputting, into a decoder, the first output data from the encoder and the discrete latent variable data from the diffusion model;

receiving, from the decoder, second output data representing an occupancy representation for the object in the environment and object state data associated with the object; and

training the encoder or the decoder based at least in part on the second output data.

18 . The method of claim 17 , wherein the second output data comprises a trajectory, a velocity, or an acceleration associated with the object.

19 . The method of claim 17 , wherein:

the map data is associated with a simulated environment,

the compressed representation represents a latent embedding of data input to the encoder, and

the object state data indicates an acceleration, a velocity, an orientation, or a position of the object.

20 . The method of claim 17 , further comprising:

performing a simulation to verify a response by a vehicle relative to the object; and

training, based at least in part on the response, the decoder to output bounding box data for the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2022
From: PRONOVOST, ETHAN MILLER
To: ZOOX, INC.
Reel/Frame 062190/0987 →
Continuity (1)
Related Publication 20240211797A1 · Jun 27, 2024
References Cited (131)
US 10019011B1 · Green et al. · 2018 [cited by applicant]
US 10086782B1 · Konrardy et al. · 2018 [cited by applicant]
US 10421453B1 · Ferguson et al. · 2019 [cited by applicant]
US 10459444B1 · Kentley-Klay · 2019 [cited by applicant]
US 10671076B1 · Kobilarov et al. · 2020 [cited by applicant]
US 10678244B2 · Forrest et al. · 2020 [cited by applicant]
US 10717004B2 · Buttner · 2020 [cited by applicant]
US 11186276B2 · Tao et al. · 2021 [cited by applicant]
US 11195418B1 · Hong et al. · 2021 [cited by applicant]
US 11200679B1 · Li et al. · 2021 [cited by applicant]
US 11370424B1 · Cohen et al. · 2022 [cited by applicant]
US 11599972B1 · Xu et al. · 2023 [cited by applicant]
US 11667301B2 · Misra et al. · 2023 [cited by applicant]
US 11731652B2 · Dolben et al. · 2023 [cited by applicant]
US 11772663B2 · Anthony · 2023 [cited by applicant]
US 11774978B2 · Song et al. · 2023 [cited by applicant]
US 11912301B1 · Hendy · 2024 [cited by examiner]
US 11960292B2 · Nayhouse et al. · 2024 [cited by applicant]
US 11975726B1 · Gu et al. · 2024 [cited by applicant]
US 12198225B2 · Bradley et al. · 2025 [cited by applicant]
US 12322068B1 · Kim et al. · 2025 [cited by applicant]
US 12361689B1 · Yazdani et al. · 2025 [cited by applicant]
US 12394311B2 · Vozar et al. · 2025 [cited by applicant]
US 20170132334A1 · Levinson et al. · 2017 [cited by applicant]
US 20180253093A1 · Augugliaro et al. · 2018 [cited by applicant]
US 20190072966A1 · Zhang et al. · 2019 [cited by applicant]
US 20190129436A1 · Sun et al. · 2019 [cited by applicant]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190152490A1 · Lan et al. · 2019 [cited by applicant]
US 20190164007A1 · Liu et al. · 2019 [cited by applicant]
US 20190303759A1 · Farabet et al. · 2019 [cited by applicant]
US 20190332875A1 · Vallespi-Gonzalez et al. · 2019 [cited by applicant]
US 20200026287A1 · Jiang et al. · 2020 [cited by applicant]
US 20200148201A1 · King et al. · 2020 [cited by applicant]
US 20200160532A1 · Urtasun · 2020 [cited by examiner]
US 20200174481A1 · Van Heukelom et al. · 2020 [cited by applicant]
US 20200180647A1 · Anthony · 2020 [cited by applicant]
US 20200216085A1 · Bobier-Tiu et al. · 2020 [cited by applicant]
US 20200225669A1 · Silva et al. · 2020 [cited by applicant]
US 20200283016A1 · Blaiotta · 2020 [cited by applicant]
US 20200324795A1 · Bojarski et al. · 2020 [cited by applicant]
US 20200380085A1 · Behrendt · 2020 [cited by applicant]
US 20200409368A1 · Caldwell et al. · 2020 [cited by applicant]
US 20200409378A1 · Benisch et al. · 2020 [cited by applicant]
US 20210026355A1 · Chen et al. · 2021 [cited by applicant]
US 20210114617A1 · Phillips et al. · 2021 [cited by applicant]
US 20210220739A1 · Zinno et al. · 2021 [cited by applicant]
US 20210286923A1 · Kristensen et al. · 2021 [cited by applicant]
US 20210286924A1 · Wyrwas et al. · 2021 [cited by applicant]
US 20210294944A1 · Nassar et al. · 2021 [cited by applicant]
US 20210341927A1 · Refaat et al. · 2021 [cited by applicant]
US 20210347382A1 · Huang et al. · 2021 [cited by applicant]
US 20220032970A1 · Vadivelu · 2022 [cited by examiner]
US 20220153309A1 · Cui et al. · 2022 [cited by applicant]
US 20220153314A1 · Suo et al. · 2022 [cited by applicant]
US 20220161811A1 · Lu et al. · 2022 [cited by applicant]
US 20220291690A1 · Goyal et al. · 2022 [cited by applicant]
US 20220315049A1 · Stenson et al. · 2022 [cited by applicant]
US 20230048926A1 · Kurbiel et al. · 2023 [cited by applicant]
US 20230109379A1 · Kreis et al. · 2023 [cited by applicant]
US 20230121388A1 · Taslim et al. · 2023 [cited by applicant]
US 20230150529A1 · Stenson et al. · 2023 [cited by applicant]
US 20230150550A1 · Shi et al. · 2023 [cited by applicant]
US 20230177819A1 · Forrest et al. · 2023 [cited by applicant]
US 20230202511A1 · Atsmon et al. · 2023 [cited by applicant]
US 20230213945A1 · Sajjan et al. · 2023 [cited by applicant]
US 20230237335A1 · Hallac · 2023 [cited by applicant]
US 20230267315A1 · Kingma · 2023 [cited by examiner]
US 20230286539A1 · Malloch et al. · 2023 [cited by applicant]
US 20230368337A1 · Karras et al. · 2023 [cited by applicant]
US 20230377099A1 · Kreis et al. · 2023 [cited by applicant]
US 20230377226A1 · Saharia et al. · 2023 [cited by applicant]
US 20230377584A1 · Pascual et al. · 2023 [cited by applicant]
US 20230394823A1 · Weng · 2023 [cited by examiner]
US 20240046422A1 · Song · 2024 [cited by applicant]
US 20240101150A1 · Pronovost · 2024 [cited by applicant]
US 20240101157A1 · Pronovost · 2024 [cited by applicant]
US 20240104698A1 · Nie et al. · 2024 [cited by applicant]
US 20240104934A1 · Pronovost · 2024 [cited by applicant]
US 20240119261A1 · Strudel et al. · 2024 [cited by applicant]
US 20240134640A1 · Han et al. · 2024 [cited by applicant]
US 20240160888A1 · Rempe et al. · 2024 [cited by applicant]
US 20240169500A1 · Zheng et al. · 2024 [cited by applicant]
US 20240184812A1 · Mcdaniel et al. · 2024 [cited by applicant]
US 20240199071A1 · Atsmon · 2024 [cited by examiner]
US 20240202577A1 · Bagschik et al. · 2024 [cited by applicant]
US 20240210942A1 · Pronovost · 2024 [cited by applicant]
US 20240211731A1 · Pronovost · 2024 [cited by applicant]
US 20240212360A1 · Pronovost · 2024 [cited by applicant]
US 20240217530A1 · Martin Bragado · 2024 [cited by applicant]
US 20240221178A1 · Hughes et al. · 2024 [cited by applicant]
US 20240273261A1 · Brehmer et al. · 2024 [cited by applicant]
US 20240296919A1 · Alesiani et al. · 2024 [cited by applicant]
US 20240394944A1 · Liu et al. · 2024 [cited by applicant]
US 20250021761A1 · Santhanam et al. · 2025 [cited by applicant]
US 20250037298A1 · Liang et al. · 2025 [cited by applicant]
US 20250058802A1 · Chen et al. · 2025 [cited by applicant]
US 20250103779A1 · Zhang et al. · 2025 [cited by applicant]
US 20250111552A1 · Yu et al. · 2025 [cited by applicant]
US 20250162150A1 · Chen et al. · 2025 [cited by applicant]
US 20250171017A1 · Chen et al. · 2025 [cited by applicant]
US 20250218139A1 · Kharbanda et al. · 2025 [cited by applicant]
US 20250225659A1 · Choi et al. · 2025 [cited by applicant]
US 20250232471A1 · Zhou et al. · 2025 [cited by applicant]
CN 113936243 · 2022 [cited by applicant]
Office Action for U.S. Appl. No. 18/087,609, mailed on Oct. 30, 2024, Pronovost, “Generating a Scenario Using a Variable Autoencoder Conditioned With a Diffusion Model”, 22 Pages. [cited by applicant]
Office Action for U.S. Appl. No. 18/087,540, dated Dec. 4, 2024, 20 pages. [cited by applicant]
Office Action for U.S. Appl. No. 18/087,540, Dated May 6, 2025, Pronovost, “Latent Variable Determination By a Diffusion Model .” 19 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/855,671, mailed on Dec. 4, 2024, Pronovost, “Conditional Trajectory Determination By a Machine Learned Model”, 17 pages. [cited by applicant]
Balakrishnan, et al. “MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction” Submitted to Cornell University on Nov. 29, 2021; 22 pages. [cited by applicant]
Esser, “Taming Transformers for High-Resolution Image Synthesis” Submitted to Cornell University on Dec. 17, 2020; 52 pages. [cited by applicant]
Gilled, et al. “GOHOME: Graph-Oriented Heatmap Output for future Motion Estimation” submitted to Corness University on Sep. 4, 2021; 8 pages. [cited by applicant]
Giris, et al. “Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction” ICLR 2022 Spotlight; Sep. 29, 2021; 26 pages. [cited by applicant]
Janjos, et al; “StarNet: Joint Action-Space Prediction with Star Graphs and Implicit Global Frame Self-Attention” Submitted to Corness University on Nov. 26, 2021; 7 pages. [cited by applicant]
Nglam, et al. “Scene Transformer: A unified architecture for predicting multiple agent trajectories” Published as a conference paper at ICLR 2022; Mar. 4, 2022, 25 pages. [cited by applicant]
Rhinehart, et al. “PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent Settings” Submitted to Cornell University on May 3, 2019; 24 pages. [cited by applicant]
Salzmann, et al. “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data” Submitted to Cornell University on Jan. 9, 2020; 23 pages. [cited by applicant]
Tang, et al. “Multiple Futures Prediction” Submitted to Cornell University on Nov. 4, 2019; 17 pages. [cited by applicant]
Yang, et al; “TPPO A Novel Trajectory Predictor with Pseudo Oracle” IEEE Dec. 29, 2021; 14 pages. [cited by applicant]
Office Action for U.S. Appl. No. 18/087,609, Dated Mar. 5, 2025, Pronovost, Generating a Scenario Using a Variable Autoencoder Conditioned With a Diffusion Model , 17 pages. [cited by applicant]
ICLR 2021 Conference Paper 2345 Authors, Official Comment on Latent Optimization Variational Autoencoder for Conditional Molecular Generation [online], Nov. 13, 2020 (last modification date)Y [retrieved on Mar. 11, 2024… [cited by applicant]
Office Action for U.S. Appl. No. 17/855,671, Dated Jun. 21, 2024, 13 pages. [cited by applicant]
PCT Search Report and Written Opinion mailed Apr. 29, 2024 for PCT Application No. PCT/US2023/084627 from PCT Summary, 11 pages. [cited by applicant]
PCT Search Report and Written Opinion mailed Apr. 18, 2024 for PCT Application No. PCT/US2023/084618 from PCT Summary, 13 pages. [cited by applicant]
Bruno Sauvalle et al., Autoencoder-based background reconstruction and foreground segmentation with background noise estimation [online], Jun. 27, 2022\Y [retrieved on Mar. 11, 2024]. Retrievedfrom<URL:https://www.resea… [cited by applicant]
Arsal Syed, Forecasting Pedestrian Trajectory Using Deep Learning, In: Unlv Theses, Dissertations, Professional Papers, and Capstones [online], 2021Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://digitalsch… [cited by applicant]
Hao Xue, Deep Learning Based Pedestrian Trajectory Prediction, In: Thesis—Doctor of Philosophy (research output) [online], 2020Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://research-repository.uwa.edu.au/… [cited by applicant]
Office Action for U.S. Appl. No. 18/087,609, mailed Oct. 30, 2024, Pronovost, “Generating a Scenario Using a Variable Autoencoder Conditioned With a Difusion Model”, 22 pages. [cited by applicant]
Office Action for U.S. Appl. No. 18/087,586, Dated Sep. 4, 2025, Pronovost,“Generating Object Data Using a Diffusion Model ,” 21 pages. [cited by applicant]
Bruno Sauvalle et al., Autoencoder-based background reconstruction and foreground segmentation with background noise estimation [online], Jun. 27, 2022 \ [retrieved on Mar. 11, 2024]. Retrieved from<URL:https://www.rese… [cited by applicant]
ICLR 2021 Conference Paper 2345 Authors, Official Comment on Latent Optimization Variational Autoencoder for Conditional Molecular Generation [online], Nov. 13, 2020 (last modification date)Y [retrieved on Mar. 11, 2024… [cited by applicant]