IP Library Granted Patent US 12,353,979
Granted Patent B2
US 12,353,979 · App. 18/087,570 · Granted Jul 8, 2025

Generating object representations using a variable autoencoder

Inventor: Ethan Miller Pronovost (Redwood City, CA)
Assignee: Zoox, Inc.
G06N3/0455G05B19/4069
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,979
App. No.
18/087,570
Filed
Dec 22, 2022
Granted
Jul 8, 2025
Kind
B2
Art Unit
2115
USPC
700/28
Abstract

Techniques for generating a representation for an object in an environment of an autonomous vehicle are described herein. For example, the techniques may include a decoder of a variable autoencoder receiving latent variable data from a diffusion model and determining an object representation such as a bounding box or a heatmap for one or more objects proximate the autonomous vehicle. The bounding box can include orientation data indicating an orientation for each of the one or more bounding boxes. The object representation(s) can be sent to a vehicle computing device for consideration during vehicle planning, which may include simulation.

Claims (50)

1. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations comprising:

inputting, as a first input into a decoder of a variable autoencoder, map data representing an environment;

inputting, as a second input into the decoder, discrete latent variable data associated with a first object and a second object in the environment, the discrete latent variable data representing a first action of the first object and a second action of the second object, the second action different than the first action;

receiving, from the decoder and based at least in part on the first input and the second input, output data representing a first bounding box for the first object and a second bounding box for the second object, the first bounding box including a first orientation and the second bounding box including a second orientation; and

at least one of:

performing, based at least in part on the output data, a simulation between a vehicle, the first object, and the second object; or

controlling, based at least in part on the output data, the vehicle in the environment relative to the first object and the second object.

2. The system of claim 1 , wherein a first number of channels associated with the output data from the decoder is greater than a second number of channels associated with the first input.

3. The system of claim 1 , the operations further comprising:

determining an object type associated with the first object or the second object; and

determining the output data based at least in part on the object type.

4. The system of claim 1 , wherein the discrete latent variable data is received from a diffusion model configured to receive input data, determine cross attention data between the first object and the second object, and output the discrete latent variable data based at least in part on the cross attention data.

5. The system of claim 4 , wherein the diffusion model determines a number of objects to include in the environment based at least in part on condition data.

6. One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform operations comprising:

inputting, into a decoder of a variable autoencoder, latent variable data associated with an object in an environment, the latent variable data representing a behavior of the object;

receiving, from the decoder and based at least in part on the latent variable data, output data representing a discrete occupancy representation for the object; and

at least one of:

performing, based at least in part on the output data, a simulation between a robotic device and the object: or

controlling, based at least in part on the output data, a robotic device in the environment.

7. The one or more non-transitory computer-readable media of claim 6 , wherein the latent variable data comprises discrete latent variable data representing at least one of an action, an intent, or an attribute of the object for use during the simulation, and the operations further comprising:

inputting map data representing the environment into the decoder; and

determining, by the decoder, the output data based at least in part on the map data.

8. The one or more non-transitory computer-readable media of claim 6 , wherein:

the object is a first object,

the behavior is a first behavior,

the latent variable data further represents a second behavior of a second object, and

the output data comprises a second occupancy representation.

9. The one or more non-transitory computer-readable media of claim 6 , wherein a first number of channels associated with the output data from the decoder is greater than a second number of channels associated with an input to an encoder of the variable autoencoder.

10. The one or more non-transitory computer-readable media of claim 6 , wherein the output data further represents an orientation of the object.

11. The one or more non-transitory computer-readable media of claim 6 , wherein the latent variable data is received from a diffusion model configured to receive input data and apply an algorithm to denoise the input data.

12. The one or more non-transitory computer-readable media of claim 11 , wherein the diffusion model determines a number of objects to include in the environment.

13. The one or more non-transitory computer-readable media of claim 6 , the operations further comprising:

determining, based at least in part on the output data, one or more of: a position, a size, an acceleration, or a velocity of the object at a future time.

14. The one or more non-transitory computer-readable media of claim 6 , wherein the occupancy representation for the object comprises a bounding box representing a two-dimensional shape or a three-dimensional shape of the object for a current time.

15. The one or more non-transitory computer-readable media of claim 6 , wherein the output data comprises a vector representation of the object.

16. The one or more non-transitory computer-readable media of claim 6 , wherein:

the decoder is trained based at least in part on an output from an encoder that is configured to receive map data and occupancy data associated with the object as input.

17. A method comprising:

inputting, into a decoder of a variable autoencoder, latent variable data associated with an object in an environment, the latent variable data representing a behavior of the object;

receiving, from the decoder and based at least in part on the latent variable data, output data representing a discrete occupancy representation for the object; and

at least one of:

performing, based at least in part on the output data, a simulation between a robotic device and the object; or

controlling, based at least in part on the output data, a robotic device in the environment.

18. The method of claim 17 , wherein the latent variable data comprises discrete latent variable data representing at least one of an action, an intent, or an attribute of the object for use during the simulation, and further comprising:

inputting map data representing the environment into the decoder; and

determining, by the decoder, the output data based at least in part on the map data.

19. The method of claim 17 , wherein a first number of channels associated with the output data from the decoder is greater than a second number of channels associated with an input to an encoder of the variable autoencoder.

20. The method of claim 17 , wherein the latent variable data is received from a diffusion model configured to receive input data and apply an algorithm to denoise the input data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2022
From: PRONOVOST, ETHAN MILLER
To: ZOOX, INC.
Reel/Frame 062190/0890 →
Continuity (1)
Related Publication 20240211731A1 · Jun 27, 2024
References Cited (62)
US 10019011B1 · Green et al. · 2018 [cited by applicant]
US 10086782B1 · Konrardy et al. · 2018 [cited by applicant]
US 10421453B1 · Ferguson et al. · 2019 [cited by applicant]
US 10459444B1 · Kentley-Klay · 2019 [cited by applicant]
US 10671076B1 · Kobilarov et al. · 2020 [cited by applicant]
US 11200679B1 · Li et al. · 2021 [cited by applicant]
US 11370424B1 · Cohen et al. · 2022 [cited by applicant]
US 11960292B2 · Nayhouse et al. · 2024 [cited by applicant]
US 20170132334A1 · Levinson et al. · 2017 [cited by applicant]
US 20190072966A1 · Zhang et al. · 2019 [cited by applicant]
US 20190129436A1 · Sun et al. · 2019 [cited by applicant]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190152490A1 · Lan et al. · 2019 [cited by applicant]
US 20190164007A1 · Liu et al. · 2019 [cited by applicant]
US 20190303759A1 · Farabet et al. · 2019 [cited by applicant]
US 20190332875A1 · Vallespi-Gonzalez et al. · 2019 [cited by applicant]
US 20200026287A1 · Jiang · 2020 [cited by examiner]
US 20200148201A1 · King et al. · 2020 [cited by applicant]
US 20200174481A1 · Van Heukelom et al. · 2020 [cited by applicant]
US 20200180647A1 · Anthony · 2020 [cited by applicant]
US 20200216085A1 · Bobier-Tiu et al. · 2020 [cited by applicant]
US 20200225669A1 · Silva et al. · 2020 [cited by applicant]
US 20200283016A1 · Blaiotta · 2020 [cited by applicant]
US 20200324795A1 · Bojarski et al. · 2020 [cited by applicant]
US 20200380085A1 · Behrendt · 2020 [cited by applicant]
US 20200409368A1 · Caldwell et al. · 2020 [cited by applicant]
US 20200409378A1 · Benisch et al. · 2020 [cited by applicant]
US 20210114617A1 · Phillips et al. · 2021 [cited by applicant]
US 20210286924A1 · Wyrwas et al. · 2021 [cited by applicant]
US 20210341927A1 · Refaat et al. · 2021 [cited by applicant]
US 20210347382A1 · Huang et al. · 2021 [cited by applicant]
US 20220153309A1 · Cui et al. · 2022 [cited by applicant]
US 20220153314A1 · Suo et al. · 2022 [cited by applicant]
US 20220161811A1 · Lu et al. · 2022 [cited by applicant]
US 20230121388A1 · Taslim et al. · 2023 [cited by applicant]
US 20240101150A1 · Pronovost · 2024 [cited by applicant]
US 20240101157A1 · Pronovost · 2024 [cited by applicant]
US 20240104934A1 · Pronovost · 2024 [cited by applicant]
US 20240210942A1 · Pronovost · 2024 [cited by applicant]
US 20240211797A1 · Pronovost · 2024 [cited by applicant]
US 20240212360A1 · Pronovost · 2024 [cited by applicant]
CN 113936243 · 2022 [cited by applicant]
Office Action for U.S. Appl. No. 18/087,540, dated Dec. 4, 2024, 20 pages. [cited by applicant]
ICLR 2021 Conference Paper 2345 Authors, Official Comment on Latent Optimization Variational Autoencoder for Conditional Molecular Generation [online], Nov. 13, 2020 (last modification date)Y [retrieved on Mar. 11, 2024… [cited by applicant]
Office Action for U.S. Appl. No. 17/855,671, Dated Jun. 21, 2024, 13 pages. [cited by applicant]
The PCT Search Report and Written Opinion mailed Apr. 29, 2024 for PCT Application No. PCT/US2023/084627 from PCT Summary, 11 pages. [cited by applicant]
The PCT Search Report and Written Opinion mailed Apr. 18, 2024 for PCT Application No. PCT/US2023/084618 from PCT Summary, 13 pages. [cited by applicant]
Bruno Sauvalle et al., Autoencoder-based background reconstruction and foreground segmentation with background noise estimation [online], Jun. 27, 2022\Y [retrieved on Mar. 11, 2024]. Retrievedfrom<URL:https://www.resea… [cited by applicant]
Arsal Syed, Forecasting Pedestrian Trajectory Using Deep Learning, In: Unlv Theses, Dissertations, Professional Papers, and Capstones [online], 2021Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://digitalsch… [cited by applicant]
Hao Xue, Deep Learning Based Pedestrian Trajectory Prediction, In: Thesis—Doctor of Philosophy (research output) [online], 2020Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://research-repository.uwa.edu.au/… [cited by applicant]
ICLR 2021 Conference Paper 2345 Authors, Official Comment on Latent Optimization Variational Autoencoder for Conditional Molecular Generation [online], Nov. 13, 2020 (last modification date)Y [retrieved on Mar. 11, 2024… [cited by applicant]
Balakrishnan, et al. “MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction” Submitted to Cornell University on Nov. 29, 2021; 22 pages. [cited by applicant]
Esser, “Taming Transformers for High-Resolution Image Synthesis” Submitted to Cornell University on Dec. 17, 2020; 52 pages. [cited by applicant]
Gilled, et al. “GOHOME: Graph-Oriented Heatmap Output for future Motion Estimation” submitted to Corness University on Sep. 4, 2021; 8 pages. [cited by applicant]
Giris, et al. “Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction” ICLR 2022 Spotlight; Sep. 29, 2021; 26 pages. [cited by applicant]
Janjos, et al; “StarNet: Joint Action-Space Prediction with Star Graphs and Implicit Global Frame Self-Attention” Submitted to Corness University on Nov. 26, 2021; 7 pages. [cited by applicant]
Nglam, et al. “Scene Transformer: A unified architecture for predicting multiple agent trajectories” Published as a conference paper at ICLR 2022; Mar. 4, 2022, 25 pages. [cited by applicant]
Rhinehart, et al. “PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings” Submitted to Cornell University on May 3, 2019; 24 pages. [cited by applicant]
Salzmann, et al. “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data” Submitted to Cornell University on Jan. 9, 2020; 23 pages. [cited by applicant]
Tang, et al. “Multiple Futures Prediction” Submitted to Cornell University on Nov. 4, 2019; 17 pages. [cited by applicant]
Yang, et al; “TPPO A Novel Trajectory Predictor with Pseudo Oracle” IEEE Dec. 29, 2021; 14 pages. [cited by applicant]
Office Action for U.S. Appl. No. 18/087,609, mailed on Oct. 30, 2024, Pronovost, “Generating a Scenario Using a Variable Autoencoder Conditioned With a Diffusion Model”, 22 Pages. [cited by applicant]