IP Library Granted Patent US 12,217,515
Granted Patent B2
US 12,217,515 · App. 17/855,696 · Granted Feb 4, 2025

Training a codebook for trajectory determination

Inventor: Ethan Miller Pronovost (Redwood City, CA)
Assignee: Zoox, Inc.
G06V20/58B60W40/10B60W2554/4049
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,515
App. No.
17/855,696
Granted
Feb 4, 2025
Kind
B2
Abstract

Techniques for training a codebook usable by a machine learned model to predict an object trajectory or scene data are described herein. For example, the techniques may include generating tokens representing discrete object behavior into a machine learned model that outputs a sequence of tokens that is usable by another machine learned model to generate the object trajectory (e.g., position data, velocity data, acceleration data, etc.) or the scene data associated with the environment. The object trajectory can be sent to a vehicle computing device for consideration during vehicle planning, which may include simulation.

Claims (75)

1. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the one or more processors to perform actions comprising:

receiving, by a training component, state data representing a previous state of an object in an environment;

receiving, by the training component, feature vectors representing a vehicle and the object in the environment;

training a codebook as a trained codebook, the training comprising:

assigning, based at least in part on the state data, a first token to represent the previous state of the object;

assigning a second token to represent a characteristic of the vehicle;

assigning a third token to represent a feature of the environment; and

mapping the feature vectors to respective tokens; and

outputting the trained codebook for use by a machine learned model configured to access tokens from the trained codebook and to arrange the first token, the second token, and the third token to represent a potential interaction between the vehicle and the object.

2. The system of claim 1 , wherein the machine learned model is a transformer model, and the actions further comprising:

causing the transformer model to arrange the tokens based at least in part on a number of tokens and an order of the tokens.

3. The system of claim 1 , wherein:

the feature vectors represent continuous variables,

the tokens represent discrete latent variables, and

mapping the feature vectors to respective tokens comprises mapping the continuous variables associated with the feature vectors to the discrete latent variables associated with the tokens.

4. The system of claim 1 , the actions further comprising:

receiving, by the training component, environmental data representing the feature of the environment,

wherein training the codebook further comprises assigning the second token to represent the feature of the environment, and

the codebook is usable by the machine learned model to cause generation of a scene with the feature in the environment.

5. The system of claim 1 , wherein the vehicle is an autonomous vehicle.

6. A method comprising:

receiving, by a training component, state data representing a previous state of an object or a vehicle in an environment;

receiving, by the training component, feature vectors representing the vehicle and the object in the environment;

training a codebook to output a trained codebook, the training comprising:

assigning, based at least in part on the state data, a first token to represent the previous state of the object;

assigning a second token to represent a characteristic of the vehicle;

assigning a third token to represent a feature of the environment; and

mapping the feature vectors to respective tokens; and

outputting the trained codebook for use by a machine learned model configured to access tokens from the trained codebook and to arrange the first token, the second token, and the third token to represent a potential interaction between the vehicle and the object.

7. The method of claim 6 , further comprising:

determining a number of tokens to include in the codebook; and

determining an order of the tokens.

8. The method of claim 6 , further comprising:

causing the machine learned model to arrange the tokens based at least in part on a number of tokens and an order of the tokens.

9. The method of claim 6 , wherein:

the feature vectors represent continuous variables,

the tokens represent discrete latent variables, and

mapping the feature vectors to respective tokens comprises mapping the continuous variables associated with the feature vectors to the discrete latent variables associated with the tokens.

10. The method of claim 6 , further comprising:

receiving, by the training component, environmental data representing the feature of the environment,

wherein training the codebook further comprises assigning the second token to represent the feature of the environment, and

the codebook is usable by the machine learned model to cause generation of a scene with the feature in the environment.

11. The method of claim 6 , wherein the state data representing the previous state of the object or the vehicle comprising one or more of: position data, orientation data, heading data, velocity data, speed data, acceleration data, yaw rate data, or turning rate data.

12. The method of claim 6 , further comprising:

inputting image data into an encoder; and

outputting, by the encoder, the feature vectors,

wherein receiving the feature vectors by the training component comprises receiving the feature vectors from the encoder.

13. The method of claim 6 , wherein:

the feature vectors represent the characteristic of the vehicle and a characteristic of the object in the environment.

14. The method of claim 6 , wherein mapping the feature vectors to respective tokens comprises:

identifying an action or a state of the vehicle associated with a first feature vector of the feature vectors;

comparing the characteristic of the vehicle associated with the first feature vector to a characteristic associated with one or more of the tokens.

15. The method of claim 6 , wherein the vehicle is an autonomous vehicle.

16. One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform actions comprising:

receiving, by a training component, state data representing a previous state of an object or a vehicle in an environment;

receiving, by the training component, feature vectors representing the vehicle and the object in the environment;

training a codebook to output a trained codebook, the training comprising:

assigning, based at least in part on the state data, a first token to represent the previous state of the object;

assigning a second token to represent a characteristic of the vehicle;

assigning a third token to represent a feature of the environment; and

mapping the feature vectors to respective tokens; and

outputting the trained codebook for use by a machine learned model configured to access tokens from the trained codebook and to arrange the first token, the second token, and the third token to represent a potential interaction between the vehicle and the object.

17. The one or more non-transitory computer-readable media of claim 16 , wherein the machine learned model is a first machine learned model, and the actions further comprising:

causing the machine learned model to arrange the tokens based at least in part on a number of tokens and an order of the tokens.

18. The one or more non-transitory computer-readable media of claim 16 , wherein:

the feature vectors represent continuous variables,

the tokens represent discrete latent variables, and

mapping the feature vectors to respective tokens comprises mapping the continuous variables associated with the feature vectors to the discrete latent variables associated with the tokens.

19. The one or more non-transitory computer-readable media of claim 16 , the actions further comprising:

receiving, by the training component, environmental data representing the feature of the environment,

wherein training the codebook further comprises assigning the second token to represent the feature of the environment, and

the codebook is usable by the machine learned model to cause generation of a scene with the feature in the environment.

20. The one or more non-transitory computer-readable media of claim 16 , wherein the vehicle is an autonomous vehicle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2022
From: PRONOVOST, ETHAN MILLER
To: INC., ZOOX
Reel/Frame 060377/0412 →
Continuity (1)
Related Publication 20240104934A1 · Mar 28, 2024
References Cited (46)
US 10019011B1 · Green et al. · 2018 [cited by applicant]
US 10421453B1 · Ferguson et al. · 2019 [cited by applicant]
US 10459444B1 · Kentley-Klay · 2019 [cited by applicant]
US 10671076B1 · Kobilarov et al. · 2020 [cited by applicant]
US 11370424B1 · Cohen et al. · 2022 [cited by applicant]
US 11960292B2 · Nayhouse et al. · 2024 [cited by applicant]
US 20170132334A1 · Levinson et al. · 2017 [cited by applicant]
US 20190072966A1 · Zhang et al. · 2019 [cited by applicant]
US 20190129436A1 · Sun et al. · 2019 [cited by applicant]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190152490A1 · Lan et al. · 2019 [cited by applicant]
US 20190164007A1 · Liu et al. · 2019 [cited by applicant]
US 20190303759A1 · Farabet et al. · 2019 [cited by applicant]
US 20190332875A1 · Vallespi-Gonzalez et al. · 2019 [cited by applicant]
US 20200148201A1 · King et al. · 2020 [cited by applicant]
US 20200174481A1 · Van Heukelom et al. · 2020 [cited by applicant]
US 20200180647A1 · Anthony · 2020 [cited by applicant]
US 20200216085A1 · Bobier-Tiu et al. · 2020 [cited by applicant]
US 20200225669A1 · Silva et al. · 2020 [cited by applicant]
US 20200283016A1 · Blaiotta · 2020 [cited by applicant]
US 20200324795A1 · Bojarski et al. · 2020 [cited by applicant]
US 20200380085A1 · Behrendt · 2020 [cited by applicant]
US 20200409368A1 · Caldwell et al. · 2020 [cited by applicant]
US 20200409378A1 · Benisch et al. · 2020 [cited by applicant]
US 20210114617A1 · Phillips et al. · 2021 [cited by applicant]
US 20210286924A1 · Wyrwas et al. · 2021 [cited by applicant]
US 20210341927A1 · Refaat et al. · 2021 [cited by applicant]
US 20210347382A1 · Huang et al. · 2021 [cited by applicant]
US 20220153314A1 · Suo · 2022 [cited by examiner]
US 20220161811A1 · Lu et al. · 2022 [cited by applicant]
US 20230121388A1 · Taslim et al. · 2023 [cited by applicant]
US 20240101150A1 · Pronovost · 2024 [cited by applicant]
US 20240101157A1 · Pronovost · 2024 [cited by applicant]
US 20240104934A1 · Pronovost · 2024 [cited by applicant]
US 20240210942A1 · Pronovost · 2024 [cited by applicant]
US 20240211731A1 · Pronovost · 2024 [cited by applicant]
US 20240211797A1 · Pronovost · 2024 [cited by applicant]
US 20240212360A1 · Pronovost · 2024 [cited by applicant]
CN 113936243 · 2022 [cited by applicant]
ICLR 2021 Conference Paper 2345 Authors, Official Comment on Latent Optimization Variational Autoencoder for Conditional Molecular Generation [online], Nov. 13, 2020 (last modification date)Y [retrieved on Mar. 11, 2024… [cited by applicant]
Office Action for U.S. Appl. No. 17/855,671, Dated Jun. 21, 2024, 13 pages. [cited by applicant]
The PCT Search Report and Written Opinion mailed Apr. 29, 2024 for PCT Application No. PCT/US2023/084627 from PCT Summary, 11 pages. [cited by applicant]
The PCT Search Report and Written Opinion mailed Apr. 18, 2024 for PCT Application No. PCT/US2023/084618 from PCT Summary, 13 pages. [cited by applicant]
Bruno Sauvalle et al., Autoencoder-based background reconstruction and foreground segmentation with background noise estimation [online], Jun. 27, 2022\Y [retrieved on Mar. 11, 2024]. Retrieved from<URL: https://www.res… [cited by applicant]
Arsal Syed, Forecasting Pedestrian Trajectory Using Deep Learning, In: Unlv Theses, Dissertations, Professional Papers, and Capstones [online], 2021Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://digitalsch… [cited by applicant]
Hao Xue, Deep Learning Based Pedestrian Trajectory Prediction, In: Thesis—Doctor of Philosophy (research output) [online], 2020Y [retrieved on Mar. 11, 2024]. Retrieved from <URL: https://research-repository.uwa.edu.au/… [cited by applicant]