IP Library › Granted Patent US 12,387,088
Granted Patent B2
US 12,387,088 · App. 18/199,865 · Granted Aug 12, 2025

Method and system for generating one or more conditionally dependent data entries

Inventors: William Harvey (Vancouver, CA); Saeid Naderiparizi (Vancouver, CA); Vaden Masrani (Vancouver, CA); Christian Dietrich Weilbach (Vancouver, CA); Frank Wood (Vancouver, CA)
Assignee: The University of British Columbia
G06N3/047G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,088
App. No.
18/199,865
Granted
Aug 12, 2025
Kind
B2
Abstract

Methods, systems, and techniques for generating one or more conditionally dependent data entries using a probabilistic generative model, and for training that model. The probabilistic generative model is trained using a plurality of data generation tasks respectively corresponding to a plurality of vectors each having a sequence of differently indexed data entries that are conditionally dependent on each other. Each of the data generation tasks involves generating at least one latent data entry selected from the sequence of data entries in response to being provided at least one index for each of the at least one latent data entry. Training can be performed by minimizing an expected value of a denoising loss over all training stages.

Claims (50)

1. A method comprising generating data using a probabilistic generative model, wherein the probabilistic generative model is trained using a plurality of different data generation tasks respectively corresponding to a plurality of vectors each comprising a sequence of differently indexed data entries that are conditionally dependent on each other, wherein each of the data generation tasks comprises generating at least one latent data entry selected from the sequence of data entries in response to being provided at least one index for each of the at least one latent data entry, and wherein during training each of the at least one latent data entry that is generated is evaluated against an actual value of the at least one latent data entry.

2. The method of claim 1 , wherein the probabilistic generative model comprises a score based diffusion model.

3. The method of claim 2 , wherein the score based diffusion model comprises a denoising diffusion probabilistic model (DDPM).

4. The method of claim 1 , wherein the differently indexed data entries are indexed at least by time.

5. The method of claim 1 , wherein the probabilistic generative model comprises at least one attention layer, wherein each of the at least one attention layer relates different ones of the differently indexed data entries according to indices or modalities of the differently indexed data entries.

6. The method of claim 5 , wherein the sequence of data entries comprises different frames of a video, wherein the at least one attention layer comprises at least one spatial attention layer and at least one temporal attention layer, wherein the at least one spatial attention layer relates different ones of the frames to each other spatially, and wherein the at least one temporal layer relates different ones of the frames to each other temporally.

7. The method of claim 1 , wherein relative position encodings are used to relate different ones of the differently indexed data entries to each other by respective indices of the differently indexed data entries.

8. The method of claim 7 , wherein the sequence of data entries comprises different frames of a video, and wherein the relative position encodings are used to relate different ones of the frames to each other temporally in the at least one temporal layer.

9. The method of claim 1 , wherein at least one of the data generation tasks is a conditional data generation task and comprises providing to the probabilistic generative model at least one observed data entry and at least one corresponding index selected from the sequence of data entries, and wherein generating the at least one latent data entry is conditioned on the at least one observed data entry and at least one corresponding index.

10. The method of claim 9 , wherein the at least one corresponding index comprises a time index, and wherein the time index is later in time than a corresponding time index of the at least one latent data entry.

11. The method of claim 9 , wherein the at least one observed data entry and the at least one latent data entry are of different modalities.

12. The method of claim 9 , wherein for the at least one of the data generation tasks that is the conditional data generation task, at least one of the at least one latent data entry that is generated is non-consecutively indexed relative to each of the at least one observed data entry.

13. The method of claim 1 , wherein for at least one of the different data generation tasks, the at least one latent data entry is unconditionally generated by requiring generating without inputting to the probabilistic generative model a particular one of the data entries on which to condition generation of the at least one latent data entry.

14. The method of claim 1 , wherein for at least one of the different data generation tasks, the data is generated unconditionally by requiring generation without providing an input data sequence to the conditional density estimator on which to condition generation of the data.

15. The method of claim 1 , wherein for at least one of the different data generation tasks, generating the data is performed conditionally and comprises:

providing to the probabilistic generative model an input data sequence comprising at least one data entry and at least one corresponding index of the same type as the at least one data entry and at least one corresponding index used to train the probabilistic generative model; and

using the probabilistic generative model to generate at least one new sample data entry over multiple stages conditioned on the at least one data entry and at least one corresponding index provided to the probabilistic generative model.

16. The method of claim 15 , wherein the conditional density estimator comprises a denoising diffusion probabilistic model (DDPM), the data entries of the input data sequence comprise different frames of a video segment, the sample data entries are sample frames and the data entries on which the sample frames are conditioned are conditional frames, and wherein following the multiple stages a consecutive series of the sample frames have been added to the video segment.

17. The method of claim 16 , wherein one of the sample frames generated during one of the stages is used as one of the conditional frames during a subsequent one of the stages.

18. The method of claim 16 , wherein one of the sample frames is conditioned on one of the conditional frames that occurs later in time than the one of the sample frames.

19. The method of claim 18 , wherein the one of the sample frames is also conditioned on one of the conditional frames that occurs earlier in time than the one of the sample frames.

20. The method of claim 16 , wherein at least some of conditional frames are used to condition the new sample frames for all of the stages.

21. The method of claim 15 , wherein the at least one new sample data entry is of a different modality than the at least one data entry on which generation of the at least one new sample data entry is conditioned.

22. A system comprising:

at least one processor;

storage, communicatively coupled to the at least one processor; and

at least one non-transitory computer readable medium communicatively coupled to the at least one processor, the at least one non-transitory computer readable medium having stored thereon computer program code that, when executed, causes the at least one processor to perform a method comprising:

generating data using a probabilistic generative model, wherein the probabilistic generative model is trained using a plurality of different data generation tasks respectively corresponding to a plurality of vectors each comprising a sequence of differently indexed data entries that are conditionally dependent on each other, and wherein each of the data generation tasks comprises generating at least one latent data entry selected from the sequence of data entries in response to being provided at least one index for each of the at least one latent data entry; and

storing the data that is generated in the storage.

23. The method of claim 22 , wherein for at least one of the data generation tasks:

the data generation task is a conditional data generation task and comprises providing to the probabilistic generative model at least one observed data entry and at least one corresponding index selected from the sequence of data entries, and wherein generating the at least one latent data entry is conditioned on the at least one observed data entry and the at least one corresponding index, and

at least one of the at least one latent data entry that is generated is non-consecutively indexed relative to each of the at least one observed data entry.

24. A non-transitory computer readable medium comprising computer program code that, when executed, causes at least one processor to perform a method comprising generating data using a probabilistic generative model, wherein the probabilistic generative model is trained using a plurality of different data generation tasks respectively corresponding to a plurality of vectors each comprising a sequence of differently indexed data entries that are conditionally dependent on each other, wherein each of the data generation tasks comprises generating at least one latent data entry selected from the sequence of data entries in response to being provided at least one index for each of the at least one latent data entry, and wherein during training each of the at least one latent data entry that is generated is evaluated against an actual value of the at least one latent data entry.

25. The non-transitory computer readable medium of claim 24 , wherein for at least one of the data generation tasks:

the data generation task is a conditional data generation task and comprises providing to the probabilistic generative model at least one observed data entry and at least one corresponding index selected from the sequence of data entries, and wherein generating the at least one latent data entry is conditioned on the at least one observed data entry and the at least one corresponding index, and

at least one of the at least one latent data entry that is generated is non-consecutively indexed relative to each of the at least one observed data entry.

26. A method comprising:

performing, over multiple stages:

inputting to a probabilistic generative model at least one index respectively corresponding to at least one latent data entry to be generated;

generating, using the probabilistic generative model, the at least one latent data entry based on the at least one index corresponding to the at least one latent data entry, wherein the at least one latent data entry is selected from a sequence of differently indexed data entries that are conditionally dependent on each other; and

determining a denoising loss based on the at least one latent data entry; and

training the probabilistic generative model based on reducing an overall denoising loss over at least some of the multiple stages, wherein the multiple stages correspond to different data generation tasks.

27. The method of claim 26 , wherein the training comprises minimizing an expected value of the overall denoising loss over all of the stages.

28. The method of claim 27 , wherein the score based diffusion model comprises a denoising diffusion probabilistic model (DDPM).

29. The method of claim 26 , wherein the probabilistic generative model comprises a score based diffusion model.

30. The method of claim 26 , wherein the probabilistic generative model comprises at least one attention layer, wherein each of the at least one attention layer relates different ones of the differently indexed data entries according to indices or modalities of the differently indexed data entries.

31. The method of claim 26 , wherein relative position encodings are used to relate different ones of the differently indexed data entries to each other by respective indices of the differently indexed data entries.

32. The method of claim 26 , wherein at least one of the data generation tasks is a conditional data generation task and comprises providing to the probabilistic generative model at least one observed data entry and at least one corresponding index selected from the sequence of data entries, and wherein generating the at least one latent data entry is conditioned on the at least one observed data entry and at least one corresponding index.

33. The method of claim 32 , wherein for the at least one of the data generation tasks that is the conditional data generation task, at least one of the at least one latent data entry that is generated is non-consecutively indexed relative to each of the at least one observed data entry.

34. The method of claim 26 , wherein for at least one of the different data generation tasks, the at least one latent data entry is unconditionally generated by requiring generating without inputting to the probabilistic generative model a particular one of the data entries on which to condition generation of the at least one latent data entry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2023
From: HARVEY, WILLIAM; NADERIPARIZI, SAEID; MASRANI, VADEN; WEILBACH, CHRISTIAN DIETRICH; WOOD, FRANK
To: THE UNIVERSITY OF BRITISH COLUMBIA
Reel/Frame 064578/0507 →
Continuity (1)
Related Publication 20240386249A1 · Nov 21, 2024
References Cited (71)
Yang, Ruihan, Prakhar Srivastava, and Stephan Mandt. “Diffusion probabilistic modeling for video generation.” arXiv preprint arXiv:2203.09481 (2022). (Year: 2022). [cited by examiner]
Kumar, Manoj, et al. “Videoflow: A flow-based generative model for video.” arXiv preprint arXiv:1903.01434 2.5 (2019): 3. (Year: 2019). [cited by examiner]
Kumar, Manoj, et al. “Videoflow: A flow-based generative model for video.” arXiv preprint arXiv:1903.01434 2.5 (2019): (Year: 2019). [cited by examiner]
Filipowicz, Wlodzimierz. “Conditional dependencies in imprecise data handling.” Procedia Computer Science 192 (2021): 80-89. (Year: 2021). [cited by examiner]
Sruthi, Panchapakesan C., Sanjay Rao, and Bruno Ribeiro. “Pitfalls of data-driven networking: A case study of latent causal confounders in video streaming.” Proceedings of the Workshop on Network Meets AI & ML. 2020. (Y… [cited by examiner]
D. Nguyen-tuong et al. “Learning Inverse Dynamics: A Comparison”. Conference: ESANN 2008, 16th European Symposium on Artificial Neural Networks, Bruges, Belgium, Apr. 23-25, 2008. [cited by applicant]
T. Osa et al. “An algorithmic perspective on imitation learning”. Foundations and Trends® in Robotics 7.1-2 (2018), pp. 1-179. [cited by applicant]
A. Radford et al. “Learning transferable visual models from natural language supervision”. International Confer¬ence on Machine Learning. PMLR. 2021, pp. 8748-8763. [cited by applicant]
J. Rawlings, E. Meadows, and K. Muske. “Nonlinear Model Predictive Control: A Tutorial and Survey”. IFAC Proceedings vols. 27.2 (1994), pp. 185-197. ISSN: 1474-6670. DOI: 10.1016/S1474-6670(17)48151-1. [cited by applicant]
G. Rong et al. “LGSVL Simulator: A High Fidelity Simulator for Autonomous Driving”. arXiv preprint arXiv:2005.03778 (2020). [cited by applicant]
C. Saharia et al. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. May 2022. DOI: 10.48550/arXiv.2205.11487. arXiv: 2205.11487 [cs]. [cited by applicant]
T. Salimans et al. “Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications”. arXiv preprint arXiv:1701.05517 (2017). [cited by applicant]
S. Shah et al. “AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles”. Field and Service Robotics. 2017. eprint: arXiv:1705.05065. URL: https://arxiv.org/abs/1705.05065. [cited by applicant]
X. Shi et al. “Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAM”. 2020 International Conference on Robotics and Automation (ICRA). 2020, pp. 3139-3145. [cited by applicant]
Y. Song and S. Ermon. “Generative modeling by estimating gradients of the data distribution”. Advances in Neural Information Processing Systems 32 (2019). [cited by applicant]
F. Torabi, G. Warnell, and P. Stone. Behavioral Cloning from Observation. May 2018. DOI: 10.48550/arXiv.1805.01954. arXiv: 1805.01954 [cs]. [cited by applicant]
W. Zhan et al. Interaction Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps. Sep. 2019. DOI: 10. 48550/arXiv. 1910 . 03088. arXiv: 1910.03088 [cs,… [cited by applicant]
Dan Bigioi, Shubhajit Basak, Hugh Jordan, Rachel McDonnell, and Peter Corcoran. Speech driven video editing via an audio-conditioned diffusion model. arXiv preprint arXiv:2301.04474, 2023. [cited by applicant]
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. arXiv preprint arXiv:2204.03638, 2022. [cited by applicant]
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. [cited by applicant]
Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo. Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. arXiv preprint a… [cited by applicant]
Christian Weilbach, William Harvey, and Frank Wood. Graphically structured diffusion models. arXiv preprint arXiv:2210.11633, 2022. [cited by applicant]
Nuha Aldausari, Arcot Sowmya, Nadine Marcus, and Gelareh Mohammadi. Video generative adversarial networks: a review. ACM Computing Surveys (CSUR), 55(2):1-25, 2022. [cited by applicant]
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H Campbell, and Sergey Levine. Stochastic variational video prediction. arXiv preprint arXiv:1710.11252, 2017. [cited by applicant]
Mohammad Babaeizadeh, Mohammad Taghi Saffar, Suraj Nair, Sergey Levine, Chelsea Finn, and Dumitru Erhan. Fitvid: Overfitting in pixel-level video prediction. arXiv preprint arXiv:2106.13195, 2021. [cited by applicant]
Rewon Child. Very deep vaes generalize autoregressive models and can outperform them on images. arXiv preprint arXiv:2011.10650, 2020. [cited by applicant]
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio. A recurrent latent variable model for sequential data. Advances in neural information processing systems, 28, 2015. [cited by applicant]
Aidan Clark, Jeff Donahue, and Karen Simonyan. Adversarial video generation on complex datasets. arXiv preprint arXiv:1907.06571, 2019. [cited by applicant]
Emily Denton and Rob Fergus. Stochastic video generation with a learned prior. In International conference on machine learning, pp. 1174-1183. PMLR, 2018. [cited by applicant]
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34, 2021. [cited by applicant]
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. Carla: An open urban driving simulator. In Conference on robot learning, pp. 1-16. PMLR, 2017. [cited by applicant]
SM Ali Eslami, Danilo Jimenez Rezende, Frederic Besse, Fabio Viola, Ari S Morcos, Marta Garnelo, Avraham Ruderman, Andrei A Rusu, Ivo Danihelka, Karol Gregor, et al. Neural scene representation and rendering. Science, 3… [cited by applicant]
Audrunas Gruslys, Remi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves. Memory efficient backpropagation through time. Advances in Neural Information Processing Systems, 29, 2016. [cited by applicant]
William H Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov. Minerl: A large-scale dataset of minecraft demonstrations. arXiv preprint arXiv:1907.13440, 2019. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. arxiv 2015. arXiv preprint arXiv:1512.03385, 2015. [cited by applicant]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840-6851, 2020. [cited by applicant]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022. [cited by applicant]
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson. Deep variational reinforcement learning for pomdps. In International Conference on Machine Learning, pp. 2117-2126. PMLR, 2018. [cited by applicant]
Zahra Kadkhodaie and Eero P Simoncelli. Solving linear inverse problems using the prior implicit in a denoiser. arXiv preprint arXiv:2007.13640, 2020. [cited by applicant]
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al. Modelbased reinforcement learning for atari. arXi… [cited by applicant]
Taesup Kim, Sungjin Ahn, and Yoshua Bengio. Variational temporal abstraction. Advances in Neural Information Processing Systems, 32, 2019. [cited by applicant]
Gautam Mittal, Jesse Engel, Curtis Hawthorne, and Ian Simon. Symbolic music generation with diffusion models. arXiv preprint arXiv:2103.16091, 2021. [cited by applicant]
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pp. 8162-8171. PMLR, 2021. [cited by applicant]
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pp. 234-241… [cited by applicant]
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, 2022. [cited by applicant]
Vaibhav Saxena, Jimmy Ba, and Danijar Hafner. Clockwork variational autoencoders. Advances in Neural Information Processing Systems, 34, 2021. [cited by applicant]
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018. [cited by applicant]
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256-2265. PMLR, 2015. [cited by applicant]
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. [cited by applicant]
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. [cited by applicant]
Corentin Tallec and Yann Ollivier. Unbiasing truncated backpropagation through time. arXiv preprint arXiv:1705.08209, 2017. [cited by applicant]
Yusuke Tashiro, Jiaming Song, Yang Song, and Stefano Ermon. Csdi: Conditional score-based diffusion models for probabilistic time series imputation. Advances in Neural Information Processing Systems, 34, 2021. [cited by applicant]
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 20… [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. [cited by applicant]
Ruben Villegas, Dumitru Erhan, Honglak Lee, et al. Hierarchical long-term video prediction without supervision. In International Conference on Machine Learning, pp. 6038-6046. PMLR, 2018. [cited by applicant]
Dirk Weissenborn, Oscar Täckström, and Jakob Uszkoreit. Scaling autoregressive video models. arXiv preprint arXiv:1906.02634, 2019. [cited by applicant]
Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao. Rethinking and improving relative position encoding for vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, … [cited by applicant]
Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. Tackling the generative learning trilemma with denoising diffusion GANs. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id… [cited by applicant]
Ruihan Yang, Prakhar Srivastava, and Stephan Mandt. Diffusion probabilistic modeling for video generation. arXiv preprint arXiv:2203.09481, 2022. [cited by applicant]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern … [cited by applicant]
B. Baker et al. Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos. Jun. 2022. arXiv:2206.11795 [cs]. [cited by applicant]
A. Bucker et al. LaTTe: Language Trajectory TransformEr. Aug. 2022. DOI: 10.48550/arXiv.2208.02918. arXiv:2208.02918 [cs]. [cited by applicant]
J. Devlin et al. “Bert: Pre-training of deep bidirectional transformers for language understanding”. arXiv preprint arXiv:1810.04805 (2018). [cited by applicant]
B. B. Elallid et al. “A Comprehensive Survey on the Application of Deep and Reinforcement Learning Approaches in Autonomous Driving”. Journal of King Saud University—Computer and Information Sciences (Apr. 2022). ISSN: … [cited by applicant]
S. Grigorescu et al. “A Survey of Deep Learning Techniques for Autonomous Driving”. Journal of Field Robotics 37.3 (Apr. 2020), pp. 362-386. ISSN: 1556-4959, 1556-4967. DOI: 10. 1002/rob.21918. arXiv: 1910.07738 [cs]. [cited by applicant]
W. Harvey et al. Flexible Diffusion Modeling of Long Videos. May 2022. DOI: 10. 48550/arXiv .2205. 11495. arXiv: 2205.11495 [cs]. [cited by applicant]
D. Hrovat et al. “The Development of Model Predictive Control in Automotive Industry: A Survey”. 2012 IEEE International Conference on Control Applications. Oct. 2012, pp. 295-302. DOI: 10.1109/CCA.2012.6402735. [cited by applicant]
Y. Huang and Y. Chen. Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies. Jul. 2020. DOI: 10.48550/arXiv.2006.06091. arXiv: 2006.06091 [cs]. [cited by applicant]
A. Hussein et al. “Imitation learning: A survey of learning methods”. ACM Computing Surveys (CSUR) 50.2 (2017), pp. 1-35. [cited by applicant]
M. Janner et al. Planning with Diffusion for Flexible Behavior Synthesis. May 2022. arXiv: 2205.09991 [cs]. [cited by applicant]