IP Library Granted Patent US 12,482,189
Granted Patent B2
US 12,482,189 · App. 16/898,110 · Granted Nov 25, 2025

Environment generation using one or more neural networks

Inventors: Seung Wook Kim (Toronto, CA); Sanja Fidler (Toronto, CA); Jonah Philion (Oak Park, IL); Antonio Torralba Barriuso (Somerville, MA)
Assignee: NVIDIA Corporation
G06T19/006G06N3/084G06N5/046G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,189
App. No.
16/898,110
Granted
Nov 25, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to generate a simulated environment. In at least one embodiment, one or more neural networks are used to generate a simulated environment based, at least in part, on stored information associated with objects within the simulated environment.

Claims (42)

1 . One or more processors, comprising:

circuitry to use one or more first neural networks to determine whether objects within a three-dimensional (3D) environment are static objects or dynamic objects and to store positional information corresponding to objects determined by the one or more first neural networks to be static in a different location in memory from positional information corresponding to objects determined by the one or more first neural networks to be dynamic, and to use one or more second neural networks to generate a rendering of the 3D environment that depicts the objects within the 3D environment by accessing the respective stored positional information in memory corresponding to static and dynamic objects.

2 . The one or more processors of claim 1 , wherein the one or more second neural networks include a generative network to receive action input and one or more prior images for an agent in the environment, and further to infer a next image for the agent in the 3D environment.

3 . The one or more processors of claim 2 , wherein the generative network includes a rendering engine to distinguish between static objects and dynamic objects in the 3D environment, and to generate the next image based at least in part upon state data determined for the one or more prior images.

4 . The one or more processors of claim 3 , wherein the circuitry is further to cause information associated with the static objects to be stored to a memory module, wherein the memory module is enabled to generate a memory map of the static objects in the 3D environment for subsequent reuse in generating the 3D environment.

5 . The one or more processors of claim 4 , wherein the generative network further includes a dynamics engine for generating state information from the action input and the one or more prior images, the state information to be provided as input to the rendering engine along with relevant information from the memory module, the relevant information to be determined based at least in part upon the memory map.

6 . The one or more processors of claim 2 , wherein the generative network is trained using an end-to-end training process with a loss function including adversarial loss terms for enforcing realistic image generation, action-conditioned generation, and temporally-consistent generation.

7 . The one or more processors of claim 1 , wherein the one or more second neural networks are to generate the rendering of the 3D environment based, at least in part, on the positional information of the one or more objects within the 3D environment stored in one or more storage locations based on whether the one or more objects are static or dynamic with respect to the 3D environment.

8 . A system comprising:

one or more processors to use one or more first neural networks to determine whether objects within a three-dimensional (3D) environment are static objects or dynamic objects and to store positional information corresponding to objects determined by the one or more first neural networks to be static in a different location in memory from positional information corresponding to objects determined by the one or more first neural networks to be dynamic, and to use one or more second neural networks to generate a rendering of the 3D environment that depicts the objects within the 3D environment by accessing the respective stored positional information in memory corresponding to static and dynamic objects.

9 . The system of claim 8 , wherein the one or more second neural networks include a generative network to receive action input and one or more prior images for an agent in the 3D environment, and further to infer a next image for the agent in the 3D environment.

10 . The system of claim 9 , wherein the generative network includes a rendering engine to distinguish between static objects and dynamic objects in the 3D rendering of the 3D environment, and to generate the next image based at least in part upon state data determined for the one or more prior images.

11 . The system of claim 10 , wherein the one or more processors are further to cause information associated with the static objects to be stored to a memory module, wherein the memory module is enabled to generate a memory map of the static objects in the environment for subsequent reuse in generating the rendering of the 3D environment.

12 . The system of claim 11 , wherein the generative network further includes a dynamics engine for generating state information from the action input and the one or more prior images, the state information to be provided as input to the rendering engine along with relevant information from the memory module, the relevant information to be determined based at least in part upon the memory map.

13 . The system of claim 9 , wherein the generative network is trained using an end-to-end training process with a loss function including adversarial loss terms for enforcing realistic image generation, action-conditioned generation, and temporally-consistent generation.

14 . A method comprising:

determining, using one or more first neural networks, whether objects within a three-dimensional (3D) environment are static objects or dynamic objects;

storing positional information corresponding to objects determined by the one or more first neural networks to be static in a different location in memory from positional information corresponding to objects determined by the one or more first neural networks to be dynamic; and

using one or more second neural networks to generate a rendering of the 3D environment that depicts the objects within the 3D environment by accessing the respective stored positional information in memory corresponding to static and dynamic objects.

15 . The method of claim 14 , wherein the one or more second neural networks include a generative network to receive action input and one or more prior images for an agent in the 3D environment, and further to infer a next image for the agent in the 3D environment.

16 . The method of claim 15 , wherein the generative network includes a rendering engine to distinguish between static objects and dynamic objects in the 3D environment, and to generate the next image based at least in part upon state data determined for the one or more prior images.

17 . The method of claim 16 , further comprising: causing information associated with the static objects to be stored to a memory module, wherein the memory module is enabled to generate a memory map of the static objects in the environment for subsequent reuse in generating the rendering of the 3D environment.

18 . The method of claim 17 , wherein the generative network further includes a dynamics engine for generating state information from the action input and the one or more prior images, the state information to be provided as input to the rendering engine along with relevant information from the memory module, the relevant information to be determined based at least in part upon the memory map.

19 . The method of claim 15 , wherein the generative network is trained using an end-to-end training process with a loss function including adversarial loss terms for enforcing realistic image generation, action-conditioned generation, and temporally-consistent generation.

20 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

determine, using one or more first neural networks, whether objects within a three-dimensional (3D) environment are static objects or dynamic objects;

cause positional information corresponding to objects determined by the one or more first neural networks to be static to be stored in a different location in memory from positional information corresponding to objects determined by the one or more first neural networks to be dynamic; and

use one or more second neural networks to generate a rendering of the 3D environment that depicts the objects within the 3D environment by accessing the respective stored positional information in memory corresponding to static and dynamic objects.

21 . The non-transitory machine-readable medium of claim 20 , wherein the one or more second neural networks include a generative network to receive action input and one or more prior images for an agent in the 3D environment, and further to infer a next image for the agent in the 3D environment.

22 . The non-transitory machine-readable medium of claim 21 , wherein the generative network includes a rendering engine to distinguish between static objects and dynamic objects in the 3D environment, and to generate the next image based at least in part upon state data determined for the one or more prior images.

23 . The non-transitory machine-readable medium of claim 22 , wherein the instructions if performed by the one or more processors further cause the one or more processors to:

cause information associated with the static objects to be stored to a memory module, wherein the memory module is enabled to generate a memory map of the static objects in the 3D environment for subsequent reuse in generating the rendering of the 3D environment.

24 . The non-transitory machine-readable medium of claim 23 , wherein the generative network further includes a dynamics engine for generating state information from the action input and the one or more prior images, the state information to be provided as input to the rendering engine along with relevant information from the memory module, the relevant information to be determined based at least in part upon the memory map.

25 . The non-transitory machine-readable medium of claim 21 , wherein the generative network is trained using an end-to-end training process with a loss function including adversarial loss terms for enforcing realistic image generation, action-conditioned generation, and temporally-consistent generation.

26 . A simulation system, comprising:

one or more processors to use one or more first neural networks to determine whether objects within a three-dimensional (3D) environment are static objects or dynamic objects and to store positional information corresponding to objects determined by the one or more first neural networks to be static in a different location in memory from positional information corresponding to objects determined by the one or more first neural networks to be dynamic, and to use one or more second neural networks to generate a rendering of the 3D environment that depicts the objects within the 3D environment by accessing the respective stored positional information in memory corresponding to static and dynamic objects and

memory for storing network parameters for the one or more first neural networks.

27 . The simulation system of claim 26 , wherein the one or more second neural networks include a generative network to receive action input and one or more prior images for an agent in the 3D environment, and further to infer a next image for the agent in the 3D environment.

28 . The simulation system of claim 27 , wherein the generative network includes a rendering engine to distinguish between static objects and dynamic objects in the 3D environment, and to generate the next image based at least in part upon state data determined for the one or more prior images.

29 . The simulation system of claim 28 , wherein the one or more processors are further to cause information associated with the static objects to be stored to a memory module, wherein the memory module is enabled to generate a memory map of the static objects in the 3D environment for subsequent reuse in generating the rendering of the 3D environment.

30 . The simulation system of claim 29 , wherein the generative network further includes a dynamics engine for generating state information from the action input and the one or more prior images, the state information to be provided as input to the rendering engine along with relevant information from the memory module, the relevant information to be determined based at least in part upon the memory map.

31 . The simulation system of claim 27 , wherein the generative network is trained using an end-to-end training process with a loss function including adversarial loss terms for enforcing realistic image generation, action-conditioned generation, and temporally-consistent generation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2020
From: KIM, SEUNG WOOK; FIDLER, SANJA; PHILION, JONAH; BARRIUSO, ANTONIO TORRALBA
To: NVIDIA CORPORATION
Reel/Frame 052978/0257 →
Continuity (1)
Related Publication 20210390778A1 · Dec 16, 2021
References Cited (135)
US 10540798B1 · Walters et al. · 2020 [cited by applicant]
US 10565458B2 · Fukuhara et al. · 2020 [cited by applicant]
US 10762389B2 · Zuev et al. · 2020 [cited by applicant]
US 10810745B2 · Kang et al. · 2020 [cited by applicant]
US 10832487B1 · Ulbricht · 2020 [cited by examiner]
US 10937169B2 · Dharur et al. · 2021 [cited by applicant]
US 11094134B1 · Fallin et al. · 2021 [cited by applicant]
US 11302012B2 · Kalra et al. · 2022 [cited by applicant]
US 11436825B2 · Lee et al. · 2022 [cited by applicant]
US 11468658B2 · Mironica · 2022 [cited by applicant]
US 11481563B2 · Wason et al. · 2022 [cited by applicant]
US 20060273792A1 · Kholmovski et al. · 2006 [cited by applicant]
US 20180293965A1 · Vembu et al. · 2018 [cited by applicant]
US 20180300940A1 · Sakthivel et al. · 2018 [cited by applicant]
US 20180315248A1 · Bastov · 2018 [cited by examiner]
US 20180350322A1 · Marcu et al. · 2018 [cited by applicant]
US 20190266449A1 · Viola et al. · 2019 [cited by applicant]
US 20190286950A1 · Kiapour et al. · 2019 [cited by applicant]
US 20190295302A1 · Fu et al. · 2019 [cited by applicant]
US 20200215695A1 · Cristache · 2020 [cited by examiner]
US 20200226807A1 · Walters · 2020 [cited by examiner]
US 20200279392A1 · Shamir · 2020 [cited by examiner]
US 20200306638A1 · Fear et al. · 2020 [cited by applicant]
US 20200310442A1 · Halder · 2020 [cited by examiner]
US 20200409380A1 · Song et al. · 2020 [cited by applicant]
US 20210141867A1 · Wason et al. · 2021 [cited by applicant]
US 20220013231A1 · Xanthis et al. · 2022 [cited by applicant]
US 20220114361A1 · Kale et al. · 2022 [cited by applicant]
CN 102938072A · 2013 [cited by applicant]
CN 107031656A · 2017 [cited by applicant]
CN 107045293A · 2017 [cited by applicant]
CN 108932534A · 2018 [cited by applicant]
CN 109410337A · 2019 [cited by applicant]
CN 111052045A · 2020 [cited by applicant]
CN 111402179A · 2020 [cited by applicant]
CN 111738046A · 2020 [cited by applicant]
EP 4001902A1 · 2022 [cited by applicant]
WO 0140897A2 · 2001 [cited by applicant]
Combined Search and Abbreviated Examination Report received in GB Application No. GB2108272.2, dated Oct. 15, 2021. [cited by applicant]
Anonymous, “Model Based Reinforcement Learning for Atari,” International Conference on Learning Representations, 2020, 25 pages. [cited by applicant]
Beer et al., “Evolving Dynamical Neural Networks for Adaptive Behavior,” Adaptive Behavior, 1(1): 1992, 32 pages. [cited by applicant]
Brock et al., “Large Scale GAN Training for High Fidelity Natural Image Synthesis,” Sep. 28, 2018, 29 pages. [cited by applicant]
Chen et al., “Infogan: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets,” In Advances in Neural Information Processing Systems, 2016, 14 pages. [cited by applicant]
Chen et al., “Learning by Cheating,” CoRL, 2019, 10 pages. [cited by applicant]
Chiappa et al., “Recurrent Environment Simulators,” Apr. 19, 2017, 61 pages. [cited by applicant]
Clark et al., “Efficient Video Generation on Complex Datasets,” Sep. 25, 2019, 21 pages. [cited by applicant]
Denton et al., “Stochastic Video Generation with a Learned Prior,” In Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Denton et al., “Unsupervised Learning of Disentangled Representations from Video,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” In Advances in Neural Information Processing Systems, Jun. 10, 2014, 9 pages. [cited by applicant]
Graves et al., “Neural Turing Machines,” Dec. 10, 2014, 26 pages. [cited by applicant]
Guzdial et al., “Game Engine Learning from Video, ” IJCAI, 2017. [cited by applicant]
Ha et al., “Recurrent World Models Facilitate Policy Evolution,” Advances in Neural Informa-tion Processing Systems, 2018, 13 pages. [cited by applicant]
Hansen et al., “Completely De-randomized Self-adaptation in Evolution Strategies,” Evolutionary computation, 9(2): 2001, 39 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, 9(8): 1997, pp. 1735-1780. [cited by applicant]
Huh et al., “Feedback Adversarial Learning: Spatial Feedback for Improving Generative Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” Mar. 2, 2015, 11 pages. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks”, IEEE Conference on Computer Vision and Pattern Recognition, 2017, 10 pages. [cited by applicant]
Jain et al., “Discrete Residual Flow for Probabilistic Pedestrian Behavior Prediction,” Oct. 17, 2019, 13 pages. [cited by applicant]
Kempka et al., “ViZDoom: A Doom-based Al Research Platform for Visual Reinforcement Learning,” IEEE Conference on Computational Intelligence and Games, 2016, 8 pages. [cited by applicant]
Kim et al., “Learning to Discover Cross-Domain Relations with Generative Adversarial Networks,” ICML, 2017, 9pages. [cited by applicant]
Kim et al., “Learning to Simulate Dynamic Environments with GameGAN,” May 25, 2020, 16 pages. [cited by applicant]
Kingma et al. “Adam: A Method for Stochastic Optimization,” arXiv:1412.6980, dated Dec. 22, 2014, 9 pages. [cited by applicant]
Larsen et al., “Autoencoding Beyond Pixels using a Learned Similarity Metric,” International Conference on Machine Learning, 2016, 9 pages. [cited by applicant]
Lee et al., “Stochastic Adversarial Video Prediction,” Apr. 4, 2018, 26 pages. [cited by applicant]
Liu et al., “Unsupervised Image-to-Image Translation Networks,” Oct. 9, 2017, 11 pages. [cited by applicant]
Mescheder et al., “Which Training Methods for GANs do actually Converge?,” Jul. 31, 2018, 39 pages. [cited by applicant]
Mirza et al., “Conditional Generative Adversarial Nets,” Nov. 6, 2014, 7 pages. [cited by applicant]
Miyato et al., cGANs with Projection Discriminator, ICLR, 2018, 23 pages. [cited by applicant]
Mnih et al., “Asynchronous Methods for Deep Reinforcement Learning,” Jun. 16, 2016, 19 pages. [cited by applicant]
Mnih et al., “Playing Atari with Deep Reinforcement Learning,” Dec. 19, 2013, 9 pages. [cited by applicant]
Oh et al., “Action-Conditional Video Prediction using Deep Networks in Atari Games,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Park et al., “Semantic Image Synthesis with Spatially Adaptive Normalization,” CVPR, 2019, 10 pages. [cited by applicant]
Paxton et al., “Combining Neural Networks and Tree Search for Task and Motion Planning in Challenging Environments,” IEEE/RSJ International Conference on Intelligent Robots and Systems, 2017, 2 pages. [cited by applicant]
Reed et al., “Generative Adversarial Text-to-Image Synthesis,” ICML, Jun. 5, 2016, 10 pages. [cited by applicant]
Salimans et al., “Improved Techniques for Training GANs,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Srivastava et al., “Unsupervised Learning of Video Representations using LSTMs,” International Conference on Machine Learning, 2015, 10 pages. [cited by applicant]
Tulyakov et al., “MoCoGAN: Decomposing Motion and Content for Video Generation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Ulyanov et al., “Instance Normalization: The Missing Ingredient for Fast Stylization,” Nov. 6, 2017, 6 pages. [cited by applicant]
Wang et al., “Video-to-Video Synthesis”, In Neural Information Processing Systems, 2018, 13 pages. [cited by applicant]
Yang et al., “Diversity-Sensitive Conditional Generative Adversarial Networks,” Jan. 25, 2019, 23 pages. [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks,” Nov. 24, 2017, 20 pages. [cited by applicant]
Formisano et al., “Spatial Independent Component Analysis of Functional Magnetic Resonance Imaging Time-series: Characterization of the Cortical Components,” Neurocomputing, 49(1-4): 2002, 14 pages. [cited by applicant]
Ji et al., “Target Detection and Classified Recognition Method Based on Optical Remote Sensing Image,” Journal of Shenyang University, Feb. 2015, 9 pages. [cited by applicant]
Office Action for Chinese Application No. 202210167572.3, mailed Apr. 27, 2025, 27 pages. [cited by applicant]
Zhao et al., “Fundamentals of Big Data Technology,” China Machine Press, 2020, 6 pages. [cited by applicant]
Babaeizadeh et al., “Stochastic Variational Video Prediction,” 2017, 12 pages. [cited by applicant]
Bellemare et al., “The Arcade Learning Environment: An Evaluation Platform for General Agents,” Journal of Artificial Intelligence Research, 2013, 27 pages. [cited by applicant]
Chen et al., “Traffic Signal Control Based on Deep Reinforcement Learning,” Modern Computers—Research and Development, Jan. 25, 2020, 5 pages. [cited by applicant]
Combined Search and Examination Report for Application No. GB2202537.3, mailed Aug. 22, 2022, 7 pages. [cited by applicant]
Deisenroth et al., “PILCO: AModel-Based and Data-Efficient Approach toPolicy Search,” Proceedings of the 28th International Conference on machine learning, 2011, 8 pages. [cited by applicant]
Devaranjan et al., “Meta-Sim2: Learning to Generate Synthetic Datasets,” ECCV, 2020, 19 pages. [cited by applicant]
Dosovitskiy et al., “CARLA: An Open Urban Driving Simulator,” 1st Conference on Robot Learning (CoRL '17), Nov. 10, 2017, 16 pages. [cited by applicant]
Dumoulin et al., “A Learned Representation for Artistic Style,” International Conference on Learning Representations, Sep. 9, 2017, 26 pages. [cited by applicant]
Finn et al., “Unsupervised Learning for Physical Interaction Through Video Prediction,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Ghiasi et al., “Exploring the Structure of a Real-Time, Arbitrary Neural Artistic Stylization Network,” BMVC, 2017, 16 pages. [cited by applicant]
Hafner et al., “Learning Latent Dynamics for Planning from Pixels,” International Conference on Machine Learning, 2019, 11 pages. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-scale Update Rule Converge to a Local Nash Equilibrium,” Advances in Neural Information Processing Systems, 2017, 12 pages. [cited by applicant]
Higgins et al., “Beta-vae: Learning Basic Visual Concepts with a Constrained Variational Framework,” ICLR, 2(5):6, 2017, 13 pages. [cited by applicant]
Hsieh et al., “Learning to Decompose and Disentangle Representations for Video Prediction,” CoRR, 2018, 14 pages. [cited by applicant]
Huang et al., “Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization,” ICCV, 2017, 10 pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
Kaiser et al., “Model-Based Reinforcement Learning for Atari,” Jun. 11, 2019, 18 pages. [cited by applicant]
Kalchbrenner et al., “Video Pixel Networks,” Oct. 3, 2016, 16 pages. [cited by applicant]
Kar et al., “Meta-Sim: Learning to Generate Synthetic Datasets, ” ICCV, 2019, 14 pages. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” CVPR, 2019, 10 pages. [cited by applicant]
Karras et al., “Analyzing and Improving the Image Quality of Stylegan,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Kim, “Convolutional Neural Networks for Sentence Classification,” 2014, 6 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” International Conference on Learning Representations 2014 (ICLR 2014), Apr. 14, 2014, 14 pages. [cited by applicant]
Kumar et al., “Videoflow: A Flow-Based Generative Model for Video,” Jun. 10, 2019, 13 pages. [cited by applicant]
Liese et al., A Study of the Simulated Evolution of the Spectral Sensitivity of Visual Agent Receptors, Artificial Life, Mar. 31, 2001, 26 pages. [cited by applicant]
Lotter et al., “Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning,” Aug. 31, 2016, 12 pages. [cited by applicant]
Mallya et al., “World-consistent Video-to-Video Synthesis,” 2020, 25 pages. [cited by applicant]
Manivasagam et al., “LiDARsim: Realistic LiDAR Simulation by Leveraging the RealWorld,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Mathieu et al., “Deep Multi-Scale Video Prediction Beyond Mean Square Error,” International Conference on Learning Representations, 2016, 14 pages. [cited by applicant]
Minderer et al., “Unsupervised learning of object structure and dynamics from videos,” Advances in Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Office Action for Chinese Application No. 202110638931.4, mailed Jun. 1, 2023, 11 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2108272.2, mailed Nov. 29, 2022, 5 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2202537.3, mailed Sep. 18, 2023, 3 pages. [cited by applicant]
Philion et al., “Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D,” 2020, 17 pages. [cited by applicant]
Ranzato et al., “Video (Language) Modeling: A Baseline for Generative Models of Natural Videos,” 2014, 15 pages. [cited by applicant]
Ruiz et al., “Learning To Simulate,” Oct. 5, 2018, retrieved Jan. 14, 2021 from https://arxiv.org/pdf/1810.02513.pdf, 12 pages. [cited by applicant]
Saito et al., “Temporal Generative Adversarial Nets with Singular Value Clipping,” IEEE International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Saito et al., “TGANv2: Efficient Training of Large Models for Video Generation with Multiple Subsampling Layers,” 2018, 12 pages. [cited by applicant]
Shaham et al., “SinGAN: Learning a Generative Model from a Single Natural Image,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 27, 2019, 11 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” Machine Learning Proceedings 1990, 9 pages. [cited by applicant]
Todorov et al., “MuJoCo : A Physics Engine for Model-Based Control,” 2012, 8 pages. [cited by applicant]
Unterthiner et al., “Towards Accurate Generative Models of Video: A New Metric & Challenges,” 2018, 16 pages. [cited by applicant]
Vondrick et al., “Generating Videos with Scene Dynamics,” 29th Conference on Neural Information Processing Systems (NIPS), 2016, pp. 1-10. [cited by applicant]
Weissenborn et al., “Scaling Autoregressive Video Models,” 2019, 22 pages. [cited by applicant]
Xia et al., “Gibson Env: Real-World Perception for Embodied Agents,” Aug. 31, 2018, 12 pages. [cited by applicant]
Yu et al., “Efficient and Information-Preserving Future Frame Prediction and Beyond,” ICLR, 2020, 14 pages. [cited by applicant]
Zhang et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” CVPR, 2018, 10 pages. [cited by applicant]