IP Library Granted Patent US 12,738,176
Granted Patent B2
US 12,738,176 · App. 17/953,894 · Granted Sep 15, 2026

Systems and methods for training a scene simulator using real and simulated agent data

Inventors: Tsun-Hsuan Wang (Cambridge, MA); Alexander Amini (Brookline, MA); Wilko Schwarting (Boston, MA); Igor Gilitschenski (Cambridge, MA); Sertac Karaman (Cambridge, MA); Daniela Rus (Weston, MA)
Assignees: Toyota Research Institute, Inc.; Toyota Jidosha Kabushiki Kaisha; Massachusetts Institute of Technology
G09B9/042G06N20/00G09B9/05
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,176
App. No.
17/953,894
Granted
Sep 15, 2026
Kind
B2
Abstract

System, methods, and other embodiments described herein relate to training a scene simulator for rendering 2D scenes using data from real and simulated agents. In one embodiment, a method includes acquiring trajectories and three-dimensional (3D) views for multiple agents from observations of real vehicles. The method also includes generating a 3D scene having the multiple agents using the 3D views and information from simulated agents. The method also includes training a scene simulator to render scene projections using the 3D scene. The method also includes outputting a 2D scene having simulated observations for a driving scene using the scene simulator.

Claims (61)

1 . A learning system, comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the processor to:

acquire trajectories and three-dimensional (3D) views for multiple agents from observations of real vehicles;

generate a 3D scene having the multiple agents by casting ambient light within the 3D views to the real vehicles associated with an infinite distance with the multiple agents and information from simulated agents, wherein the ambient light is associated with an average color of the 3D scene;

train a scene simulator with a policy to render scene projections using the 3D scene with a foreground mask, and a reinforcement model computes the policy and increases a reward for a driving task associated with an ego vehicle within a driving scene that is using the policy; and

output a 2D scene having simulated observations for the driving scene using the scene simulator.

2 . The learning system of claim 1 , further including instructions to:

compute, by the reinforcement model, the reward for the driving task using a simulated scene estimated by the scene simulator;

form the policy that increases the reward, wherein the reinforcement model outputs a control signal for an ego vehicle within the driving scene using the policy; and

train the scene simulator according to the policy.

3 . The learning system of claim 2 , wherein the instructions to compute the reward further include instructions to:

compose the simulated scene using representations of visual observations for a driving scenario, wherein the visual observations are associated with data from an ego vehicle and represent a simulated state.

4 . The learning system of claim 2 , wherein the control signal is one of a steering angle and a throttle amount and the driving task is one of vehicle following, avoidance, and overtaking a vehicle.

5 . The learning system of claim 2 , further including instructions to increase the reward in response to a vehicle avoiding an object or tracking a lane during a driving scenario.

6 . The learning system of claim 1 , wherein the instructions to generate the 3D scene further include instructions to:

cast directional light within the 3D views according to the real vehicles; and

process the 3D scene with the 3D views and the foreground mask to further train the scene simulator.

7 . The learning system of claim 1 , further including instructions to:

select vehicle bodies for the multiple agents randomly from a mesh library; and

mesh the vehicle bodies of the multiple agents for the 3D scene using distances between the multiple agents.

8 . The learning system of claim 7 , further including instructions to:

transform the vehicle bodies in the 3D scene to determine collisions between the multiple agents; and

train the scene simulator according to the collisions for estimating 2D views.

9 . A non-transitory computer-readable medium comprising: instructions that when executed by a processor cause the processor to:

acquire trajectories and three-dimensional (3D) views for multiple agents from observations of real vehicles;

generate a 3D scene having the multiple agents by casting ambient light within the 3D views to the real vehicles associated with an infinite distance with the multiple agents and information from simulated agents, wherein the ambient light is associated with an average color of the 3D scene;

train a scene simulator with a policy to render scene projections using the 3D scene with a foreground mask, and a reinforcement model computes the policy and increases a reward for a driving task associated with an ego vehicle within a driving scene that is using the policy; and

output a 2D scene having simulated observations for the driving scene using the scene simulator.

10 . The non-transitory computer-readable medium of claim 9 , further including instructions to:

compute, by the reinforcement model, the reward for the driving task using a simulated scene estimated by the scene simulator;

form the policy that increases the reward, wherein the reinforcement model outputs a control signal for an ego vehicle within the driving scene using the policy; and

train the scene simulator according to the policy.

11 . The non-transitory computer-readable medium of claim 10 , wherein the instructions to compute the reward further include instructions to:

compose the simulated scene using representations of visual observations for a driving scenario, wherein the visual observations are associated with data from an ego vehicle and represent a simulated state.

12 . The non-transitory computer-readable medium of claim 10 , wherein the instructions to generate the 3D scene further include instructions to:

cast directional light within the 3D views according to the real vehicles; and

process the 3D scene with the 3D views and the foreground mask to further train the scene simulator.

13 . A method comprising:

acquiring trajectories and three-dimensional (3D) views for multiple agents from observations of real vehicles;

generating a 3D scene having the multiple agents by casting ambient light within the 3D views to the real vehicles associated with an infinite distance with the multiple agents and information from simulated agents, wherein the ambient light is associated with an average color of the 3D scene;

training a scene simulator with a policy to render scene projections using the 3D scene with a foreground mask, and a reinforcement model computes the policy and increases a reward for a driving task associated with an ego vehicle within a driving scene that is using the policy; and

outputting a 2D scene having simulated observations for the driving scene using the scene simulator.

14 . The method of claim 13 , further comprising:

computing, by the reinforcement model, the reward for the driving task using a simulated scene estimated by the scene simulator;

forming the policy that increases the reward, wherein the reinforcement model outputs a control signal for an ego vehicle within the driving scene using the policy; and

training the scene simulator according to the policy.

15 . The method of claim 14 , wherein computing the reward further includes:

composing the simulated scene using representations of visual observations for a driving scenario, wherein the visual observations are associated with data from an ego vehicle and represent a simulated state.

16 . The method of claim 14 , wherein the control signal is one of a steering angle and a throttle amount and the driving task is one of vehicle following, avoidance, and overtaking a vehicle.

17 . The method of claim 14 , further comprising:

increasing the reward in response to a vehicle avoiding an object or tracking a lane during a driving scenario.

18 . The method of claim 13 , wherein generating the 3D scene further includes:

casting directional light within the 3D views according to the real vehicles; and

processing the 3D scene with the 3D views and the foreground mask to further train the scene simulator.

19 . The method of claim 13 , further comprising:

selecting vehicle bodies for the multiple agents randomly from a mesh library; and

meshing the vehicle bodies of the multiple agents for the 3D scene using distances between the multiple agents.

20 . The method of claim 19 , further comprising:

transforming the vehicle bodies in the 3D scene to determine collisions between the multiple agents; and

training the scene simulator according to the collisions for estimating 2D views.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: WANG, TSUN-HSUAN; AMINI, ALEXANDER; SCHWARTING, WILKO; KARAMAN, SERTAC; RUS, DANIELA
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 061457/0747 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: GILITSCHENSKI, IGOR
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 061457/0944 →
Continuity (1)
Related Publication 20240119857A1 · Apr 11, 2024
References Cited (80)
US 10643320B2 · Lee et al. · 2020 [cited by applicant]
US 12037013B1 · Linscott · 2024 [cited by examiner]
US 12112259B1 · Dirac · 2024 [cited by examiner]
US 20160314224A1 · Wei · 2016 [cited by examiner]
US 20180275658A1 · Iandola · 2018 [cited by examiner]
US 20190094875A1 · Schulter · 2019 [cited by examiner]
US 20190228571A1 · Atsmon · 2019 [cited by examiner]
US 20190303759A1 · Farabet et al. · 2019 [cited by applicant]
US 20200301799A1 · Manivasagam · 2020 [cited by examiner]
US 20200339109A1 · Hong · 2020 [cited by examiner]
US 20210312244A1 · Asa · 2021 [cited by examiner]
US 20220135086A1 · Mahjourian · 2022 [cited by examiner]
US 20220388532A1 · Yangel · 2022 [cited by examiner]
US 20240051575A1 · Lee · 2024 [cited by examiner]
Chen et al., “GeoSim: Realistic Video Simulation via Geometry-Aware Composition for Self-Driving,” Computer Vision Foundation, 2021, p. 7230-7240. [cited by applicant]
Liu et al., “Learning End-to-End Multimodal Sensor Policies for Autonomous Navigation,” 1st Conference on Robot Learning, 2017, 13 pages. [cited by applicant]
Xiao et al., “Multimodal End-to-End Autonomous Driving,” arXiv:1906.03199v2 [cs.CV], Oct. 25, 2020, pp. 1-10. [cited by applicant]
Ly et al., “Learning to Drive by Imitation: An Overview of Deep Behavior Cloning Methods,” IEEE Transactions on Intelligent Vehicles, vol. 6, No. 2, Jun. 2021, pp. 1-3. [cited by applicant]
Wymann et al., “TORCS: The Open Racing Car Simulator,” Mar. 12, 2015, pp. 1-5. [cited by applicant]
Codevilla et al., “End-to-end driving via conditional imitation learning,” in Proceedings of the International Conference on Robotics and Automation (ICRA), 2018, pp. 4693-4700. [cited by applicant]
Kendall et al., “Learning to drive in a day,” arXiv preprint 1807.00412, 2018, pp. 8248-8254. [cited by applicant]
Amini et al., “Variational End-to-End Navigation and Localization,” in Proceedings of the International Conference on Robotics and Automation (ICRA), 2019, pp. 8958-8964. [cited by applicant]
Bojarski et al., “End to end learning for self-driving cars,” arXiv preprint:1604.07316, 2016, pp. 1-9. [cited by applicant]
Xu et al., “End-to-end learning of driving models from large-scale video datasets,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2174-2182. [cited by applicant]
Hubschneider et al., “Calibrating uncertainty models for steering angle estimation,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE, 2019, pp. 1511-1518. [cited by applicant]
Rhinehart et al., “Deep imitative models for flexible inference, planning, and control,” arXiv preprint arXiv:1810.06544, 2018, pp. 1-19. [cited by applicant]
A. Mehta, “Learning end-to-end autonomous driving using guided auxiliary supervision,” in Proceedings of the 11th Indian Conference on Computer Vision, Graphics and Image Processing, 2018, pp. 1-12. [cited by applicant]
Pan et al., “Virtual to Real Reinforcement Learning for Autonomous Driving,” in Proceedings of the British Machine Vision Conference (BMVC), 2017, pp. 1-13. [cited by applicant]
Bewley et al., “Learning to Drive from Simulation without Real World Labels,” in Proceedings of the International Conference on Robotics and Automation (ICRA), 2019, pp. 4818-4824. [cited by applicant]
Amini et al., “Learning Robust Control Policies for End-to-End Autonomous Driving From Data-Driven Simulation,” IEEE Robotics and Automation Letters, vol. 5, No. 2, 2020, pp. 1143-1150. [cited by applicant]
Unknown, “Bullet Real-Time Physics Simulation” http://pybullet.org, last accessed on Sep. 25, 2022, 9 pages. [cited by applicant]
Tassa et al., “dm control: Software and tasks for continuous control,” arXiv preprint: 2006.12983, 2020, pp. 1-34. [cited by applicant]
Todorov et al., “MuJoCo: A physics engine for model-based control,” in Proceedings of the International Conference on Intelligent Robots and Systems (IROS), 2012, pp. 5026-5033. [cited by applicant]
Dosovitskiy et al., “Carla: An open urban driving simulator,” in Proceedings of the Conference on Robot Learning (CoRL), 2017, pp. 1-16. [cited by applicant]
Shah et al., “AirSim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and Service Robotics (FSR), 2017, pp. 1-14. [cited by applicant]
Gan et al., “Threedworld: A platform for interactive multi-modal physical simulation,” arXiv preprint arXiv:2007.04954, 2020, pp. 1-23. [cited by applicant]
Guerra et al., “FlightGoggles: Photorealistic Sensor Simulation for Perception-driven Robotics using Photogrammetry and Virtual Reality,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS… [cited by applicant]
Loquercio et al., “Deep Drone Racing: From Simulation to Reality With Domain Randomization,” Transactions on Robotics, vol. 36, No. 1, 2020, pp. 1-14. [cited by applicant]
Tobin et al., “Domain randomization for transferring deep neural networks from simulation to the real world,” in Intelligent Robots and Systems (IROS), 2017 IEEE/RSJ International Conference on. IEEE, 2017, pp. 23-30. [cited by applicant]
Abu Alhaija et al., “Augmented Reality Meets Computer Vision: Efficient Data Generation for Urban Driving Scenes,” International Journal of Computer Vision, vol. 126, No. 9, 2018, pp. 961-972. [cited by applicant]
Remez et al., “Learning to Segment via Cut-and-Paste,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 37-52. [cited by applicant]
Menapace et al., “Playable Video Generation,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 10061-10070. [cited by applicant]
Li et al., “AADS: Augmented autonomous driving simulation using data-driven algorithms,” Science Robotics, vol. 4, No. 28, 2019, pp. 1-12. [cited by applicant]
Chen et al., “GeoSim: Photorealistic Image Simulation with Geometry-Aware Composition,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. [cited by applicant]
Manivasagam et al., “LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11167-11176. [cited by applicant]
Wang et al., “V2VNet: Vehicle-to-Vehicle Communication for Joint Perception and Prediction,” in Proceedings of the European Conference on Computer Vision (ECCV), 2020, pp. 605-621. [cited by applicant]
Xia et al., “Gibson Env: Real-World Perception for Embodied Agents,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 9068-9079. [cited by applicant]
Savva et al., “Habitat: A Platform for Embodied AI Research,” in Proceedings of the International Conference on Computer Vision (ICCV), 2019, pp. 9339-9347. [cited by applicant]
Deitke et al., “RoboTHOR: An Open Simulation-to-Real Embodied AI Platform,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 3164-3174. [cited by applicant]
Alcala et al., “Autonomous racing using linear parameter varying-model predictive control (LPV-MPC),” Control Engineering Practice, vol. 95, 2020, pp. 1-8. [cited by applicant]
Carrau et al., “Efficient implementation of Randomized MPC for miniature race cars,” in European Control Conference (ECC), 2016, pp. 957-962. [cited by applicant]
Galceran et al., “Multipolicy decision-making for autonomous driving via changepoint-based behavior prediction: Theory and experiment,” Autonomous Robots, vol. 41, No. 6, 2017, pp. 1367-1382. [cited by applicant]
Liniger et al., “Real-time control for autonomous racing based on viability theory,” IEEE Transactions on Control Systems Technology, vol. 27, No. 2, 2019, pp. 464-478. [cited by applicant]
Schwarting et al., “Planning and decision-making for autonomous vehicles,” Annual Review of Control, Robotics, and Autonomous Systems, 2018, pp. 187-210. [cited by applicant]
Bansal et al., “ChauffeurNet: Learning to drive by imitating the best and synthesizing the worst,” in Proceedings of Robotics: Science and Systems (RSS), 2019, pp. 1-20. [cited by applicant]
Lechner et al., “Neural circuit policies enabling auditable autonomy,” Nature Machine Intelligence, vol. 2, No. 10, pp. 642-652, 2020, pp. 642-652. [cited by applicant]
Hawke et al., “Urban Driving with Conditional Imitation Learning,” in Proceedings of the International Conference on Robotics and Automation (ICRA), 2020, pp. 251-257. [cited by applicant]
Zhu et al., “Autonomous Robot Navigation Based on Multi-Camera Perception,” in Proceedings of the International Conference on Intelligent Robots and Systems (IROS), 2020, pp. 5879-5885. [cited by applicant]
Brunnbauer et al., “Model-based versus Model-free Deep Reinforcement Learning for Autonomous Racing Cars,” In Proceedings of the International Conference on Robotics and Automation (ICRA), 2021, pp. 1-8. [cited by applicant]
Sauer et al., “Conditional Affordance Learning for Driving in Urban Environments,” in Proceedings of the Conference on Robot Learning (CoRL), 2018, pp. 237-252. [cited by applicant]
Amini et al., “Variational autoencoder for end-to-end control of autonomous driving with novelty detection and training de-biasing,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IE… [cited by applicant]
Song et al., “Autonomous Overtaking in Gran Turismo Sport Using Curriculum Reinforcement Learning,” in Proceedings of the International Conference on Robotics and Automation (ICRA), 2021, pp. 9403-9409. [cited by applicant]
Schwarting et al., “Deep Latent Competition: Learning to Race Using Visual Control Policies in Latent Space,” in Proceedingss of the Conference on Robot Learning (CoRL), 2020, pp. 1-16. [cited by applicant]
Toromanoff et al., “End-to-End Model-Free Reinforcement Learning for Urban Driving Using Implicit Affordances,” In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7153-7162. [cited by applicant]
Amini et al., “Vista 2.0: An open, data-driven simulator for multimodal sensing and policy learning for autonomous vehicles,” arXiv preprint arXiv:2111.12083, 2021, pp. 2419-2426. [cited by applicant]
Sofiiuk et al., “Foreground-aware semantic representations for image harmonization,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 1620-1629. [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms,” arXiv prepring:1707.06347, 2017, pp. 1-12. [cited by applicant]
Guan et al., “Centralized cooperation for connected and automated vehicles at intersections by proximal policy optimization,” IEEE Transactions on Vehicular Technology, vol. 69, No. 11, 2020, pp. 12597-12608. [cited by applicant]
Bohn et al., “Deep reinforcement learning attitude control of fixed-wing UAVs using proximal policy optimization,” in 2019 International Conference on Unmanned Aircraft Systems (ICUAS). IEEE, 2019, pp. 523-533. [cited by applicant]
Codevilla et al., “On offline evaluation of vision-based driving models,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 236-251. [cited by applicant]
Schwarting et al., “Stochastic dynamic games in belief space,” arxiv preprint: 1909.06963, 2019, pp. 1-16. [cited by applicant]
Williams et al., “Autonomous racing with AutoRally vehicles and differential games,” arXiv preprint: 1707.04540, 2017, pp. 1-8. [cited by applicant]
Liniger et al., “A non-cooperative game approach to autonomous racing,” IEEE Transactions on Control Systems Technology, vol. 28, No. 3, 2020, pp. 884-897. [cited by applicant]
Schwarting et al., “Social behavior for autonomous vehicles,” Proceedings of the National Academy of Sciences, vol. 116, No. 50, 2019, pp. 24972-24978. [cited by applicant]
Chen et al., “Learning by cheating,” in Conference on Robot Learning (CoRL), 2020, pp. 66-75. [cited by applicant]
Liang et al., “CIRL: Controllable Imitative Reinforcement Learning for Vision-based Self-driving,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 584-599. [cited by applicant]
D.A. Pomerleau, “ALVINN: An autonomous land vehicle in a neural network,” in Advances in Neural Information Processing Systems (NeurIPS), 1989, pp. 305-313. [cited by applicant]
Pfeiffer et al., “From perception to decision: A data-driven approach to end-to-end motion planning for autonomous ground robots,” in Proceedings of the International Conference on Robotics and Automation (ICRA), 2017, … [cited by applicant]
Amini et al., “Learning steering bounds for parallel autonomous systems,” in 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 4717-4724. [cited by applicant]
Toromanoff et al., “End to end vehicle lateral control using a single fisheye camera,” in Proceedings of the International Conference on Intelligent Robots and Systems (IROS), 2018, pp. 3613-3619. [cited by applicant]