IP Library Granted Patent US 12,429,878
Granted Patent B1
US 12,429,878 · App. 17/340,870 · Granted Sep 30, 2025

Systems and methods for dynamic object removal from three-dimensional data

Inventors: Ricson Cheng (Bridgewater, NJ); Adam Wlodzimierz Harley (Pittsburgh, PA); Justin Liang (Toronto, CA); Xinchen Yan (San Mateo, CA); Raquel Urtasun (Toronto, CA); Mehmet Ersin Yumer (San Francisco, CA)
Assignee: AURORA OPERATIONS, INC.
G05D1/0219G01S17/89G05D1/0221G05D1/0248G06N20/00G06V20/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,429,878
App. No.
17/340,870
Granted
Sep 30, 2025
Kind
B1
Abstract

Systems and methods for generating simulation data based on real-world environments are provided. A method includes obtaining multi-modal sensor data indicative of a dynamic object within an environment of a robotic platform. The multi-modal sensor data is associated with a plurality of timesteps including a first timestep and a second timestep. The method includes providing the multi-modal sensor data indicative of the dynamic object within the environment as an input to a machine-learned dynamic object removal model. And, the method includes receiving as an output of the machine-learned dynamic object removal model, in response to receipt of the multi-modal sensor data, a scene representation indicative of at least a portion of the environment including a reconstructed region based at least in part on removal of the dynamic object and multiple levels of granularity. The scene representation is used as a template for generating different simulations within the depicted environment.

Claims (49)

1. A computing system, comprising:

one or more processors; and

one or more computer-readable medium storing instructions that when executed by the one or more processors cause the computing system to perform operations, the operations comprising:

obtaining multi-modal sensor data indicative of a dynamic object within an environment of an autonomous vehicle, wherein the multi-modal sensor data is associated with a plurality of timesteps including a first timestep and a second timestep;

providing the multi-modal sensor data indicative of the dynamic object within the environment as an input to a machine-learned dynamic object removal model;

generating, using the machine-learned dynamic object removal model and based at least in part on the multi-modal sensor data, an intermediate scene representation comprising a coarse reconstructed region based at least in part on removal of the dynamic object, wherein the coarse reconstructed region comprises inpainted data describing features of the environment that were occluded by the dynamic object; and

generating, using the machine-learned dynamic object removal model and based at least in part on the intermediate scene representation, a scene representation output indicative of at least a portion of the environment comprising a refined reconstructed region based at least in part on removal of the dynamic object.

2. The computing system of claim 1 , wherein the intermediate scene representation comprises inpainted pixel data.

3. The computing system of claim 2 , wherein the intermediate scene representation comprises inpainted depth data.

4. The computing system of claim 2 , wherein the intermediate scene representation comprises inpainted semantic segmentation data.

5. The computing system of claim 1 , wherein the machine-learned object removal model comprises:

a first network configured to generate the intermediate scene representation; and

a second network configured to generate the scene representation output.

6. The computing system of claim 5 , wherein the first network and the second network are trained end-to-end.

7. The computing system of claim 1 , wherein the multi-modal sensor data comprises a plurality of image frames, and wherein the machine-learned object removal model is configured to:

select a subset of image frames from the plurality of image frames based at least in part on a comparison of pixels centered around an area associated with the coarse reconstructed region across the subset of image frames.

8. The computing system of claim 1 , wherein obtaining the multi-modal sensor data comprises:

obtaining, through one or more first sensors and one or more second sensors, sensor data indicative of the dynamic object within the environment, wherein the one or more first sensors are a different type of sensor than the one or more second sensors; and

generating the multi-modal sensor data based at least in part on the sensor data, the multi-modal sensor data comprising a three-dimensional reconstruction of at least a portion of the environment with the dynamic object.

9. The computing system of claim 1 , wherein the refined reconstructed region comprises a static background previously occluded by the dynamic object.

10. The computing system of claim 1 , wherein the refined reconstructed region reduces a shadow associated with the dynamic object.

11. The computing system of claim 1 , wherein the operations further comprise:

generating simulation data based at least in part on the scene representation output.

12. The computing system of claim 11 , wherein the simulation data comprises:

a simulated environment that is based at least in part on the scene representation output; and

one or more simulated dynamic objects designed to move within the simulated environment.

13. The computing system of claim 12 , wherein the operations further comprise:

training a machine-learned model of an autonomous vehicle computing system using the simulation data.

14. An autonomous vehicle, comprising:

a plurality of sensors comprising at least one first sensor and at least one second sensor, the at least one first sensor being a different type of sensor than the at least one second sensor;

one or more processors; and

one or more computer-readable medium storing instructions that when executed by the one or more processors cause the autonomous vehicle to perform operations, the operations comprising:

obtaining, through the at least one first sensor and the at least one second sensor, multi-modal sensor data indicative of a dynamic object within an environment, wherein the multi-modal sensor data is associated with a plurality of timesteps including a first timestep and a second timestep;

providing the multi-modal sensor data indicative of the dynamic object within the environment as an input to a machine-learned dynamic object removal model;

generating, using the machine-learned dynamic object removal model and based at least in part on the multi-modal sensor data, an intermediate scene representation comprising a coarse reconstructed region based at least in part on removal of the dynamic object, wherein the coarse reconstructed region comprises inpainted data describing features of the environment that were occluded by the dynamic object; and

generating, using the machine-learned dynamic object removal model and based at least in part on the intermediate scene representation, a scene representation output indicative of at least a portion of the environment comprising a refined reconstructed region based at least in part on removal of the dynamic object.

15. The autonomous vehicle of claim 14 , wherein the intermediate scene representation comprises inpainted pixel data.

16. The autonomous vehicle of claim 15 , wherein the intermediate scene representation comprises inpainted depth data.

17. The autonomous vehicle of claim 14 , wherein the machine-learned object removal model comprises:

a first network configured to generate the intermediate scene representation; and

a second network configured to generate the scene representation output.

18. The autonomous vehicle of claim 17 , wherein the first network and the second network are trained end-to-end.

19. A computer-implemented method, comprising:

obtaining multi-modal sensor data indicative of a dynamic object within an environment of a robotic platform, wherein the multi-modal sensor data is associated with a plurality of timesteps including a first timestep and a second timestep;

providing the multi-modal sensor data indicative of the dynamic object within the environment as an input to a machine-learned dynamic object removal model;

generating, using the machine-learned dynamic object removal model and based at least in part on the multi-modal sensor data, an intermediate scene representation indicative of at least a portion of the environment comprising a coarse reconstructed region based at least in part on removal of the dynamic object, wherein the coarse reconstructed region comprises inpainted data describing features of the environment that were occluded by the dynamic object; and

generating, using the machine-learned dynamic object removal model and based at least in part on the intermediate scene representation, a scene representation output indicative of at least a portion of the environment comprising a refined reconstructed region based at least in part on removal of the dynamic object.

20. The computer-implemented method of claim 19 , wherein the robotic platform comprises an autonomous vehicle, and the environment is a surrounding environment of the autonomous vehicle, and wherein the method further comprises:

generating simulation data based at least in part on the first scene representation, wherein the simulation data comprises a simulated environment for simulating autonomous vehicle operation, the simulated environment comprising a scene that is based at least in part on the first scene representation and comprises one or more simulated dynamic objects.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2022
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 058962/0140 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: LIANG, JUSTIN; CHENG, RICSON; YAN, XINCHEN; HARLEY, ADAM WLODZIMIERZ; YUMER, ERSIN
To: UATC, LLC
Reel/Frame 058795/0888 →
EMPLOYMENT AGREEMENT Recorded Jan 24, 2022
From: SOTIL, RAQUEL URTASUN
To: UBER TECHNOLOGIES, INC.
Reel/Frame 058826/0936 →
Continuity (1)
Provisional Application 63035577 · Jun 5, 2020
References Cited (60)
US 20200218961A1 · Kanazawa · 2020 [cited by examiner]
US 20210261148A1 · Li · 2021 [cited by examiner]
Arjovsky et al., “Wasserstein GAN”, arXiv:1701.07875v3, Dec. 6, 2017, 32 pages. [cited by applicant]
Barnes et al., “PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing”, ACM Transactions on Graphics, 2009, vol. 28, 11 pages. [cited by applicant]
Behley et al., “SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences”, arXiv:1904.01416v3, Aug. 16, 2019, 17 pages. [cited by applicant]
Bousmalis et al., “Using Simulation and Domain Adaption to Improve Efficiency of Deep Robotic Grasping”, IEEE International Conference on Robotics and Automation, May 21-25, 2018, Brisbane, Australia, pp. 4243-4250. [cited by applicant]
Casser et al., “Depth Prediction without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos”, AAAI Conference on Artificial Intelligence, 2019, vol. 33, pp. 8001-8008. [cited by applicant]
Chang et al., “ShapeNet: An Information-Rich Model Repository”, arXiv:1512.03012v1, Dec. 9, 2015, 11 pages. [cited by applicant]
Delage et al., “A dynamic Bayesian network model for autonomous 3d reconstruction from a single indoor image”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 17-22, 2006, New York, NY,… [cited by applicant]
Denton et al., “Semi-Supervised Learning with Context-Conditional Generative Adversarial Networks”, arXiv:1611.06430v1, Nov. 19, 2016, 10 pages. [cited by applicant]
Dosovitskiy et al., “CARLA: An Open Urban Driving Simulator”, Conference on Robot Learning, Nov. 13-15, 2017, Mountain View, CA, 16 pages. [cited by applicant]
Efros et al., “Texture Synthesis by Non-parametric Sampling”, IEEE International Conference on Computer Vision, Sep. 20-27, 1999, Kerkyra, Greece, 6 pages. [cited by applicant]
Eslami et al., “Neural scene representation and rendering”, Science, Jun. 15, 2018, vol. 360, 8 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets”, Conference on Neural Information Processing Systems, Dec. 8-13, 2014, Montreal, Canada, 9 pages. [cited by applicant]
Gulrajani et al., “Improved Training of Wasserstein GANs”, Conference on Neural Information Processing Systems, Dec. 4-9, 2017, Long Beach, CA, 11 pages. [cited by applicant]
Harley et al., “Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D Mapping”, International Conference on Learning Representations, Apr. 26-May 1, 2020, Virtual, 19 pages. [cited by applicant]
Hinton et al., “Transforming Auto-Encoders”, International Conference on Artificial Neural Networks, Jun. 14-17, 2011, Espoo, Finland, pp. 44-51. [cited by applicant]
Hong et al., “Learning Hierarchical Semantic Image Manipulation through Structured Representations”, Conference on Neural Information Processing Systems, Dec. 2-8, 2018, Montreal, Canada, 11 pages. [cited by applicant]
Iizuka et al., “Globally and Locally Consistent Image Completion”, ACM Transactions on Graphics, 2017, vol. 36, No. 4, pp. 107:1-107:14. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu, Hawaii, pp. 1125-1134. [cited by applicant]
Kar et al., “Meta-Sim: Learning to Generate Synthetic Datasets”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 4551-4560. [cited by applicant]
Kim et al., “Deep Video Inpainting”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 5792-5801. [cited by applicant]
Kolve et al., “AI2-THOR: An Interactive 3D Environment for Visual AI”, arXiv:1712.05474v3, Mar. 15, 2019, 4 pages. [cited by applicant]
Lee et al., “Context-Aware Synthesis and Placement of Object Instances”, 32 [cited by applicant]
Lee et al., “Copy-and-Paste Networks for Deep Video Inpainting”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 4413-4421. [cited by applicant]
Lee et al., “Geometric Reasoning for Single Image Structure Recovery, IEEE Computer Society Conference on Computer Vision and Pattern Recognition”, Jun. 20-25, 2009, Miami Beach, FL, pp. 2136-2143. [cited by applicant]
Lee et al., “Inserting Videos into Videos”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, CA, pp. 10061-10070. [cited by applicant]
Li et al., “Generative Face Completion”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu, HI, pp. 3911-3919. [cited by applicant]
Liu et al., “Image Inpainting for Irregular Holes Using Partial Convolutions”, European Conference on Computer Vision. Sep. 8-14, 2018, Munich, Germany, 16 pages. [cited by applicant]
Malik et al., “The three R's of computer vision: Recognition, reconstruction and reorganization”, Pattern Recognition Letters 72, 2016, pp. 4-14. [cited by applicant]
Newson et al., “Towards fast, generic video inpainting, CVMP2013: The 10 [cited by applicant]
Oh et al., “Onion-Peel Networks for Deep Video Completion”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 4403-4412. [cited by applicant]
Park et al., “Probabilistic Surfel Fusion for Dense LiDAR Mapping”, International Conference on Computer Vision. Oct. 22-29, 2017, Venice Italy, pp. 2418-2426. [cited by applicant]
Pathak et al., “Context Encoders: Feature Learning and Inpainting”, Conference on Computer Vision and Pattern Recognition, Jun. 26-Jul. 1, 2016, Las Vegas, NV, pp. 2536-2544. [cited by applicant]
Ros et al., “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 27-30, 2016… [cited by applicant]
Sangkloy et al., “Scribbler: Controlling Deep Image Synthesis with Sketch and Color”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu HI, pp. 5400-5409. [cited by applicant]
Savva et al., “Habitat: A Platform for Embodied AI Research”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 9339-9347. [cited by applicant]
Savva et al., “MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments”, arXiv:1712.03931v1, Dec. 11, 2017, 14 pages. [cited by applicant]
Shechtman et al., “Space-Time Super-Resolution”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 27, No. 4, Apr. 2005, pp. 531-545. [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, arXiv:1409.1556v6, Apr. 10, 2015, 14 pages. [cited by applicant]
Sitzmann et al., “DeepVoxels: Learning Persistent 3D Feature Embeddings”, Conference on Computer Vision and Pattern Recognition, Jun. 16-Jun. 20, 2019, Long Beach, CA, pp. 2437-2446. [cited by applicant]
Song et al., “Im2Pano3D: Extrapolating 360 Structure and Semantics Beyond the Field of View”, Conference on Computer Vision and Patter Recognition, Jun. 18-Jun. 22, 2018, Salt Lake City, UT, pp. 3847-3856. [cited by applicant]
Thonat et al., “Multi-View Inpainting for Image-Based Scene Editing and Rendering”, 2016 Fourth International Conference on 3D Vision, Oct. 25-28, 2016, Stanford, CA, pp. 351-359. [cited by applicant]
Tran et al., “A Closer Look at Spatiotemporal Convolutions for Action Recognition”, Conference on Computer Vision and Patter Recognition, Jun. 18-Jun. 22, 2018, Salt Lake City, UT, pp. 6450-6459. [cited by applicant]
Ulyanov et al., “Deep Image Prior”, Conference on Computer Vision and Patter Recognition, Jun. 18-Jun. 22, 2018, Salt Lake City, UT, pp. 9446-9454. [cited by applicant]
Wang et al., “Image Quality Assessment: From Error Visibility to Structural Similarity”, IEEE Transactions on Image Processing, vol. 13, No. 4, Apr. 2004, pp. 600-612. [cited by applicant]
Wang et al., “Video-to-Video Synthesis”, Conference on Neural Information Processing Systems, Dec. 2-8, 2018, Montreal, Canada, 13 pages. [cited by applicant]
Wexler et al., “Space-Time Completion of Video”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 29, No. 3, Mar. 2007, pp. 463-476. [cited by applicant]
Xia et al., “Gibson Env: Real-World Perception for Embodied Agents”, Conference on Computer Vision and Patter Recognition, Jun. 18-Jun. 22, 2018, Salt Lake City, UT, pp. 9068-9079. [cited by applicant]
Xian et al., “TextureGAN: Controlling Deep Image Synthesis with Texture Patches”, Conference on Computer Vision and Patter Recognition, Jun. 18-Jun. 22, 2018, Salt Lake City, UT, pp. 8456-8465. [cited by applicant]
XU eta l., “Deep Flow-Guided Video Inpainting”, Conference on Computer Vision and Pattern Recognition, Jun. 16-Jun. 20, 2019, Long Beach, CA, pp. 3723-3732. [cited by applicant]
Yan et al., “Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Rpresentations”, 2018 IEEE International Conference on Robotics and Automation, May 21-25, 2018, Brisbane, Australia, pp. 3766-3773. [cited by applicant]
Yang et al., “High-Resolution Image Inpainting using Multi-Scale Neural Patch Synthesis”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu HI, pp. 6721-6729. [cited by applicant]
Yao et al., “3D-Aware Scene Manipulation via Inverse Graphics”, Conference on Neural Information Processing Systems, Dec. 2-8, 2018, Montreal, Canada, 12 pages. [cited by applicant]
Yeh et al., “Semantic Image Inpainting with Deep Generative Models”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu HI, pp. 5485-5493. [cited by applicant]
Yu et al., “Free-Form Image Inpainting with Gated Convolution”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 4471-4480. [cited by applicant]
Yu et al., “Generative Image Inpainting with Contextual Attention”, Conference on Computer Vision and Patter Recognition, Jun. 18-Jun. 22, 2018, Salt Lake City, UT, pp. 5505-5514. [cited by applicant]
Zhang et al., “An Internal Learning Approach to Video Inpainting”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 2720-2729. [cited by applicant]
Zhang et al., “AutoRemover: Automatic Object Removal for Autonomous Driving Videos”, arXiv:1911.12588v1, Nov. 28, 2019, 9 pages. [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu HI, pp. 2223-2232. [cited by applicant]
Cited By (1)
US 12,651,318