IP Library Granted Patent US 12,544,913
Granted Patent B2
US 12,544,913 · App. 18/605,452 · Granted Feb 10, 2026

Domain adaptation using simulation to simulation transfer

Inventors: Paul Wohlhart (Sunnyvale, CA); Stephen James (Santa Clara, CA); Mrinal Kalakrishnan (Palo Alto, CA); Konstantinos Bousmalis (London, GB)
Assignee: GDM Holding LLC
B25J9/161B25J9/163B25J9/1671B25J9/1697G05B13/027G06F18/2148G06F18/217G06F18/2431G06N3/045G06N3/08G06T7/50G06V10/764G06V10/776G06V10/82G06V20/10G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,544,913
App. No.
18/605,452
Granted
Feb 10, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a generator neural network to adapt input images.

Claims (33)

1 . A method, comprising:

obtaining a plurality of simulation training inputs, each simulation training input comprising (i) a canonical simulated training image of a canonical simulation of a real-world environment and (ii) a corresponding randomized simulated training image of a randomized simulation of the real-world environment generated by randomizing one or more characteristics of a scene depicted in the canonical simulated training image;

training a generator neural network based on the plurality of simulation training inputs to adapt images of the randomized simulation of the real-world environment into images of the canonical simulation of the real-world environment; and

providing the trained generator neural network.

2 . The method of claim 1 , wherein the corresponding randomized simulated training image of each simulation training input is generated by randomizing one or more characteristics of a scene depicted in the canonical simulated training image.

3 . The method of claim 1 , wherein the trained generator neural network is trained to optimize an objective function that includes one or more terms that encourage adapted images generated by the generator neural network by processing randomized simulated training images to be similar to corresponding canonical simulated training images.

4 . The method of claim 3 , wherein the trained generator neural network is trained jointly with a canonical-randomized discriminator neural network having a plurality of canonical-randomized discriminator parameters,

wherein the canonical-randomized discriminator neural network is configured to process input images in accordance with the canonical-randomized discriminator parameters to classify each input image as either being an adapted image or a canonical simulated image, and

wherein the objective function includes a term that penalizes the generator neural network for generating training adapted images that are accurately classified by the canonical-randomized discriminator neural network.

5 . The method of claim 4 , wherein the trained generator neural network is trained by repeatedly alternating between the following:

performing an iteration of a machine learning training technique that adjusts generator parameters to minimize the term by determining gradients with respect to the generator parameters for a batch of simulation training inputs; and

performing an iteration of a machine learning training technique that adjusts the canonical-randomized discriminator parameters to maximize the term by determining gradients with respect to the canonical-randomized discriminator parameters for a batch of simulation training inputs.

6 . The method of claim 1 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages visual similarity between adapted training images generated by the generator neural network and corresponding randomized simulation images.

7 . The method of claim 1 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages semantic similarity between predicted segmentation masks generated by the generator neural network and corresponding ground truth segmentation masks.

8 . The method of claim 1 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages similarity between predicted depth maps generated by the generator neural network and corresponding ground truth depth maps.

9 . The method of claim 1 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages adapted images generated by the generator neural network by processing real-world training images to appear to be images of the canonical simulation while maintaining semantics of the real-world training images.

10 . The method of claim 1 , wherein the images of the randomized simulation of the real-world environment have randomized lighting characteristics of a scene.

11 . The method of claim 1 , wherein the images of the randomized simulation of the real-world environment have randomized texture characteristics of a scene.

12 . The method of claim 1 , wherein the images of the randomized simulation of the real-world environment have randomized object properties of objects in a scene.

13 . A method, comprising:

receiving an input image of a real-word environment;

processing the input image of the real-world environment using a trained generator neural network to generate an adapted image from the input image, the trained generator neural network having been trained to adapt images of a randomized simulation of the real-world environment into images of a canonical simulation of the real-world environment; and

providing the adapted image.

14 . The method of claim 13 , wherein the trained generator neural network is trained to optimize an objective function that includes one or more terms that encourage adapted images generated by the generator neural network by processing randomized simulated training images to be similar to corresponding canonical simulated training images.

15 . The method of claim 13 , wherein the trained generator neural network is trained to optimize an objective function that includes one or more terms that encourage adapted images generated by the generator neural network by processing randomized simulated training images to be similar to corresponding canonical simulated training images.

16 . The method of claim 13 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages semantic similarity between predicted segmentation masks generated by the generator neural network and corresponding ground truth segmentation masks.

17 . The method of claim 13 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages similarity between predicted depth maps generated by the generator neural network and corresponding ground truth depth maps.

18 . The method of claim 13 , wherein the trained generator neural network is trained using an objective function that includes a term that encourages adapted images generated by the generator neural network by processing real-world training images to appear to be images of the canonical simulation while maintaining semantics of the real-world training images.

19 . The method of claim 13 , further comprising providing the adapted image as input to a control policy for a robotic agent.

20 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

receiving an input image of a real-word environment;

processing the input image of the real-world environment using a trained generator neural network to generate an adapted image from the input image, the trained generator neural network having been trained to adapt images of a randomized simulation of the real-world environment into images of a canonical simulation of the real-world environment; and

providing the adapted image.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071465/0754 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2024
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 068927/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2024
From: WOHLHART, PAUL; JAMES, STEPHEN; KALAKRISHNAN, MRINAL; BOUSMALIS, KONSTANTINOS
To: X DEVELOPMENT LLC
Reel/Frame 066789/0020 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2024
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 066799/0867 →
Priority Claims (1)
GR 20180100527 · Nov 23, 2018 · national
Continuity (3)
Continuation 17656137 · Mar 23, 2022
Continuation 16692509 · Nov 22, 2019
Related Publication 20240238967A1 · Jul 18, 2024
References Cited (71)
US 20200070352A1 · Larzaro-Gredilla · 2020 [cited by examiner]
US 20200122321A1 · Khansari Zadeh · 2020 [cited by examiner]
US 20200174490A1 · Ogale · 2020 [cited by examiner]
Antonova et al., “Reinforcement Learning for Pivoting Task,” https://arxiv.org/abs/1703.00472, Mar. 2017, 7 pages. [cited by applicant]
Bohg et al., “Data-Driven Grasp Synthesis—A Survey,” IEEE Transactions on Robotics, Apr. 2014, pp. 289-309, vol. 30, No. 2. [cited by applicant]
Bousmalis et al., “Domain Separation Networks,” Advances in Neural Information Processing Systems, Dec. 2016, 9 pages. [cited by applicant]
Bousmalis et al., “Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 95-104. [cited by applicant]
Bousmalis et al., “Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping,” 2018 IEEE International Conference on Robotics and Automation (ICRA), May 2018, pp. 4243-4250. [cited by applicant]
Brook et al., “Collaborative Grasp Planning with Multiple Object Representations,” 2011 IEEE International Conference on Robotics and Automation, May 2011, pp. 2851-2858. [cited by applicant]
Caseiro et al., “Beyond the shortest path: Unsupervised Domain Adaptation by Sampling Subspaces Along the Spline Flow,” 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2015, pp. 3846-3854. [cited by applicant]
Chang et al., “ShapeNet: An Information-Rich 3D Model Repository,” https://arxiv.org/abs/1512.03012, Dec. 2015, 11 pages. [cited by applicant]
Chebotar et al., “Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience,” https://arxiv.org/abs/1810.05687v3, Oct. 2018, 9 pages. [cited by applicant]
Chen et al., “Photographic Image Synthesis with Cascaded Refinement Networks,” 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 2017, pp. 1520-1529. [cited by applicant]
Ciocarlie et al., “Towards Reliable Grasping and Manipulation in Household Environments,” Experimental Robotics, 2014, pp. 241-252. [cited by applicant]
Csurka, Gabriela, “Domain Adaptation for Visual Applications: A Comprehensive Survey,” https://arxiv.org/abs/1702.0537 4v2, Mar. 2017, 46 pages. [cited by applicant]
Ganin et al., “Domain-Adversarial Training of Neural Networks,” The Journal of Machine Learning Research, Apr. 2016, 35 pages. [cited by applicant]
Gong et al., “Geodesic Flow Kernel for Unsupervised Domain Adaptation,” 2012 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2012, pp. 2066-2073. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” Advances in Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Gopalan et al., “Domain Adaptation for Object Recognition: An Unsupervised Approach,” 2011 International Conference on Computer Vision, Nov. 2011, pp. 999-1006. [cited by applicant]
Gualtieri et al., “High precision grasp pose detection in dense clutter,” 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 2016, pp. 598-605. [cited by applicant]
Hernandez et al., “Team Delft's Robot Winner of the Amazon Picking Challenge 2016,” Robot World Cup, 2016, pp. 613-624. [cited by applicant]
Hinterstoisser et al., “Multimodal Templates for Real-Time Detection of Texture-less Objects in Heavily Cluttered Scenes,” 2011 International Conference on Computer Vision, Nov. 2011, pp. 858-865. [cited by applicant]
Hoffman et al., “CyCADA: Cycle-Consistent Adversarial Domain Adaptation,” In Intl. Conference on Machine Learning, Jul. 2018, 10 pages. [cited by applicant]
Ioffee et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” https://arxiv.org/abs/1502.03167v3, Mar. 2015, 11 pages. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 5967-5976. [cited by applicant]
James et al., “3D Simulation for Robot Arm Control with Deep Q-Learning,” https://arxiv.org/abs/1609.03759v2, Dec. 2016, 6 pages. [cited by applicant]
James et al., “Task-Embedded Control Networks for Few-Shot Imitation Learning,” https://arxiv.org/abs/1810.03237, Oct. 2018, 13 pages. [cited by applicant]
James et al., “Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task,” https://arxiv.org/abs/1707.02267v2, Oct. 2017, 10 pages. [cited by applicant]
Kalashnikov et al., “QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,” https://arxiv.org/abs/1806.10293v3, Nov. 2018, 23 pages. [cited by applicant]
Kappler et al., “Leveraging Big Data for Grasp Planning,” 2015 IEEE International Conference on Robotics and Automation (ICRA), May 2015, pp. 4304-4311. [cited by applicant]
Kehoe et al., “Cloud-Based Robot Grasping with the Google Object Recognition Engine,” 2013 IEEE International Conference on Robotics and Automation, May 2013, pp. 4263-4270. [cited by applicant]
Kim et al., “Learning to Discover Cross-Domain Relations with Generative Adversarial Networks,” Intl. Conference on Machine Learning, Aug. 2017, 9 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems, 2012, 9 pages. [cited by applicant]
Larsen et al., “Autoencoding beyond pixels using a learned similarity metric,” Intl. Conference on Machine Learning, 2016, 9 pages. [cited by applicant]
Lenz et al., “Deep learning for detecting robotic grasps,” The International Journal of Robotics Research, Mar. 2015, pp. 705-724, vol. 34 No. 4-5. [cited by applicant]
Levine et al., “Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection,” https://arxiv.org/abs/1603.02199v4, last revised Aug. 2016, 12 pages. [cited by applicant]
Lillicrap et al., “Continuous control with deep reinforcement learning,” https://arxiv.org/abs/1509.0297lvl, Sep. 2015, 14 pages. [cited by applicant]
Long et al., “Learning Transferable Features with Deep Adaptation Networks,” https://arxiv.org/abs/1502.0279lvl, Feb. 2015, 10 pages. [cited by applicant]
Mahler et al., “Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics,” https://arxiv.org/abs/l 703.09312v3, last revised Aug. 2017, 12 pages. [cited by applicant]
Matas et al., “Sim-to-Real Reinforcement Learning for Deformable Object Manipulation,” https://arxiv.org/abs/1806.0785lv2, last revised Oct. 2018, 10 pages. [cited by applicant]
Mnih et al., “Asynchronous Methods for Deep Reinforcement Learning,” Intl. Conference on Machine Learning, Jun. 2016, 10 pages. [cited by applicant]
Mordatch et al., “Ensemble-CIO: Full-Body Dynamic Motion Planning that Transfers to Physical Humanoids,” 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 2015, pp. 5307-5314. [cited by applicant]
Morrison et al., “Cartman: The Low-Cost Cartesian Manipulator that Won the Amazon Robotics Challenge,” 2018 IEEE International Conference on Robotics and Automation (ICRA), May 2018, pp. 7757-7764. [cited by applicant]
Patel et al., “Visual Domain Adaptation: A survey of recent advances,” IEEE Signal Processing Magazine, May 2015, pp. 53-69, vol. 32, No. 3. [cited by applicant]
Peng et al., “Sim-to-Real Transfer of Robotic Control with Dynamics Randomization,” https://arxiv.org/abs/l710.06537v3, last revised Mar. 2018, 8 pages. [cited by applicant]
Pinto et al., “Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours,” 2016 IEEE International Conference on Robotics and Automation (ICRA), May 2016, pp. 3406-3413. [cited by applicant]
Prattichizzo et al., “Grasping,” Springer Handbook of Robotics, 2008, pp. 671-700. [cited by applicant]
Rajeswaran et al., “EPOpt: Learning Robust Neural Network Policies Using Model Ensembles,” https://arxiv.org/abs/1610.01283v4, last revised Mar. 2017, 15 pages. [cited by applicant]
Rodriguez et al., “From caging to grasping,” The International Journal of Robotics Research, Apr. 2012, 16 pages. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” Intl. Conference on Medical Image Computing and Computer Assisted Intervention, 2015, pp. 234-241. [cited by applicant]
Rusu et al., “Sim-to-Real Robot Learning from Pixels with Progressive Nets,” Conference on Robot Learning, Oct. 2017, 9 pages. [cited by applicant]
Sadeghi et al., “CAD2RL: Real Single-Image Flight without a Single Real Image,” https://arxiv.org/abs/161 l.0420lv4, last revised Jun. 2017, 12 pages. [cited by applicant]
Sadeghi et al., “Sim2Real Viewpoint Invariant Visual Servoing by Recurrent Control,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, 4691-4699. [cited by applicant]
Saxena et al., “Robotic Grasping of Novel Objects using Vision,” The Intl. Journal of Robotics Research, Feb. 2008, pp. 157-173. [cited by applicant]
Shrivastava et al., “Learning from Simulated and Unsupervised Images through Adversarial Training,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 2242-2251. [cited by applicant]
Shu et al., “A DIR T-T Approach to Unsupervised Domain Adaptation,” https://arxiv.org/abs/1802.08735v2, last revised Mar. 2018, 19 pages. [cited by applicant]
Stein et al., “GeneSIS-RT: Generating Synthetic Images for Training Secondary Real-World Tasks,” 2018 IEEE International Conference on Robotics and Automation (ICRA), May 2018, pp. 7151-7158. [cited by applicant]
Sun et al., “Return of Frustratingly Easy Domain Adaptation,” Association for the Advancement of Artificial Intelligence, Mar. 2016, pp. 2058-2065. [cited by applicant]
Sunderhauf et al., “The limits and potentials of deep learning for robotics,” The Intl. Journal of Robotics Research, Apr. 2018, pp. 405-420, vol. 37, No. 4-5. [cited by applicant]
Taigman et al., “Unsupervised Cross-Domain Image Generation,” Intl. Conference on Learning Representations, retrieved from URL <https://pdfs.semanticscholar.org/b4ee/64022cc3ccdl4c7f9d4935c59bl6456067d3.pdf>, 2017, 6 pa… [cited by applicant]
Ten Pas et al., “Grasp Pose Detection in Point Clouds,” The International Journal of Robotics Research, Oct. 2017, pp. 1455-1473 vol. 36, No. 13-14. [cited by applicant]
Tobin et al., “Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 2017, pp. 23-30. [cited by applicant]
Tzeng et al., “Adapting Deep Visuomotor Representations with Weak Pairwise Constraints,” https://arxiv.org/abs/1511.07lllv5, last revised May 2017, 16 pages. [cited by applicant]
Ulyanov et al., “Instance Normalization: The Missing Ingredient for Fast Stylization,” https://arxiv.org/abs/1607.08022vl, Jul. 2016, 6 pages. [cited by applicant]
Viereck et al., “Learning a visuomotor controller for real world robotic grasping using simulated depth images,” Conference on Robot Learning, Oct. 2017, 10 pages. [cited by applicant]
Yi et al., “DualGAN: Unsupervised Dual Learning for Image-to-Image Translation,” 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 2017, pp. 2868-2876. [cited by applicant]
Yoo et al., “Pixel-Level Domain Transfer,” European Conference on Computer Vision, Mar. 2016, pp. 517-532. [cited by applicant]
Yu et al., “Preparing for the Unknown: Learning a Universal Policy with Online System Identification,” https://arxiv.org/abs/1702.02453v3, last revised May 2017, 10 pages. [cited by applicant]
Zeng et al., “Learning Synergies Between Pushing and Grasping with Self-Supervised Deep Reinforcement Learning,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 2018, pp. 4238-4245. [cited by applicant]
Zhang et al., “YR-Goggles for Robots: Real-to-Sim Domain Adaptation for Visual Control,” IEEE Robotics and Automation Letters, Apr. 2019, pp. 1148-1155. [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks,” https://arxiv.org/abs/l703.10593v3, Nov. 2017, 20 pages. [cited by applicant]