IP Library Granted Patent US 12,430,564
Granted Patent B2
US 12,430,564 · App. 17/684,245 · Granted Sep 30, 2025

Fine-tuning policies to facilitate chaining

Inventors: Yuke Zhu (Austin, TX); Anima Anandkumar (Pasadena, CA); Youngwoon Lee (Los Angeles, CA)
Assignee: NVIDIA CORPORATION
G06N3/092G05B19/41865G05B19/41885G05B19/41895
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,564
App. No.
17/684,245
Granted
Sep 30, 2025
Kind
B2
Abstract

A manipulation task may include operations performed by one or more manipulation entities on one or more objects. This manipulation task may be broken down into a plurality of sequential sub-tasks (policies). These policies may be fine-tuned so that a terminal state distribution of a given policy matches an initial state distribution of another policy that immediately follows the given policy within the plurality of policies. The fine-tuned plurality of policies may then be chained together and implemented within a manipulation environment.

Claims (29)

1. A method comprising, at a device:

determining an initial state distribution of a second state-action policy, the initial state distribution including possible states of an environment immediately before the second state-action policy is implemented;

fine-tuning a first state-action policy to match a terminal state distribution of the first state-action policy to the initial state distribution of the second state-action policy, the terminal state distribution including possible states of the environment immediately after the first state-action policy is implemented;

implementing the fine-tuned first state-action policy and the second state-action policy in sequence, wherein a terminal state of the environment resulting from implementation of the first state-action policy is provided as an initial state of the environment to the second state-action policy.

2. The method of claim 1 , wherein the first state-action policy and the second state-action policy each describes one or more manipulations performed by one or more manipulation entities on one or more objects.

3. The method of claim 2 , wherein the one or more manipulation entities include one or more robotic manipulation devices.

4. The method of claim 2 , wherein the one or more manipulation entities include one or more vehicle manipulation devices.

5. The method of claim 2 , wherein the one or more objects include one or more components of a product being assembled.

6. The method of claim 2 , wherein the one or more objects include one or more components of a vehicle being controlled.

7. The method of claim 2 , wherein the initial state distribution of the second state-action policy identifies all possible states of the one or more manipulation entities and the one or more objects being manipulated immediately before the second state-action policy is implemented.

8. The method of claim 2 , wherein the terminal state distribution of the first state-action policy identifies all possible states of the one or more manipulation entities and the one or more objects being manipulated immediately after the first state-action policy is implemented.

9. The method of claim 1 , wherein the first state-action policy is adjusted so that the terminal state distribution of the first policy is within a predetermined threshold of the initial state distribution of the second state-action policy.

10. The method of claim 1 , wherein the fine-tuning is performed within a simulation of the environment.

11. The method of claim 1 , wherein the fine-tuning is performed within the environment.

12. The method of claim 1 , comprising implementing the chained policies within the environment.

13. A system comprising:

a hardware processor of a device that is configured to:

determine an initial state distribution of a second state-action policy, the initial state distribution including possible states of an environment immediately before the second state-action policy is implemented;

fine-tune a first state-action policy to match a terminal state distribution of the first state-action policy to the initial state distribution of the second state-action policy, the terminal state distribution including possible states of the environment immediately after the first state-action policy is implemented;

implement the fine-tuned first state-action policy and the second state-action policy in sequence, wherein a terminal state of the environment resulting from implementation of the first state-action policy is provided as an initial state of the environment to the second state-action policy.

14. The system of claim 13 , wherein the first state-action policy and the second state-action policy each describes one or more manipulations performed by one or more manipulation entities on one or more objects.

15. The system of claim 14 , wherein the one or more manipulation entities include one or more robotic manipulation devices.

16. The system of claim 14 , wherein the one or more manipulation entities include one or more vehicle manipulation devices.

17. The system of claim 14 , wherein the one or more objects include one or more components of a product being assembled.

18. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor of a device, causes the processor to cause the device to:

determine an initial state distribution of a second state-action policy, the initial state distribution including possible states of an environment immediately before the second state-action policy is implemented;

fine-tune a first state-action policy to match a terminal state distribution of the first state-action policy to the initial state distribution of the second state-action policy, the terminal state distribution including possible states of the environment immediately after the first state-action policy is implemented;

implement the fine-tuned first state-action policy and the second state-action policy in sequence, wherein a terminal state of the environment resulting from implementation of the first state-action policy is provided as an initial state of the environment to the second state-action policy.

19. The non-transitory computer-readable medium of claim 18 , wherein the first state-action policy is adjusted so that the terminal state distribution of the first policy is within a predetermined threshold of the initial state distribution of the second state-action policy.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2022
From: ZHU, YUKE; ANANDKUMAR, ANIMA; LEE, YOUNGWOON
To: NVIDIA CORPORATION
Reel/Frame 059785/0750 →
Continuity (1)
Related Publication 20230280726A1 · Sep 7, 2023
References Cited (64)
US 11157010B1 · Narang · 2021 [cited by examiner]
US 11657266B2 · Yang · 2023 [cited by examiner]
US 12154029B2 · Schaul · 2024 [cited by examiner]
US 20180165603A1 · Van Seijen · 2018 [cited by examiner]
US 20180218267A1 · Halim · 2018 [cited by examiner]
US 20200334565A1 · Tresp · 2020 [cited by examiner]
US 20210341904A1 · Woehlke · 2021 [cited by examiner]
US 20220245503A1 · Li · 2022 [cited by examiner]
W. Han, S. Levine, and P. Abbeel. “Learning compound multi-step controllers under unknown dynamics”. In: 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE. 2015, pp. 6435-6442 (Year: … [cited by examiner]
Zhu et al., “Reinforcement and Imitation Learning for Diverse Visuomotor Skills,” Robotics: Science and Systems, 2018, 10 pages, retrieved http://www.roboticsproceedings.org/rss14/p09.pdf. [cited by applicant]
Peng et al., “AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control,” ACM Transactions on Graphics, vol. 40, No. 4, Aug. 2021, pp. 1:1-1:20. [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms,” arXiv, 2017, 12 pages, retrieved from https://arxiv.org/pdf/1707.06347.pdf. [cited by applicant]
Pertsch et al., “Accelerating Reinforcement Learning with Learned Skill Priors,” 4th Conference on Robot Learning (CoRL), 2020, pp. 1-17. [cited by applicant]
Sutton, R., “Temporal Credit Assignment in Reinforcement Learning,” PhD thesis, University of Massachusetts, Feb. 1984, 223 pages. [cited by applicant]
Levine et al., “End-to-End Training of Deep Visuomotor Policies,” Journal of Machine Learning Research, vol. 17, 2016, pp. 1-40. [cited by applicant]
Suarez-Ruiz et al., “A Framework for Fine Robotic Assembly,” IEEE International Conference on Robotics and Automation (ICRA), 2016, 8 pages, retrieved from https://www.semanticscholar.org/paper/A-framework-for-fine-robo… [cited by applicant]
Rajeswaran et al., “Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations,” Robotics: Science and Systems (RSS), 2018, 9 pages, retrieved from https://arxiv.org/abs/1709.10087. [cited by applicant]
Jain et al., “Learning Deep Visuomotor Policies for Dexterous Hand Manipulation,” International Conference on Robotics and Automation (ICRA), 2019, 7 pages, retrieved from https://www.semanticscholar.org/paper/Learning-… [cited by applicant]
Andrychowicz et al., “What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale Study,” ICLR, 2021, pp. 1-10. [cited by applicant]
Clegg et al., “Learning to Dress: Synthesizing Human Dressing Motion via Deep Reinforcement Learning,” ACM Transactions on Graphics, vol. 37, No. 6, Nov. 2018, pp. 179:1-179:10. [cited by applicant]
Lee et al., “Composing Complex Skills by Learning Transition Policies,” International Conference on Learning Representations, 2019, pp. 1-19. [cited by applicant]
Peng et al., “MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies,” 33rd Conference on Neural Information Processing Systems (NeurIPS), 2019, pp. 1-12. [cited by applicant]
Lee et al., “Learning to Coordinate Manipulation Skills via Skill Behavior Diversification,” ICLR, 2020, pp. 1-15. [cited by applicant]
Ghosh et al., “Divide-and-Conquer Reinforcement Learning,” ICLR, 2018, pp. 1-14. [cited by applicant]
Konidaris et al., “Skill Discovery in Continuous Reinforcement Learning Domains using Skill Chaining,” Proceedings of the 22nd International Conference on Neural Information Processing Systems, Dec. 2009, 9 pages, retri… [cited by applicant]
Bagaria et al., “Option discovery using deep skill chaining,” ICLR, 2020, pp. 1-21. [cited by applicant]
Pastor et al., “Learning and Generalization of Motor Skills by Learning from Demonstration,”—IEEE International Conference on Robotics and Automation, May 2009, 7 pages. [cited by applicant]
Kober et al., “Movement Templates for Learning of Hitting and Batting,” Proceedings—IEEE International Conference on Robotics and Automation, May 2010, 7 pages. [cited by applicant]
Muelling et al., “Learning to Select and Generalize Striking Movements in Robot Table Tennis,” AAAI Technical Report FS-12-07, Robots Learning Interactively from Human Teachers, 8 pages. [cited by applicant]
Hausmann et al., “Learning an Embedding Space for Transferable Robot Skills,” ICLR, 2018, pp. 1-16. [cited by applicant]
Lee et al., “Ikea Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks,” IEEE International Conference on Robotics and Automation, 2021, 7 pages, retrieved from https://clvrai.com/assets/research/l… [cited by applicant]
Lillicrap et al., “Continuous Control with Deep Reinforcement Learning,” ICLR, 2016, pp. 1-14. [cited by applicant]
Schulman et al., “Trust Region Policy Optimization,” Proceedings of the 31st International Conference on Machine Learning, 2015, 9 pages. [cited by applicant]
Haarnoja et al., “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,” Proceedings of the 35 th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Riedmiller et al., “Learning by Playing—Solving Sparse Reward Tasks from Scratch,” arXiv, 2018, 18 pages, retrieved from https://arxiv.org/pdf/1802.10567.pdf. [cited by applicant]
Pomerleau, D., “Alvinn: An Autonomous Land Vehicle in a Neural Network,” Advances in Neural Information Processing Systems, 1988, pp. 305-313, retrieved from https://proceedings.neurips.cc/paper/1988/file/812b4ba287f5ee… [cited by applicant]
Schaal et al., “Learning From Demonstration,” Advances in Neural Information Processing Systems, 1996, pp. 1040-1046. [cited by applicant]
Finn et al., “Deep Spatial Autoencoders for Visuomotor Learning,” IEEE International Conference on Robotics and Automation, 2016, 9 pages, retrieved https://arxiv.org/abs/1509.06113. [cited by applicant]
Pathak et al., “Zero-Shot Visual Imitation,” International Conference on Learning Representations, 2018, pp. 1-16. [cited by applicant]
Ng et al., “Algorithms for Inverse Reinforcement Learning,” International Conference on Machine Learning, 2000, pp. 1-8. [cited by applicant]
Abbeel et al., “Apprenticeship Learning via Inverse Reinforcement Learning,” Proceedings of the 21 st International Conference on Machine Learning, 2004, 8 pages. [cited by applicant]
Ziebart et al., “Maximum Entropy Inverse Reinforcement Learning,” AAAI Conference on Artificial Intelligence, 2008, 6 pages, retrieved from http://www.cs.cmu.edu/˜bziebart/publications/maximum-entropy-inverse-reinforcem… [cited by applicant]
Ho et al., “Generative Adversarial Imitation Learning,” 30th Conference on Neural Information Processing Systems, 2016, pp. 1-9. [cited by applicant]
Fu et al., “Learning robust rewards with adversarial inverse reinforcement learning,” International Conference on Learning Representations, 2018, pp. 1-15. [cited by applicant]
Kostrikov et al., “Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning,” International Conference on Learning Representations, 2019, pp. 1-15. [cited by applicant]
Zolna et al., “Task-Relevant Adversarial Imitation Learning,” 4th Conference on Robot Learning, 2020, pp. 1-17. [cited by applicant]
Sutton et al., “Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning,” Artificial Intelligence, vol. 112, 1999, pp. 181-211. [cited by applicant]
Schmidhuber, J., “Towards Compositional Learning in Dynamic Networks,” TUM, Institut fur Informatik, 1990, 11 pages. [cited by applicant]
Bacon et al., “The Option-Critic Architecture,” Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, 2017, pp. 1726-1734. [cited by applicant]
Nachum et al., “Data-Efficient Hierarchical Reinforcement Learning,” 32nd Conference on Neural Information Processing Systems, 2018, pp. 1-11. [cited by applicant]
Levy et al., “Learning Multi-Level Hierarchies with Hindsight,” International Conference on Learning Representations, 2019, pp. 1-16. [cited by applicant]
Konidaris et al., “Robot Learning from Demonstration by Constructing Skill Trees,” The International Journal of Robotics Research, May 2012, pp. 1-33. [cited by applicant]
Kipf et al., “CompILE: Compositional Imitation Learning and Execution,” Proceedings of the 36 th International Conference on Machine Learning, 2019, pp. 1-11. [cited by applicant]
Lu et al., “Learning Task Decomposition with Ordered Memory Policy Network,” International Conference on Learning Representations, 2021, pp. 1-21. [cited by applicant]
Frans et al., “Meta Learning Shared Hierarchies,” International Conference on Learning Representations, 2018, pp. 1-11. [cited by applicant]
Kulkarni et al., “Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation,” 30th Conference on Neural Information Processing Systems, 2016, pp. 1-9. [cited by applicant]
Oh et al., “Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning,” Proceedings of the 34 th International Conference on Machine Learning, 2017, 10 pages. [cited by applicant]
Merel et al., “Hierarchical Visuomotor Control of Humanoids,” International Conference on Learning Representations, 2019, pp. 1-19. [cited by applicant]
Andreas et al., “Modular Multitask Reinforcement Learning with Policy Sketches,” International Conference on Machine Learning, 2017, pp. 1-12. [cited by applicant]
Kase et al., “Transferable Task Execution from Pixels through Deep Planning Domain Learning,” IEEE International Conference on Robotics and Automation, 2020, 7 pages, retrieved from https://www.semanticscholar.org/paper… [cited by applicant]
Driess et al., “Learning Geometric Reasoning and Control for Long-Horizon Tasks from Visual Input,” IEEE Iternational Conference on Robotics and Automation, May 2021, 9 pages. [cited by applicant]
Niekum et al., “Incremental Semantically Grounded Learning from Demonstration,” Robotics: Science and Systems, 2013, 8 pages, retrieved from http://www.roboticsproceedings.org/rss09/p48.pdf. [cited by applicant]
Knepper et al., “IkeaBot: An Autonomous Multi-Robot Coordinated Furniture Assembly System,” IEEE International Conference on Robotics and Automation, May 2013, pp. 855-862. [cited by applicant]
Suarez-Ruiz et al., “Can robots assemble an Ikea chair?” Science Robotics, Apr. 2018, pp. 1-8. [cited by applicant]