IP Library › Granted Patent US 12,576,515
Granted Patent B2
US 12,576,515 · App. 18/018,421 · Granted Mar 17, 2026

Off-line learning for robot control using a reward prediction model

Inventors: Konrad Zolna (London, GB); Scott Ellison Reed (Atlanta, GA)
Assignee: GDM Holding LLC
B25J9/161B25J9/163G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,576,515
App. No.
18/018,421
Granted
Mar 17, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for off-line learning using a reward prediction model. One of the methods includes obtaining robot experience data; training, on a first subset of the robot experience data, a reward prediction model that receives a reward input comprising an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation; processing experiences in the robot experience data using the trained reward prediction model to generate a respective reward prediction for each of the processed experiences; and training a policy neural network on (i) the processed experiences and (ii) the respective reward predictions for the processed experiences.

Claims (49)

1 . A method comprising:

obtaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise (i) an observation characterizing a state of the environment and (ii) an action performed by a respective robot in response to the observation, wherein the experiences comprise:

expert experiences from episodes of a particular task being performed by an expert agent, and

unlabeled experiences;

training, on a first subset of the robot experience data, a reward prediction model that receives a reward input comprising an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation, wherein training the reward prediction model comprises optimizing an objective function that:

includes a first term that encourages the reward prediction model to assign, to observations from expert experiences, a first reward value that indicates that the particular task was completed successfully after the environment was in the state characterized by the observation, and

includes a second term that encourages the reward prediction model to assign, to observations from unlabeled experiences, a second reward value that indicates that the particular task was not completed successfully after the environment was in the state characterized by the observation;

processing experiences in the robot experience data using the trained reward prediction model to generate a respective reward prediction for each of the processed experiences; and

training a policy neural network on (i) the processed experiences and (ii) the respective reward predictions for the processed experiences, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task.

2 . The method of claim 1 , further comprising:

controlling a robot using the trained policy neural network while the robot performs the particular task.

3 . The method of claim 1 , further comprising:

providing data specifying the trained policy neural network for use in controlling a robot while the robot performs the particular task.

4 . The method of claim 1 , wherein the first subset includes the expert experiences and a proper subset of the unlabeled experiences.

5 . The method of claim 1 , wherein the objective function includes a third term that encourages the reward prediction model to assign, to observations from expert experiences, the second reward value.

6 . The method of claim 5 , wherein the first and second terms have a different sign from the third term in the objective function.

7 . The method of claim 1 , wherein the objective function includes a fourth term that penalizes the reward prediction model for correctly distinguishing expert experiences from unlabeled experiences based on a first predetermined number of observations of an episode of the particular task being performed by an expert agent.

8 . The method of claim 1 , wherein training the policy neural network comprises training the policy neural network on (i) the experiences and (ii) the respective reward predictions for the experiences using an off-line reinforcement learning technique.

9 . The method of claim 8 , wherein the off-line reinforcement learning technique is Critic-Regularized Regression (CRR).

10 . The method of claim 1 , wherein the off-line reinforcement learning technique is an off-line actor-critic technique.

11 . The method of claim 1 , wherein training the reward prediction model comprises applying data augmentation to the experiences in the robot experience data that are used for the training of the reward prediction model.

12 . The method of claim 1 , wherein at least some of experiences of the first subset of the robotic experience data relate to a real-world environment.

13 . The method of any claim 1 , further comprising:

controlling a robot using the trained policy neural network while the robot performs the particular task, wherein controlling the robot comprises obtaining observations from one or more sensors sensing a real-world environment, providing the observations to the trained policy neural network, and using an output of the trained policy neural network to select actions to control the robot to perform the particular task.

14 . A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise (i) an observation characterizing a state of the environment and (ii) an action performed by a respective robot in response to the observation, wherein the experiences comprise:

expert experiences from episodes of a particular task being performed by an expert agent, and

unlabeled experiences;

training, on a first subset of the robot experience data, a reward prediction model that receives a reward input comprising an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation, wherein training the reward prediction model comprises optimizing an objective function that:

includes a first term that encourages the reward prediction model to assign, to observations from expert experiences, a first reward value that indicates that the particular task was completed successfully after the environment was in the state characterized by the observation, and

includes a second term that encourages the reward prediction model to assign, to observations from unlabeled experiences, a second reward value that indicates that the particular task was not completed successfully after the environment was in the state characterized by the observation;

processing experiences in the robot experience data using the trained reward prediction model to generate a respective reward prediction for each of the processed experiences; and

training a policy neural network on (i) the processed experiences and (ii) the respective reward predictions for the processed experiences, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task.

15 . The system of claim 14 , the operations further comprising:

controlling a robot using the trained policy neural network while the robot performs the particular task.

16 . The system of claim 14 , the operations further comprising:

providing data specifying the trained policy neural network for use in controlling a robot while the robot performs the particular task.

17 . The system of claim 14 , wherein the first subset includes the expert experiences and a proper subset of the unlabeled experiences.

18 . The system of claim 14 , wherein the objective function includes a third term that encourages the reward prediction model to assign, to observations from expert experiences, the second reward value.

19 . The system of claim 18 , wherein the first and second terms have a different sign from the third term in the objective function.

20 . One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

obtaining robot experience data characterizing robot interactions with an environment, the robot experience data comprising a plurality of experiences that each comprise (i) an observation characterizing a state of the environment and (ii) an action performed by a respective robot in response to the observation, wherein the experiences comprise:

expert experiences from episodes of a particular task being performed by an expert agent, and

unlabeled experiences;

training, on a first subset of the robot experience data, a reward prediction model that receives a reward input comprising an input observation and generates as output a reward prediction that is a prediction of a task-specific reward for the particular task that should be assigned to the input observation, wherein training the reward prediction model comprises optimizing an objective function that:

includes a first term that encourages the reward prediction model to assign, to observations from expert experiences, a first reward value that indicates that the particular task was completed successfully after the environment was in the state characterized by the observation, and

includes a second term that encourages the reward prediction model to assign, to observations from unlabeled experiences, a second reward value that indicates that the particular task was not completed successfully after the environment was in the state characterized by the observation;

processing experiences in the robot experience data using the trained reward prediction model to generate a respective reward prediction for each of the processed experiences; and

training a policy neural network on (i) the processed experiences and (ii) the respective reward predictions for the processed experiences, wherein the policy neural network is configured to receive a network input comprising an observation and to generate a policy output that defines a control policy for a robot performing the particular task.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2023
From: ZOLNA, KONRAD; REED, SCOTT ELLISON
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 063932/0129 →
Continuity (2)
Provisional Application 63057850 · Jul 28, 2020
Related Publication 20230256593A1 · Aug 17, 2023
References Cited (60)
US 8908044B2 · Luo · 2014 [cited by examiner]
US 9147338B2 · Hunter · 2015 [cited by examiner]
US 11712799B2 · Cabi · 2023 [cited by examiner]
US 12061481B2 · Toshev · 2024 [cited by examiner]
US 20130155252A1 · Luo · 2013 [cited by examiner]
US 20170302951A1 · Joshi · 2017 [cited by examiner]
US 20200104680A1 · Reed et al. · 2020 [cited by applicant]
CN 110119844A · 2019 [cited by applicant]
Moridian et al., Learning Navigation Tasks from Demonstration for Semi-Autonomous Remote Operation of Mobile Robots, 2018, IEEE, p. 1-8 (Year: 2018). [cited by examiner]
Oudeyer et al., Intrinsic Motivation Systems for Autonomous Mental Development, 2007, IEEE, p. 265-279 (Year: 2007). [cited by examiner]
Pfeiffer et al., Reinforced Imitation: Sample Efficient Deep Reinforcement Learning for Mapless Navigation by Leveraging Prior Demonstrations, 2018, IEEE, p. 4423-4430 (Year: 2018). [cited by examiner]
Khan et al., Learning Sample-Efficient Target Reaching for Mobile Robots, 2018, IEEE, p. 3080-303087 (Year: 2018). [cited by examiner]
Hafez et al., Efficient Intrinsically Motivated Robotic Grasping with Learning-Adaptive Imagination in Latent Space, 2019, IEEE, p. 240-246 (Year: 2019). [cited by examiner]
Buehler et al., Online inference of human belief for cooperative robots, 2018, IEEE, p. 409-415 (Year: 2018). [cited by examiner]
Office Action in Chinese Appln. No. 202180049339.9, mailed on May 6, 2025, 26 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 21751553.5, mailed on Apr. 23, 2025, 8 pages. [cited by applicant]
Weiming et al., “Semi-supervised dynamic soft measurement modeling method based on recurrent neural networks,” Journal of Electronic Measurement and Instrumentation, Nov. 15, 2019, 33(11):7-13 (with machine translation). [cited by applicant]
Decision to Grant Patent in Japanese Appln. No. 2023-506025, dated Aug. 5, 2024, 5 pages (with English translation). [cited by applicant]
Abbeel et al., “Apprenticeship leaming via inverse reinforcement learning,” ICML '04: Proceedings of the twenty-first international conference on Machine learning, Jul. 4, 2004, 8 pages. [cited by applicant]
Baram et al., “End-to-end differentiable adversarial imitation learning,” Proceedings of the 34th International Conference on Machine Learning, 2017, 70:390-399. [cited by applicant]
Barth-Maron et al., “Distributed distributional deterministic policy gradients,” CoRR, Apr. 23, 2018, arxiv.org/abs/1804.08617, 16 pages. [cited by applicant]
Cabi et al., “Scaling data-driven robotics with reward sketching and batch reinforcement learning,” CoRR, Sep. 26, 2019, arxiv.org/abs/1909.12200, 11 pages. [cited by applicant]
Chen et al., “BAIL: Best-action imitation learning for batch deep reinforcement learning,” 34th Conference on Neural Information Processing Systems, 2020, 11 pages. [cited by applicant]
Elkan et al., “Leaming classifiers from only positive and unlabeled data,” KDD '08: Proceedings of the 14th ACM SIGKDD intemational conference on Knowledge discovery and data mining, Aug. 2008, pp. 213-220. [cited by applicant]
Finn et al., “Guided cost learning: Deep inverse optimal control via policy optimization,” Proceedings of The 33rd International Conference on Machine Leaming, 2016, 48:49-58. [cited by applicant]
Fu et al., “D4RL: Datasets for Deep Data-Driven Reinforcement Learning,” CoRR, Jun. 27, 2020, arXiv:2004.07219v3, 15 pages. [cited by applicant]
Fu et al., “Learning robust rewards with adversarial inverse reinforcement learning,” CoRR, Oct. 30, 2017, arxiv.org/abs/1710.11248, 15 pages. [cited by applicant]
Fujimoto et al., “Off-policy deep reinforcement learning without exploration,” Proceedings of the 36th International Conference on Machine Learning, 2019, 97:2052-2062. [cited by applicant]
Gülçehre et al., “RL unplugged: Benchmarks for offline reinforcement learning,” CoRR, Jun. 24, 2020, arxiv.org/abs/2006.13888, 12 pages. [cited by applicant]
Ho et al., “Generative adversarial imitation learning,” Advances in Neural Information Processing Systems 29 (NIPS 2016), 2016, 9 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/EP2021/071079, dated Feb. 9, 2023, 11 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2021/071079, dated Nov. 5, 2021, 18 pages. [cited by applicant]
Lange et al., “Batch reinforcement learning,” Reinforcement learning, Mar. 5, 2012, 12:45-72. [cited by applicant]
Levine et al., “Offline reinforcement learning: Tutorial, review, and perspectives on open problems,” CoRR, May 4, 2020, arXiv:2005.01643, 43 pages. [cited by applicant]
Li et al., “InfoGAIL .; Interpretable imitation learning from visual demon strations,” Advances in Neural Information Processing Systems 30 (NIPS 2017), 2017, 11 pages. [cited by applicant]
Merel et al., “Learning human behaviors from motion capture by adversarial imitation,” CoRR, Jul. 7, 2017, arXiv:1707.02201, 12 pages. [cited by applicant]
Nair et al., “Overcoming exploration in reinforcement learning with demonstrations,” 2018 IEEE International Conference on Robotics and Automation (ICRA), May 21-25, 2018, 8 pages. [cited by applicant]
Ng et al., “Algorithms for inverse reinforcement learning,” ICML, Jun. 29, 2000, 1:2. [cited by applicant]
Office Action in Japanese Appln. No. 2023-506025, dated Mar. 4, 2024, 6 pages (with English translation). [cited by applicant]
Osa et al., “An algorithmic perspective on imitation learning,” Foundations and Trends in Robotics, Mar. 26, 2018, 7(1-2):1-79. [cited by applicant]
Paine et al., “Hyperparameter selection for offline reinforcement learning,” CoRR, Jul. 17, 2020, arXiv:2007.09055, 19 pages. [cited by applicant]
Paine et al., “Making efficient use of demonstrations to solve hard exploration problems,” CoRR, Sep. 3, 2019, arxiv.org/abs/1909.01387, 22 pages. [cited by applicant]
Peng et al., “Advantage-weighted regression: Simple and scalable off-policy reinforcement learning,” CoRR, Oct. 1, 2019, arXiv:1910.00177, 19 pages. [cited by applicant]
Pohlen et al., “Observe and look further: Achieving consistent performance on Atari,” CoRR, May 29, 2018, arXiv:1805.11593, 19 pages. [cited by applicant]
Pomerleau, “ALVINN: An autonomous land vehicle in a neural network,” Advances in Neural Information Processing Systems 1 (NIPS 1988), 1988, pp. 305-313. [cited by applicant]
Rajeswaran et al., “Learning complex dexterous manipulation with deep reinforcement leaming and demonstrations,” CoRR, Sep. 28, 2017, arxiv.org/abs/1709.10087, 9 pages. [cited by applicant]
Reddy et al., “SQIL: imitation learning via reinforcement learning with sparse rewards,” CoRR, May 27, 2019, arxiv.org/abs/1905.11108, 14 pages. [cited by applicant]
Ross et al., “A reduction of imitation learning and structured prediction to No. regret online leaming,” Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 2011, 15:627-635. [cited by applicant]
Siegel et al., “Keep doing what worked: Behavior modelling priors for offline reinforcement learning,” CoRR, Feb. 19, 2020, arxiv.org/abs/2002.08396, 21 pages. [cited by applicant]
Vecerik et al., “A practical approach to insertion with variable socket position using deep reinforcement learning,” 2019 International Conference on Robotics and Automation, May 20-24, 2019, 7 pages. [cited by applicant]
Vecerik et al., “Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards,” CoRR, Jul. 27, 2017, arxiv.org/abs/1707.08817, 10 pages. [cited by applicant]
Wang et al., “Critic regularized regression,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020, 11 pages. [cited by applicant]
Wang et al., “Exponentially weighted imitation learning for batched historical data,” 32nd Conference on Neural Information Processing Systems, 2018, 10 pages. [cited by applicant]
Xu et al., “Positive-unlabeled reward learning, ” CoRR, Nov. 1, 2019, arxiv.org/abs/1911.00459, 18 pages. [cited by applicant]
Zhu et al., “Reinforcement and imitation learning for diverse visuomotor skills,” CoRR, Feb. 26, 2018, arxiv.org/abs/1802.09564, 12 pages. [cited by applicant]
Zolna et al., “Combating false negatives in adversarial imitation learning, ” CoRR, Feb. 2, 2020, arxiv.org/abs/2002.00412, 9 pages. [cited by applicant]
Zolna et al., “Offline Learning from Demonstrations and Unlabeled Experience,” CoRR, Nov. 27, 2020, arXiv:2011.13885v1, 13 pages. [cited by applicant]
Zolna et al., “Reinforced imitation in heterogeneous action space,” CoRR, Apr. 6, 2019, arXiv:1904.03438, 17 pages. [cited by applicant]
Zolna et al., “Task-relevant adversarial imitation learning,” CoRR, Oct. 2, 2019, arXiv:1910.01077, 17 pages. [cited by applicant]
Office Action in Korean Appln. No. 10-2023-7002829, mailed on Oct. 24, 2025, 18 pages (with English translation). [cited by applicant]
Cited By (1)
US 12,709,003