IP Library › Granted Patent US 12,246,450
Granted Patent B2
US 12,246,450 · App. 17/652,983 · Granted Mar 11, 2025

Device and method to improve learning of a policy for robots

Inventors: Felix Berkenkamp (Munich, DE); Lukas Froehlich (Freiburg, DE); Maksym Lefarov (Stuttgart, DE); Andreas Doerr (Stuttgart, DE)
Assignee: Robert Bosch GmbH
B25J9/163G06F17/17
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,246,450
App. No.
17/652,983
Granted
Mar 11, 2025
Kind
B2
Abstract

A computer-implemented method for for learning a policy. The method includes: recording at least an episode of interactions of the agent with its environment following policy and adding the recorded episode to a set of training data; optimizing a transition dynamics model based on the training data such that the transition dynamics model predicts the next states of the environment depending on the states and actions contained in the training data; optimizing policy parameters based on the training data and the transition dynamics model by optimizing a reward. In the method, the transition dynamics model comprises a first model characterizing the global model and a second model characterizing a correction model, which is configured to correct outputs of the first model.

Claims (72)

1. A computer-implemented method for operating an agent depending on a learned policy obtained by:

initializing a policy and a transition dynamics model which predicts a next state of an environment and/or of the agent; and

repeating the following steps until a termination condition is fulfilled:

recording at least an episode of interactions of the agent with the environment following the policy and adding the recorded episode to a set of training data,

optimizing the transition dynamics model based on the training data such that the transition dynamics model predicts next states of the environment depending on states and actions contained in the training data, and

optimizing policy parameters of the policy based on the training data and the transition dynamics model by optimizing a reward over at least one episode by following the policy;

wherein the transition dynamics model includes a first model representing a learned model of the environment and a correction model which is configured to correct errors of the first model;

wherein the method includes:

sensing the environment using a sensor of the agent;

determining a current state depending on the sensed environment;

determining, using the learning policy, an action for the agent depending on the current state;

carrying out the determined action, by the agent, wherein the correction model is optimized such that, when actions are selected as done when recording the episodes of the training data, then a sequence of states predicted by the transition dynamics model will be equal to the recorded states of the training data.

2. The method according to claim 1 , wherein the correction model is selected by minimizing a difference between an output of the correction model and a difference between the recorded state of the training data and the predicted state by the first model.

3. A computer-implemented method for operating an agent depending on a learned policy obtained by:

initializing a policy and a transition dynamics model which predicts a next state of an environment and/or of the agent; and

repeating the following steps until a termination condition is fulfilled:

recording at least an episode of interactions of the agent with the environment following the policy and adding the recorded episode to a set of training data,

optimizing the transition dynamics model based on the training data such that the transition dynamics model predicts next states of the environment depending on states and actions contained in the training data, and

optimizing policy parameters of the policy based on the training data and the transition dynamics model by optimizing a reward over at least one episode by following the policy;

wherein the transition dynamics model includes a first model representing a learned model of the environment and a correction model which is configured to correct errors of the first model;

wherein the method includes:

sensing the environment using a sensor of the agent;

determining a current state depending on the sensed environment;

determining, using the learning policy, an action for the agent depending on the current state;

carrying out the determined action, by the agent, wherein the correction model is dependent on a state or a time, wherein the time characterizes a time span elapsed since a beginning of a respective episode.

4. The method according to claim 3 , wherein the environment is deterministic and the correction model is dependent on the time.

5. A computer-implemented method for operating an agent depending on a learned policy obtained by:

initializing a policy and a transition dynamics model which predicts a next state of an environment and/or of the agent; and

repeating the following steps until a termination condition is fulfilled:

recording at least an episode of interactions of the agent with the environment following the policy and adding the recorded episode to a set of training data,

optimizing the transition dynamics model based on the training data such that the transition dynamics model predicts next states of the environment depending on states and actions contained in the training data, and

optimizing policy parameters of the policy based on the training data and the transition dynamics model by optimizing a reward over at least one episode by following the policy;

wherein the transition dynamics model includes a first model representing a learned model of the environment and a correction model which is configured to correct errors of the first model;

wherein the method includes:

sensing the environment using a sensor of the agent;

determining a current state depending on the sensed environment;

determining, using the learning policy, an action for the agent depending on the current state;

carrying out the determined action, by the agent, wherein the correction model is a probabilistic function, and wherein the probabilistic function is optimized by approximate inference.

6. A computer-implemented method for operating an agent depending on a learned policy obtained by:

initializing a policy and a transition dynamics model which predicts a next state of an environment and/or of the agent; and

repeating the following steps until a termination condition is fulfilled:

recording at least an episode of interactions of the agent with the environment following the policy and adding the recorded episode to a set of training data,

optimizing the transition dynamics model based on the training data such that the transition dynamics model predicts next states of the environment depending on states and actions contained in the training data, and

optimizing policy parameters of the policy based on the training data and the transition dynamics model by optimizing a reward over at least one episode by following the policy;

wherein the transition dynamics model includes a first model representing a learned model of the environment and a correction model which is configured to correct errors of the first model;

wherein the method includes:

sensing the environment using a sensor of the agent;

determining a current state depending on the sensed environment;

determining, using the learning policy, an action for the agent depending on the current state;

carrying out the determined action, by the agent, wherein the correction model is optimized jointly with the first model.

7. The method according to claim 6 , wherein the agent is an at least partially autonomous robot and/or a manufacturing machine and/or an access control system.

8. The method according to claim 6 , carrying out the determined action, by the agent, wherein for optimizing the transition dynamics model, after optimizing the first model on the training data, the correction model is selected such that error of the first model is minimized for actions selected from the policy on the training data.

9. A machine-readable storage medium on which is stored a computer program for learning a policy for an agent, the computer program, when executed by a computer, causing the computer to perform the following steps:

initializing a policy and a transition dynamics model which predicts a next state of an environment and/or of the agent; and

repeating the following steps until a termination condition is fulfilled:

recording at least an episode of interactions of the agent with the environment following the policy and adding the recorded episode to a set of training data,

optimizing the transition dynamics model based on the training data such that the transition dynamics model predicts next states of the environment depending on states and actions contained in the training data, and

optimizing policy parameters of the policy based on the training data and the transition dynamics model by optimizing a reward over at least one episode by following the policy;

wherein the transition dynamics model includes a first model representing a learned model of the environment and a correction model which is configured to correct errors of the first model, wherein the computer program, when executed by the computer, causing the computer to perform the following further steps:

sensing the environment using a sensor of the agent;

determining a current state depending on the sensed environment;

determining, using the learning policy, an action for the agent depending on the current state;

carrying out the determined action, by the agent wherein the correction model is optimized jointly with the first model.

10. A control system for operating an actuator, control system comprising:

a policy trained by:

initializing a policy and a transition dynamics model which predicts a next state of an environment and/or of the agent; and

repeating the following steps until a termination condition is fulfilled:

recording at least an episode of interactions of the agent with the environment following the policy and adding the recorded episode to a set of training data,

optimizing the transition dynamics model based on the training data such that the transition dynamics model predicts next states of the environment depending on states and actions contained in the training data, and

optimizing policy parameters of the policy based on the training data and the transition dynamics model by optimizing a reward over at least one episode by following the policy;

wherein the transition dynamics model includes a first model representing a learned model of the environment and a correction model which is configured to correct errors of the first model;

wherein the control system is configured to operate the actuator in accordance with an output of the policy, wherein the correction model is optimized jointly with the first model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2022
From: BERKENKAMP, FELIX; FROEHLICH, LUKAS; LEFAROV, MAKSYM; DOERR, ANDREAS
To: ROBERT BOSCH GMBH
Reel/Frame 060744/0496 →
Priority Claims (1)
EP 21162920 · Mar 16, 2021 · regional
Continuity (1)
Related Publication 20220297290A1 · Sep 22, 2022
References Cited (14)
US 10558929B2 · Alam · 2020 [cited by examiner]
US 20160258782A1 · Sadjadi · 2016 [cited by examiner]
US 20190258254A1 · Kadin · 2019 [cited by examiner]
US 20210178600A1 · Jha · 2021 [cited by examiner]
US 20220261630A1 · Fulton · 2022 [cited by examiner]
Janner et al., “When to Trust Your Model: Model-Based Policy Optimization,” 33rd Conference on Neural Information Processing Systems (NEURIPS2019), vol. 32, 2019, pp. 1-12. <https://dl.acm.org/doi/pdf/10.5555/3454287.34… [cited by applicant]
Doerr et al., “Optimizing Long-Term Predictions for Model-Based Policy Search,” 1st Conference on Robot Learning (CORL 2017), vol. 78, 2017, pp. 1-12. <http://proceedings.mlr.press/v78/doerr17a/doerr17a.pdf> Downloaded … [cited by applicant]
Bristow et al., “A Survey of Iterative Learning Control,” IEEE Control Systems Magazine, vol. 26, No. 3, 2006, pp. 96-114. <https://www.researchgate.net/publication/3207727_A_survey_of_iterative_learning> Downloaded Feb… [cited by applicant]
Haarnoja et al., “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning With a Stochastic Actor,” Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 201… [cited by applicant]
Heess et al., “Learning Continuous Control Policies By Stochastic Value Gradients,” Advances in Neural Information Processing Systems, vol. 28, 2015, pp. 1-9. <https://proceedings.neurips.cc/paper/2015/file/148510031349… [cited by applicant]
Lee et al., “Bayesian Residual Policy Optimization: Scalable Bayesian Reinforcement Learning With Clairvoyant Experts,” Cornell University Library, 2020, pp. 1-12. [cited by applicant]
Silver et al., “Residual Policy Learning,” Cornell University Library, 2019, pp. 1-12. [cited by applicant]
Johannink et al., “Residual Reinforcement Learning for Robot Control,” 2019 International Conference on Robotics and Automation (ICRA), Palais Des Congres De Montreal, Montreal, Canada, 2019, pp. 6023-6029. [cited by applicant]
Ma et al., “Efficient Insertion Control for Precision Assembly Based on Demonstration Learning and Reinforcement Learning,” IEEE Transactions on Industrial Informatics, vol. 17, No. 7, 2021, pp. 4492-4502. [cited by applicant]