IP Library Granted Patent US 12,246,449
Granted Patent B2
US 12,246,449 · App. 17/447,553 · Granted Mar 11, 2025

Device and method for controlling a robotic device

Inventor: Fabian Otto (Tuebingen, DE)
Assignee: ROBERT BOSCH GMBH
B25J9/163B25J9/161G05B13/027G05B2219/39001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,246,449
App. No.
17/447,553
Granted
Mar 11, 2025
Kind
B2
Abstract

A device and a method for controlling a robotic device are described. The method includes: carrying out a sequence of actions by the robotic device using a robot control model; ascertaining an updated policy using the carried-out sequence of actions; projecting the updated policy onto a projected policy in such a way that for each state of the plurality of states of the projected policy: a similarity value according to a similarity metric between the projected policy and the updated policy is maximized, and a similarity value according to the similarity metric between the projected policy and the initial policy is greater than a predefined threshold value; adapting the robot control model for implementing the projected policy; and controlling the robotic device, using the adapted robot control model.

Claims (67)

1. A method for controlling a robotic device, comprising:

carrying out a sequence of actions by the robotic device using a robot control model, the carrying out of each action of the sequence of actions including:

ascertaining an action for a present state of a plurality of states of the robotic device with the aid of the robot control model, using an initial policy,

carrying out the ascertained action by the robotic device, and

ascertaining the state of the robotic device resulting from the carried-out action;

ascertaining an updated policy using the carried-out sequence of actions;

projecting the updated policy onto a projected policy in such a way that for each state of a plurality of states of the projected policy:

a similarity value according to a similarity metric between the projected policy and the updated policy is maximized, and

the similarity value according to the similarity metric between the projected policy and the initial policy is greater than a predefined threshold value;

adapting the robot control model for implementing the projected policy; and

controlling the robotic device, using the adapted robot control model, wherein the projected policy is ascertained by a numerical optimizer that ascertains a first optimized Lagrange multiplier and a second optimized Lagrange multiplier for a canonical parameter and a cumulant-generating function of a multivariate normal distribution.

2. The method as recited in claim 1 , wherein the projection of the updated policy onto the projected policy includes:

projecting the updated policy onto the projected policy in such a way that for each state of the plurality of states of the projected policy:

the similarity value according to the similarity metric between the projected policy and the updated policy is maximized,

the similarity value according to the similarity metric between the projected policy and the initial policy is greater than the predefined threshold value, and

an entropy of the projected policy is greater than or equal to a predefined entropy threshold value.

3. The method as recited in claim 1 , wherein:

the initial policy includes an initial multivariate normal distribution of the plurality of actions;

the updated policy includes an updated multivariate normal distribution of the plurality of actions;

the projected policy includes a projected multivariate normal distribution of the plurality of actions;

the projection of the updated policy onto the projected policy includes:

projecting the updated policy onto the projected policy in such a way that for each state of the plurality of states of the projected policy:

a similarity value according to the similarity metric between the projected multivariate normal distribution and the updated multivariate normal distribution is maximized, and

the similarity value according to the similarity metric between the projected multivariate normal distribution and the initial multivariate normal distribution is greater than the predefined threshold value.

4. The method as recited in claim 3 , wherein the projection of the updated policy onto the projected policy includes:

projecting the updated policy onto the projected policy in such a way that for each state of the plurality of states of the projected policy:

the similarity value according to the similarity metric between the projected multivariate normal distribution and the updated multivariate normal distribution is maximized, and

the similarity value according to the similarity metric between the projected multivariate normal distribution and the initial multivariate normal distribution is greater than the predefined threshold value; and

an entropy of the projected multivariate normal distribution is greater than or equal to a predefined entropy threshold value.

5. The method as recited in claim 3 , wherein the projection of the updated policy onto the projected policy in such a way that for each state of the plurality of states of the projected policy, the similarity value according to the similarity metric between the projected multivariate normal distribution and the updated multivariate normal distribution is maximized, and the similarity value according to the similarity metric between the projected multivariate normal distribution and the initial multivariate normal distribution is greater than the predefined threshold value, includes:

ascertaining the projected multivariate normal distribution using the initial multivariate normal distribution, the updated multivariate normal distribution, and the predefined threshold value with the aid of the Mahalanobis distance and the Frobenius norm.

6. The method as recited in claim 5 , wherein the ascertainment of the projected multivariate normal distribution includes a Lagrange multiplier method.

7. The method as recited in claim 3 , wherein the projection of the updated policy onto the projected policy in such a way that for each state of the plurality of states of the projected policy, the similarity value according to the similarity metric between the projected multivariate normal distribution and the updated multivariate normal distribution is maximized, and the similarity value according to the similarity metric between the projected multivariate normal distribution and the initial multivariate normal distribution is greater than the predefined threshold value, includes:

ascertaining the projected multivariate normal distribution using the initial multivariate normal distribution, the updated multivariate normal distribution, and the predefined threshold value using a Wasserstein distance.

8. The method as recited in claim 3 , wherein the projection of the updated policy onto the projected policy in such a way that for each state of the plurality of states of the projected policy, the similarity value according to the similarity metric between the projected multivariate normal distribution and the updated multivariate normal distribution is maximized, and the similarity value according to the similarity metric between the projected multivariate normal distribution and the initial multivariate normal distribution is greater than the predefined threshold value, includes:

ascertaining the projected multivariate normal distribution using the initial multivariate normal distribution, the updated multivariate normal distribution, and the predefined threshold value using the numerical optimizer.

9. The method as recited in claim 8 , wherein the numerical optimizer ascertains the projected multivariate normal distribution, using a Kullback-Leibler divergence.

10. The method as recited in claim 1 , wherein the robot control model is a neural network, and the projection of the updated policy onto the projected policy is implemented as one or multiple layers in the neural network.

11. The method as recited in claim 1 , wherein the adaptation of the robot control model for implementing the projected policy includes an adaptation of the robot control model using a gradient method.

12. The method as recited in claim 1 , wherein the control of the robotic device using the adapted robot control model includes:

carrying out one or multiple actions by the robotic device, using the adapted robot control model; and

updating the policy using a regression, using the carried-out one or multiple actions.

13. The method as recited in claim 1 , wherein the control of the robotic device using the adapted robot control model includes:

carrying out one or multiple actions by the robotic device, using the adapted robot control model;

updating the policy, using the carried-out one or multiple actions, in such a way that a difference between an expected reward and a similarity value according to the similarity metric between the projected policy and the updated policy is maximized.

14. A device configured to control a robotic device, the device configured to:

carry out a sequence of actions by the robotic device using a robot control model, the carrying out of each action of the sequence of actions including:

ascertaining an action for a present state of a plurality of states of the robotic device with the aid of the robot control model, using an initial policy,

carrying out the ascertained action by the robotic device, and

ascertaining the state of the robotic device resulting from the carried-out action;

ascertain an updated policy using the carried-out sequence of actions;

project the updated policy onto a projected policy in such a way that for each state of a plurality of states of the projected policy:

a similarity value according to a similarity metric between the projected policy and the updated policy is maximized, and

the similarity value according to the similarity metric between the projected policy and the initial policy is greater than a predefined threshold value;

adapt the robot control model for implementing the projected policy; and

control the robotic device, using the adapted robot control model, wherein the projected policy is ascertained by a numerical optimizer that ascertains a first optimized Lagrange multiplier and a second optimized Lagrange multiplier for a canonical parameter and a cumulant-generating function of a multivariate normal distribution.

15. A non-transitory nonvolatile memory medium that stores program instructions for controlling a robotic device, the program instructions, when executed by a computer, causing the computer to perform the following steps:

carrying out a sequence of actions by the robotic device using a robot control model, the carrying out of each action of the sequence of actions including:

ascertaining an action for a present state of a plurality of states of the robotic device with the aid of the robot control model, using an initial policy,

carrying out the ascertained action by the robotic device, and

ascertaining the state of the robotic device resulting from the carried-out action;

ascertaining an updated policy using the carried-out sequence of actions;

projecting the updated policy onto a projected policy in such a way that for each state of a plurality of states of the projected policy:

a similarity value according to a similarity metric between the projected policy and the updated policy is maximized, and

the similarity value according to the similarity metric between the projected policy and the initial policy is greater than a predefined threshold value;

adapting the robot control model for implementing the projected policy; and

controlling the robotic device, using the adapted robot control model, wherein the projected policy is ascertained by a numerical optimizer that ascertains a first optimized Lagrange multiplier and a second optimized Lagrange multiplier for a canonical parameter and a cumulant-generating function of a multivariate normal distribution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: OTTO, FABIAN
To: ROBERT BOSCH GMBH
Reel/Frame 059267/0058 →
Priority Claims (1)
DE 102020211648.2 · Sep 17, 2020 · national
Continuity (1)
Related Publication 20220080586A1 · Mar 17, 2022
References Cited (21)
US 5319738A · Shima et al. · 1994 [cited by applicant]
US 8019713B2 · Gupta et al. · 2011 [cited by applicant]
US 9707680B1 · Jules et al. · 2017 [cited by applicant]
US 10751879B2 · Li et al. · 2020 [cited by applicant]
US 10786900B1 · Bohez · 2020 [cited by examiner]
US 20060184491A1 · Gupta et al. · 2006 [cited by applicant]
US 20150217449A1 · Meier · 2015 [cited by examiner]
US 20200130177A1 · Kolouri et al. · 2020 [cited by applicant]
DE 4440859A1 · 1996 [cited by applicant]
DE 112010000775T5 · 2013 [cited by applicant]
DE 102018201949A1 · 2019 [cited by applicant]
DE 102019131385A1 · 2020 [cited by applicant]
Yang, Tsung-Yen, et al. “Projection-Based Constrained Policy Optimization.” International Conference on Learning Representations. Published 2019. [cited by examiner]
Abdullah, Mohammed Amin, et al. “Wasserstein robust reinforcement learning.” arXiv preprint arXiv:1907.13196. Published 2019. [cited by examiner]
Law, Marc T., et al. “Closed-form training of mahalanobis distance for supervised clustering.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Published 2016. [cited by examiner]
Blind Review Yang et al., “Projection-Based Constrained Policy Optimization,” Accessed Jun. 6, 2024. [cited by examiner]
OpenReview page detailing the original publication date of Yang et al., accessed Jun. 6, 2024. [cited by examiner]
Schulman et al., “Trust Region Policy Optimization,” Proceedings of the 31st International Conference on Machine Learning, JMLR: W&Cp, vol. 37, 2015, pp. 1-9. <http://proceedings.mlr.press/v37/schulman15.pdf> Downloaded… [cited by applicant]
Abdolmaleki et al., “Model-Based Relative Entropy Stochastic Search,” Advances in Neural Information Processing Systems, 2015, pp. 1-9. <https://papers.nips.cc/paper/2015/file/36ac8e558ac7690b6f44e2cb5ef93322-Paper.pdf>… [cited by applicant]
Akrour et al., “Projections for Approximate Policy Iteration Algorithms,” Proceedings of the 36th International Conference on Machine Learning, PMLR, vol. 97, 2019, pp. 1-10. <http://proceedings.mlr.press/v97/akrour19a/… [cited by applicant]
Amos et al., “Optnet: Differentiable Optimization as a Layer in Neural Networks,” Proceedings of the 34th International Conference on Machine Learning, PMLR, vol. 70, 2017, pp. 1-10. <https://dl.acm.org/doi/pdf/10.5555/… [cited by applicant]