IP Library › Granted Patent US 12,189,349
Granted Patent B2
US 12,189,349 · App. 17/319,442 · Granted Jan 7, 2025

System and method of efficient, continuous, and safe learning using first principles and constraints

Inventors: Lifeng Liu (Boston, MA); Yingxuan Zhu (Plano, TX); Jun Zhang (Plano, TX); Xiaotian Yin (Plano, TX); Jian Li (Plano, TX); Yongxiang Tao (Shanghai, CN); Dayao Liang (Shanghai, CN)
Assignee: Huawei Technologies Co., Ltd.
G05B13/0265B60W50/06G06F16/24578B60W2050/0062B60W60/00B60W60/001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,349
App. No.
17/319,442
Filed
May 13, 2021
Granted
Jan 7, 2025
Kind
B2
Art Unit
2171
USPC
700/47
Abstract

A computer implemented method for self-learning of a control system. The method includes creating an initial knowledge base. The method learns first principles using the knowledge base. The method creates initial control commands derived from the knowledge base. The method generates constraints for the control commands. The method performs constrained reinforcement learning by executing the control commands with the constraints and observing feedback to improve the control commands. The method enriches the knowledge base based on the feedback.

Claims (69)

1. A computer implemented method for self-learning of a control system, the method comprising:

creating a knowledge base comprising data obtained from real-world experiments;

learning first principles using the knowledge base, wherein the first principles are foundational principles that cannot be deduced from other principles, and wherein learning the first principles is based on kinematic parameters of motion of an object, estimated structural parameters of the object, and a relationship between the kinematic parameters and control parameters of the object based on the estimated structural parameters of the object;

creating control commands for controlling the object based on the first principles derived from the knowledge base;

generating constraints for the control commands;

performing constrained reinforcement learning by executing the control commands with the constraints,

utilizing feedback from the constrained reinforcement learning to improve the control commands; and

enriching the knowledge base based on the feedback.

2. The method of claim 1 , wherein the object is a vehicle, and wherein the first principles comprises first principals of vehicle motion.

3. The method of claim 1 , wherein creating the control commands derived from the knowledge base comprises receiving a current state, a target state, and a reference to the knowledge base as input parameters.

4. The method of claim 3 , wherein creating the control commands derived from the knowledge base comprises:

breaking a control task down into separate components according to the current state and the target state;

creating a query for each of the separate components of the control task;

retrieving a corresponding query result from the knowledge base for each query to generate a plurality of query results; and

combining the query results to generate a control command for the control task.

5. The method of claim 4 , wherein the query results are combined according to corresponding weights assigned to the separate components of the control task.

6. The method of claim 1 , wherein the constraints for the control commands include hard constraints that specify conditions that cannot be exceeded.

7. The method of claim 1 , wherein the constraints for the control commands include soft constraints that specify preferable conditions.

8. The method of claim 1 , wherein generating constraints for the control commands comprises:

generating a first subset of constraint items based on a state of an operating environment;

generating a second subset of constraint items for a set of filtered moving objects based on a state of a host machine and a target state;

generating a third subset of constraint items for a set of filtered stationary obstacles; and

combining the first subset of constraint items, the second subset of constraint items, and the third subset of constraint items.

9. The method of claim 1 , wherein performing the constrained reinforcement learning to improve the control commands comprises:

decomposing the control commands into multiple categories and dimensions to enable learning of the control commands in each category and dimension separately;

generating control command candidates for a control command based on a current state;

applying the constraints to the control command candidates;

refining the control commands based on past experiences when the current state is a learned state; and

adapting results from learned states to new environments when the current state is not the learned state.

10. The method of claim 1 , wherein enriching the knowledge base based on the feedback comprises refining a manifestation of dynamics and kinematics models of the knowledge base.

11. A system comprising

a memory storage unit comprising instructions; and

one or more processors in communication with the memory storage unit, wherein the one or more processors execute the instructions to:

create a knowledge base comprising data obtained from real-world experiments;

learn first principles using the knowledge base, wherein the first principles are foundational principles that cannot be deduce from other principles, and wherein learning the first principles is based on kinematic parameters of motion of an object, estimated structural parameters of the object, and a relationship between the kinematic parameters and control parameters of the object based on the estimated structural parameters of the object;

create control commands for controlling the object based on the first principles derived from the knowledge base;

generate constraints for the control commands;

perform constrained reinforcement learning by executing the control commands with the constraints,

utilize feedback from the constrained reinforcement learning to improve the control commands; and

enrich the knowledge base based on the feedback.

12. The system of claim 11 , wherein the object is a vehicle, and wherein the first principles comprises first principals of vehicle motion.

13. The system of claim 11 , wherein the one or more processors further execute the instructions to receive a current state, a target state, and a reference to the knowledge base as input parameters.

14. The system of claim 13 , wherein the one or more processors further execute the instructions to:

break a control task down into separate components according to the current state and the target state;

create a query for each of the separate components of the control task;

retrieve a corresponding query result from the knowledge base for each query to generate a plurality of query results; and

combine the query results to generate a control command for the control task.

15. The system of claim 14 , wherein the one or more processors further execute the instructions to combine the query results according to corresponding weights assigned to the separate components of the control task.

16. The system of claim 11 , wherein the constraints for the control commands include hard constraints that specify conditions that cannot be exceeded.

17. The system of claim 11 , wherein the constraints for the control commands include soft constraints that specify preferable conditions.

18. The system of claim 11 , wherein the one or more processors further execute the instructions to:

generate a first subset of constraint items based on a state of an operating environment;

generate a second subset of constraint items for a set of filtered moving objects based on a state of a host machine and a target state;

generate a third subset of constraint items for a set of filtered stationary obstacles; and

combine the first subset of constraint items, the second subset of constraint items, and the third subset of constraint items.

19. The system of claim 11 , wherein the one or more processors further execute the instructions to:

decompose the control commands into multiple categories and dimensions to enable learning of the control commands in each category and dimension separately;

generate control command candidates for a control command based on a current state;

apply the constraints to the control command candidates;

refine the control commands based on past experiences when the current state is a learned state; and

adapt results from learned states to new environments when the current state is not the learned state.

20. A non-transitory computer readable medium storing computer instructions, that when executed by one or more processors, cause the one or more processors to perform the steps of:

creating a knowledge base comprising data obtained from real-world experiments;

learning first principles using the knowledge base, wherein the first principles are foundational principles that cannot be deduce from other principles, and wherein learning the first principles is based on kinematic parameters of motion of an object, estimated structural parameters of the object, and a relationship between the kinematic parameters and control parameters of the object based on the estimated structural parameters of the object;

creating control commands for controlling the object based on the first principles derived from the knowledge base;

generating constraints for the control commands;

performing constrained reinforcement learning by executing the control commands with the constraints,

utilizing feedback from the constrained reinforcement learning to improve the control commands; and

enriching the knowledge base based on the feedback.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2021
From: FUTUREWEI TECHNOLOGIES, INC.
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 057099/0553 →
Continuity (4)
Continuation PCTCN2019083895 · Apr 23, 2019
Provisional Application 62768467 · Nov 16, 2018
Related Publication 20210341886A1 · Nov 4, 2021
Related Publication 20220155732A9 · May 19, 2022
References Cited (25)
US 10640111B1 · Gutmann · 2020 [cited by examiner]
US 10679497B1 · Konrardy · 2020 [cited by examiner]
US 20140063232A1 · Fairfield · 2014 [cited by examiner]
US 20180074493A1 · Prokhorov · 2018 [cited by examiner]
US 20180189647A1 · Calvo · 2018 [cited by examiner]
US 20190204842A1 · Jafari Tafti · 2019 [cited by examiner]
US 20200033868A1 · Palanisamy · 2020 [cited by examiner]
US 20200346665A1 · Araujo et al. · 2020 [cited by examiner]
CN 106842925A · 2017 [cited by applicant]
CN 106873585A · 2017 [cited by applicant]
CN 107194612A · 2017 [cited by applicant]
CN 107506830A · 2017 [cited by applicant]
WO 2018139993A1 · 2018 [cited by applicant]
Mnih, V., et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, Feb. 26, 2015, 13 pages. [cited by applicant]
Arulkumaran, K., et al., “A Brief Survey of Deep Reinforcement Learning,” IEEE Signal Processing magazine, Special Issue on Deep Learning for Image Understanding, Sep. 28, 2017, 16 pages. [cited by applicant]
Silver, D., et al., “Deterministic Policy Gradient Algorithms,” International Conference on Machine Learning, Beijing, China, 2014, 9 pages. [cited by applicant]
Lillicrap, T., et al., “Continuous Control with Deep Reinforcement Learning,” International Conference on Learning Representations, Published as conference paper at ICLR, 2016, 14 pages. [cited by applicant]
Emami, P., “Deep Deterministic Policy Gradients in TensorFlow,” Updates on my machine learning research, summaries of papers, and blog posts, retrieved from the internet: http://pemami4911.github.io/blog/2016/08/21/ddpg… [cited by applicant]
Levine, S., et al., “Guided Policy Search,” Proceedings of the 30th International Conference on Machine Learning, PMLR 28(3):1-9, Atlanta, Georgia, USA, 2013, 2 pages. [cited by applicant]
Levine, S., et al., “Learning Contact-Rich Manipulation Skills with Guided Policy Search,” IEEE International Conference on Robotics and Automation, 2015, 3 pages. [cited by applicant]
Parisotto, E., et al., “Actor-Mimic Deep Multitask and Transfer Reinforcement Learning,” International Conference on Learning Representations, Published as a conference paper at ICLR, Feb. 22, 2016, 16 pages. [cited by applicant]
Kahn, G., et al., “PLATO: Policy Learning using Adaptive Trajectory Optimization,” IEEE International Conference on Robotics and Automation, Mar. 2, 2016, 13 pages. [cited by applicant]
Schulman, J., et al., “Trust Region Policy Optimization,” International Conference on Machine Learning, Lille, France, 2015, 9 pages. [cited by applicant]
Wang, T., “Trust Region Policy Optimization,” Machine Learning Group, University of Toronto, retrieved from the Internet: http://www.cs.toronto.edu/˜tingwuwang/trpo.pdf, 21 pages. [cited by applicant]
Kurin, V., “Introduction to Imitation Learning,” retrieved from the internet: https://blog.statsbot.co/introduction-to-mitation-learning-32334c3b1e7a, 14 pages. [cited by applicant]