IP Library Granted Patent US 12,208,521
Granted Patent B1
US 12,208,521 · App. 17/580,552 · Granted Jan 28, 2025

System and method for robot learning from human demonstrations with formal logic

Inventors: Aniruddh Gopinath Puranic (Los Angeles, CA); Jyotirmoy Deshmukh (Los Angeles, CA); Stefanos Nikolaidis (Los Angeles, CA)
Assignee: UNIVERSITY OF SOUTHERN CALIFORNIA
B25J9/163B25J9/0081B25J9/161B25J9/1658
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,208,521
App. No.
17/580,552
Granted
Jan 28, 2025
Kind
B1
Abstract

Systems and methods are provided for learning control policies that are safe, robust, and interpretable in developing robotic systems. In one example, a signal temporal logic is used to evaluate and rank the quality of demonstrations. The temporal logic based specifications allow generation of non-Markovian rewards, and also define causal dependencies between tasks such as sequential task specifications.

Claims (47)

1. A system for reinforcement learning, the system comprising:

one or more processors;

a computer-readable medium storing executable instructions that, when executed, cause the system to perform operations comprising:

receive a set of demonstrations, the set of demonstrations obtained via interaction of an agent with an environment determined according to sensor input from one or more sensors of the agent;

receive a set of specifications, the set of specifications providing descriptions of one or more tasks and/or one or more objectives;

convert the set of specifications into a signal temporal logic formal language in real-time;

evaluate the set of demonstrations based on the set of specifications in the signal temporal logic formal language;

generate a robustness value for each demonstration in the set of demonstrations based on the evaluation;

infer rewards for each demonstration based on the robustness value;

learn a control policy based on the inferred rewards; and

provide one or more control signals to one or more actuators of the agent based on the control policy.

2. The system of claim 1 , wherein the computer-readable medium stores further instructions that when executed cause the system to:

verify the learned control policy based on the set of specifications to determine a final policy.

3. The system of claim 1 , wherein the computer-readable medium stores further instructions that when executed cause the system to:

for each demonstration comprising a set of state and action pairs, generate a state reward corresponding to each state; and

generate a candidate reward for each demonstration based on the state reward for each state in the set of states for each demonstration.

4. The system of claim 3 , wherein inferring rewards for each demonstration comprises ranking each demonstration based on the robustness value, and determining a learner reward for the agent based on the ranks and corresponding candidate rewards for each demonstration.

5. The system of claim 1 , wherein the control policy is determined based on a reinforcement learning algorithm.

6. The system of claim 1 , wherein the set of specifications are provided in natural language.

7. The system of claim 1 , wherein the agent is selected from the group consisting of a cyber-physical system, a robotic system, an autonomous vehicle, an insulin delivery system, and a drone.

8. The system of claim 1 , wherein each specification in the set of specifications is represented as a directed acyclic graph (DAG).

9. A system, comprising:

one or more sensors configured to acquire environmental data of an environment interacting with the system;

one or more processors;

a computer-readable medium storing executable instructions that, when executed, cause the system to perform operations comprising:

evaluate a current state of the system according to the environmental data from the one or more sensors; and

determine an action to be performed by the system based on the current state according to a control policy;

wherein the control policy is learned based on inferred rewards from a plurality of demonstrations; and

wherein the plurality of demonstrations are evaluated and ranked based on a robustness value of each demonstration using a set of specifications in a formal language; and

provide control signals to one or more actuators based on the control policy.

10. The system of claim 9 , wherein the formal language is selected from the group consisting of is selected from the group consisting of a temporal logic, a Signal Temporal Logic (STL), a Linear Temporal Logic (LTL), and a Computation Tree Logic (CTL).

11. The system of claim 9 , wherein the inferred rewards are based on uncertainties in the environment; and wherein the inferred reward increases with increase in uncertainty.

12. A method for performing reinforcement learning, the method comprising:

receiving a set of demonstrations, the set of demonstrations obtained via interaction of an agent with an environment;

receiving a set of specifications, the set of specifications providing descriptions of one or more tasks and/or one or more objectives;

converting the set of specifications into a temporal logic formal language;

evaluating the set of demonstrations based on the set of specifications in the temporal logic formal language;

generating a robustness value for each demonstration in the set of demonstrations based on the evaluation;

inferring rewards for each demonstration based on the robustness value;

learning a control policy based on the inferred rewards; and

storing the control policy; and

provide one or more control signals to one or more actuators of the agent based on the control policy.

13. The method of claim 12 , wherein the temporal logic is selected from the group consisting of Signal Temporal Logic (STL), Linear Temporal Logic (LTL), and Computation Tree Logic (CTL).

14. The method of claim 12 , further comprising, for each demonstration comprising a set of state and action pairs, generating a state reward corresponding to each state.

15. The method of claim 14 , further comprising, generating a candidate reward for each demonstration based on the state reward for each state in the set of states.

16. The method of claim 15 , wherein inferring rewards for each demonstration comprises ranking each demonstration based on the robustness value, and determining a learner reward for the agent based on the ranks and corresponding candidate rewards for each demonstration.

17. The method of claim 15 , wherein the agent is selected from the group consisting of a cyber-physical system, a robotic system, an autonomous vehicle, an insulin delivery system, and a drone.

Assignments (2)
CONFIRMATORY LICENSE Recorded Feb 12, 2025
From: UNIVERSITY OF SOUTHERN CALIFORNIA
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 070189/0563 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: PURANIC, ANIRUDDH GOPINATH; DESHMUKH, JYOTIRMOY; NIKOLAIDIS, STEFANOS
To: UNIVERSITY OF SOUTHERN CALIFORNIA
Reel/Frame 059081/0036 →
Continuity (1)
Provisional Application 63139540 · Jan 20, 2021
References Cited (22)
US 9050200B2 · Digiovanna · 2015 [cited by examiner]
US 11034019B2 · Tellex · 2021 [cited by examiner]
US 20130218335A1 · Barajas · 2013 [cited by examiner]
US 20180272535A1 · Ogawa · 2018 [cited by examiner]
US 20190126472A1 · Tunyasuvunakool · 2019 [cited by examiner]
US 20190272465A1 · Kimura · 2019 [cited by examiner]
US 20200023514A1 · Tellex · 2020 [cited by examiner]
US 20200226467A1 · Fainekos · 2020 [cited by examiner]
US 20200276703A1 · Chebotar · 2020 [cited by examiner]
US 20210334657A1 · Jordan · 2021 [cited by examiner]
US 20220197306A1 · Cella · 2022 [cited by examiner]
US 20230031545A1 · Oleynik · 2023 [cited by examiner]
US 20230241772A1 · Schillinger · 2023 [cited by examiner]
US 20240173855A1 · Thon · 2024 [cited by examiner]
CN 109726813A · 2019 [cited by examiner]
DE 102019134794B4 · 2021 [cited by examiner]
DE-102019134794-B4 translation (Year: 2021). [cited by examiner]
CN-109726813-A translation (Year: 2019). [cited by examiner]
Elaborating on Learned Demonstrations with Temporal Logic Specifications (Year: 2020). [cited by examiner]
Robust Model Predictive Control for Signal Temporal Logic Synthesis (Year: 2015). [cited by examiner]
A Policy Search Method for Temporal Logic Specified Reinforcement Learning Task (Year: 2018). [cited by examiner]
Learning From Demonstrations Using Signal Temporal Logic in Stochastic and Continuous Domains (Year: 2021). [cited by examiner]
Cited By (1)
US 12,665,094