IP Library Granted Patent US 12,346,815
Granted Patent B2
US 12,346,815 · App. 18/484,805 · Granted Jul 1, 2025

Meta imitation learning with structured skill discovery

Inventors: Wenchao Yu (Plainsboro, NJ); Wei Cheng (Princeton Junction, NJ); Haifeng Chen (West Windsor, NJ); Yiwei Sun (State College, PA)
Assignee: NEC Corporation
G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,815
App. No.
18/484,805
Granted
Jul 1, 2025
Kind
B2
Abstract

A method for acquiring skills through imitation learning by employing a meta imitation learning framework with structured skill discovery (MILD) is presented. The method includes learning behaviors or tasks, by an agent, from demonstrations: by learning to decompose the demonstrations into segments, via a segmentation component, the segments corresponding to skills that are transferrable across different tasks, learning relationships between the skills that are transferrable across the different tasks, employing, via a graph generator, a graph neural network for learning implicit structures of the skills from the demonstrations to define structured skills, and generating policies from the structured skills to allow the agent to acquire the structured skills for application to one or more target tasks.

Claims (207)

1. A method for acquiring skills of medication through imitation learning, the method comprising:

learning behaviors or tasks, by an agent, from medical treatment demonstrations for a given disease by:

learning to decompose the demonstrations into segments, via a segmentation component, the segments corresponding to skills that are transferrable across different tasks;

learning relationships between the skills;

employing, via a graph generator, a graph neural network for learning implicit structures of the skills from the demonstrations to define structured skills; and

generating policies for medication from the structured skills to allow the agent to acquire the structured skills for application to one or more target tasks by optimizing an objective function:

min

π

θ

(

p

π

E

p

π

θ

)

-

H

[

c

]

+

H

[

c

"\[LeftBracketingBar]"

X

,

g

]

+

H

[

X

"\[LeftBracketingBar]"

c

]

-

H

[

g

]

+

H

[

g

"\[LeftBracketingBar]"

s

,

c

]

 wherein p π θ is a generated policy with parameters π θ , p π E is an expert policy, is a distance function, H is a Shannon entropy function, c is a set of skills, X is an implicit structure, g is a segmentation, and s is a state; and

providing a medication to a patient in accordance with the generated policies to treat the given disease.

2. The method of claim 1 , wherein the demonstrations are expert trajectories generated by an expert policy for medication.

3. The method of claim 2 , wherein each of the expert trajectories includes state-action pairs.

4. The method of claim 3 , wherein the skills correspond to subsequences of the state-action pairs extracted from the expert trajectories.

5. The method of claim 1 , wherein the relationship between the skills is learned by employing a graph generator decoder to predict a graph probability distribution over latent interaction of the skills and a graph decoder to generate skills conditioned on a graph structure.

6. A non-transitory computer-readable storage medium comprising a computer-readable program for acquiring skills of medication through imitation learning, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of:

learning behaviors or tasks, by an agent, from medical treatment demonstrations of a given disease by:

learning to decompose the demonstrations into segments, via a segmentation component, the segments corresponding to skills that are transferrable across different tasks;

learning relationships between the skills;

employing, via a graph generator, a graph neural network for learning implicit structures of the skills from the demonstrations to define structured skills; and

generating policies for medication from the structured skills to allow the agent to acquire the structured skills for application to one or more target tasks by optimizing an objective function:

min

π

θ

(

p

π

E

p

π

θ

)

-

H

[

c

]

+

H

[

c

"\[LeftBracketingBar]"

X

,

g

]

+

H

[

X

"\[LeftBracketingBar]"

c

]

-

H

[

g

]

+

H

[

g

"\[LeftBracketingBar]"

s

,

c

]

 wherein p π θ is a generated policy with parameters π θ , p π E is an expert policy, is a distance function, H is a Shannon entropy function, c is a set of skills, X is an implicit structure, g is a segmentation, and s is a state; and

providing a medication to a patient in accordance with the generated policies to treat the given disease.

7. The non-transitory computer-readable storage medium of claim 6 , wherein the demonstrations are expert trajectories generated by an expert policy for medication.

8. The non-transitory computer-readable storage medium of claim 7 , wherein each of the expert trajectories includes state-action pairs.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the skills correspond to subsequences of the state-action pairs extracted from the expert trajectories.

10. The non-transitory computer-readable storage medium of claim 6 , wherein the relationship between the skills is learned by employing a graph generator decoder to predict a graph probability distribution over latent interaction of the skills and a graph decoder to generate skills conditioned on a graph structure.

11. A system for acquiring skills of medication through imitation learning, the system comprising:

an imitation component to minimize a measure of discrepancy between a learned policy and an expert policy for medication by optimizing an objective function:

min

π

θ

(

p

π

E

p

π

θ

)

-

H

[

c

]

+

H

[

c

"\[LeftBracketingBar]"

X

,

g

]

+

H

[

X

"\[LeftBracketingBar]"

c

]

-

H

[

g

]

+

H

[

g

"\[LeftBracketingBar]"

s

,

c

]

wherein p π θ is a generated policy with parameters π θ , p π E is an expert policy, is a distance function, H is a Shannon entropy function, c is a set of skills, X is an implicit structure g is a segmentation, and s is a state;

a graph neural network to learn the implicit structure of skills from medical treatment demonstrations to define structured skills;

a meta controller to learn predictable skills;

a segmentation component to learn to decompose the demonstrations into segments corresponding to skills that are transferrable across different tasks concurrently with learning relationships between the skills; and

a treatment component to provide a medication to a patient in accordance with the learned policies to treat the given disease.

12. The system of claim 11 , wherein the demonstrations are expert trajectories generated by an expert policy for medication.

13. The system of claim 12 , wherein each of the expert trajectories includes state-action pairs.

14. The system of claim 13 , wherein the skills correspond to subsequences of the state-action pairs extracted from the expert trajectories.

15. The system of claim 11 , wherein the relationship between the skills is learned by employing a graph generator decoder to predict a graph probability distribution over latent interaction of the skills and a graph decoder to generate skills conditioned on a graph structure.

16. The system of claim 11 , wherein the learned policy for medication is optimized based on gradients of an imitation learning loss with respect to expert demonstrations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 071095/0825 →
Continuity (4)
Continuation 17391427 · Aug 2, 2021
Provisional Application 63084035 · Sep 28, 2020
Provisional Application 63067009 · Aug 18, 2020
Related Publication 20240037400A1 · Feb 1, 2024
References Cited (4)
US 20220092441A1 · Zhu · 2022 [cited by examiner]
Bacciu et al, A Gentle Introduction to Deep Learning for Graphs, Jun. 15, 2020, https://arxiv.org/pdf/1912.12693 (Year: 2020). [cited by examiner]
Hester et al, Deep Q-learning from Demonstrations, Nov. 22, 2017, https://arxiv.org/pdf/1704.03732 (Year: 2017). [cited by examiner]
Luo et al, Grouped Spatial-Temporal Aggregation for Efficient Action Recognition, Sep. 28, 2019, https://arxiv.org/pdf/1909.13130 (Year: 2019). [cited by examiner]