IP Library Granted Patent US 12,333,434
Granted Patent B2
US 12,333,434 · App. 18/484,816 · Granted Jun 17, 2025

Meta imitation learning with structured skill discovery

Inventors: Wenchao Yu (Plainsboro, NJ); Wei Cheng (Princeton Junction, NJ); Haifeng Chen (West Windsor, NJ); Yiwei Sun (State College, PA)
Assignee: NEC Corporation
G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,434
App. No.
18/484,816
Granted
Jun 17, 2025
Kind
B2
Abstract

A method for acquiring skills through imitation learning by employing a meta imitation learning framework with structured skill discovery (MILD) is presented. The method includes learning behaviors or tasks, by an agent, from demonstrations: by learning to decompose the demonstrations into segments, via a segmentation component, the segments corresponding to skills that are transferrable across different tasks, learning relationships between the skills that are transferrable across the different tasks, employing, via a graph generator, a graph neural network for learning implicit structures of the skills from the demonstrations to define structured skills, and generating policies from the structured skills to allow the agent to acquire the structured skills for application to one or more target tasks.

Claims (198)

1. A method for acquiring skills through imitation learning, the method comprising:

learning behaviors or tasks, by an agent, from state-action pairs of medical treatment for a given disease by:

learning to decompose the state-action pairs into segments, via a segmentation component, the segments corresponding to skills that are transferrable across different tasks;

learning relationships between the skills;

employing, via a graph generator, a graph neural network for learning implicit structures of the skills from the state-action pairs to define structured skills; and

generating policies from the structured skills to allow the agent to acquire the structured skills for application to one or more target tasks by optimizing an objective function:

min

π

θ

(

p

π

E

p

π

θ

)

-

H

[

c

]

+

H

[

c

"\[LeftBracketingBar]"

X

,

g

]

+

H

[

X

"\[LeftBracketingBar]"

c

]

-

H

[

g

]

+

H

[

g

"\[LeftBracketingBar]"

s

,

c

]

wherein p πθ is a generated policy with parameters π θ , p π E is an expert policy, is a distance function, H is a Shannon entropy function, c is a set of skills, X is an implicit structure, g is a segmentation, and s is a state; and

providing a medication to a patient in accordance with the generated policies to treat the given disease.

2. The method of claim 1 , wherein the relationship between the skills is learned by employing a graph generator decoder to predict a graph probability distribution over latent interaction of the skills and a graph decoder to generate skills conditioned on a graph structure.

3. A non-transitory computer-readable storage medium comprising a computer-readable program for acquiring skills through imitation learning, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of:

learning behaviors or tasks, by an agent, from state-action pairs of medical treatment for a given disease by:

learning to decompose the state-action pairs into segments, via a segmentation component, the segments corresponding to skills that are transferrable across different tasks;

learning relationships between the skills;

employing, via a graph generator, a graph neural network for learning implicit structures of the skills from the state-action pairs to define structured skills; and

generating policies from the structured skills to allow the agent to acquire the structured skills for application to one or more target tasks by optimizing an objective function:

min

π

θ

(

p

π

E

p

π

θ

)

-

H

[

c

]

+

H

[

c

"\[LeftBracketingBar]"

X

,

g

]

+

H

[

X

"\[LeftBracketingBar]"

c

]

-

H

[

g

]

+

H

[

g

"\[LeftBracketingBar]"

s

,

c

]

wherein p π θ is a generated policy with parameters π θ , p π E is an expert policy, is a distance function, H is a Shannon entropy function, c is a set of skills, X is an implicit structure, g is a segmentation, and s is a state; and

providing a medication to a patient in accordance with the generated policies to treat the given disease.

4. The non-transitory computer-readable storage medium of claim 3 , wherein the relationship between the skills is learned by employing a graph generator decoder to predict a graph probability distribution over latent interaction of the skills and a graph decoder to generate skills conditioned on a graph structure.

5. A system for acquiring skills through imitation learning, the system comprising:

an imitation component to minimize a measure of discrepancy between a learned policy and an expert policy by optimizing an objective function:

min

π

θ

(

p

π

E

p

π

θ

)

-

H

[

c

]

+

H

[

c

"\[LeftBracketingBar]"

X

,

g

]

+

H

[

X

"\[LeftBracketingBar]"

c

]

-

H

[

g

]

+

H

[

g

"\[LeftBracketingBar]"

s

,

c

]

wherein p π θ is a generated policy with parameters π θ , p π E is an expert policy, is a distance function, H is a Shannon entropy function, c is a set of skills, X is an implicit structure, g is a segmentation, and s is a state;

a graph neural network to learn the implicit structure of skills from state-action pairs of medical treatment to define structured skills;

a meta controller to learn predictable skills;

a segmentation component to learn to decompose the state-action pairs into segments corresponding to skills that are transferrable across different tasks concurrently with learning relationships between the skills; and

a treatment component to provide a medication to a patient in accordance with the learned policies to treat the given disease.

6. The system of claim 5 , wherein the relationship between the skills is learned by employing a graph generator decoder to predict a graph probability distribution over latent interaction of the skills and a graph decoder to generate skills conditioned on a graph structure.

7. The system of claim 5 , wherein the learned policy is optimized based on gradients of an imitation learning loss with respect to expert state-action pairs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 071095/0825 →
Continuity (4)
Continuation 17391427 · Aug 2, 2021
Provisional Application 63084035 · Sep 28, 2020
Provisional Application 63067009 · Aug 18, 2020
Related Publication 20240046092A1 · Feb 8, 2024
References Cited (5)
US 20200082940A1 · Qiao · 2020 [cited by examiner]
US 20220058482A1 · Yu · 2022 [cited by examiner]
Luo et al, Grouped Spatial-Temporal Aggregation for Efficient Action Recognition, Sep. 28, 2019, https://arxiv.org/pdf/1909.13130 (Year: 2019). [cited by examiner]
Hester et al, Deep Q-learning from Demonstrations, Nov. 22, 2017, https://arxiv.org/pdf/1704.03732 (Year: 2017). [cited by examiner]
Bacciu et al, A Gentle Introduction to Deep Learning for Graphs, Jun. 15, 2020, https://arxiv.org/pdf/1912.12693 (Year: 2020). [cited by examiner]