IP Library Granted Patent US 12,380,360
Granted Patent B2
US 12,380,360 · App. 17/323,475 · Granted Aug 5, 2025

Interpretable imitation learning via prototypical option discovery for decision making

Inventors: Wenchao Yu (Plainsboro, NJ); Haifeng Chen (West Windsor, NJ); Wei Cheng (Princeton Junction, NJ)
Assignee: NEC Corporation
G06N20/00G06F16/951G06N20/10G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,360
App. No.
17/323,475
Granted
Aug 5, 2025
Kind
B2
Abstract

A method for learning prototypical options for interpretable imitation learning is presented. The method includes initializing options by bottleneck state discovery, each of the options presented by an instance of trajectories generated by experts, applying segmentation embedding learning to extract features to represent current states in segmentations by dividing the trajectories into a set of segmentations, learning prototypical options for each segment of the set of segmentations to mimic expert policies by minimizing loss of a policy and projecting prototypes to the current states, training option policy with imitation learning techniques to learn a conditional policy, generating interpretable policies by comparing the current states in the segmentations to one or more prototypical option embeddings, and taking an action based on the interpretable policies generated.

Claims (195)

1. A method for learning prototypical options for interpretable imitation learning, the method comprising:

initializing options by bottleneck state discovery, each of the options presented by an instance of trajectories generated by experts;

applying segmentation embedding learning to extract features to represent current states in segmentations by dividing the trajectories into a set of segmentations;

learning prototypical options for each segment of the set of segmentations to mimic expert policies by minimizing loss of a policy and projecting prototypes to the current states;

learning prototypical option embedding using an objective function:

option

=

-

λ

1

*

I

L

loss

+

λ

2

*

i

=

1

K

min

m

=

1

M

=

f

ϕ

(

s

v

m

:

v

m

)

-

e

i

2

2

)

+

λ

3

*

i

=

1

K

.

j

=

i

+

1

K

max

(

0

,

d

min

-

e

i

-

e

j

)

where L IL loss is an imitation learning loss, f φ is, the second term is a segment representation function for a segment s ν′ m ,ν m from segment ν m to segment ν m′ , e i and e j are embedded prototypes, K is a number of prototypes, M is a number of segments, d min is a threshold value, and λ 1 , λ 2 , and λ 3 , are weighting parameters;

training option policy with imitation learning techniques to learn a conditional policy;

generating interpretable policies by comparing the current states in the segmentations to one or more prototypical option embeddings;

generating dosage options for a patient based on the interpretable policies;

displaying the dosage options on a user interface for a user; and

taking an action based on the dosage options.

2. The method of claim 1 , wherein option initialization includes identifying states from the current states that connect different densely connected regions in a state space.

3. The method of claim 2 , wherein a soft attention mechanism is employed to obtain important states with particular attention weights.

4. The method of claim 3 , wherein the important states are found with density-based spatial clustering of applications with noise (DBSCAN).

5. The method of claim 1 , wherein the bottleneck state discovery divides the trajectories generated by the experts into disjoint segments of variable length by a density-based clustering method.

6. The method of claim 1 , wherein each of the options includes an intra-option policy, a termination condition, an initiation state set, and an option prototype.

7. The method of claim 6 , wherein the option prototype is defined by a sub-trajectory generated by the experts.

8. The method of claim 1 , wherein each of the one or more prototypical option embeddings is assigned with a respective closest segment embedding in a training set.

9. The method of claim 1 , wherein the loss is a least square loss.

10. The method of claim 1 , wherein a diversity regularization term is employed to penalize one or more of the prototypical options that are close to each other.

11. A non-transitory computer-readable storage medium comprising a computer-readable program for learning prototypical options for interpretable imitation learning, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of:

initializing options by bottleneck state discovery, each of the options presented by an instance of trajectories generated by experts;

applying segmentation embedding learning to extract features to represent current states in segmentations by dividing the trajectories into a set of segmentations;

learning prototypical options for each segment of the set of segmentations to mimic expert policies by minimizing loss of a policy and projecting prototypes to the current states; learning prototypical option embedding using an objective function:

option

=

-

λ

1

*

I

L

loss

+

λ

2

*

i

=

1

K

min

m

=

1

M

=

f

ϕ

(

s

v

m

:

v

m

)

-

e

i

2

2

)

+

λ

3

*

i

=

1

K

.

j

=

i

+

1

K

max

(

0

,

d

min

-

e

i

-

e

j

)

where L IL loss is an imitation learning loss, f φ is, the second term is a segment representation function for a segment s ν′ m ,ν m from segment ν m to segment ν m′ , e i and e j are embedded prototypes, K is a number of prototypes, M is a number of segments, d min is a threshold value, and λ 1 , λ 2 , and λ 3 are weighting parameters:

training option policy with imitation learning techniques to learn a conditional policy;

generating interpretable policies by comparing the current states in the segmentations to one or more prototypical option embeddings;

generating dosage options for a patient based on the interpretable policies;

displaying the dosage options on a user interface for a user; and

taking an action based on the interpretable policies generated dosage options.

12. The non-transitory computer-readable storage medium of claim 11 , wherein option initialization includes identifying states from the current states that connect different densely connected regions in a state space.

13. The non-transitory computer-readable storage medium of claim 12 , wherein a soft attention mechanism is employed to obtain important states with particular attention weights.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the important states are found with density-based spatial clustering of applications with noise (DBSCAN).

15. The non-transitory computer-readable storage medium of claim 11 , wherein the bottleneck state discovery divides the trajectories generated by the experts into disjoint segments of variable length by a density-based clustering method.

16. The non-transitory computer-readable storage medium of claim 11 , wherein each of the options includes an intra-option policy, a termination condition, an initiation state set, and an option prototype.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the option prototype is defined by a sub-trajectory generated by the experts.

18. The non-transitory computer-readable storage medium of claim 11 , wherein each of the one or more prototypical option embeddings is assigned with a respective closest segment embedding in a training set.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 071486/0094 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2021
From: YU, WENCHAO; CHEN, HAIFENG; CHENG, WEI
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 056276/0245 →
Continuity (3)
Provisional Application 63029754 · May 26, 2020
Provisional Application 63033304 · Jun 2, 2020
Related Publication 20210374612A1 · Dec 2, 2021
References Cited (24)
US 20190324795A1 · Gao · 2019 [cited by examiner]
US 20200334093A1 · Dubey · 2020 [cited by examiner]
US 20210295171A1 · Kamenev · 2021 [cited by examiner]
CA 2872831A1 · 2012 [cited by examiner]
CN 105393264A · 2016 [cited by examiner]
CN 105893256A · 2016 [cited by examiner]
CN 108805877B · 2019 [cited by examiner]
CN 110491171A · 2019 [cited by examiner]
CN 111712862A · 2020 [cited by examiner]
CN 111950950A · 2020 [cited by examiner]
CN 109739585B · 2022 [cited by examiner]
EP 3462385A1 · 2019 [cited by examiner]
EP 2504776B1 · 2019 [cited by examiner]
JP 2017142549A · 2017 [cited by examiner]
JP 7390126B2 · 2023 [cited by examiner]
KR 20130049201A · 2013 [cited by examiner]
WO WO2020162680A1 · 2020 [cited by examiner]
WO WO2020235693A1 · 2020 [cited by examiner]
Abbeel et al., “Apprenticeship Learning via Inverse Reinforcement Learning”, Proceedings of the 21st International Conference on Machine Learning. Jul. 5-9, 2004. pp. 1-8. [cited by applicant]
Eysenbach et al., “Diversity is All You Need: Learning Skills Without a Reward Function”, arXiv:1802.06070v6 [cs.AI]. Oct. 9, 2018. pp. 1-22. [cited by applicant]
Ho et al., “Generative Adversarial Imitation Learning”, arXiv:1606.03476v1 [cs.LG]. Jun. 10, 2016. pp. 1-14. [cited by applicant]
Li et al., “InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations”, arXiv:1703.08840v2 [cs.LG]. Nov. 14, 2017. pp. 1-14. [cited by applicant]
Ming et al., “Interpretable and Steerable Sequence Learning via Prototypes”, 25th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Aug. 4-8, 2019. pp. 1-11. [cited by applicant]
Tomar et al., “Successor Options: An Option Discovery Framework for Reinforcement Learning”, associarXiv:1905.05731v1 [cs.LG]. May 14, 2019. pp. 1-7. [cited by applicant]