IP Library › Granted Patent US 12,217,137
Granted Patent B1
US 12,217,137 · App. 17/039,447 · Granted Feb 4, 2025

Meta-Q learning

Inventors: Rasool Fakoor (San Jose, CA); Alexander Johannes Smola (Sunnyvale, CA); Stefano Soatto (Pasadena, CA); Pratik Anil Chaudhari (Pasadena, CA)
Assignee: Amazon Technologies, Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,137
App. No.
17/039,447
Granted
Feb 4, 2025
Kind
B1
Abstract

Techniques for Meta-Q-Learning (MQL) are described. A method of MQL may include receiving a request from an agent to perform adaptation based at least on task data associated with a new task collected by the agent, identifying a subset of meta-training data corresponding to the task data in a replay buffer, and adapting a policy using the subset of meta-training data and the task data to generate an adapted policy, wherein the adapted policy is used identify a next action for the agent to perform.

Claims (53)

1. A computer-implemented method comprising:

receiving a set of training tasks;

initializing a replay buffer and parameters for a policy;

for each training task in the set of training tasks:

obtaining training data for the training task by interacting with an environment;

updating a context variable based on the training data;

storing the training data in the replay buffer; and

updating the parameters for the policy based on the training data to create a meta-trained policy;

performing adaptation on the meta-trained policy using task data and a subset of the training data to generate an adapted meta-trained policy, wherein the subset of the training data is identified using the context variable and a propensity score, wherein the propensity score indicates a similarity between the task data and at least one training task from the set of training tasks.

2. The computer-implemented method of claim 1 , wherein the context variable represents past action and reward data associated with the set of training tasks.

3. The computer-implemented method of claim 1 , wherein the propensity score is estimated using a classifier.

4. A computer-implemented method comprising:

performing, by an agent, adaptation based at least on task data associated with a new task collected by the agent;

identifying, by the agent in a replay buffer using a propensity score, a subset of meta-training data corresponding to the task data, wherein the propensity score indicates a similarity between the new task and at least one training task from a set of training tasks;

adapting a policy using the subset of meta-training data and the task data to generate an adapted policy; and

identifying a next action for the agent to perform using the adapted policy.

5. The computer-implemented method of claim 4 , further comprising meta-training an initial policy using a set of training tasks, wherein the initial policy is trained to maximize an average reward across the set of training tasks.

6. The computer-implemented method of claim 5 , wherein meta-training an initial policy using the set of training tasks, wherein the initial policy is trained to maximize an average reward across the set of training tasks, further comprises for each training task in the set of training tasks:

interacting with a training environment to obtain task data;

updating a context based on the task data;

storing the task data in the replay buffer; and

updating parameters for the policy based on the meta-training data to create a meta-trained policy.

7. The computer-implemented method of claim 6 , wherein identifying a subset of meta-training data corresponding to the task data in a replay buffer further comprises:

sampling the meta-training data from the replay buffer, the meta-training data associated with a plurality of tasks from a multi-task training phase;

determining the propensity score between the sampled meta-training data from the replay buffer and the task data; and

determining the subset of the meta-training data using the propensity score.

8. The computer-implemented method of claim 7 , wherein the propensity score is estimated based at least on the context.

9. The computer-implemented method of claim 8 , wherein the context represents past trajectory of the agent.

10. The computer-implemented method of claim 6 , wherein the context implements a machine learning model that encodes a recent subset of a past trajectory of the agent into an embedding and compares the embedding to embeddings associated with the set of training tasks to identify a most similar task.

11. The computer-implemented method of claim 4 , further comprising performing the next action based on a current state of the agent and the adapted policy.

12. The computer-implemented method of claim 4 , further comprising sampling the task data from a temporary buffer, the task data generated based at least on prior interactions with an environment by the agent while performing a new task.

13. A system comprising:

a first one or more electronic devices to implement a storage service in a multi-tenant provider network; and

a second one or more electronic devices to implement a Meta-Q-Learning (MQL) agent in the multi-tenant provider network, the MQL agent including instructions that upon execution by one or more processors cause the MQL agent to:

perform adaptation based at least on task data associated with a new task collected by the MQL agent;

identify, in a replay buffer of the storage service using a propensity score, a subset of meta-training data corresponding to the task data, wherein the propensity score indicates a similarity between the new task and at least one training task from a set of training tasks;

adapt a policy using the subset of meta-training data and the task data to generate an adapted policy; and

identify a next action for the MQL agent to perform using the adapted policy.

14. The system of claim 13 , wherein the instructions, when executed, further cause the MQL agent to meta-train an initial policy using a set of training tasks, wherein the initial policy is trained to maximize an average reward across the set of training tasks.

15. The system of claim 14 , wherein to meta-train an initial policy using the set of training tasks, wherein the initial policy is trained to maximize an average reward across the set of training tasks, the instructions, when executed, further cause the MQL agent to:

for each training task in the set of training tasks:

interact with a training environment to obtain task data;

update a context based on the task data;

store the task data in the replay buffer; and

update parameters for the policy based on the meta-training data to create a meta-trained policy.

16. The system of claim 15 , wherein to identify a subset of meta-training data corresponding to the task data in a replay buffer in the storage service, the instructions, when executed, further cause the MQL agent to:

sample the meta-training data from the replay buffer, the meta-training data associated with a plurality of tasks from a multi-task training phase;

determine the propensity score between the sampled meta-training data from the replay buffer and the task data; and

determine the subset of the meta-training data using the propensity score.

17. The system of claim 16 , wherein the propensity score is estimated based at least on the context.

18. The system of claim 17 , wherein the context represents past trajectory of the MQL agent.

19. The system of claim 13 , wherein the instructions, when executed, further cause the MQL agent to identify the next action to perform based at least on a current state of the MQL agent and the adapted policy.

20. The system of claim 13 , wherein the instructions, when executed, further cause the MQL agent to sample the task data from a temporary buffer, the task data generated based at least on prior interactions with an environment by the agent while performing the new task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: FAKOOR, RASOOL; SMOLA, ALEXANDER JOHANNES; SOATTO, STEFANO; CHAUDHARI, PRATIK ANIL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 054552/0429 →
References Cited (22)
US 10963754B1 · Ravichandran · 2021 [cited by examiner]
US 11263222B2 · Misra · 2022 [cited by examiner]
US 11501167B2 · Camilo Gamboa Higuera · 2022 [cited by examiner]
US 11526812B2 · Devlin · 2022 [cited by examiner]
US 11685045B1 · Herzog · 2023 [cited by examiner]
US 20180336480A1 · Chang · 2018 [cited by examiner]
US 20190180302A1 · Ventrice · 2019 [cited by examiner]
US 20190232488A1 · Levine · 2019 [cited by examiner]
US 20190354859A1 · Xu · 2019 [cited by examiner]
US 20200250493A1 · Simmons-Edler · 2020 [cited by examiner]
US 20200311585A1 · Helenius · 2020 [cited by examiner]
US 20210174205A1 · Rosa · 2021 [cited by examiner]
US 20210205988A1 · James · 2021 [cited by examiner]
US 20220013230A1 · Wu · 2022 [cited by examiner]
US 20220019878A1 · Li · 2022 [cited by examiner]
US 20220036179A1 · Garg · 2022 [cited by examiner]
US 20220105624A1 · Kalakrishnan · 2022 [cited by examiner]
US 20220161423A1 · Perez · 2022 [cited by examiner]
US 20220327814A1 · Han · 2022 [cited by examiner]
US 20230214649A1 · Jeong · 2023 [cited by examiner]
Rasool Fakoor, Pratik Chaudhari, Stefano Soatto, and Alexander J. Smola. Meta-q-learning. In ICLR, Apr. 4, 2020. (Year: 2020). [cited by examiner]
Siraj Raval, “Q-learning”, working screenshot video available online at [https://www.youtube.com/watch?v=aCEvtRtNO-M], published on 2017 (Year: 2017). [cited by examiner]
Cited By (3)
US 12,387,138 US 12,567,255 US 12,619,912