IP Library Granted Patent US 11,580,445
Granted Patent B2
US 11,580,445 · App. 16/653,890 · Granted Feb 14, 2023

Efficient off-policy credit assignment

Inventors: Hao Liu (San Francisco, CA); Richard Socher (Menlo Park, CA); Caiming Xiong (Mountain View, CA)
Assignee: salesforce.com, inc.
G06N20/00G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,445
App. No.
16/653,890
Filed
Oct 15, 2019
Granted
Feb 14, 2023
Kind
B2
Art Unit
2124
USPC
706/12
Abstract

Systems and methods are provided for efficient off-policy credit assignment (ECA) in reinforcement learning. ECA allows principled credit assignment for off-policy samples, and therefore improves sample efficiency and asymptotic performance. One aspect of ECA is to formulate the optimization of expected return as approximate inference, where policy is approximating a learned prior distribution, which leads to a principled way of utilizing off-policy samples. Other features are also provided.

Claims (41)

1. A method for efficient off-policy credit assignment in reinforcement learning, the method comprising:

receiving, via a communication interface, a training dataset comprising a plurality of contexts and sample results generated from one or more agents implementing one or more actions,

wherein the one or more actions are generated by a learning model according to one or more policies that are unrelated to a current policy in response to a context from the plurality of contexts,

wherein the received sample results comprise a successful sample result resulting in a high reward for the one or more agents, and an unsuccessful sample result resulting in a low reward for the one or more agents;

computing, via a processor, a learned prior distribution of the one or more actions based on an average of the one or more policies that generate the one or more actions over the training dataset;

computing one or more adaptive weights of samples based at least in part on the learned prior distribution and the received sample results including both the successful sample result and the unsuccessful sample result;

generating an estimate of a gradient using the one or more adaptive weights; and

training the learning model by updating the current policy using the estimate of the gradient.

2. The method of claim 1 , wherein computing one or more adaptive weights comprises generating at least one of a high-reward result credit and a zero-reward result credit.

3. The method of claim 2 , wherein generating an estimate of a gradient comprises generating a high-reward score function gradient using the high-reward result credit to weight successful sample results.

4. The method of claim 2 , wherein generating an estimate of a gradient comprises generating a zero-reward score function gradient using the zero-reward result credit to weight unsuccessful sample results.

5. The method of claim 1 , wherein the learning model is applied to a task of semantic parsing.

6. The method of claim 5 , wherein the sample results correspond to executable software code generated by the learning model based on natural language instructions text.

7. A system for efficient off-policy credit assignment in reinforcement learning, the system comprising:

a memory storing machine executable code; and

one or more processors coupled to the memory and configurable to execute the machine executable code to cause the one or more processors to:

receive a training dataset comprising a plurality of contexts and sample results generated from one or more agents implementing one or more actions,

wherein the one or more actions are generated by a learning model according to one or more policies that are unrelated to a current policy in response to a context from the plurality of contexts,

wherein the received sample results comprise a successful sample result resulting in a high reward for the one or more agents, and an unsuccessful sample result resulting in a low reward for the one or more agents;

compute a learned prior distribution of the one or more actions based on an average of the one or more policies that generate the one or more actions over the training dataset;

compute one or more adaptive weights of samples based at least in part on the learned prior distribution and the received sample results including both the successful sample result and the unsuccessful sample result;

generate an estimate of a gradient using the one or more adaptive weights; and

train the learning model by updating the current policy using the estimate of the gradient.

8. The system of claim 7 , wherein the one or more processors are configurable to generate at least one of a high-reward result credit and a zero-reward result credit.

9. The system of claim 8 , wherein the one or more processors are configurable to generate a high-reward score function gradient using the high-reward result credit to weight successful sample results.

10. The system of claim 8 , wherein the one or more processors are configurable to generate a zero-reward score function gradient using the zero-reward result credit to weight unsuccessful sample results.

11. The system of claim 7 , wherein the learning model is applied to a task of semantic parsing.

12. The system of claim 11 , wherein the sample results correspond to executable software code generated by the learning model based on natural language instructions text.

13. A non-transitory machine-readable medium comprising executable code which when executed by one or more processors associated with a computer are adapted to cause the one or more processors to perform a method for efficient off-policy credit assignment in reinforcement learning, the method comprising:

receiving a training dataset comprising a plurality of contexts and sample results generated from one or more agents implementing one or more actions,

wherein the one or more actions are generated by a learning model according to one or more policies that are unrelated to a current policy in response to a context from the plurality of contexts,

wherein the received sample results comprise a successful sample result resulting in a high reward for the one or more agents, and an unsuccessful sample result resulting in a low reward for the one or more agents;

computing a learned prior distribution of the one or more actions based on an average of the one or more policies that generate the one or more actions over the training dataset;

computing one or more adaptive weights of samples based at least in part on the learned prior distribution and the received sample results including both the successful sample result and the unsuccessful sample result;

generating an estimate of a gradient using the one or more adaptive weights; and

training the learning model by updating the current policy using the estimate of the gradient.

14. The non-transitory machine-readable medium of claim 13 , wherein computing one or more adaptive weights comprises generating at least one of a high-reward result credit and a zero-reward result credit.

15. The non-transitory machine-readable medium of claim 14 , wherein generating an estimate of a gradient comprises generating a high-reward score function gradient using the high-reward result credit to weight successful sample results.

16. The non-transitory machine-readable medium of claim 14 , wherein generating an estimate of a gradient comprises generating a zero-reward score function gradient using the zero-reward result credit to weight unsuccessful sample results.

17. The non-transitory machine-readable medium of claim 13 , wherein the learning model is applied to a task of semantic parsing.

18. The non-transitory machine-readable medium of claim 17 , wherein the sample results correspond to executable software code generated by the learning model based on natural language instructions text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2019
From: LIU, HAO; SOCHER, RICHARD; XIONG, CAIMING
To: SALESFORCE.COM, INC.
Reel/Frame 050971/0256 →
Continuity (4)
Provisional Application 62852258 · May 23, 2019
Provisional Application 62849007 · May 16, 2019
Provisional Application 62813937 · Mar 5, 2019
Related Publication 20200285993A1 · Sep 10, 2020
Cited By (2)
US 12,639,189 US 12,705,477