IP Library Granted Patent US 11,922,323
Granted Patent B2
US 11,922,323 · App. 16/395,083 · Granted Mar 5, 2024

Meta-reinforcement learning gradient estimation with variance reduction

Inventor: Hao Liu (Palo Alto, CA)
Assignee: Salesforce, Inc.
G06N3/088G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,922,323
App. No.
16/395,083
Granted
Mar 5, 2024
Kind
B2
Abstract

A method for deep reinforcement learning using a neural network model includes receiving a distribution including a plurality of related tasks. Parameters for the reinforcement learning neural network model is trained based on gradient estimation associated with the parameters using samples associated with the plurality of related tasks. Control variates are incorporated into the gradient estimation by automatic differentiation.

Claims (71)

1. A method for training a reinforcement learning neural network model, comprising:

receiving a distribution including a plurality of related tasks; and

training parameters for a reinforcement learning neural network model based on a gradient estimation associated with the parameters using samples associated with the plurality of related tasks,

wherein control variates are incorporated into the gradient estimation using a combination of a first gradient estimator and a second gradient estimator,

wherein the first gradient estimator uses the control variates and the second gradient estimator does not use the control variates,

wherein the control variates are represented by a baseline function depending on the tasks and corresponding states of the tasks,

wherein the gradient estimation is generated using Hessian estimation, and

wherein a product of two factors for the Hessian estimation based on the first gradient estimator and the second gradient estimator is reduced by subtracting the baseline function from one of the two factors.

2. The method of claim 1 ,

wherein the control variates are incorporated into the gradient estimation by automatic differentiation to reduce variance without introducing bias into the gradient estimation.

3. The method of claim 2 , wherein the automatic differentiation includes twice auto-differentiating for Hessian estimation for gradient estimation.

4. The method of claim 1 , wherein the method further comprises:

training the control variates to generate trained control variates; and

incorporating the trained control variates into the gradient estimation.

5. The method of claim 4 , wherein the training the control variates includes:

for each task of the plurality of related tasks, training a separate control variate.

6. The method of claim 5 , wherein the training the control variates includes:

training the control variates across the plurality of related tasks.

7. The method of claim 6 , wherein the training the control variates across the plurality of related tasks includes:

receiving a plurality of episodes corresponding to the plurality of related tasks;

adapting meta control variates parameters for control variates using a first portion of the episodes associated with a first portion of the tasks to generate first adapted meta control variates parameters;

generating first meta control variates estimates associated with a second portion of the tasks based on the first adapted meta control variates parameters;

adapting the meta control variates parameters using a second portion of the episodes associated with the second portion of the tasks to generate second adapted meta control variates parameters;

generating second meta control variates estimates associated with the first portion of the tasks based on the second adapted meta control variates parameters; and

updating the meta control variates parameters using the first and second meta control variates estimates.

8. A non-transitory machine-readable medium comprising executable code which when executed by one or more processors associated with a computing device are adapted to cause the one or more processors to perform a method comprising:

receiving a distribution including a plurality of related tasks; and

training parameters for a reinforcement learning neural network model based on gradient estimation associated with the parameters using samples associated with the plurality of related tasks,

wherein control variates are incorporated into the gradient estimation using a combination of a first gradient estimator and a second gradient estimator,

wherein the first gradient estimator uses the control variates and the second gradient estimator does not use the control variates,

wherein the control variates are represented by a baseline function depending on the tasks and corresponding states of the tasks,

wherein the gradient estimation is generated using Hessian estimation, and

wherein a product of two factors for the Hessian estimation based on the first gradient estimator and the second gradient estimator is reduced by subtracting the baseline function from one of the two factors.

9. The non-transitory machine-readable medium of claim 8 ,

wherein the control variates are incorporated into the gradient estimation by automatic differentiation to reduce variance without introducing bias into the gradient estimation.

10. The non-transitory machine-readable medium of claim 9 , wherein the automatic differentiation includes twice auto-differentiating for the Hessian estimation for gradient estimation.

11. The non-transitory machine-readable medium of claim 8 , wherein the method further comprises:

training the control variates to generate trained control variates; and

incorporating the trained control variates into the gradient estimation.

12. The non-transitory machine-readable medium of claim 11 , wherein the training the control variates includes:

for each task of the plurality of related tasks, training a separate control variate.

13. The non-transitory machine-readable medium of claim 11 , wherein the training the control variates includes:

training the control variates across the plurality of related tasks.

14. The non-transitory machine-readable medium of claim 13 , wherein the training the control variates across the plurality of related tasks includes:

receiving a plurality of episodes corresponding to the plurality of related tasks;

adapting meta control variates parameters for control variates using a first portion of the episodes associated with a first portion of the tasks to generate first adapted meta control variates parameters;

generating first meta control variates estimates associated with a second portion of the tasks based on the first adapted meta control variates parameters;

adapting the meta control variates parameters using a second portion of the episodes associated with the second portion of the tasks to generate second adapted meta control variates parameters;

generating second meta control variates estimates associated with the first portion of the tasks based on the second adapted meta control variates parameters; and

updating the meta control variates parameters using the first and second meta control variates estimates.

15. A computing device comprising:

a memory; and

one or more processors coupled to the memory;

wherein the one or more processors are configured to:

receive a distribution including a plurality of related tasks; and

train parameters for a reinforcement learning neural network model based on gradient estimation associated with the parameters using samples associated with the plurality of related tasks,

wherein control variates are incorporated into the gradient estimation using a combination of a first gradient estimator and a second gradient estimator,

wherein the first gradient estimator uses the control variates and the second gradient estimator does not use the control variates,

wherein the control variates are represented by a baseline function depending on the tasks and corresponding states of the tasks,

wherein the gradient estimation is generated using Hessian estimation, and

wherein a product of two factors for the Hessian estimation based on the first gradient estimator and the second gradient estimator is reduced by subtracting the baseline function from one of the two factors.

16. The computing device of claim 15 ,

wherein the control variates are incorporated into the gradient estimation by automatic differentiation to reduce variance without introducing bias into the gradient estimation.

17. The computing device of claim 16 , wherein the automatic differentiation includes twice auto-differentiating for the Hessian estimation for the gradient estimation.

18. The computing device of claim 15 , wherein the one or more processors are further configured to:

train the control variates to generate trained control variates; and

incorporate the trained control variates into the gradient estimation.

19. The computing device of claim 18 , wherein the training the control variates includes:

for each task of the plurality of related tasks, training a separate control variate.

20. The computing device of claim 18 , wherein the training the control variates includes:

training the control variates across the plurality of related tasks.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0416 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2019
From: LIU, HAO
To: SALESFORCE.COM, INC.
Reel/Frame 049398/0976 →
Continuity (3)
Provisional Application 62796000 · Jan 23, 2019
Provisional Application 62793524 · Jan 17, 2019
Related Publication 20200234113A1 · Jul 23, 2020