IP Library Patent Application 18662481
Patent Application
App. No. 18/662,481

REINFORCEMENT LEARNING USING BASELINE AND POLICY NEURAL NETWORKS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/662,481
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for asynchronous deep reinforcement learning. One of the systems includes a plurality of workers, wherein each worker is configured to operate independently of each other worker, and wherein each worker is associated with a respective actor that interacts with a respective replica of the environment during the training of the deep neural network.

Claims (232)

1 .- 20 . (canceled)

21 . A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

training a deep neural network used to control an agent that interacts with an environment by performing actions selected from a set of actions, the deep neural network comprising a policy neural network having a plurality of policy parameters and a baseline neural network having a plurality of baseline parameters, the training comprising:

generating training trajectories using a plurality of worker computing units, wherein each worker computing unit is configured to operate independently of each other worker computing unit, wherein each worker computing unit is associated with a respective actor that interacts with a respective replica of the environment during the training of the deep neural network, and wherein each worker computing unit repeatedly performs operations comprising:

determining current values of the parameters of the deep neural network, comprising:

receiving, from a shared memory accessible by each of the worker computing units, current values of the baseline parameters of the baseline neural network; and

receiving, from the shared memory, current values of the policy parameters of the policy neural network;

until the environment replica interacted with by the actor associated with the worker computing unit transitions into a state that satisfies particular criteria:

receiving observations characterizing states of the environment replica and, in response to each observation, selecting an action to be performed by the actor associated with the worker computing unit in response to the observation in accordance with a respective action selection policy for the worker computing unit, wherein the selecting comprises, for each observation:

processing the observation using the policy neural network and in accordance with the current values of the parameters of the policy neural network to generate a respective score for each action in the predetermined set of actions in accordance with the current values of the parameters of the policy neural network, and

selecting the current action to be performed by the actor using the respective scores;

wherein the observations characterizing states of the environment replica and the actions selected to be performed in response to the observations collectively define one or more training trajectories;

training the deep neural network on the training trajectories generated using the plurality of worker computing units, comprising, for each training trajectory:

for each observation received before a last observation in the training trajectory:

generating, based on the current values of the parameters of the baseline network, a baseline score representing an estimated reward received by the agent starting from a state characterized by the observation;

identifying an actual reward resulting from the actor performing the selected action win response to the observation;

determining an actual long-term reward from at least the actual reward and a baseline score for the observation characterizing the terminal state; and

performing an iteration of a reinforcement learning technique to determine respective current gradients for the baseline neural network and policy neural network based on the baseline score and the actual long-term reward;

updating the parameters of the deep neural network using the current gradients.

22 . The system of claim 21 , wherein each worker executes independently of each other worker on the same computer.

23 . The system of claim 21 , wherein the respective action selection policy for each worker is different from the respective action selection policy for each other worker.

24 . The system of claim 21 , wherein determining the actual long-term reward comprises:

for the last observation in the training trajectory, setting the actual long-term reward based on the actual reward for the observation and the baseline score for the observation; and

for each other observation before the last observation in the training trajectory, setting the actual long-term reward for the observation as the actual reward plus a discounted sum of the actual long-term rewards for observations received after the observation.

25 . The system of claim 21 , wherein the gradient update for the policy neural network is determined by:

θ

log

π

(

a

t

|

s

t

;

θ

)

?

(

R

t

-

b

t

)

?

indicates text missing or illegible when filed

wherein π(a t |s t ; θ′) represents the score for the action performed in response to the observation s t determined in accordance with the current values of the parameters of the policy neural network, R t is the actual reward and b t is the baseline score.

26 . The system of claim 25 , wherein the gradient update for the baseline neural network is determined by:

(

R

t

-

b

(

s

t

;

θ

b

)

)

2

/

θ

b

wherein b(s t ; θ′ b ) is the baseline score for the observation s t .

27 . The system of claim 21 , wherein the operations further comprise, after training the deep neural network, using the deep neural network to control a mechanical agent interacting with a real-world environment.

28 . A method performed by one or more computers, the method comprising:

training a deep neural network used to control an agent that interacts with an environment by performing actions selected from a set of actions, the deep neural network comprising a policy neural network having a plurality of policy parameters and a baseline neural network having a plurality of baseline parameters, the training comprising:

generating training trajectories using a plurality of worker computing units, wherein each worker computing unit is configured to operate independently of each other worker computing unit, wherein each worker computing unit is associated with a respective actor that interacts with a respective replica of the environment during the training of the deep neural network, and wherein each worker computing unit repeatedly performs operations comprising:

determining current values of the parameters of the deep neural network, comprising:

receiving, from a shared memory accessible by each of the worker computing units, current values of the baseline parameters of the baseline neural network; and

receiving, from the shared memory, current values of the policy parameters of the policy neural network;

until the environment replica interacted with by the actor associated with the worker computing unit transitions into a state that satisfies particular criteria:

receiving observations characterizing states of the environment replica and, in response to each observation, selecting an action to be performed by the actor associated with the worker computing unit in response to the observation in accordance with a respective action selection policy for the worker computing unit, wherein the selecting comprises, for each observation:

processing the observation using the policy neural network and in accordance with the current values of the parameters of the policy neural network to generate a respective score for each action in the predetermined set of actions in accordance with the current values of the parameters of the policy neural network, and

selecting the current action to be performed by the actor using the respective scores;

wherein the observations characterizing states of the environment replica and the actions selected to be performed in response to the observations collectively define one or more training trajectories;

training the deep neural network on the training trajectories generated using the plurality of worker computing units, comprising, for each training trajectory:

for each observation received before a last observation in the training trajectory:

generating, based on the current values of the parameters of the baseline network, a baseline score representing an estimated reward received by the agent starting from a state characterized by the observation;

identifying an actual reward resulting from the actor performing the selected action win response to the observation;

determining an actual long-term reward from at least the actual reward and a baseline score for the observation characterizing the terminal state; and

performing an iteration of a reinforcement learning technique to determine respective current gradients for the baseline neural network and policy neural network based on the baseline score and the actual long-term reward;

updating the parameters of the deep neural network using the current gradients.

29 . The method of claim 28 , wherein each worker executes independently of each other worker on the same computer.

30 . The method of claim 28 , wherein the respective action selection policy for each worker is different from the respective action selection policy for each other worker.

31 . The method of claim 28 , wherein determining the actual long-term reward comprises:

for the last observation in the training trajectory, setting the actual long-term reward based on the actual reward for the observation and the baseline score for the observation; and

for each other observation before the last observation in the training trajectory, setting the actual long-term reward for the observation as the actual reward plus a discounted sum of the actual long-term rewards for observations received after the observation.

32 . The method of claim 28 , wherein the gradient update for the policy neural network is determined by:

θ

log

π

(

a

t

|

s

t

;

θ

)

?

(

R

t

-

b

t

)

?

indicates text missing or illegible when filed

wherein π(a t |s t ; θ′) represents the score for the action performed in response to the observation s t determined in accordance with the current values of the parameters of the policy neural network, R t is the actual reward and b t is the baseline score.

33 . The method of claim 32 , wherein the gradient update for the baseline neural network is determined by:

(

R

t

-

b

(

s

t

;

θ

b

)

)

2

/

θ

b

wherein b(s t ;θ′ b ) is the baseline score for the observation s t .

34 . The method of claim 28 , wherein the operations further comprise, after training the deep neural network, using the deep neural network to control a mechanical agent interacting with a real-world environment.

35 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

training a deep neural network used to control an agent that interacts with an environment by performing actions selected from a set of actions, the deep neural network comprising a policy neural network having a plurality of policy parameters and a baseline neural network having a plurality of baseline parameters, the training comprising:

generating training trajectories using a plurality of worker computing units, wherein each worker computing unit is configured to operate independently of each other worker computing unit, wherein each worker computing unit is associated with a respective actor that interacts with a respective replica of the environment during the training of the deep neural network, and wherein each worker computing unit repeatedly performs operations comprising:

determining current values of the parameters of the deep neural network, comprising:

receiving, from a shared memory accessible by each of the worker computing units, current values of the baseline parameters of the baseline neural network; and

receiving, from the shared memory, current values of the policy parameters of the policy neural network;

until the environment replica interacted with by the actor associated with the worker computing unit transitions into a state that satisfies particular criteria:

receiving observations characterizing states of the environment replica and, in response to each observation, selecting an action to be performed by the actor associated with the worker computing unit in response to the observation in accordance with a respective action selection policy for the worker computing unit, wherein the selecting comprises, for each observation:

processing the observation using the policy neural network and in accordance with the current values of the parameters of the policy neural network to generate a respective score for each action in the predetermined set of actions in accordance with the current values of the parameters of the policy neural network, and

selecting the current action to be performed by the actor using the respective scores;

wherein the observations characterizing states of the environment replica and the actions selected to be performed in response to the observations collectively define one or more training trajectories;

training the deep neural network on the training trajectories generated using the plurality of worker computing units, comprising, for each training trajectory:

for each observation received before a last observation in the training trajectory:

generating, based on the current values of the parameters of the baseline network, a baseline score representing an estimated reward received by the agent starting from a state characterized by the observation;

identifying an actual reward resulting from the actor performing the selected action win response to the observation;

determining an actual long-term reward from at least the actual reward and a baseline score for the observation characterizing the terminal state; and

performing an iteration of a reinforcement learning technique to determine respective current gradients for the baseline neural network and policy neural network based on the baseline score and the actual long-term reward;

updating the parameters of the deep neural network using the current gradients.

36 . The one or more non-transitory computer storage media of claim 35 , wherein each worker executes independently of each other worker on the same computer.

37 . The one or more non-transitory computer storage media of claim 35 , wherein the respective action selection policy for each worker is different from the respective action selection policy for each other worker.

38 . The one or more non-transitory computer storage media of claim 35 , wherein determining the actual long-term reward comprises:

for the last observation in the training trajectory, setting the actual long-term reward based on the actual reward for the observation and the baseline score for the observation; and

for each other observation before the last observation in the training trajectory, setting the actual long-term reward for the observation as the actual reward plus a discounted sum of the actual long-term rewards for observations received after the observation.

39 . The one or more non-transitory computer storage media of claim 35 , wherein the gradient update for the policy neural network is determined by:

θ

log

π

(

a

t

s

t

;

θ

)

?

(

R

t

-

b

t

)

?

indicates text missing or illegible when filed

wherein π(a t |s t ; θ′) represents the score for the action performed in response to the observation s t determined in accordance with the current values of the parameters of the policy neural network, R t is the actual reward and b t is the baseline score.

40 . The one or more non-transitory computer storage media of claim 39 , wherein the gradient update for the baseline neural network is determined by:

(

R

t

-

b

(

s

t

;

θ

b

)

)

2

/

θ

b

wherein b(s t ; θ′ b ) is the baseline score for the observation s t .

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2025
From: MNIH, VOLODYMYR; BADIA, ADRIÀ PUIGDOMÈNECH; GRAVES, ALEXANDER BENJAMIN; HARLEY, TIMOTHY JAMES ALEXANDER; SILVER, DAVID; KAVUKCUOGLU, KORAY
To: GOOGLE INC.
Reel/Frame 070647/0235 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2025
From: GOOGLE INC.
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 070647/0344 →