IP Library Granted Patent US 12705485
Granted Patent B2
US 12705485 · App. 18/919,108 · Granted Aug 11, 2026

Training action selection neural networks using look-ahead search

Inventors: Karen Simonyan (London, GB); David Silver (Hitchin, GB); Julian Schrittwieser (London, GB)
Assignee: GDM Holding LLC
G06N3/08G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705485
App. No.
18/919,108
Granted
Aug 11, 2026
Kind
B2
Abstract

Methods, systems and apparatus, including computer programs encoded on computer storage media, for training an action selection neural network. One of the methods includes receiving an observation characterizing a current state of the environment; determining a target network output for the observation by performing a look ahead search of possible future states of the environment starting from the current state until the environment reaches a possible future state that satisfies one or more termination criteria, wherein the look ahead search is guided by the neural network in accordance with current values of the network parameters; selecting an action to be performed by the agent in response to the observation using the target network output generated by performing the look ahead search; and storing, in an exploration history data store, the target network output in association with the observation for use in updating the current values of the network parameters.

Claims (43)

1 . A method performed by one or more computers and of selecting, using a neural network, actions to be performed in an attempt to achieve a specified result,

wherein the neural network has a plurality of network parameters and is configured to receive an input observation characterizing a state of an environment and to process the input observation in accordance with the network parameters to generate a network output that comprises an action selection output that defines an action selection policy for selecting an action to be performed in response to the input observation, and

wherein the method comprises:

receiving a current observation characterizing a current state of the environment;

determining a target action selection output for the current observation by performing, using the neural network and in accordance with values of the network parameters, a look ahead search of possible future states of the environment starting from the current state until the environment reaches a possible future state that satisfies one or more termination criteria, wherein the look ahead search is a tree search of a state tree having nodes representing states of the environment starting from a root node that represents the current state; and

selecting an action to be performed in response to the current observation using the target action selection output generated by performing the look ahead search.

2 . The method of claim 1 , wherein performing the look ahead search comprises evaluating leaf nodes of the state tree encountered during the look ahead search using the neural network and in accordance with the values of the network parameters.

3 . The method of claim 2 , wherein evaluating leaf nodes of the state tree encountered during the look ahead search using the trained neural network and in accordance with the values of the network parameters comprises, for each leaf node:

adding one or more new edges from the leaf node of the state tree;

processing a new observation characterizing a new state of the environment that is characterized by the leaf node using the neural network and in accordance with the values of the network parameters to generate a new action selection output; and

generating, using the new action selection output, a respective prior probability for each of the one or more new edges.

4 . The method of claim 1 , wherein the values of the network parameters have been determined by training the neural network using target network outputs determined by performing look ahead searches using the neural network.

5 . The method of claim 1 , wherein performing the look ahead search comprises determining a respective visit count for each of a plurality of outgoing edges from the root node, each outgoing edge representing a respective action to be performed by the agent.

6 . The method of claim 5 , wherein the target action selection output comprises a respective probability for each action that is represented by an outgoing edge from the root node, and wherein determining the target action selection output comprises determining the target action selection output from the respective visit counts for the outgoing edges.

7 . The method of claim 6 , wherein performing the look ahead search comprises traversing the state tree starting from the root node until encountering a leaf node by selecting edges to be traversed using adjusted action scores for edges in the state tree.

8 . The method of claim 6 , wherein selecting an action to be performed in response to the current observation using the target action selection output generated by performing the look ahead search comprises:

sampling an action using the respective probabilities for the actions.

9 . The method of claim 6 , wherein determining the target action selection output from the respective visit counts for the outgoing edges comprises applying a softmax over the respective visit counts for the outgoing edges.

10 . The method of claim 1 , further comprising: storing the current observation and the target action selection output for use in training the neural network.

11 . A method performed by one or more computers and of selecting, using a neural network, actions to be performed by an agent interacting with an environment to perform a task in an attempt to achieve a specified result,

wherein the neural network has a plurality of network parameters and is configured to receive an input characterizing a state of the environment and to process the input in accordance with the network parameters to generate a network output, and

wherein the method comprises:

receiving a current observation characterizing a current state of the environment;

determining a target action selection output for the current observation by performing, using the neural network and in accordance with values of the network parameters, a look ahead search of possible future states of the environment starting from the current state until the environment reaches a possible future state that satisfies one or more termination criteria, wherein the look ahead search is a tree search of a state tree having nodes representing states of the environment starting from a root node that represents the current state; and

selecting an action to be performed by the agent in response to the current observation using the target action selection output generated by performing the look ahead search.

12 . The method of claim 11 , wherein performing the look ahead search comprises evaluating leaf nodes of the state tree encountered during the look ahead search using the neural network and in accordance with the values of the network parameters.

13 . The method of claim 12 , wherein evaluating leaf nodes of the state tree encountered during the look ahead search using the trained neural network and in accordance with the values of the network parameters comprises, for each leaf node:

adding one or more new edges from the leaf node of the state tree;

processing a new observation characterizing a new state of the environment that is characterized by the leaf node using the neural network and in accordance with the values of the network parameters to generate a new action selection output; and

generating, using the new action selection output, a respective prior probability for each of the one or more new edges.

14 . The method of claim 11 , wherein the values of the network parameters have been determined by training the neural network using target network outputs determined by performing look ahead searches using the neural network.

15 . The method of claim 11 , wherein performing the look ahead search comprises determining a respective visit count for each of a plurality of outgoing edges from the root node, each outgoing edge representing a respective action to be performed by the agent.

16 . The method of claim 15 , wherein the target action selection output comprises a respective probability for each action that is represented by an outgoing edge from the root node, and wherein determining the target action selection output comprises determining the target action selection output from the respective visit counts for the outgoing edges.

17 . The method of claim 16 , wherein performing the look ahead search comprises traversing the state tree starting from the root node until encountering a leaf node by selecting edges to be traversed using adjusted action scores for edges in the state tree.

18 . The method of claim 16 , wherein selecting an action to be performed in response to the current observation using the target action selection output generated by performing the look ahead search comprises:

sampling an action using the respective probabilities for the actions.

19 . The method of claim 16 , wherein determining the target action selection output from the respective visit counts for the outgoing edges comprises applying a softmax over the respective visit counts for the outgoing edges.

20 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising selecting, using a neural network, actions to be performed in an attempt to achieve a specified result,

wherein the neural network has a plurality of network parameters and is configured to receive an input observation characterizing a state of an environment and to process the input observation in accordance with the network parameters to generate a network output that comprises an action selection output that defines an action selection policy for selecting an action to be performed in response to the input observation, and

wherein the method comprises:

receiving a current observation characterizing a current state of the environment;

determining a target action selection output for the current observation by performing, using the neural network and in accordance with values of the network parameters, a look ahead search of possible future states of the environment starting from the current state until the environment reaches a possible future state that satisfies one or more termination criteria, wherein the look ahead search is a tree search of a state tree having nodes representing states of the environment starting from a root node that represents the current state; and

selecting an action to be performed in response to the current observation using the target action selection output generated by performing the look ahead search.