IP Library Patent Application 17763914
Patent Application
App. No. 17/763,914

CONTROLLING AGENTS USING CAUSALLY CORRECT ENVIRONMENT MODELS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/763,914
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using an environment model to simulate state transitions of an environment being interacted with by an agent that is controlled using a policy neural network. One of the methods includes initializing an internal representation of a state of the environment at a current time point; repeatedly performing the following operations: receiving an action to be performed by the agent; generating, based on the internal representation, a predicted latent representation that is a prediction of a latent representation that would have been generated by the policy neural network by processing an observation characterizing the state of the environment corresponding to the internal representation; and updating the internal representation to simulate a state transition caused by the agent performing the received action by processing the predicted latent representation and the received action using the environment model.

Claims (52)

1 . A computer-implemented method of using an environment model to simulate state transitions of an environment being interacted with by an agent that is controlled using a policy neural network, wherein the policy neural network is configured to receive an observation characterizing a state of the environment, update a belief representation of the state of the environment, generate a latent representation from the belief representation, and generate an output specifying an action to be performed by the agent from the latent representation, and wherein the method comprises:

initializing an internal representation of a state of the environment at a current time point;

repeatedly performing the following operations:

receiving an action to be performed by the agent;

generating, based on the internal representation, a predicted latent representation that is a prediction of a latent representation that would have been generated by the policy neural network by processing an observation characterizing the state of the environment corresponding to the internal representation; and

updating the internal representation to simulate a state transition caused by the agent performing the received action by processing the predicted latent representation and the received action using the environment model.

2 . The method of claim 1 , further comprising:

generating, from the internal representation of the state of the environment, a target to be provided for use in controlling the agent.

3 . The method of claim 1 , wherein initializing an internal representation of a state of the environment at a current time point comprises:

receiving, by the policy neural network, an observation characterizing the state of the environment at the current time point;

updating, by the policy neural network and based on processing the received observation, a belief representation of the state of the environment; and

initializing the internal representation based on the belief representation of the state of the environment.

4 . The method of claim 1 , wherein updating the internal representation does not include processing the observation to be provided to the policy neural network that characterizes the state of the environment.

5 . The method of claim 1 , further comprising:

selecting, based on a result of repeatedly performing the operations, an action to be performed by the agent in the environment at the current time point.

6 . The method of claim 1 , further comprising:

processing, by the policy neural network, the belief representation of the state of the environment and the action that is performed by the agent to update the belief representation of the state of the environment at a future time point that is after the current time point.

7 . The method of claim 6 , wherein updating the belief representation of the state of the environment at the future time point further comprises processing an observation that characterizes the state of the environment at the future time point.

8 . The method of claim 1 , wherein the latent representation corresponds to one or more layers of the policy neural network after updating the belief representation of the state of the environment.

9 . The method of claim 8 , wherein the one or more layers comprise an input layer of the policy neural network after updating the belief representation of the state of the environment.

10 . The method of claim 1 , wherein the latent representation corresponds to respective probabilities generated by the policy neural network for controlling the agent to perform different actions.

11 . The method of claim 1 , wherein the latent representation corresponds to an intended action to be performed by the agent before selecting actions under exploration.

12 . The method of claim 1 , wherein the policy neural network and the environment model are each a respective neural network having a plurality of network parameters.

13 . The method of claim 12 , wherein the policy neural network and the environment model are each a recurrent neural network.

14 . The method of claim 1 , wherein the environment specified in the latent representation of the state of the environment corresponds to a partial view of the environment being interacted with by the agent.

15 . The method of claim 1 , wherein:

generating the latent representation from the belief representation comprises:

sampling from a distribution of a plurality of variables that describe the latent representation, the distribution being generated by the policy neural network and being conditioned on the belief representation; and

generating the predicted latent representation that is a prediction of the latent representation comprises:

sampling from a distribution of a plurality of variables that describe the latent representation, the distribution being generated by the environment model and being conditioned on the internal representation.

16 . The method of claim 1 , further comprising:

iteratively training the environment model on training data to determine trained values of the model parameters, wherein the training data includes observation received by the agent during interaction with the environment.

17 . The method of claim 16 , wherein training the environment model comprises, at each training iteration:

generating, by the environment model, a training predicted latent representation;

evaluating an objective function measuring a difference between the training predicted latent representation and the actual latent representation that is generated by the policy neural network; and

updating, based on a computed gradient of the objective function, corresponding values of the environment model parameters.

18 . (canceled)

19 . (canceled)

20 . A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for using an environment model to simulate state transitions of an environment being interacted with by an agent that is controlled using a policy neural network, wherein the policy neural network is configured to receive an observation characterizing a state of the environment, update a belief representation of the state of the environment, generate a latent representation from the belief representation, and generate an output specifying an action to be performed by the agent from the latent representation, and wherein the operations comprise:

initializing an internal representation of a state of the environment at a current time point;

repeatedly performing the following operations:

receiving an action to be performed by the agent;

generating, based on the internal representation, a predicted latent representation that is a prediction of a latent representation that would have been generated by the policy neural network by processing an observation characterizing the state of the environment corresponding to the internal representation; and

updating the internal representation to simulate a state transition caused by the agent performing the received action by processing the predicted latent representation and the received action using the environment model.

21 . The system of claim 20 , wherein the operations further comprise:

selecting, based on a result of repeatedly performing the operations, an action to be performed by the agent in the environment at the current time point.

21 . One or more non-transitory computer storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for using an environment model to simulate state transitions of an environment being interacted with by an agent that is controlled using a policy neural network, wherein the policy neural network is configured to receive an observation characterizing a state of the environment, update a belief representation of the state of the environment, generate a latent representation from the belief representation, and generate an output specifying an action to be performed by the agent from the latent representation, and wherein the operations comprise:

initializing an internal representation of a state of the environment at a current time point;

repeatedly performing the following operations:

receiving an action to be performed by the agent;

generating, based on the internal representation, a predicted latent representation that is a prediction of a latent representation that would have been generated by the policy neural network by processing an observation characterizing the state of the environment corresponding to the internal representation; and

updating the internal representation to simulate a state transition caused by the agent performing the received action by processing the predicted latent representation and the received action using the environment model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2022
From: DANIHELKA, IVO; REZENDE, DANILO JIMENEZ; GREGOR, KAROL; PAPAMAKARIOS, GEORGIOS; WEBER, THEOPHANE GUILLAUME
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 059428/0784 →