IP Library › Granted Patent US 12,265,795
Granted Patent B2
US 12,265,795 · App. 18/649,774 · Granted Apr 1, 2025

Action selection based on environment observations and textual instructions

Inventors: Karl Moritz Hermann (Berlin, DE); Philip Blunsom (Oxford, GB); Felix George Hill (London, GB)
Assignee: DeepMind Technologies Limited
G06F40/30G06F17/16G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,795
App. No.
18/649,774
Granted
Apr 1, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting actions to be performed by an agent interacting with an environment. In one aspect, a system includes a language encoder model that is configured to receive a text string in a particular natural language, and process the text string to generate a text embedding of the text string. The system includes an observation encoder neural network that is configured to receive an observation characterizing a state of the environment, and process the observation to generate an observation embedding of the observation. The system includes a subsystem that is configured to obtain a current text embedding of a current text string and a current observation embedding of a current observation. The subsystem is configured to select an action to be performed by the agent in response to the current observation.

Claims (69)

1. A method performed by one or more computers for selecting actions to be performed by an agent interacting with an environment, the method comprising:

at each of a plurality of time steps:

receiving a current text string in a natural language that expresses information about a current task being performed by the agent;

receiving a current observation characterizing a current state of the environment;

processing an input comprising the current text string and the current observation using a neural network to generate an action selection output, the processing comprising:

combining, by the neural network and in accordance with values of a set of neural network parameters, the current text string and the current observation to produce a combined embedding; and

generating, by the neural network and in accordance with the values of the set of neural network parameters, the action selection output based on the combined embedding; and

selecting an action to be performed by the agent at the time step based on the action selection output;

wherein the neural network has been trained from end-to-end using a machine learning training technique.

2. The method of claim 1 , further comprising:

receiving, at each of the plurality of time steps, a current reward as a result of the agent performing the action in response to the current observation; and

training the neural network from end-to-end using reinforcement learning based on the rewards received over the plurality of time steps.

3. The method of claim 1 , wherein combining, by the neural network and in accordance with the values of the set of neural network parameters, the current text string and the current observation to produce the combined embedding comprises:

processing the current text string using a language encoder model of the neural network to generate a current text embedding of the current text string;

processing the current observation using an observation encoder neural network of the neural network to generate a current observation embedding of the current observation; and

combining the current observation embedding and the current text embedding to generate the combined embedding.

4. The method of claim 1 , wherein generating, by the neural network and in accordance with the values of the set of neural network parameters, the action selection output based on the combined embedding comprises:

processing the combined embedding using an action selection neural network of the neural network to generate the action selection output.

5. The method of claim 3 , wherein the language encoder model is a recurrent neural network.

6. The method of claim 3 , wherein the language encoder model is a bag-of-words encoder.

7. The method of claim 3 , wherein the current observation embedding is a feature matrix of the current observation, and wherein the current text embedding is a feature vector of the current text string.

8. The method of claim 7 , wherein combining the current observation embedding and the current text embedding comprises:

flattening the feature matrix of the current observation; and

concatenating the flattened feature matrix and the feature vector of the current text string.

9. The method of claim 1 , wherein at each of the plurality of time steps, the current text string is a natural language instruction for the agent for performing the current task.

10. The method of claim 1 , wherein at each of the plurality of time steps:

the action selection output defines a probability distribution over possible actions to be performed by the agent; and

selecting the action to be performed by the agent comprises:

sampling an action from the probability distribution or selecting an action having a highest probability according to the probability distribution.

11. The method of claim 1 , wherein at each of the plurality of time steps:

the action selection output comprises, for each of a plurality of possible actions to be performed by the agent, a respective Q value that is an estimate of a return resulting from the agent performing the possible action in response to the current observation; and

selecting the action to be performed by the agent comprises:

selecting an action having a highest Q value.

12. The method of claim 1 , wherein at each of the plurality of time steps:

the action selection output identifies a best possible action to be performed by the agent in response to the current observation; and

selecting the action to be performed by the agent comprises:

selecting the best possible action.

13. The method of claim 1 , wherein the current text string is the same for each observation received during the performance of the current task.

14. The method of claim 1 , wherein the current text string is different from a preceding text string received during the performance of the current task.

15. A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for selecting actions to be performed by an agent interacting with an environment, the operations comprising:

at each of a plurality of time steps:

receiving a current text string in a natural language that expresses information about a current task being performed by the agent;

receiving a current observation characterizing a current state of the environment;

processing an input comprising the current text string and the current observation using a neural network to generate an action selection output, the processing comprising:

combining, by the neural network and in accordance with values of a set of neural network parameters, the current text string and the current observation to produce a combined embedding; and

generating, by the neural network and in accordance with the values of the set of neural network parameters, the action selection output based on the combined embedding; and

selecting an action to be performed by the agent at the time step based on the action selection output;

wherein the neural network has been trained from end-to-end using a machine learning training technique.

16. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for selecting actions to be performed by an agent interacting with an environment, the operations comprising:

at each of a plurality of time steps:

receiving a current text string in a natural language that expresses information about a current task being performed by the agent;

receiving a current observation characterizing a current state of the environment;

processing an input comprising the current text string and the current observation using a neural network to generate an action selection output, the processing comprising:

combining, by the neural network and in accordance with values of a set of neural network parameters, the current text string and the current observation to produce a combined embedding; and

generating, by the neural network and in accordance with the values of the set of neural network parameters, the action selection output based on the combined embedding; and

selecting an action to be performed by the agent at the time step based on the action selection output;

wherein the neural network has been trained from end-to-end using a machine learning training technique.

17. The non-transitory computer storage media of claim 16 , wherein the operations further comprise:

receiving, at each of the plurality of time steps, a current reward as a result of the agent performing the action in response to the current observation; and

training the neural network from end-to-end using reinforcement learning based on the rewards received over the plurality of time steps.

18. The non-transitory computer storage media of claim 16 , wherein combining, by the neural network and in accordance with the values of the set of neural network parameters, the current text string and the current observation to produce the combined embedding comprises:

processing the current text string using a language encoder model of the neural network to generate a current text embedding of the current text string;

processing the current observation using an observation encoder neural network of the neural network to generate a current observation embedding of the current observation; and

combining the current observation embedding and the current text embedding to generate the combined embedding.

19. The non-transitory computer storage media of claim 17 , wherein generating, by the neural network and in accordance with the values of the set of neural network parameters, the action selection output based on the combined embedding comprises:

processing the combined embedding using an action selection neural network of the neural network to generate the action selection output.

20. The non-transitory computer storage media of claim 18 , wherein the language encoder model is a recurrent neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: HERMANN, KARL MORITZ; BLUNSOM, PHILIP; HILL, FELIX GEORGE
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 067357/0654 →
Continuity (4)
Continuation 17744921 · May 16, 2022
Continuation 16497602
Provisional Application 62515458 · Jun 5, 2017
Related Publication 20240320438A1 · Sep 26, 2024
References Cited (80)
US 5680557A · Karamchetty · 1997 [cited by examiner]
US 5963966A · Mitchell · 1999 [cited by examiner]
US 9754221B1 · Nagaraja · 2017 [cited by examiner]
US 10645073B1 · Agarmore · 2020 [cited by examiner]
US 11791914B2 · Cella · 2023 [cited by examiner]
US 20120044250A1 · Landers · 2012 [cited by examiner]
US 20130145241A1 · Salama · 2013 [cited by examiner]
US 20140355861A1 · Nirenberg · 2014 [cited by examiner]
US 20170124432A1 · Chen · 2017 [cited by examiner]
US 20170337682A1 · Liao · 2017 [cited by examiner]
US 20170372696A1 · Lee · 2017 [cited by examiner]
US 20180012159A1 · Kozloski · 2018 [cited by examiner]
US 20180060301A1 · Li · 2018 [cited by examiner]
US 20180061074A1 · Yamamichi · 2018 [cited by examiner]
US 20180129742A1 · Li · 2018 [cited by examiner]
US 20180129938A1 · Xiong · 2018 [cited by examiner]
US 20180329887A1 · Bull · 2018 [cited by examiner]
US 20180329998A1 · Thomson · 2018 [cited by examiner]
US 20210110115A1 · Hermann · 2021 [cited by examiner]
US 20210390270A1 · Fei · 2021 [cited by examiner]
US 20220020355A1 · Ming · 2022 [cited by examiner]
US 20220318516A1 · Hermann · 2022 [cited by examiner]
CN 103430232 · 2013 [cited by applicant]
CN 106056213 · 2016 [cited by applicant]
WO WO2018224471 · 2018 [cited by applicant]
Arumugam et al., “Accurately and efficiently interpreting human-robot instructions of varying granularities,” arXiv, Jun. 2018, 10 pages. [cited by applicant]
Beattie et al., “DeepMind Lab,” arXiv, Dec. 2016, 11 pages. [cited by applicant]
Berant et al., “Semantic Parsing on Freebase from Question-Answer Pairs,” Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Oct. 2013, 12 pages. [cited by applicant]
Chen et al., “Learning to Sportscast: A Test of Grounded Language Acquisition,” Proceedings of the 25th international conference on Machine learning, Jul. 2008, 8 pages. [cited by applicant]
Chomsky, “A review of BF Skinner's Verbal Behavior,” Readings in philosophy of psychology 1, 1980, 29 pages. [cited by applicant]
De Anda et al., “Lexical access in the second year: a cross-linguistic study of monolingual and bilingual vocabulary development,” San Diego Linguistic Papers 6 (2016), 2016, 16 pages. [cited by applicant]
Doumas et al., “A Theory of the Discovery and Predication of Relational Concepts,” Psychological review 115, Jan. 2008, 43 pages. [cited by applicant]
Extended Search Report in European Appln. No. 23198528.4, dated Jan. 16, 2024, 8 pages. [cited by applicant]
Fernald et al., “Blue car, red car: Developing efficiency in online interpretation of adjective-noun phrases,” Cognitive psychology 60.3 (2010), May 2010, 34 pages. [cited by applicant]
Frank et al., “Social and discourse contributions to the determination of reference in cross-situational word learning,” Language Learning and Development 9.1 (2013), Jan. 2013, 24 pages. [cited by applicant]
Harnad, “The symbol grounding problem,” Physica D: Nonlinear Phenomena 42, 12 pages. [cited by applicant]
Hemachandra et al., “Learning Spatial-Semantic Representations from Natural Language Descriptions and Scene Classifications,” 2014 IEEE International Conference on Robotics and Automation (ICRA), May 2014, 8 pages. [cited by applicant]
Hermann et al., “Grounded Language Learning in a Simulated 3D World,” arXiv, Jun. 2017, 22 pages. [cited by applicant]
Hochreiter et al., “Long short-term memory,” Neural computation, Nov. 1997, 32 pages. [cited by applicant]
Jaderberg et al., “Reinforcement Learning with Unsupervised Auxiliary Tasks,” arXiv, Nov. 2016, 14 pages. [cited by applicant]
Kaplan et al., “Beating atari with natural language guided reinforcement learning,” arXiv, Apr. 2017, 13 pages. [cited by applicant]
Krening et al., “Learning from explanations using sentiment and advice in RL,” IEEE Transactions on Cognitive and Developmental Systems 9.1, Nov. 2016, 12 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in neural information processing systems, Dec. 2012, 9 pages. [cited by applicant]
LeCun et al., “Backpropagation applied to handwritten zip code recognition,” Neural computation, Dec. 1989, 11 pages. [cited by applicant]
McClelland et al, “The appeal of parallel distributed processing,” IEEE, 1988, 1:43 pages. [cited by applicant]
McMurray, “Defusing the childhood vocabulary explosion,” Science 317.5838 (2007), Aug. 2007, 1 page. [cited by applicant]
Mikolov et al., “A roadmap towards machine intelligence,” arXiv, Feb. 2016, 36 pages. [cited by applicant]
Mirowski et al., “Learning to Navigate in Complex Environments,” arXiv, Jan. 2017, 16 pages. [cited by applicant]
Mnih et al., “Asynchronous methods for deep reinforcement learning,” International conference on machine learning, Jun. 2016, 10 pages. [cited by applicant]
Mnih et al., “Human-level control through deep reinforcement learning,” Nature 518, Feb. 2015, 13 pages. [cited by applicant]
Narasimhan et al., “Language understanding for text-based games using deep reinforcement learning,” arXiv, Sep. 2015, 11 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201880026852.4, dated Oct. 21, 2022, 14 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 18729406.1, dated Jun. 30, 2021, 13 pages. [cited by applicant]
Oh et al., “Action-Conditional Video Prediction using Deep Networks in Atari Games,” arXiv, Dec. 2015, 26 pages. [cited by applicant]
PCT International Preliminary Report on Patentability in International Appln. No. PCT/EP2018/064703, dated Dec. 10, 2019, 16 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Appln. No. PCT/EP2018/06470, mailed on Sep. 7, 2018, 22 pages. [cited by applicant]
Plunkett, “Lexical segmentation and vocabulary growth in early language acquisition,” Journal of Child Language 20.1, Feb. 1993, 18 pages. [cited by applicant]
Poulin-Dubois et al., “Lexical access and vocabulary development in very young bilinguals,” International Journal of Bilingualism, Feb. 2013, 17(1):57-70. [cited by applicant]
Quinn et al., “Evidence for representations of perceptually similar natural categories by 3-month-old and 4-month-old infants,” Perception, Apr. 1993, 22:463-475. [cited by applicant]
Rowe, “Child-directed speech: Relation to socioeconomic status, knowledge of child development and child vocabulary skill,” Journal of child language 35.1 (2008), Feb. 2008, 22 pages. [cited by applicant]
Roy et al., “Learning words from sights and sounds: A computational model,” Cognitive science, Jan. 2002, 113-146. [cited by applicant]
Searle, “Minds, brains, and programs,” Behavioral and brain sciences 3.3 (1980), Sep. 1980, 19 pages. [cited by applicant]
Shelhamer et al., “Loss is its own Reward: Self-Supervision for Reinforcement Learning,” arXiv, Mar. 2017, 9 pages. [cited by applicant]
Silberer et al., “Grounded models of semantic representation,” Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, Jul. 2012, 1423-14… [cited by applicant]
Siskind, “Grounding Language in Perception,” Artificial Intelligence Review, Sep. 1994, 21 pages. [cited by applicant]
Siskind, “Grounding the lexical semantics of verbs in visual perception using force dynamics and event logic,” Journal of artificial intelligence research 15, Aug. 2001, 31-90. [cited by applicant]
Smith et al., “Infants rapidly learn word-referent mappings via cross-situational statistics,” Cognition 106.3 (2008), Mar. 2008, 12 pages. [cited by applicant]
Smith et al., “Naming in young children: A dumb attentional mechanism?” Cognition 60, Aug. 1996, 29 pages. [cited by applicant]
Steels, “The symbol grounding problem has been solved. so what's next,” Symbols and embodiment: Debates on meaning and cognition, Aug. 2007, 21 pages. [cited by applicant]
Thomason et al., “Learning to interpret natural language commands through human-robot dialog,” Twenty-Fourth International Joint Conference on Artificial Intelligence, Jun. 2015, 1923-1929. [cited by applicant]
Vendrov et al., “Order-embeddings of images and language,” arXiv, Mar. 2016, 12 pages. [cited by applicant]
Vosniadou et al., “Mental models of the earth: A study of conceptual change in childhood,” Cognitive psychology 24, Oct. 1992, 535-585. [cited by applicant]
Walter et al., “A framework for learning semantic maps from grounded natural language descriptions,” International Journal of Robotics Research 33.9, Aug. 2014, 23 pages. [cited by applicant]
Wang et al., “Learning language games through interaction,” arXiv, Jun. 2016, 11 pages. [cited by applicant]
Weisleder et al., “Talking to children matters: Early language experience strengthens processing and builds vocabulary,” Psychological science 24.11 (2013), Nov. 2013, 14 pages. [cited by applicant]
Winograd, “Understanding natural language,” Cognitive psychology 3.1 (1972), Jan. 1972, 191 pages. [cited by applicant]
Xu et al., “Show, attend and tell: Neural image caption generation with visual attention,” International conference on machine learning, Jun. 2015, 10 pages. [cited by applicant]
Yu et al., “A Deep Compositional Framework for Human-like Language Acquisition in Virtual Environment,” arXiv, May 2017, 16 pages. [cited by applicant]
Yu et al., “Grounded language learning from video described with sentences,” Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), Aug. 2013, 11 pages. [cited by applicant]
Zettlemoyer et al., “Learning to map sentences to logical form: Structured classification with probabilistic categorial grammars,” arXiv, Jul. 2012, 9 pages. [cited by applicant]