IP Library Granted Patent US 12,468,779
Granted Patent B1
US 12,468,779 · App. 18/656,462 · Granted Nov 11, 2025

Training action-selection neural networks from demonstrations using multiple losses

Inventor: Todd Andrew Hester (Seattle, WA)
Assignee: GDM Holding LLC
G06F18/2148G06F17/11G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,779
App. No.
18/656,462
Granted
Nov 11, 2025
Kind
B1
Abstract

A method of training an action selection neural network to perform a demonstrated task using a supervised learning technique. The action selection neural network is configured to receive demonstration data comprising actions to perform the task and rewards received for performing the actions. The action selection neural network has auxiliary prediction task neural networks on one or more of its intermediate outputs. The action selection policy neural network is trained using multiple combined losses, concurrently with the auxiliary prediction task neural networks.

Claims (39)

1 . A method of training a neural network system, the method comprising:

training an action selection neural network to perform a task,

wherein the action selection neural network is configured to receive inputs comprising observations of an environment and to process the inputs to generate action selection outputs indicating actions to perform the task, and

during training of the action selection neural network:

training an auxiliary prediction task neural network, wherein the auxiliary prediction task neural network is configured to receive an intermediate output from the action selection neural network and to generate a prediction output which indicates a predicted characteristic of the task,

wherein training the auxiliary prediction task neural network comprises training the auxiliary prediction task neural network and the action selection neural network using demonstration data for the task by backpropagating gradients determined from an auxiliary learning loss function through the auxiliary prediction task neural network and into the action selection neural network to bring the predicted characteristic closer to a corresponding observed characteristic of the task from the demonstration data.

2 . A method as claimed in claim 1 , further comprising, during training of the action selection neural network:

training the action selection neural network to select actions to be performed by an agent interacting with the environment to perform the task using a reinforcement learning technique.

3 . A method as claimed in claim 1 wherein the auxiliary prediction task neural network is configured to predict a demonstrated action from the demonstration data at a subsequent observation to a current observation, and wherein the predicted characteristic comprises a predicted demonstrated action from the demonstration data.

4 . A method as claimed in claim 1 wherein training the action selection neural network to perform the task comprises training the action selection neural network using both a supervised learning technique and a reinforcement learning technique.

5 . A method as claimed in claim 1 wherein the auxiliary prediction task neural network is configured to predict one or more Q-values.

6 . A method as claimed in claim 5 wherein the one or more Q-values comprise a time-discounted Q-value characterizing a future state of the environment.

7 . A method as claimed in claim 1 wherein the auxiliary prediction task neural network is configured to predict a reward from the environment at a subsequent observation to a current observation.

8 . A method as claimed in claim 1 wherein the action selection neural network has a policy output to determine actions to be performed by an agent once trained, wherein the policy output defines, each possible action, a probability distribution over a set of possible returns, and wherein training the action selection neural network comprises estimating the probability distribution over the set of possible returns for each of the possible actions.

9 . A method as claimed in claim 1 wherein the auxiliary prediction task neural network is configured to predict termination of the demonstrated task.

10 . A method as claimed in claim 1 wherein one or both of action selection network parameters and auxiliary prediction task network parameters comprise neural network weights including noise characterizing parameters, and wherein training the action selection neural network further comprises adjusting values of the noise characterizing parameters.

11 . A method as claimed in claim 1 further comprising using the trained neural network system to perform the task.

12 . A method as claimed in claim 11 wherein the trained neural network system learns solely from the demonstration data prior to using the trained neural network system to perform the task.

13 . A method as claimed in claim 1 comprising generating the demonstration data by controlling or manipulating a robot or other mechanical agent, and training the neural network system using the demonstration data to control the robot or other mechanical agent to perform the task.

14 . A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of training a neural network system using demonstrations, the operations comprising:

training an action selection neural network to perform a task,

wherein the action selection neural network is configured to receive inputs comprising observations of an environment and to process the inputs to generate action selection outputs indicating actions to perform the task, and

during training of the action selection neural network:

training an auxiliary prediction task neural network, wherein the auxiliary prediction task neural network is configured to receive an intermediate output from the action selection neural network and to generate a prediction output which indicates a predicted characteristic of the task,

wherein training the auxiliary prediction task neural network comprises training the auxiliary prediction task neural network and the action selection neural network using demonstration data for the task by backpropagating gradients determined from an auxiliary learning loss function through the auxiliary prediction task neural network and into the action selection neural network to bring the predicted characteristic closer to a corresponding observed characteristic of the task from the demonstration data.

15 . A system as claimed in claim 14 , wherein the operations further comprise, during training of the action selection neural network:

training the action selection neural network to select actions to be performed by an agent interacting with the environment to perform the task using a reinforcement learning technique.

16 . A system as claimed in claim 14 wherein the auxiliary prediction task neural network is configured to predict a demonstrated action from the demonstration data at a subsequent observation to a current observation, and wherein the predicted characteristic comprises a predicted demonstrated action from the demonstration data.

17 . A system as claimed in claim 14 wherein training the action selection neural network to perform the task comprises training the action selection neural network using both a supervised learning technique and a reinforcement learning technique.

18 . A system as claimed in claim 14 wherein the auxiliary prediction task neural network is configured to predict one or more Q-values.

19 . A system as claimed in claim 18 wherein the one or more Q-values comprise a time-discounted Q-value characterizing a future state of the environment.

20 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of training a neural network system using demonstrations, the operations comprising:

training an action selection neural network to perform a task,

wherein the action selection neural network is configured to receive inputs comprising observations of an environment and to process the inputs to generate action selection outputs indicating actions to perform the task, and

during training of the action selection neural network:

training an auxiliary prediction task neural network, wherein the auxiliary prediction task neural network is configured to receive an intermediate output from the action selection neural network and to generate a prediction output which indicates a predicted characteristic of the task,

wherein training the auxiliary prediction task neural network comprises training the auxiliary prediction task neural network and the action selection neural network using demonstration data for the task by backpropagating gradients determined from an auxiliary learning loss function through the auxiliary prediction task neural network and into the action selection neural network to bring the predicted characteristic closer to a corresponding observed characteristic of the task from the demonstration data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: HESTER, TODD ANDREW
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 067366/0059 →
Continuity (3)
Continuation 18120912 · Mar 13, 2023
Continuation 16174148 · Oct 29, 2018
Provisional Application 62578367 · Oct 27, 2017
References Cited (62)
US 9536191B1 · Arel et al. · 2017 [cited by applicant]
US 9754221B1 · Nagaraja · 2017 [cited by applicant]
US 10282662B2 · Schaul et al. · 2019 [cited by applicant]
US 10650310B2 · Schaul et al. · 2020 [cited by applicant]
US 10748039B2 · Lonescu et al. · 2020 [cited by applicant]
US 10867242B2 · Graepel et al. · 2020 [cited by applicant]
US 11551165B1 · Reynders, III · 2023 [cited by examiner]
US 11604941B1 · Hester · 2023 [cited by examiner]
US 12008077B1 · Hester · 2024 [cited by examiner]
US 20150100530A1 · Mnih et al. · 2015 [cited by applicant]
US 20160292568A1 · Schaul · 2016 [cited by examiner]
US 20170140269A1 · Schaul · 2017 [cited by examiner]
US 20170337682A1 · Liao · 2017 [cited by examiner]
US 20180032863A1 · Graepel et al. · 2018 [cited by applicant]
US 20180032864A1 · Graepel et al. · 2018 [cited by applicant]
US 20180260707A1 · Schaul et al. · 2018 [cited by applicant]
US 20190014488A1 · Tan et al. · 2019 [cited by applicant]
US 20190126472A1 · Tunyasuvunakool · 2019 [cited by examiner]
US 20190251437A1 · Finn et al. · 2019 [cited by applicant]
US 20190258938A1 · Mnih · 2019 [cited by examiner]
US 20210216822A1 · Paik et al. · 2021 [cited by applicant]
Abbeel et al. “An application of reinforcement learning to aerobatic helicopter flight,” Advances in Neural Information Processing Systems, Dec. 2007, 8 pages. [cited by applicant]
Bellemare et al. “The arcade learning environment: An evaluation platform for general agents.” Journal of Artificial Intelligence Research, 47, Jun. 2013, 5 pages. [cited by applicant]
Brys et al. “Reinforcement Learning from demonstration through shaping,” International Joint Conference on Artificial Intelligence, Jul. 2015, 7 pages. [cited by applicant]
Cederborg et al. “Policy shaping with human teachers,” International Joint Conference on Artificial Intelligence, Jul. 2015, 7 pages. [cited by applicant]
Chmali et al. “Direct policy iteration from demonstrations,” International Joint Conference on Artificial Intelligence, Jul. 2015, 7 pages. [cited by applicant]
De la Cruz Jr et al. “Pre-training Neural Networks with Human Demonstrations for Deep Reinforcement Learning,” CoRR, Submitted on Sep. 12, 2017, arXiv:1709.04083v1, 8 pages. [cited by applicant]
Duan et al. “One-shot imitation learning,” CoRR, Submitted on Dec. 4, 2017, arXiv 1703.07326v3, 27 pages. [cited by applicant]
Finn et al. “Guided cost learning: Deep inverse optimal control via policy optimization,” International Conference on Machine Learning, Jun. 2016, 10 pages. [cited by applicant]
Fortunato et al. “Noisy networks for exploration,” CoRR, Submitted on Feb. 15, 2018, arXiv:1706.10295v2, 21 pages. [cited by applicant]
Hasselt et al. “Deep reinforcement learning with double Q-learning,” AAAI Conference on Artificial Intelligence, Feb. 2016, 7 pages. [cited by applicant]
Hester et al. “Learning from Demonstrations for Real World Reinforcement Learning,” CoRR, Submitted on Apr. 12, 2017, arXiv:1704.03732v1, 11 pages. [cited by applicant]
Hester et al. “TEXPLORE: Real-time sample-efficient reinforcement learning for robots,” Machine Learning, 90(3), Mar. 2013, 45 pages. [cited by applicant]
Ho et al. “Generative adversarial imitation learning,” Advances in Neural Information Processing Systems, Dec. 2016, 9 pages. [cited by applicant]
Hosu et al. “Playing Atari games with deep reinforcement learning and human check-point replay,” CoRR, Submitted on Jul. 18, 2016, arXiv:1607.05077v1, 6 pages. [cited by applicant]
Jaderberg et al. “Reinforcement learning with unsupervised auxiliary tasks,” CoRR, Submitted on Nov. 16, 2016, arXiv:1611.05397v1, 14 pages. [cited by applicant]
Kim et al. “Learning from limited demonstrations,” Advances in Neural Information Processing Systems, Dec. 2013, 9 pages. [cited by applicant]
Kingma et al. “Adam: A method for stochastic optimization,” CoRR, Submitted on Jan. 30, 2017, arXiv:1412.6980v9, 15 pages. [cited by applicant]
Kurin et al. “The atari grand challenge dataset,” CoRR, Submitted on May 31, 2017, arXiv:1705.10998v1, 10 pages. [cited by applicant]
Lakshminarayanan et al. “Reinforcement learning with few expert demonstrations,” NIPS Workshop on Deep Learning for Action and Interaction, Dec. 2016, 8 pages. [cited by applicant]
LeCun et al. “Deep Learning,” Nature 521(7553) May 2015, 10 pages. [cited by applicant]
Levine et al. “End-to-end training of deep visuomotor policies,” Journal of Machine Learning, Jan. 2016, 40 pages. [cited by applicant]
Lipton et al. “Efficient exploration for dialog policy learning with deep BBQ network and replay buffer spiking,” CoRR, Submitted on Aug. 17, 2016, arXiv:1608.05081v1, 16 pages. [cited by applicant]
Mnih et al. “Asynchronous methods for deep reinforcement learning”, International Conference on Machine Learning, Jun. 2016, 10 pages. [cited by applicant]
Mnih et al. “Human-level control through deep reinforcement learning,” Nature, 518, Feb. 2015, 13 pages. [cited by applicant]
Piot et al. “Boosted and Reward-regularized Classification for Apprenticeship Learning,” International Conference on Autonomous Agents and Multiagent Systems, May 2014, 8 pages. [cited by applicant]
Piot et al. “Boosted bellman residual minimization handling expert demonstrations,” European Conference on Machine Learning, Sep. 2014, 17 pages. [cited by applicant]
Ross et al. “A reduction of imitation learning and structured prediction to No. regret online learning” International Conference on Artificial Intelligence and Statistics, Apr. 2011, 9 pages. [cited by applicant]
Schaal “Learning from demonstration,” December NIPS, 1997, 7 pages. [cited by applicant]
Schaul et al. “Prioritized experience replay,” CoRR, Submitted on May 2016, arXiv:1511.05952v4, 21 pages. [cited by applicant]
Shani et al. “An mdp-based recommender system,” Journal of Machine Learning Research, 6, Sep. 2005, 31 pages. [cited by applicant]
Silver et al. “Mastering the game of Go with deep neural networks and tree search,” Nature, 529, Jan. 2016, 20 pages. [cited by applicant]
Srivastava et al. “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, 15(1), Jun. 2014, 30 pages. [cited by applicant]
Suay et al. “Learning from demonstration for shaping through inverse reinforcement learning,” International Conference on Autonomous Agents and Multiagent Systems, May 2016, 9 pages. [cited by applicant]
Subramanian et al. “Exploration from demonstration for interactive reinforcement learning,” International Conference on Autonomous Agents and Multiagent Systems, May 2016, 10 pages. [cited by applicant]
Sun et al. “Deep aggrevated: Differentiable imitation learning for sequential prediction,” CoRR, Submitted on Mar. 3, 2017, arXiv:1703.01030v1, 17 pages. [cited by applicant]
Sutton et al. “Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction,” The 10th International Conference on Autonomous Agents and Multiagent Systems, vol. 2, May 2011… [cited by applicant]
Syed et al. “A game-theoretic approach to apprenticeship learning,” NIPS, Dec. 2007, 8 pages. [cited by applicant]
Syed et al. “Apprenticeship learning using linear programming,” International Conference on Machine Learning, Jul. 2008, 8 pages. [cited by applicant]
Taylor et al. “Integrating reinforcement learning with human demonstrations of varying ability” International Conference on Autonomous Agents and Multiagent Systems, May 2011, 8 pages. [cited by applicant]
Wang et al. “Dueling network architectures for deep reinforcement learning,” International Conference on Machines Learning, Jun. 2016, 9 pages. [cited by applicant]
Watter et al. “Embed to control: A locally linear latent dynamics model for control from raw images,” NIPS, Dec. 2015, 9 pages. [cited by applicant]