IP Library Granted Patent US 11,709,462
Granted Patent B2
US 11,709,462 · App. 15/894,688 · Granted Jul 25, 2023

Safe and efficient training of a control agent

Inventors: Haoxiang Li (San Jose, CA); Yinan Zhang (West Lebanon, NH)
Assignee: ADOBE INC.
G05B13/027G05B17/02G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,709,462
App. No.
15/894,688
Granted
Jul 25, 2023
Kind
B2
Abstract

The training of a learning agent to provide real-time control of an object is disclosed. Training of the learning agent and training of a corresponding pioneer agent are iteratively alternated. The training of the learning and pioneer agents is under the supervision of a supervisor agent. The training of the learning agent provides feedback for subsequent training of the pioneer agent. The training of the pioneer agent provides feedback for subsequent training of the learning agent. During the training, a supervisor coefficient modulates the influence of the supervisor agent. As agents are trained, the influence of the supervisor agent is decayed. The training of the learning agent, under a first level of supervisor influence, includes real-time control of the object. The subsequent training of the pioneer agent, under a reduced level of supervisor influence, includes replay of training data accumulated during the real-time control of the object.

Claims (23)

1. A computer-readable storage medium having instructions stored thereon that, when executed by a processor of a computing device, cause the computing device to perform operations comprising:

employing a first agent to provide real-time control of an object within an environment based on a combination of a learning signal generated by the first agent and a supervisor signal generated by a supervisor agent, wherein the supervisor signal is weighted by a supervisor coefficient;

accumulating training data over one or more iterations of the first agent providing real-time control of the object within the environment;

updating the supervisor coefficient;

training a second agent to provide simulated control of the object within the environment based on at least a portion of the accumulated training data and a combination of a pioneer signal generated by the second agent and the supervisor signal that is weighted by the updated supervisor coefficient; and

updating the first agent based on the trained second agent.

2. The computer-readable storage medium of claim 1 , wherein the operations further comprise:

updating a critic network based on the accumulated training data;

determining an actor-value function based on the critic network; and

updating the first agent based on the actor-value function.

3. The computer-readable storage medium of claim 1 , wherein at least the first agent and the second agent are implemented in deep neural networks.

4. The computer-readable storage medium of claim 1 , wherein the operations further comprise:

employing the updated first agent to provide further real-time control the object within the environment, wherein the combination of the learning signal and the supervisor signal is weighted by the updated supervisor coefficient.

5. The one or more computer-readable storage media of claim 1 , wherein the operations further comprise:

employing a learning policy of the first agent to generate the learning signal based on a current state of the environment;

employing a supervisor policy of the supervisor agent to generate the supervisor signal based on the current state of the environment;

selecting an action based on the combination of the learning signal and the supervisor signal;

causing the object to execute the action;

observing a transition from current state to a new state of the environment, wherein the transition from the current state to the new state is in response to the action executed by the object;

observing a reward in response to the action executed by the object; and

including the current state, the action, the reward, and the new state as associated data within the accumulated training data.

6. The one or more computer-readable storage media of claim 1 , wherein training the second agent includes comparing a difference between the combination of the pioneer signal and the supervisor signal and the combination of the learning signal and the supervisor signal, and wherein the combination of the learning signal and the supervisor signal is embedded within the accumulated training data.

7. The one or more computer-readable storage media of claim 1 , wherein updating the supervisor coefficient includes decreasing a value of the supervisor coefficient.

Assignments (2)
CHANGE OF NAME Recorded Nov 29, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047687/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2018
From: LI, HAOXIANG; ZHANG, YINAN
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 045824/0498 →
Continuity (1)
Related Publication 20190250568A1 · Aug 15, 2019