IP Library Patent Application 15634811
Patent Application
App. No. 15/634,811

SCALABILITY OF REINFORCEMENT LEARNING BY SEPARATION OF CONCERNS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
15/634,811
Abstract

Aspects provided herein are relevant to machine learning techniques, including decomposing single-agent reinforcement learning problems into simpler problems addressed by multiple agents. Actions proposed by the multiple agents are then aggregated using an aggregator, which selects an action to take with respect to an environment. Aspects provided herein are also relevant to a hybrid reward model.

Claims (39)

1 . A method comprising:

receiving a single-agent task having a set of states and a set of environment actions; and

decomposing the single-agent task by:

instantiating a plurality of non-cooperating agents, each agent having a defined output set and a reward function associated with an aspect of the single-agent task, wherein each agent is configured to choose an output from its defined output set; and

defining an aggregator that selects an environment action from the set of environment actions based, in part, on the chosen output from each agent.

2 . The method of claim 1 , wherein the defined output set of at least one agent of the plurality of agents comprises an output associated with an environment action and an output associated with a communication action.

3 . The method of claim 1 , wherein the defined output set of at least one agent of the plurality of agents comprises only outputs associated with communication actions.

4 . The method of claim 1 , wherein the defined output set of at least one agent of the plurality of agents comprises outputs only associated with environment actions.

5 . The method of claim 1 , wherein each agent of the plurality of agents sees a subset of states smaller than the set of states.

6 . The method of claim 1 , further comprising:

determining that there is a cyclic relationship within the plurality of agents; and

responsive to determining that there is a cyclic relationship, converting the cyclic relationship into an acyclic relationship.

7 . The method of claim 6 , wherein converting the cyclic relationship into an acyclic relationship comprises instantiating at least two trainer agents, each trainer agent associated with an agent of the plurality of agents.

8 . The method of claim 7 , further comprising:

pre-training agents having a trainer agent with their respective trainer agents;

after pre-training, freezing weights of the pre-trained agents; and

after freezing the weights, training additional agents of the plurality of agents.

9 . The method of claim 1 , wherein the aggregator is configured to aggregate using a technique selected from the group consisting of: majority voting, rank voting, and Q-value generalized means maximizer.

10 . The method of claim 1 , further comprising training the plurality of agents with respect to the task.

11 . A computer-implemented method comprising:

generating a plurality of agents, each agent associated with a different aspect of a task, wherein the task defines an environment and a set of environment actions that can be taken with respect to the environment;

using each agent of the plurality of agents to:

observe at least a portion of the environment of the task; and

generate an output based, in part, on the observation; and

choosing an environment action from the set of environment actions based, in part, on the outputs generated by the agents.

12 . The computer-implemented method of claim 11 , wherein each output is selected from a set of outputs defined for each agent.

13 . The computer-implemented method of claim 11 , wherein choosing the environment action comprises using a technique selected from the group consisting of: majority voting, rank voting, and Q-value generalized means maximizer.

14 . The computer-implemented method of claim 11 , further comprising performing the chosen environment action.

15 . The computer-implemented method of claim 11 , wherein the plurality of agents are non-cooperative.

16 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:

generate a plurality of agents, each agent associated with a different aspect of a same task, wherein the task defines an environment and a set of environment actions that can be taken with respect to the environment;

use each agent of the plurality of agents to:

observe at least a portion of the environment of the task; and

generate an output based, in part, on the observation; and

choose an action from the set of environment actions based, in part, on the output from the agents.

17 . The non-transitory computer readable medium of claim 16 , wherein the output comprises an output associated with an action selected from a subset of the set of environment actions.

18 . The non-transitory computer readable medium of claim 16 , wherein choosing the action comprises using a technique selected from the group consisting of: majority voting, rank voting, and Q-value generalized means maximizer.

19 . The non-transitory computer readable medium of claim 16 , wherein the instructions further cause the processor to perform the chosen environment action.

20 . The non-transitory computer readable medium of claim 16 , wherein the plurality of agents are non-cooperative.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2017
From: VAN SEIJEN, HARM HENDRIK; FATEMI BOOSHEHRI, SEYED MEHDI; LAROCHE, ROMAIN MICHEL HENRI; ROMOFF, JOSHUA SAMUEL
To: MICROSOFT TECHNOLOGY LICENSING, LLC,
Reel/Frame 042830/0716 →