IP Library Granted Patent US 11,562,251
Granted Patent B2
US 11,562,251 · App. 16/533,575 · Granted Jan 24, 2023

Learning world graphs to accelerate hierarchical reinforcement learning

Inventors: Wenling Shang (San Francisco, CA); Alexander Richard Trott (San Francisco, CA); Stephan Tao Zheng (Redwood City, CA)
Assignee: Salesforce.com, Inc.
G06N3/088G05D1/0221G06N3/0445
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,251
App. No.
16/533,575
Granted
Jan 24, 2023
Kind
B2
Abstract

Systems and methods are provided for learning world graphs to accelerate hierarchical reinforcement learning (HRL) for the training of a machine learning system. The systems and methods employ or implement a two-stage framework or approach that includes (1) unsupervised world graph discovery, and (2) accelerated hierarchical reinforcement learning by integrating the graph.

Claims (26)

1. A system for training a machine learning system, the system comprising:

a communication interface that receives environment data, the environment data relating to an environment in which the machine learning system may operate;

a memory containing machine readable medium storing machine executable code; and

one or more processors coupled to the memory and configurable to execute the machine executable code to:

generate, by implementing a recurrent differentiable binary latent model, from the environment data a graph abstraction for the environment, the graph abstraction comprising a plurality of nodes and edges, wherein nodes represent points of interest in the environment and edges represent traversals between the nodes; and

perform hierarchical reinforcement learning using the graph abstraction to train the machine learning system.

2. The system of claim 1 , wherein the one or more processors configurable to execute the machine executable code discover one or more pivotal states in the environment.

3. The system of claim 2 , wherein the one or more processors configurable to execute the machine executable code generate edge connections for the graph abstraction using the one or more pivotal states.

4. The system of claim 1 , wherein the one or more processors configurable to execute the machine executable code implement a goal-conditioned agent to sample goals in a random walk of the graph abstraction.

5. The system of claim 4 , wherein knowledge gained in by the goal-conditioned agent in the random walk of the graph abstraction is transferred to subsequent tasks for the machine learning system.

6. The system of claim 1 , wherein the recurrent differentiable binary latent model infers a sequence of binary latent variables to discover one or more pivotal states in the environment.

7. The system of claim 1 , wherein the one or more processors configurable to execute the machine executable code execute a Wide-then-Narrow Instruction.

8. The system of claim 7 , wherein options for the Wide-then-Narrow Instruction are limited to pivotal states discovered during the generation of the graph abstraction.

9. The system of claim 1 , wherein the one or more processors configurable to execute the machine executable code implement a Feudal Network to perform the hierarchical reinforcement learning.

10. A method for training a machine learning system comprising:

receiving, at one or more processors, environment data, the environment data relating to an environment in which the machine learning system may operate;

generating, by implementing a recurrent differentiable binary latent model, from the environment data, at the one or more processors, a graph abstraction for the environment, the graph abstraction comprising a plurality of nodes and edges, wherein nodes represent points of interest in the environment and edges represent traversals between the nodes; and

performing hierarchical reinforcement learning, at the one or more processors, using the graph abstraction to train the machine learning system.

11. The method of claim 10 , wherein generating the graph abstraction for the environment comprises discovering one or more pivotal states in the environment.

12. The method of claim 11 , wherein generating the graph abstraction for the environment comprises generating edge connections for the graph abstraction using the one or more pivotal states.

13. The method of claim 10 , wherein generating the graph abstraction for the environment comprises employing a goal-conditioned agent to sample goals in a random walk of the graph abstraction.

14. The method of claim 13 , wherein performing the hierarchical reinforcement learning comprises transferring knowledge gained by the goal-conditioned agent in the random walk of the graph abstraction to subsequent tasks for the machine learning system.

15. The method of claim 10 , wherein generating the graph abstraction for the environment comprises inferring a sequence of binary latent variables to discover one or more pivotal states in the environment.

16. The method of claim 10 , wherein performing the hierarchical reinforcement learning comprises executing a Wide-then-Narrow Instruction.

17. The method of claim 16 , wherein options for the Wide-then-Narrow Instruction are limited to pivotal states discovered during the generation of the graph abstraction.

18. The method of claim 10 , wherein a Feudal Network is used to perform the hierarchical reinforcement learning.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0427 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 26, 2019
From: SHANG, WENLING; TROTT, ALEXANDER RICHARD; ZHENG, STEPHAN
To: SALESFORCE.COM, INC.
Reel/Frame 051372/0714 →