IP Library › Granted Patent US 11,720,792
Granted Patent B2
US 11,720,792 · App. 16/944,420 · Granted Aug 8, 2023

Devices and methods for reinforcement learning visualization using immersive environments

Inventors: Matthew Edmund Taylor (Edmonton, CA); Bilal Kartal (Edmonton, CA); Pablo Francisco Hernandez Leal (Edmonton, CA); Nathan Douglas (Calgary, CA); Dianna Yim (Calgary, CA); Frank Maurer (Calgary, CA)
Assignee: ROYAL BANK OF CANADA
G06N3/08G06N3/006G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,792
App. No.
16/944,420
Granted
Aug 8, 2023
Kind
B2
Abstract

Disclosed are systems, methods, and devices for generating a visualization of a deep reinforcement learning (DRL) process. State data is received, reflective of states of an environment explored by an DRL agent, each state corresponding to a time step. For each given state, saliency metrics are calculated by processing the state data, each metric measuring saliency of a feature at the time step corresponding to the given state. A graphical visualization is generated, having at least two dimensions in which: each feature of the environment is graphically represented along a first axis; and each time step is represented along a second axis; and a plurality of graphical markers representing corresponding saliency metrics, each graphical marker having a size commensurate with the magnitude of the particular saliency metric represented, and a location along the first and second axes corresponding to the feature and time step for the particular saliency metric.

Claims (43)

1. A computer-implemented method for generating a visualization of a deep reinforcement learning process, the method comprising:

receiving state data reflective of a plurality of states of an environment explored by a deep reinforcement learning agent, each of the states corresponding to one of a plurality of successive time steps;

for each given state of the plurality of states, calculating a plurality of saliency metrics by processing the state data, each of the metrics measuring saliency of one of a plurality of features of the environment at the time step corresponding to the given state;

generating a graphical visualization having:

at least two dimensions in which: each of the features of the environment is graphically represented along a first axis of said dimensions; and each of the time steps is represented along a second axis of said dimensions;

a plurality of graphical markers representing a corresponding one of the saliency metrics, each of the graphical markers having a size commensurate with the magnitude of the particular saliency metric represented by the graphical marker, and a location along the first and second axes corresponding to the feature and time step for the particular saliency metric.

2. The computer-implemented method of claim 1 , further comprising:

sending signals representing the generated graphical visualization for display at a head-mounted display.

3. The computer-implemented method of claim 1 , further comprising:

receiving signals representing user input from a handheld controller of a virtual reality, augmented reality, or mixed reality system.

4. The computer-implemented method of claim 1 , wherein the plurality of saliency metrics includes perturbation-based saliency metrics.

5. The computer-implemented method of claim 1 , wherein the plurality of saliency metrics includes gradient-based saliency metrics.

6. The computer-implemented method of claim 1 , wherein said calculating the plurality of saliency metrics includes determining a change in a policy of the deep reinforcement learning agent.

7. The computer-implemented method of claim 1 , wherein the graphical visualization has three dimensions, and each of the features is represented along two axes of said dimensions.

8. The computer-implemented method of claim 7 , wherein the graphical markers are organized into a plurality of levels, with each of the levels representative of one of the time steps.

9. The computer-implemented method of claim 7 , wherein each of the plurality of graphical markers has a three-dimensional shape.

10. The computer-implemented method of claim 1 , further comprising:

receiving a user selection of one of the time steps;

updating the graphical visualization to modify an appearance of the graphical markers representing saliency metrics for the selected time step.

11. The computer-implemented method of claim 1 , wherein the graphical visualization further comprises a graphical representation of the environment.

12. The computer-implemented method of claim 11 , wherein the graphical representation of the environment includes a virtual room.

13. The computer-implemented method of claim 1 , the environment includes a platform for trading securities and the agent is a trading agent.

14. The computer-implemented method of claim 13 , wherein said calculating the plurality of saliency metrics includes calculating metrics reflective of a level of aggression of the trading agent.

15. A computing device for generating a visualization of a deep reinforcement learning process, the device comprising:

at least one processor;

memory in communication with the at least one processor, and

software code stored in the memory, which when executed by the at least one processor causes the computing device to:

receive state data reflective of a plurality of states of an environment explored by a deep reinforcement learning agent, each of the states corresponding to a successive time step;

for each given state of the plurality of states, calculate a plurality of saliency metrics by processing the state data, each of the metrics measuring saliency of one of a plurality of features of the environment at the time step corresponding to the given state;

generate a graphical visualization having:

at least two dimensions in which: each of the features of the environment is graphically represented along a first axis of the dimensions; and

each of the time steps is represented along a second axis of the dimensions;

a plurality of graphical markers representing a corresponding one of the saliency metrics, each of the graphical markers having a size commensurate with the magnitude of the particular saliency metric represented by the graphical marker, and a location along the first and second axes corresponding to the feature and time step for the particular saliency metric.

16. The device of claim 15 , further comprising a display interface for interconnection with a head-mounted display.

17. The device of claim 15 , further comprising an I/O interface for interconnection with a handheld controller.

18. The device of claim 16 , wherein the at least one processor further causes the computing device to transmit signals corresponding to the generated graphical visualization to the head-mounted display by way of the display interface.

19. The device of claim 16 , wherein the head-mounted display is part of a virtual reality, augmented reality, or mixed reality system.

20. A non-transitory computer-readable medium having stored thereon machine interpretable instructions which, when executed by a processor, cause the processor to perform a computer implemented method for generating a visualization of a deep reinforcement learning process, the method comprising:

receiving state data reflective of a plurality of states of an environment explored by a deep reinforcement learning agent, each of the states corresponding to one of a plurality of successive time steps;

for each given state of the plurality of states, calculating a plurality of saliency metrics by processing the state data, each of the metrics measuring saliency of one of a plurality of features of the environment at the time step corresponding to the given state;

generating a graphical visualization having:

at least two dimensions in which: each of the features of the environment is graphically represented along a first axis of said dimensions; and each of the time steps is represented along a second axis of said dimensions;

a plurality of graphical markers representing a corresponding one of the saliency metrics, each of the graphical markers having a size commensurate with the magnitude of the particular saliency metric represented by the graphical marker, and a location along the first and second axes corresponding to the feature and time step for the particular saliency metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2020
From: TAYLOR, MATTHEW EDMUND; KARTAL, BILAL; HERNANDEZ LEAL, PABLO FRANCISCO; DOUGLAS, NATHAN; YIM, DIANNA; MAURER, FRANK
To: ROYAL BANK OF CANADA
Reel/Frame 054017/0667 →
Continuity (2)
Provisional Application 62881033 · Jul 31, 2019
Related Publication 20210034974A1 · Feb 4, 2021
Cited By (1)
US 12,236,326