IP Library › Granted Patent US 11,823,062
Granted Patent B1
US 11,823,062 · App. 18/124,559 · Granted Nov 21, 2023

Unsupervised reinforcement learning method and apparatus based on Wasserstein distance

Inventors: Xiangyang Ji (Beijing, CN); Shuncheng He (Beijing, CN); Yuhang Jiang (Beijing, CN)
Assignee: TSINGHUA UNIVERSITY
G06N3/092G06N3/088G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,823,062
App. No.
18/124,559
Granted
Nov 21, 2023
Kind
B1
Abstract

The present disclosure discloses an unsupervised reinforcement learning method and apparatus based on Wasserstein distance. The method includes: obtaining a state distribution in a trajectory obtained with guidance of a current policy of an agent; calculating a Wasserstein distance between the state distribution and a state distribution in a trajectory obtained with another historical policy, and calculating a pseudo reward of the agent based on the Wasserstein distance, replacing a reward fed back from an environment in a target reinforcement learning framework with the pseudo reward, and guiding the current policy of the agent to keep a large distance from the other historical policy. The method uses Wasserstein distance to encourage an algorithm in an unsupervised reinforcement learning framework to obtain diverse policies and skills through training.

Claims (23)

1. An unsupervised reinforcement learning method based on Wasserstein distance, comprising:

obtaining a state distribution in a trajectory obtained with guidance of a current policy of an agent;

calculating a Wasserstein distance between the state distribution and a state distribution in a trajectory obtained with another historical policy; and

calculating a pseudo reward of the agent based on the Wasserstein distance, replacing a reward fed back from an environment in a target reinforcement learning framework with the pseudo reward, and guiding the current policy of the agent to keep a maximum distance from the other historical policy.

2. The method according to claim 1 , wherein said calculating the pseudo reward of the agent based on the Wasserstein distance comprises:

making, based on a state variable obtained from a current observation of the agent, a decision using a policy model of the agent to obtain an action variable, and interacting with the environment to obtain the pseudo reward.

3. The method according to claim 1 , further comprising, subsequent to calculating the pseudo reward of the agent:

optimizing, based on a deep reinforcement learning framework, a policy model of the agent using gradient back propagation.

4. The method according to claim 1 , wherein the Wasserstein distance is a dual form estimation.

5. The method according to claim 2 , wherein the Wasserstein distance is a dual form estimation.

6. The method according to claim 3 , wherein the Wasserstein distance is a primal form estimation.

7. The method according to claim 3 , wherein the Wasserstein distance is a primal form estimation, and wherein an average distribution of state variables obtained from all policies other than the current policy is used as a target distribution from which a maximum distance needs to be kept.

8. An unsupervised reinforcement learning apparatus based on Wasserstein distance, comprising a processor and a memory storing an executable program, wherein the executable program, when performed by the processor, implements:

obtaining a state distribution in a trajectory obtained with guidance of a current policy of an agent;

calculating a Wasserstein distance between the state distribution and a state distribution in a trajectory obtained with another historical policy; and

calculating a pseudo reward of the agent based on the Wasserstein distance, replacing a reward fed back from an environment in a target reinforcement learning framework with the pseudo reward, and guiding the current policy of the agent to keep a maximum distance from the other historical policy.

9. The apparatus according to claim 8 , wherein the executable program, when performed by the processor, further implements: making, based on a state variable obtained from a current observation of the agent, a decision using a policy model of the agent to obtain an action variable, and interacting with the environment to obtain the pseudo reward.

10. The apparatus according to claim 8 , wherein the executable program, when performed by the processor, further implements:

optimizing, based on a deep reinforcement learning framework, a policy model of the agent using gradient back propagation, subsequent to calculating the pseudo reward of the agent.

11. The apparatus according to claim 8 , wherein the Wasserstein distance is a dual form estimation.

12. The apparatus according to claim 9 , wherein the Wasserstein distance is a dual form estimation.

13. The apparatus according to claim 10 , wherein the Wasserstein distance is a primal form estimation.

14. The apparatus according to claim 10 , wherein the Wasserstein distance is a primal form estimation, and wherein an average distribution of state variables obtained from all policies other than the current policy is used as a target distribution from which a maximum distance needs to be kept.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2023
From: JI, XIANGYANG; HE, SHUNCHENG; JIANG, YUHANG
To: TSINGHUA UNIVERSITY
Reel/Frame 064929/0592 →