IP Library › Granted Patent US 11,783,232
Granted Patent B2
US 11,783,232 · App. 18/077,697 · Granted Oct 10, 2023

System and method for multi-agent reinforcement learning in a multi-agent environment

Inventors: David F. Isele (San Jose, CA); Kikuo Fujimura (Palo Alto, CA); Anahita Mohseni-Kabir (Pittsburgh, PA)
Assignee: HONDA MOTOR CO., LTD.
G06N20/00G06F30/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,783,232
App. No.
18/077,697
Granted
Oct 10, 2023
Kind
B2
Abstract

A system and method for multi-agent reinforcement learning in a multi-agent environment that include receiving data associated with the multi-agent environment in which an ego agent and a target agent are traveling. The system and method also include learning single agent policies that are respectively associated with the ego agent and the target agent based on the data associated with the multi-agent environment. The system and method additionally include learning a multi-agent policy as an interactive policy that enables the ego agent and the target agent to account for one another while traveling to respective goals within the multi-agent environment based on the single agent policies. The system and method further include implementing the multi-agent policy to control at least one of: the ego agent and the target agent to operate within the multi-agent environment.

Claims (33)

1. A computer-implemented method for multi-agent reinforcement learning in a multi-agent environment, comprising:

receiving data associated with the multi-agent environment in which an ego agent and a target agent are traveling;

learning single agent policies that are respectively associated with the ego agent and the target agent based on the data associated with the multi-agent environment;

learning a multi-agent policy as an interactive policy that enables the ego agent and the target agent to account for one another while traveling to respective goals within the multi-agent environment based on the single agent policies; and

implementing the multi-agent policy to control at least one of: the ego agent and the target agent to operate within the multi-agent environment.

2. The computer-implemented method of claim 1 , wherein receiving data associated with the multi-agent environment includes receiving image data and LiDAR data from at least one of the: ego agent and the target agent, wherein the image data and the LiDAR data are fused to determine fused environmental data associated with the multi-agent environment.

3. The computer-implemented method of claim 1 , wherein learning the single agent policies includes inputting respective observations and the respective goals of the ego agent and the target agent into a single agent actor critic model that uses an actor-critic policy gradient algorithm.

4. The computer-implemented method of claim 3 , wherein learning the single agent policies includes executing at least one iteration of a Markov Decision Process where at least one critic evaluates at least one action that is taken by the ego agent and the target agent to determine at least one reward and at least one state.

5. The computer-implemented method of claim 4 , wherein the at least one reward and the at least one state are analyzed to determine the single agent policies according to an individual goal-specific reward function, wherein the at least one iteration of the Markov Decision Process is implemented with the individual goal-specific reward function to provide a policy gradient for at least one of: the ego agent and the target agent to reach respective goals in an independent manner.

6. The computer-implemented method of claim 5 , wherein learning the multi-agent policy includes evaluating the single agent policies and passing inputs passed to the single agent actor critic model through a multi-agent actor critic model.

7. The computer-implemented method of claim 6 , wherein learning the multi-agent policy includes combining the single agent policies with an output of the multi-agent actor critic model to learn the multi-agent policy, wherein the multi-agent policy is determined according to a modification of the individual goal-specific reward function to a cooperative goal-specific reward function.

8. The computer-implemented method of claim 7 , wherein a neural network is trained at a time step with the multi-agent policy by updating a multi-agent dataset of the neural network with data pertaining to the multi-agent policy.

9. The computer-implemented method of claim 8 , wherein implementing the multi-agent policy includes analyzing the multi-agent dataset to implement the multi-agent policy to operate at least one of: the ego agent and the target agent to reach the respective goals in a cooperative manner.

10. A system for multi-agent reinforcement learning in a multi-agent environment, comprising:

a memory storing instructions when executed by a processor cause the processor to:

receive data associated with the multi-agent environment in which an ego agent and a target agent are traveling;

learn single agent policies that are respectively associated with the ego agent and the target agent based on the data associated with the multi-agent environment;

learn a multi-agent policy as an interactive policy that enables the ego agent and the target agent to account for one another while traveling to respective goals within the multi-agent environment based on the single agent policies; and

implement the multi-agent policy to control at least one of: the ego agent and the target agent to operate within the multi-agent environment.

11. The system of claim 10 , wherein receiving data associated with the multi-agent environment includes receiving image data and LiDAR data from at least one of the: ego agent and the target agent, wherein the image data and the LiDAR data are fused to determine fused environmental data associated with the multi-agent environment.

12. The system of claim 10 , wherein learning the single agent policies includes inputting respective observations and the respective goals of the ego agent and the target agent into a single agent actor critic model that uses an actor-critic policy gradient algorithm.

13. The system of claim 12 , wherein learning the single agent policies includes executing at least one iteration of a Markov Decision Process where at least one critic evaluates at least one action that is taken by the ego agent and the target agent to determine at least one reward and at least one state.

14. The system of claim 13 , wherein the at least one reward and the at least one state are analyzed to determine the single agent policies according to an individual goal-specific reward function, wherein the at least one iteration of the Markov Decision Process is implemented with the individual goal-specific reward function to provide a policy gradient for at least one of: the ego agent and the target agent to reach respective goals in an independent manner.

15. The system of claim 14 , wherein learning the multi-agent policy includes evaluating the single agent policies and passing inputs passed to the single agent actor critic model through a multi-agent actor critic model.

16. The system of claim 14 , wherein learning the multi-agent policy includes combining the single agent policies with an output of the multi-agent actor critic model to learn the multi-agent policy, wherein the multi-agent policy is determined according to a modification of the individual goal-specific reward function to a cooperative goal-specific reward function.

17. The system of claim 16 , wherein a neural network is trained at a time step with the multi-agent policy by updating a multi-agent dataset of the neural network with data pertaining to the multi-agent policy.

18. The system of claim 17 , wherein implementing the multi-agent policy includes analyzing the multi-agent dataset to implement the multi-agent policy to operate at least one of: the ego agent and the target agent to reach the respective goals in a cooperative manner.

19. A non-transitory computer readable storage medium storing instructions that when executed by a computer, which includes a processor perform a method, the method comprising:

receiving data associated with the multi-agent environment in which an ego agent and a target agent are traveling;

learning single agent policies that are respectively associated with the ego agent and the target agent based on the data associated with the multi-agent environment;

learning a multi-agent policy as an interactive policy that enables the ego agent and the target agent to account for one another while traveling to respective goals within the multi-agent environment based on the single agent policies; and

implementing the multi-agent policy to control at least one of: the ego agent and the target agent to operate within the multi-agent environment.

20. The non-transitory computer readable storage medium of claim 19 , wherein learning the multi-agent policy includes combining the single agent policies with an output of a multi-agent actor critic model to learn the multi-agent policy, wherein the multi-agent policy is determined according to a modification of an individual goal-specific reward function to a cooperative goal-specific reward function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2022
From: ISELE, DAVID F.; FUJIMURA, KIKUO; MOHSENI-KABIR, ANAHITA
To: HONDA MOTOR CO., LTD.
Reel/Frame 062029/0076 →
Continuity (3)
Continuation 16390224 · Apr 22, 2019
Provisional Application 62731426 · Sep 14, 2018
Related Publication 20230104513A1 · Apr 6, 2023
Cited By (1)
US 12,554,620