IP Library › Granted Patent US 11,321,635
Granted Patent B2
US 11,321,635 · App. 16/425,648 · Granted May 3, 2022

Method for performing multi-agent reinforcement learning in the presence of unreliable communications via distributed consensus

Inventors: Michael W. Walton (San Diego, CA); Benjamin J. Migliori (San Diego, CA); John Reeder (San Diego, CA)
Assignee: United States of America as represented by the Secretary of the Navy
G06N20/00G06F9/5027H04L67/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,321,635
App. No.
16/425,648
Granted
May 3, 2022
Kind
B2
Abstract

A system is provided for performing a predetermined function within a total area of operation, wherein the system includes a plurality of autonomous agents. Each autonomous agent is able to detect respective local parameters. Each autonomous agent uses a Kalman filter component to establish an environment state based a plurality of state measurements over time. The output of the Kalman filter component within a respective agent is applied to reinforcement learning by an actor-critic task controller, within the respective agent, to determine a subsequent action to be performed by the respective agent in accordance with a reward function. Each agent includes a Kalman consensus filter that addresses errors of the plurality of state measurements over time.

Claims (41)

1. A system for performing a predetermined function within a total area of operation, said system comprising:

a first autonomous agent comprising a first agent detector, a first agent communication component and a first agent controller, said first agent detector being operable to detect a first agent parameter within a first agent area and to generate a first agent parameter signal based on the detected first agent parameter, said first agent controller comprising a task controller and a first Kalman consensus filter and being operable to instruct said first autonomous agent to perform an initial first agent task and to perform a subsequent first agent task, wherein the first Kalman consensus filter comprises a Kalman filter component and a distributed average consensus component;

a second autonomous agent comprising a second agent detector, a second agent communication component and a second agent controller, said second agent detector being operable to detect a second agent parameter within a second agent area and to generate a second agent parameter signal based on the detected second agent parameter, said second agent communication component being operable to transmit the second agent parameter signal to said first agent communication component, said second agent controller being operable to instruct said second autonomous agent to perform an initial second agent task and to perform a subsequent second agent task;

a third autonomous agent comprising a third agent detector, a third agent communication component and a third agent controller, said third agent detector being operable to detect a third agent parameter within a third agent area and to generate a third agent parameter signal based on the detected third agent parameter, said third agent communication component being operable to transmit the third agent parameter signal to said first agent communication component and to said second agent communication component, said third agent controller being operable to instruct said third autonomous agent to perform an initial third agent task and to perform a subsequent third agent task; and

wherein said first agent communication component is operable to transmit the first agent parameter signal to said second agent communication component and to said third agent communication component,

wherein said second agent communication component is further operable to transmit the second agent parameter signal to said third agent communication component, and operable to transmit the received third agent parameter signal to the first agent communication component,

wherein said first agent controller is operable to instruct said first autonomous agent to perform said subsequent first agent task based on the first agent parameter signal, the second agent parameter signal, the third agent parameter signal and a predetermined reward function using reinforcement learning and a first Kalman consensus filter,

wherein said second agent controller is operable to instruct said second autonomous agent to perform said subsequent second agent task based on the first agent parameter signal, the second agent parameter signal, the third agent parameter signal and the predetermined reward function using reinforcement learning and a second Kalman consensus filter, wherein said third agent controller is operable to instruct said third autonomous agent to perform said subsequent third agent task based on the first agent parameter signal, the second agent parameter signal, the third agent parameter signal and the predetermined reward function using reinforcement learning and a third Kalman consensus filter, and also operable to transmit the received second agent parameter signal to the first agent communication component,

wherein the first Kalman consensus filter is operable to output a second agent parameter consensus based on the second agent parameter signal received from the second agent communication component and the received second agent parameter signal received from the third agent communication component,

wherein the task controller is configured to instruct the first autonomous agent to perform the subsequent first agent task based on the second agent parameter consensus,

wherein the Kalman filter component is configured to generate a local state signal, and the distributed average consensus component is configured to generate the second agent parameter consensus based on a weighted average of a first product of the second agent parameter signal received from the second agent communication component and a first weighting factor of the second agent parameter signal received from the second agent communication component and a second product of the received second agent parameter signal received from the third agent communication component and a second weighting factor of the received second agent parameter signal received from the third agent communication component,

wherein the first weighting factor is based on distance,

wherein the first agent area is less than and within the total area of operation, and

wherein the second agent area is less than and within the total area of operation.

2. The system of claim 1 ,

wherein said first agent detector comprises a detector selected from a group of detectors comprising a position detector, a velocity detector, an acceleration detector and combinations thereof,

wherein the position detector is operable to detect a position selected from a group of positions comprising a first agent position of the first autonomous agent, a second agent position of said second autonomous agent when said second autonomous agent is within the first agent area and a combination thereof,

wherein the velocity detector is operable to detect a velocity selected from a group of velocities comprising a first agent velocity of the first autonomous agent, a second agent velocity of said second autonomous agent when said second autonomous agent is within the first agent area and a combination thereof, and

wherein the acceleration detector is operable to detect an acceleration selected from a group of accelerations comprising a first agent acceleration of the first autonomous agent, a second agent acceleration of said second autonomous agent when said second autonomous agent is within the first agent area and a combination thereof.

3. A method performing a predetermined function within a total area of operation, the method comprising:

providing a first autonomous agent comprising a first agent detector, a first agent communication component and a first agent controller, the first agent detector being operable to detect a first agent parameter within a first agent area and to generate a first agent parameter signal based on the detected first agent parameter, the first agent controller being operable to instruct the first autonomous agent to perform an initial first agent task and to perform a subsequent first agent task;

providing a second autonomous agent comprising a second agent detector, a second agent communication component and a second agent controller, the second agent detector being operable to detect a second agent parameter within a second agent area and to generate a second agent parameter signal based on the detected second agent parameter, the second agent communication component being operable to transmit the second agent parameter signal to the first agent communication component, the second agent controller being operable to instruct the second autonomous agent to perform an initial second agent task and to perform a subsequent second agent task; and

providing a third autonomous agent comprising a third agent detector, a third agent communication component and a third agent controller, the third agent detector being operable to detect a third agent parameter within a third agent area and to generate a third agent parameter signal based on the detected third agent parameter, the third agent communication component being operable to transmit the third agent parameter signal to the first agent communication component and to the second agent communication component, the third agent controller being operable to instruct the third autonomous agent to perform an initial third agent task and to perform a subsequent third agent task;

transmitting, from the first agent communication component, a first agent parameter signal to the second agent communication component and to the third agent communication component;

transmitting, from the second agent communication component, the second agent parameter signal to the third agent communication component;

instructing, via the first agent controller, the first autonomous agent to perform the subsequent first agent task based on the first agent parameter signal, the second agent parameter signal, the third agent parameter signal and a predetermined reward function using reinforcement learning and a first Kalman consensus filter;

instructing, via the second agent controller, the second autonomous agent to perform the subsequent second agent task based on the first agent parameter signal, the second agent parameter signal, the third agent parameter signal and the predetermined reward function using reinforcement learning and a second Kalman consensus filter; and

instructing, via the third agent controller, the third autonomous agent to perform the subsequent third agent task based on the first agent parameter signal, the second agent parameter signal, the third agent parameter signal and the predetermined reward function using reinforcement learning and a third Kalman consensus filter, transmitting, via the second agent communication component, the received third agent parameter signal to the first agent communication component;

transmitting, via the third agent communication component, the received second agent parameter signal to the first agent communication component;

outputting, via a first Kalman consensus filter of the first agent controller that comprises a task controller and the first Kalman consensus filter, output a second agent parameter consensus based on the second agent parameter signal received from the second agent communication component and the received second agent parameter signal received from the third agent communication component;

instructing, via the task controller, the first autonomous agent to perform the subsequent first agent task based on the second agent parameter consensus;

generating, via a Kalman filter component of the first Kalman consensus filter that comprises the Kalman filter component and a distributed average consensus component, a local state signal; and

generating, via the distributed average consensus component, the second agent parameter consensus based on a weighted average of a first product of the second agent parameter signal received from the second agent communication component and a first weighting factor of the second agent parameter signal received from the second agent communication component and a second product of the received second agent parameter signal received from the third agent communication component and the second weighting factor of the received second agent parameter signal received from the third agent communication component,

wherein the first weighting factor is based on distance,

wherein the first agent area is less than and within the total area of operation, and

wherein the second agent area is less than and within the total area of operation.

4. The method of claim 3 ,

wherein the first agent detector comprises a detector selected from a group of detectors comprising a position detector, a velocity detector, an acceleration detector and combinations thereof,

wherein the position detector is operable to detect a position selected from a group of positions comprising a first agent position of the first autonomous agent, a second agent position of the second autonomous agent when the second autonomous agent is within the first agent area and a combination thereof,

wherein the velocity detector is operable to detect a velocity selected from a group of velocities comprising a first agent velocity of the first autonomous agent, a second agent velocity of the second autonomous agent when the second autonomous agent is within the first agent area and a combination thereof, and

wherein the acceleration detector is operable to detect an acceleration selected from a group of accelerations comprising a first agent acceleration of the first autonomous agent, a second agent acceleration of the second autonomous agent when the second autonomous agent is within the first agent area and a combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2019
From: WALTON, MICHAEL W.; MIGLIORI, BENJAMIN J.; REEDER, JOHN
To: UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY OF THE NAVY
Reel/Frame 049324/0670 →
Continuity (1)
Related Publication 20200380401A1 · Dec 3, 2020