IP Library Granted Patent US 11,560,146
Granted Patent B2
US 11,560,146 · App. 16/778,444 · Granted Jan 24, 2023

Interpreting data of reinforcement learning agent controller

Inventors: Subramanya Nageshrao (Mountain View, CA); Bruno Sielly Jales Costa (Santa Clara, CA); Dimitar Petrov Filev (Novi, MI)
Assignee: Ford Global Technologies, LLC
B60W30/14G05D1/0088G06K9/6218G06N3/08G06V20/56G05D2201/0213
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,560,146
App. No.
16/778,444
Granted
Jan 24, 2023
Kind
B2
Abstract

The present disclosure describes systems and methods that include calculating, via a reinforcement learning agent (RLA) controller, a plurality of state-action values based on sensor data representing an observed state, wherein the RLA controller utilizes a deep neural network (DNN) and generating, via a fuzzy controller, a plurality of linear models mapping the plurality of state-action values to the sensor data.

Claims (24)

1. A method, comprising:

calculating, via a reinforcement learning agent (RLA) controller, a plurality of state-action values based on sensor data representing an observed state, wherein the RLA controller utilizes a deep neural network (DNN);

wherein the RLA controller uses, to calculate the state-action values, (i) a first reward that is a function of a user-selected maximum velocity for a vehicle, (ii) a second reward that is a function of a headway parameter representing a distance between the vehicle and a second vehicle, and (iii) a third reward that is a function of a change in acceleration of the vehicle and a predetermined allowable acceleration for the vehicle;

generating, as output from a fuzzy controller arranged in series with the RLA controller to receive as input the state-action values output by the RLA controller, a plurality of linear models mapping the plurality of state-action values to the sensor data; and

actuating an agent based on at least one of the plurality of state-action values or the plurality of linear models;

wherein the agent includes the vehicle and wherein actuating the agent further comprises adjusting a speed of the vehicle based on at least one of the plurality of state-action values or the plurality of linear models.

2. The method of claim 1 , wherein the plurality of state-action values correspond to an optimal policy generated during reinforcement learning training.

3. The method of claim 1 , wherein the vehicle is an autonomous vehicle.

4. The method of claim 1 , wherein the plurality of linear models comprise a set of IF-THEN rules mapping the plurality of state-action values to the sensor data.

5. The method of claim 1 , wherein the fuzzy controller uses an Evolving Takagi-Sugeno (ETS) model to generate the plurality of linear models.

6. The method of claim 1 , further comprising: determining, via the fuzzy controller, one or more data clusters corresponding to the sensor data, wherein each of the one or more data clusters comprises a focal point and a radius.

7. A system, comprising:

at least one processor; and

at least one memory, wherein the at least one memory stores instructions executable by the at least one processor such that the at least one processor is programmed to:

calculate, via a reinforcement learning agent (RLA) controller that utilizes a deep neural network, a plurality of state-action values based on sensor data representing an observed state;

wherein the RLA controller uses, to calculate the state-action values, (i) a first reward that is a function of a user-selected maximum velocity for a vehicle, (ii) a second reward that is a function of a headway parameter representing a distance between the vehicle and a second vehicle, and (iii) a third reward that is a function of a change in acceleration of the vehicle and a predetermined allowable acceleration for the vehicle;

generate, as output from a fuzzy controller arranged in series with the RLA controller to receive as input the state-action values output by the RLA controller, a plurality of linear models mapping the plurality of state-action values to the sensor data; and

actuate an agent based on at least one of the plurality of state-action values or the plurality of linear models;

wherein the agent includes the vehicle and wherein actuating the agent further comprises adjusting a speed of the vehicle based on at least one of the plurality of state-action values or the plurality of linear models.

8. The system of claim 7 , wherein the plurality of state-action values correspond to an optimal policy generated during reinforcement learning training.

9. The system of claim 7 , wherein the vehicle is an autonomous vehicle.

10. The system of claim 7 , wherein the plurality of linear models comprise a set of IF-THEN rules mapping the plurality of state-action values to the sensor data.

11. The system of claim 7 , wherein the processor is further programmed to generate the plurality of linear models using an Evolving Takagi-Sugeno (ETS) model.

12. The system of claim 7 , wherein the processor is further programmed to determine one or more data clusters corresponding to the sensor data, wherein each of the one or more data clusters comprises a focal point and a radius.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: NAGESHRAO, SUBRAMANYA; JALES COSTA, BRUNO SIELLY; FILEV, DIMITAR PETROV
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 051685/0072 →
Continuity (2)
Provisional Application 62824015 · Mar 26, 2019
Related Publication 20200307577A1 · Oct 1, 2020
Cited By (1)
US 12,282,337