IP Library Granted Patent US 11,993,291
Granted Patent B2
US 11,993,291 · App. 17/071,115 · Granted May 28, 2024

Neural networks for navigation of autonomous vehicles based upon predicted human intents

Inventor: Mel McCurrie (Cambridge, MA)
Assignee: Perceptive Automata, Inc.
B60W60/00276B60W60/0027G06N3/045G06V10/255G06V10/82G06V20/58G06V20/584G06V40/103B60W2420/403B60W2554/402B60W2554/4041B60W2554/4045G06V10/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,993,291
App. No.
17/071,115
Granted
May 28, 2024
Kind
B2
Abstract

A system uses neural networks to determine intents of traffic entities (e.g., pedestrians, bicycles, vehicles) in an environment surrounding a vehicle (e.g., an autonomous vehicle) and generates commands to control the vehicle based on the determined intents. The system receives images of the environment captured by sensors on the vehicle, and processes the images using neural network models to determine overall intents or predicted actions of the one or more traffic entities within the images. The system generates commands to control the vehicle based on the determined overall intents of the traffic entities.

Claims (43)

1. A method comprising:

receiving a plurality of images, each image corresponding to a video frame captured by one or more sensors of a vehicle moving through traffic;

processing the plurality of images using a first neural network model configured to generate, for a first image of the plurality of images, a first feature map indicating, for each pixel of the first image, a feature vector corresponding to an intent associated with the pixel;

using a second neural network model, identifying one or more traffic entities within the first image based upon the first image and the first feature map of the first image, wherein a traffic entity represents an object moving through the traffic;

determining, for each of the identified one or more traffic entities, an overall intent of the traffic entity, based upon an aggregation of the feature vectors corresponding to pixels encompassed by the traffic entity, wherein the overall intent represents an action that the traffic entity is expected to perform in the traffic; and

generating one or more commands to control the vehicle based upon the determined overall intents of the one or more traffic entities.

2. The method of claim 1 , wherein a feature vector associated with a pixel corresponds to a plurality of intents associated with the pixel, each intent associated with a predicted value representative of a statistical distribution of the intent and an uncertainty value associated with the predicted value.

3. The method of claim 1 , wherein identifying the one or more traffic entities using the second neural network model, further comprises:

performing object recognition on the first image to generate a bounding box around each of the one or more traffic entities in the first image, the bounding box encompassing a plurality of pixels representative of the traffic entity.

4. The method of claim 1 , wherein the overall intent of the traffic entity is based on a relationship with one or more other traffic entities in the first image.

5. The method of claim 4 , wherein the relationship is based on relative positions of the traffic entity with the one or more other traffic entities in the first image.

6. The method of claim 1 , wherein the overall intent of the object is based on a presence or absence of other traffic entities in the first image.

7. The method of claim 1 , further comprising:

generating, using the first neural network model, a second feature map for a second image that includes a traffic entity from the one or more traffic entities within the first image, the second feature map indicating a feature vector for each pixel encompassed by the traffic entity within the second image.

8. The method of claim 7 , further comprising:

identifying, using the second neural network model, the traffic entity within the second image based upon the second image and the second feature map of the second image; and

determining an updated overall intent of the traffic entity based upon the aggregation of the feature vectors corresponding to the pixels encompassed by the traffic entity within the second image.

9. A non-transitory computer readable medium storing instructions that when executed by a processor cause the processor to perform steps comprising:

receiving a plurality of images, each image corresponding to a video frame captured by one or more sensors of a vehicle moving through traffic;

processing the plurality of images using a first neural network model configured to generate, for a first image of the plurality of images, a first feature map indicating, for each pixel of the first image, a feature vector corresponding to an intent associated with the pixel;

using a second neural network model, identifying one or more traffic entities within the first image based upon the first image and the first feature map of the first image, wherein a traffic entity represents an object moving through the traffic;

determining, for each of the identified one or more traffic entities, an overall intent of the traffic entity, based upon an aggregation of the feature vectors corresponding to pixels encompassed by the traffic entity, wherein the overall intent represents an action that the traffic entity is expected to perform in the traffic; and

generating one or more commands to control the vehicle based upon the determined overall intents of the one or more traffic entities.

10. The non-transitory computer readable medium of claim 9 , wherein a feature vector associated with a pixel corresponds to a plurality of intents associated with the pixel, each intent associated with a predicted value representative of a statistical distribution of the intent and an uncertainty value associated with the predicted value.

11. The non-transitory computer readable medium of claim 9 , wherein identifying the one or more traffic entities using the second neural network model, further comprises:

performing object recognition on the first image to generate a bounding box around each of the one or more traffic entities in the first image, the bounding box encompassing a plurality of pixels representative of the traffic entity.

12. The non-transitory computer readable medium of claim 9 , wherein the overall intent of the traffic entity is based on a relationship with one or more other traffic entities in the first image.

13. The non-transitory computer readable medium of claim 12 , wherein the relationship is based on relative positions of the traffic entity with the one or more other traffic entities in the first image.

14. The non-transitory computer readable medium of claim 9 , wherein the overall intent of the traffic entity is based on a presence or absence of other traffic entities in the first image.

15. The non-transitory computer readable medium of claim 9 , further storing instructions that cause the processor to perform the step of:

generating, using the first neural network model, a second feature map for a second image that includes a traffic entity from the one or more traffic entities within the first image, the second feature map indicating a feature vector for each pixel encompassed by the traffic entity within the second image.

16. The non-transitory computer readable medium of claim 15 , further storing instructions that cause the processor to perform the steps of:

identifying, using the second neural network model, the traffic entity within the second image based upon the second image and the second feature map of the second image; and

determining an updated overall intent of the object based upon the aggregation of the feature vectors corresponding to the pixels encompassed by the traffic entity within the second image.

17. A system comprising:

a hardware processor; and

a non-transitory computer readable medium storing instructions that when executed by the hardware processor cause the processor to perform steps comprising:

receiving a plurality of images, each image corresponding to a video frame captured by one or more sensors of a vehicle moving through traffic;

processing the plurality of images using a first neural network model configured to generate, for a first image of the plurality of images, a first feature map indicating, for each pixel of the first image, a feature vector corresponding to an intent associated with the pixel;

using a second neural network model, identifying one or more traffic entities within the first image based upon the first image and the first feature map of the first image, wherein a traffic entity represents an object moving through the traffic;

determining, for each of the identified one or more traffic entities, an overall intent of the traffic entity, based upon an aggregation of the feature vectors corresponding to pixels encompassed by the traffic entity, wherein the overall intent represents an action that the traffic entity is expected to perform in the traffic; and

generating one or more commands to control the vehicle based upon the determined overall intents of the one or more traffic entities.

18. The system of claim 17 , wherein a feature vector associated with a pixel corresponds to a plurality of intents associated with the pixel, each intent associated with a predicted value representative of a statistical distribution of the intent and an uncertainty value associated with the predicted value.

Assignments (3)
PATENT SECURITY AGREEMENT Recorded Mar 25, 2025
From: PERCEPTIVE AUTOMATA LLC
To: PICCADILLY PATENT FUNDING LLC, AS SECURITY HOLDER
Reel/Frame 070614/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2025
From: PERCEPTIVE AUTOMATA, INC.
To: PERCEPTIVE AUTOMATA LLC
Reel/Frame 070267/0727 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2020
From: MCCURRIE, MEL
To: PERCEPTIVE AUTOMATA, INC.
Reel/Frame 054714/0659 →
Continuity (2)
Provisional Application 62916727 · Oct 17, 2019
Related Publication 20210114627A1 · Apr 22, 2021