IP Library Granted Patent US 12,158,762
Granted Patent B1
US 12,158,762 · App. 18/244,770 · Granted Dec 3, 2024

Visual language models for perception

Inventors: Stephen O'Hara (Fort Collins, CO); Ariel Quinn (Boulder, CO)
Assignee: AURORA OPERATIONS, INC.
G05D1/0248G05D1/0219G06F40/284G06F40/56G06V20/582G06V20/588
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,158,762
App. No.
18/244,770
Granted
Dec 3, 2024
Kind
B1
Abstract

A method is provided, that includes: receiving camera data from a perception system of an autonomous vehicle; and providing the camera data to a visual language model, where the visual language model includes a mapping of a corpus of images and a corpus of text to a common parameter space. The method further includes: receiving from the visual language model an output corresponding to one or more text tokens; accessing a configuration file comprising a plurality of text tokens representing a plurality of objects or events of interest to the autonomous vehicle; and identifying a respective object or event of interest in an environment of the autonomous vehicle by determining that a text token of the output matches a respective one of the plurality of text tokens in the configuration file. The autonomous vehicle can then be controlled based at least in part on the respective object or event of interest.

Claims (42)

1. A computer-implemented method, comprising:

receiving camera data from a perception system of an autonomous vehicle;

providing the camera data to a visual language model, the visual language model including an image encoder and a text decoder and at least comprising a mapping of a corpus of images and a corpus of text to a common parameter space;

processing, as input, an image from the camera data, using the image encoder of the visual language model, to generate an image embedding for the image in the common parameter space;

processing, as input, the image embedding for the image, using the text decoder of the visual language model, to generate an output comprising one or more text tokens;

accessing a configuration file comprising a plurality of text tokens describing a plurality of objects or events of interest to the autonomous vehicle;

identifying a respective object or event of interest in an environment of the autonomous vehicle by determining that a text token of the output of the text decoder matches a respective one of the plurality of text tokens in the configuration file that represents the respective object or event; and

controlling the autonomous vehicle based at least in part on the respective object or event of interest.

2. The method of claim 1 , wherein the image captures the respective object or event of interest in the environment of the autonomous vehicle.

3. The method of claim 2 , wherein the image capturing the respective object or event of interest in the environment of the autonomous vehicle is not part of the corpus of images.

4. The method of claim 1 , wherein the visual language model includes one or more neural networks.

5. The method of claim 1 , wherein the image encoder embeds the corpus of images into the common parameter space.

6. The method of claim 1 , wherein the corpus of text includes a natural language description for each image of the corpus of images.

7. The method of claim 6 , wherein the natural language description for each image of the corpus of images describes an object or event.

8. The method of claim 1 , wherein the respective object or event of interest includes a road sign, a road condition, or a moving object within the environment of the autonomous vehicle.

9. The method of claim 1 , wherein the perception system of the autonomous vehicle includes one or more vision sensors capturing the image data, or one or more non-vision sensors.

10. An autonomous vehicle control system for controlling an autonomous vehicle in an environment, the autonomous vehicle control system comprising:

one or more processors; and

memory storing instructions that, when executed by one or more of the processors, cause the autonomous vehicle control system to:

receive sensor data from a perception system of the autonomous vehicle;

provide the sensor data to a visual language model, the visual language model including an image encoder and a text decoder and at least comprising a mapping of a corpus of images and a corpus of text to a common parameter space;

process, as input, an image from the sensor data, using the image encoder of the visual language model, to generate an image embedding for the image in the common parameter space;

process, as input, the image embedding for the image, using the text decoder of the visual language model, to generate an output comprising one or more text tokens;

access a configuration file comprising a plurality of text tokens representing a plurality of objects or events of interest to the autonomous vehicle;

identify a respective object or event of interest in an environment of the autonomous vehicle by determining that a text token of the output of the text decoder matches a respective one of the plurality of text tokens in the configuration file that represents the respective object or event; and

control the autonomous vehicle based at least in part on the respective object or event of interest.

11. The system of claim 10 , wherein the image captures the respective object or event of interest in the environment of the autonomous vehicle, and wherein the image is an RGB image, a LIDAR image, or a radar image.

12. The system of claim 11 , wherein the image capturing the respective object or event of interest in the environment of the autonomous vehicle is not part of the corpus of images.

13. The system of claim 10 , wherein the visual language model includes one or more neural networks.

14. The system of claim 10 , wherein the image encoder embeds the corpus of images into the common parameter space.

15. The system of claim 10 , wherein the corpus of text includes a natural language description for each image of the corpus of images.

16. The system of claim 15 , wherein the natural language description for each of the corpus of images describes an object or event.

17. The system of claim 10 , wherein the respective object or event of interest includes a road sign, a road condition, or a moving object within the environment of the autonomous vehicle.

18. The system of claim 10 , wherein the perception system of the autonomous vehicle includes one or more vision sensors that capture the image-sensor data, or one or more non-vision sensors.

19. A computer-implemented method, comprising:

receiving camera data from a perception system of an autonomous vehicle, the camera data including an image depicting an environment of the autonomous vehicle;

determining, from the camera data, whether the environment of the autonomous vehicle includes any object or event that is of interest to the autonomous vehicle, wherein the determining includes:

providing the camera data to a visual language model that includes an image encoder and a text decoder, and

processing, using the image encoder of the visual language model, the image from the camera data as input, to generate an image embedding for the image,

processing the image embedding for the image as input, using the text decoder of the visual language model, to generate an output from the text decoder of the visual language model that comprises one or more text tokens indicating whether an object or event of interest to the autonomous vehicle exists within the environment of the autonomous vehicle; and

in response to the output of the text decoder of the visual language model indicating that the object or event of interest to the autonomous vehicle exists within the environment of the autonomous vehicle, controlling the autonomous vehicle based at least in part on the respective object or event of interest to the autonomous vehicle.

20. The method of claim 1 , wherein the text decoder includes a text generation machine learning model that generates text based on text embeddings and/or image embeddings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2023
From: O'HARA, STEPHEN; QUINN, ARIEL
To: AURORA OPERATIONS, INC.
Reel/Frame 065697/0354 →
Cited By (7)
US 12,423,984 US 12,430,849 US 12,456,299 US 12,528,507 US 12,576,880 US 12,731,081 US 12,731,400