IP Library › Granted Patent US 12,183,204
Granted Patent B2
US 12,183,204 · App. 17/542,880 · Granted Dec 31, 2024

Trajectory prediction on top-down scenes and associated model

Inventors: Xi Joey Hong (Campbell, CA); Benjamin John Sapp (San Francisco, CA); James William Vaisey Philbin (Palo Alto, CA); Kai Zhenyu Wang (Foster City, CA)
Assignee: Zoox, Inc.
G08G1/164B60W30/0956G05D1/0221G05D1/0276G06N3/08G06N20/00G06T7/292G08G1/166G06T2207/10032G06T2207/30236G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,183,204
App. No.
17/542,880
Granted
Dec 31, 2024
Kind
B2
Abstract

Techniques are discussed for determining prediction probabilities of an object based on a top-down representation of an environment. Data representing objects in an environment can be captured. Aspects of the environment can be represented as map data. A multi-channel image representing a top-down view of object(s) in the environment can be generated based on the data representing the objects and map data. The multi-channel image can be used to train a machine learned model by minimizing an error between predictions from the machine learned model and a captured trajectory associated with the object. Once trained, the machine learned model can be used to generate prediction probabilities of objects in an environment, and the vehicle can be controlled based on such prediction probabilities.

Claims (44)

1. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising:

receiving map data associated with an environment;

receiving object data associated with an object in the environment;

generating, based at least in part on the map data and the object data, a multi-channel data structure representing the environment from a top-down perspective, wherein a channel of the multi-channel data structure includes a portion of the object data, wherein the portion of the object data comprises at least one of a semantic label, a bounding box associated with the object, a velocity associated with the object, or an acceleration associated with the object;

inputting the multi-channel data structure into a machine learned model;

receiving, from the machine learned model and based at least in part on the multi-channel data structure, a prediction probability associated with the object; and

controlling, based at least in part on the prediction probability, a vehicle to traverse the environment.

2. The system of claim 1 , wherein at least some of the object data is included in another channel of the multi-channel data structure.

3. The system of claim 1 , wherein the object data comprises semantic information indicative of at least one of a bounding box associated with the object, movement information associated with the object, or a classification associated with the object.

4. The system of claim 1 , wherein a classification associated with the object is represented as a color in the channel, and wherein an intensity of the color is indicative of a magnitude of an attribute of the object.

5. The system of claim 1 , wherein another channel of the multi-channel data structure includes semantic information associated with the environment, the semantic information comprising at least one of road network information or a traffic light status.

6. The system of claim 1 , wherein the portion of the object data included in the channel of the multi-channel data structure is encoded with an attribute indicating at least one of a color or a location associated with the object.

7. The system of claim 1 , wherein the prediction probability comprises at least one of:

a multi modal Gaussian trajectory; or

an occupancy grid associated with a future time, wherein a cell of the occupancy grid is indicative of a probability of the object being in a region associated with the cell at the future time.

8. The system of claim 1 , wherein the multi-channel data structure further comprises one or more additional channels including semantic information associated with at least one of the environment or objects in the environment.

9. A method comprising:

receiving map data associated with an environment;

receiving object data associated with an object in the environment;

generating, based at least in part on the map data and the object data, a multi-channel data structure representing the environment from a top-down perspective, wherein a channel of the multi-channel data structure includes a portion of the object data, wherein the portion of the object data comprises at least one of a semantic label, a bounding box associated with the object, a velocity associated with the object, or an acceleration associated with the object;

inputting the multi-channel data structure into a machine learned model;

receiving, from the machine learned model and based at least in part on the multi-channel data structure, a prediction probability associated with the object; and

controlling, based at least in part on the prediction probability, a vehicle to traverse the environment.

10. The method of claim 9 , wherein at least some of the object data is included in another channel of the multi-channel data structure.

11. The method of claim 9 , wherein the object data comprises semantic information indicative of at least one of a bounding box associated with the object, movement information associated with the object, or a classification associated with the object.

12. The method of claim 9 , wherein a classification associated with the object is represented as a color in the channel, and wherein an intensity of the color is indicative of a magnitude of an attribute of the object.

13. The method of claim 9 , wherein another channel of the multi-channel data structure includes semantic information associated with the environment, the semantic information comprising at least one of road network information or a traffic light status.

14. The method of claim 9 , wherein the portion of the object data included in the channel of the multi-channel data structure is encoded with an attribute indicating at least one of a color or a location associated with the object.

15. The method of claim 9 , wherein the prediction probability comprises at least one of:

a multi modal Gaussian trajectory; or

an occupancy grid associated with a future time, wherein a cell of the occupancy grid is indicative of a probability of the object being in a region associated with the cell at the future time.

16. The method of claim 9 , wherein the multi-channel data structure further comprises one or more additional channels including semantic information associated with at least one of the environment or objects in the environment.

17. A non-transitory computer-readable medium storing instructions that, when executed, cause one or more processors to perform operations comprising:

receiving map data associated with an environment;

receiving object data associated with an object in the environment;

generating, based at least in part on the map data and the object data, a multi-channel data structure representing the environment from a top-down perspective, wherein a channel of the multi-channel data structure includes a portion of the object data, wherein the portion of the object data comprises at least one of a semantic label, a bounding box associated with the object, a velocity associated with the object, or an acceleration associated with the object;

inputting the multi-channel data structure into a machine learned model;

receiving, from the machine learned model and based at least in part on the multi-channel data structure, a prediction probability associated with the object; and

controlling, based at least in part on the prediction probability, a vehicle to traverse the environment.

18. The non-transitory computer-readable medium of claim 17 , wherein at least some of the object data is included in another channel of the multi-channel data structure.

19. The non-transitory computer-readable medium of claim 17 , wherein the object data is indicative of at least one of a bounding box associated with the object, movement information associated with the object, or a classification associated with the object.

20. The non-transitory computer-readable medium of claim 17 , wherein another channel of the multi-channel data structure includes semantic information associated with the environment, the semantic information comprising at least one of road network information or a traffic light status.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2021
From: HONG, XI JOEY; SAPP, BENJAMIN JOHN; PHILBIN, JAMES WILLIAM VAISEY; WANG, KAI ZHENYU
To: ZOOX, INC.
Reel/Frame 058316/0234 →
Continuity (3)
Continuation 16420050 · May 22, 2019
Continuation In Part 16151607 · Oct 4, 2018
Related Publication 20220092983A1 · Mar 24, 2022