IP Library › Granted Patent US 12,269,462
Granted Patent B1
US 12,269,462 · App. 16/820,378 · Granted Apr 8, 2025

Spatial prediction

Inventors: Gowtham Garimella (Burlingame, CA); Marin Kobilarov (Mountain View, CA); Kai Zhenyu Wang (Foster City, CA)
Assignee: Zoox, Inc.
B60W30/06B62D15/0285G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,269,462
App. No.
16/820,378
Filed
Mar 16, 2020
Granted
Apr 8, 2025
Kind
B1
Art Unit
3661
USPC
701/70
Abstract

Techniques relating to determining regions based on intents of objects are described. In an example, a computing device onboard a first vehicle can receive sensor data associated with an environment of the first vehicle. The computing device can determine, based on the sensor data, a region associated with a second vehicle proximate the first vehicle that is to be occupied by the second vehicle while the vehicle performs a maneuver. Further, the computing device can determine an instruction for controlling the first vehicle based at least in part on the region.

Claims (84)

1. A method comprising:

receiving sensor data associated with an environment of a first vehicle;

determining, based at least in part on the sensor data, a type of intent of a second vehicle proximate the first vehicle, wherein the intent is associated with a maneuver to be performed by the second vehicle during a future time, the maneuver comprising a plurality of trajectories within a first region in the environment;

determining, based at least in part on the intent, an entirety of time that the second vehicle performs the maneuver;

determining that the type of intent satisfies a temporal condition;

determining, based at least in part on the sensor data and the type of intent satisfying the temporal condition, a second region of the environment around and encompassing the second vehicle, wherein the second region includes areas in the environment that are likely to be occupied by the second vehicle for the entirety of time and within the first region;

determining, based at least in part on the sensor data and the second region of the environment, a blocked region corresponding to the first region that the first vehicle should not enter; and

determining an instruction for controlling the first vehicle based at least in part on the blocked region.

2. The method as claim 1 recites, further comprising:

generating, based at least in part on the sensor data, a multi-channel image comprising a top-down representation of the environment, a channel of the multi-channel image being associated with a feature of the second vehicle or the environment; and

determining the second region of the environment around the second vehicle based at least in part on analyzing the multi-channel image using a convolutional neural network.

3. The method as claim 1 recites, wherein the instruction for controlling the first vehicle comprises an instruction to cause the first vehicle to at least one of:

decelerate;

stop;

perform a lane change; or

send a signal to a remote operator.

4. The method as claim 1 recites, wherein determining the second region of the environment around the second vehicle is performed at a first time, the method further comprising:

receiving, at a second time, updated sensor data;

determining, based at least in part on the updated sensor data, an updated second region of the environment around the second vehicle that is to be occupied by the second vehicle; and

determining an updated instruction for controlling the first vehicle based at least in part on the updated second region of the environment.

5. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions, that when executed by the one or more processors, cause the system to perform operations comprising:

receiving sensor data associated with an environment of a vehicle;

determining, based at least in part on the sensor data, a type of action performed at a future time by an object proximate the vehicle;

determining, based at least in part on the action, an entirety of time that the object performs the action;

determining that the type of action satisfies a temporal condition;

receiving, based at least in part on the sensor data and the type of action satisfying the temporal condition, a first region of the environment encompassing the object and that the object proximate the vehicle is likely to occupy at the future time for the entirety of time that the object performs the action, the action comprising a plurality of trajectories within the first region;

determining, based at least in part on the sensor data and the first region of the environment, a blocked region corresponding to the first region that the vehicle should not enter; and

determining an instruction for controlling the vehicle based at least in part on the blocked region.

6. The system as claim 5 recites, the operations further comprising:

determining, based at least in part on the sensor data, a confidence score associated with an intent of the object; and

determining that the confidence score meets or exceeds a threshold, wherein the first region of the environment is determined based at least in part on the intent, and wherein the action is associated with the intent.

7. The system as claim 5 recites, the operations further comprising:

determining, based at least in part on the sensor data, a plurality of intents of the object, wherein a first intent of the plurality of intents is associated with a first confidence score and a second intent of the plurality of intents is associated with a second confidence score;

determining that the first confidence score and the second confidence score meet or exceed a threshold;

determining a third region of the environment associated with the object that is to be occupied by the object while the object performs a first action associated with the first intent; and

determining a fourth region of the environment associated with the object that is to be occupied by the object while the object performs a second action associated with the second intent, wherein the first region is determined based on at least one of the third region or the fourth region.

8. The system as claim 5 recites, the operations further comprising:

generating, based at least in part on the sensor data, an input comprising a top-down representation of the environment, wherein the top-down representation comprises a multi-channel image;

inputting the input into a convolutional neural network; and

receiving the first region of the environment associated with the object from the convolutional neural network.

9. The system as claim 8 recites, the operations further comprising selecting the convolutional neural network based at least in part on the action.

10. The system as claim 5 recites, the operations further comprising determining an occupancy grid, wherein a tile associated with the occupancy grid is associated with a confidence score indicative of whether a portion of the environment is likely to be occupied by the object, and wherein the first region of the environment is defined based at least in part on the occupancy grid.

11. The system as claim 10 recites, the operations further comprising inputting the occupancy grid into a planner component associated with the vehicle to determine the instruction for controlling the vehicle.

12. The system as claim 5 recites, wherein the vehicle is an autonomous vehicle and wherein the instruction for controlling the vehicle comprises an instruction to cause the vehicle to at least one of:

decelerate;

stop;

perform a lane change; or

send a signal to a remote operator.

13. The system as claim 5 recites, wherein determining the first region of the environment associated with the object is performed at a first time, the operations further comprising:

receiving, at a second time, updated sensor data;

determining, based at least in part on the updated sensor data and an associated intent, an updated first region of the environment associated with the object that is to be occupied by the object; and

determining an updated instruction for controlling the vehicle based at least in part on the updated first region of the environment.

14. One or more non-transitory computer-readable media storing instructions, that when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving sensor data associated with an environment of a vehicle;

determining, based at least in part on the sensor data, an action performed at a future time by an object proximate the vehicle, wherein the action is a parallel parking maneuver;

determining, based at least in part on the action, an entirety of time that the object performs the action;

receiving, based at least in part on the sensor data, a first region of the environment encompassing the object and that the object proximate the vehicle is likely to occupy at the future time for the entirety of time that the object performs the action, the action comprising a plurality of trajectories within the first region;

determining, based at least in part on the sensor data and the first region of the environment, a blocked region corresponding to the first region that the vehicle should not enter; and

determining an instruction for controlling the vehicle based at least in part on the blocked region.

15. The one or more non-transitory computer-readable media as claim 14 recites, the operations further comprising:

generating, based at least in part on the sensor data, an input comprising a top-down representation of the environment, wherein the top-down representation comprises a multi-channel image;

inputting the input into a convolutional neural network, wherein the convolutional neural network is selected based at least in part on the action; and

receiving the first region of the environment associated with the object from the convolutional neural network.

16. The one or more non-transitory computer-readable media as claim 15 recites, wherein the convolutional neural network is trained based at least in part on:

receiving log data;

detecting, based at least in part on the log data, a second object that performs the action;

determining a start time associated with the action and a stop time associated with the action;

determining a third region of the environment occupied by the second object during a period of time between the start time and the stop time; and

training the convolutional neural network based at least in part on the third region of the environment.

17. The one or more non-transitory computer-readable media as claim 14 recites, wherein the first region of the environment is determined based at least in part on the action and at least one other action, and wherein each action is associated with an intent that is associated with a confidence score that meets or exceeds a threshold.

18. The one or more non-transitory computer-readable media as claim 14 recites, the operations further comprising determining an occupancy grid, wherein a tile associated with the occupancy grid is associated with a confidence score indicative of whether a portion of the environment is likely to be occupied by the object, and wherein the first region of the environment is defined based at least in part on the occupancy grid.

19. The method of claim 1 , wherein:

the plurality of trajectories includes at least one trajectory for a current time and at least one subsequent trajectory for the future time, and

the instruction for controlling the first vehicle includes determining a trajectory for the first vehicle at the current time that avoids the blocked region.

20. The method of claim 1 , wherein:

the maneuver includes at least one of:

parallel parking,

a three-point turn,

a U-turn,

a K-turn, or

an N-point turn, and

the plurality of trajectories occur within the first region.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2020
From: GARIMELLA, GOWTHAM; KOBILAROV, MARIN; WANG, KAI ZHENYU
To: ZOOX, INC.
Reel/Frame 052659/0034 →
References Cited (25)
US 9248834B1 · Ferguson · 2016 [cited by examiner]
US 9517767B1 · Kentley · 2016 [cited by examiner]
US 9612123B1 · Levinson et al. · 2017 [cited by applicant]
US 10353390B2 · Linscott et al. · 2019 [cited by applicant]
US 10421453B1 · Ferguson · 2019 [cited by examiner]
US 20170120803A1 · Kentley · 2017 [cited by examiner]
US 20190101924A1 · Styler · 2019 [cited by examiner]
US 20190243371A1 · Nister · 2019 [cited by examiner]
US 20190382007A1 · Casas · 2019 [cited by examiner]
US 20200074266A1 · Peake · 2020 [cited by examiner]
US 20200086863A1 · Rosman · 2020 [cited by examiner]
US 20210031762A1 · Matsunaga · 2021 [cited by examiner]
US 20210064040A1 · Yadmellat · 2021 [cited by examiner]
US 20210188316A1 · Marchetti-Bowick · 2021 [cited by examiner]
US 20210197813A1 · Houston · 2021 [cited by examiner]
US 20210253132A1 · Coimbra De Andrade · 2021 [cited by examiner]
US 20210261116A1 · Hosokawa · 2021 [cited by examiner]
US 20210276587A1 · Urtasun · 2021 [cited by examiner]
H.-S. Jeon, D.-S. Kum and W.-Y. Jeong, “Traffic Scene Prediction via Deep Learning: Introduction of Multi-Channel Occupancy Grid Map as a Scene Representation,” 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, C… [cited by examiner]
U.S. Appl. No. 15/982,658, filed May 17, 2018, Lee, et al. “Vehicle Lighting State Determination”, 40 pages. [cited by applicant]
U.S. Appl. No. 16/420,050, filed May 22, 2019, Hong et al., “Trajectory Prediction on Top-Down Scenes and Associated Model”, 60 pages. [cited by applicant]
U.S. Appl. No. 16/504,147, filed Jul. 5, 2019, Garimella et al., “Prediction on Top-Down Scenes Based on Action Data”, 51 pages. [cited by applicant]
U.S. Appl. No. 16/709,263, filed Dec. 19, 2019, Thalman et al., “Determining Bias of Vehicle Axles”, 39 pages. [cited by applicant]
U.S. Appl. No. 16/803,644, filed Feb. 27, 2020, Haggblade, et al., “Perpendicular Cut-In Training”, 40 pages. [cited by applicant]
U.S. Appl. No. 16/803,705, filed Feb. 27, 2020, Haggblade, et al., “Perpendicular Cut-In Detection”, 39 pages. [cited by applicant]
Cited By (5)
US 12,352,857 US 12,466,394 US 12,472,984 US 12,565,192 US 12,637,102