IP Library › Granted Patent US 12,051,276
Granted Patent B1
US 12,051,276 · App. 18/205,651 · Granted Jul 30, 2024

Machine-learned model training for pedestrian attribute and gesture detection

Inventors: Oytun Ulutan (Buena Park, CA); Xin Wang (Sunnyvale, CA); Kratarth Goel (Albany, CA); Vasiliy Karasev (San Francisco, CA); Sarah Tariq (Palo Alto, CA); Yi Xu (Pasadena, CA)
Assignee: Zoox, Inc.
G06V40/28G06F18/2148G06F18/217G06F18/24G06T7/70G06V20/582G06V40/103G06V40/23B60W60/001B60W2420/403B60W2540/041G05D1/0088G05D1/0231G06N20/00G06T2207/20081G06T2207/30196G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,051,276
App. No.
18/205,651
Granted
Jul 30, 2024
Kind
B1
Abstract

Techniques for detecting attributes and/or gestures associated with pedestrians in an environment are described herein. The techniques may include receiving sensor data associated with a pedestrian in an environment of a vehicle and inputting the sensor data into a machine-learned model that is configured to determine a gesture and/or an attribute of the pedestrian. Based on the input data, an output may be received from the machine-learned model that indicates the gesture and/or the attribute of the pedestrian and the vehicle may be controlled based at least in part on the gesture and/or the attribute of the pedestrian. The techniques may also include training the machine-learned model to detect the attribute and/or the gesture of the pedestrian.

Claims (34)

1. A vehicle comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions that, when executed, cause the one or more processors to perform operations comprising:

receiving sensor data associated with a person in an environment of the vehicle;

determining, based at least in part on analyzing the sensor data using a machine-learned model, that the person is controlling a traffic sign, the traffic sign indicative of a traffic rule that is to be followed by the vehicle;

determining, based at least in part on a gesture performed by the person, a state associated with the traffic sign, wherein the state associated with the traffic sign is at least partially indicated by an orientation of the traffic sign; and

based at least in part on the state indicating that the traffic rule is in effect, causing the vehicle to perform an action associated with the traffic rule.

2. The vehicle of claim 1 , wherein analyzing the sensor data comprises analyzing a series of frames of the sensor data captured over a period of time, each frame of the series of frames representing the person at a different instance of time during the period of time.

3. The vehicle of claim 1 , the operations further comprising determining, based at least in part on analyzing the sensor data using the machine-learned model, an attribute associated with the person, wherein causing the vehicle to perform the action is further based at least in part on the attribute.

4. The vehicle of claim 3 , wherein the attribute associated with the person comprises at least one of a classification of the person, a pose of the person, or an action of the person.

5. The vehicle of claim 1 , wherein causing the vehicle to perform the action comprises providing, to a planner component of the vehicle, an indication of a type associated with the traffic sign and the state associated with the traffic sign.

6. The vehicle of claim 1 , wherein the gesture performed by the person comprises at least one of the person raising the traffic sign or the person lowering the traffic sign.

7. The vehicle of claim 1 , wherein the gesture performed by the person is determined based at least in part on analyzing of the sensor data using the machine-learned model.

8. The vehicle of claim 1 , wherein the sensor data comprises at least one of image data, lidar data, radar data, or time of flight data.

9. A method comprising:

receiving sensor data associated with a person in an environment of a vehicle;

determining, based at least in part on analyzing the sensor data using a machine-learned model, that the person is controlling a traffic sign, the traffic sign indicative of a traffic rule that is to be followed by the vehicle;

determining, based at least in part on a gesture performed by the person, a state associated with the traffic sign, wherein the state associated with the traffic sign is at least partially indicated by an orientation of the traffic sign; and

based at least in part on the state indicating that the traffic rule is in effect, causing the vehicle to perform an action associated with the traffic rule.

10. The method of claim 9 , wherein analyzing the sensor data comprises analyzing a series of frames of the sensor data captured over a period of time, each frame of the series of frames representing the person at a different instance of time during the period of time.

11. The method of claim 9 , further comprising determining, based at least in part on analyzing the sensor data using the machine-learned model, an attribute associated with the person, wherein causing the vehicle to perform the action is further based at least in part on the attribute.

12. The method of claim 11 , wherein the attribute associated with the person comprises at least one of a classification of the person, a pose of the person, or an action of the person.

13. The method of claim 9 , wherein causing the vehicle to perform the action comprises providing, to a planner component of the vehicle, an indication of a type associated with the traffic sign and the state associated with the traffic sign.

14. The method of claim 9 , wherein the gesture performed by the person is determined based at least in part on analyzing the sensor data using the machine-learned model.

15. The method of claim 9 , wherein the sensor data comprises at least one of image data, lidar data, radar data, or time of flight data.

16. The method of claim 9 , wherein the action includes a change in trajectory in accordance with the traffic rule.

17. One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform operations comprising:

receiving sensor data associated with a person in an environment of a vehicle;

determining, based at least in part on analyzing the sensor data using a machine-learned model, that the person is controlling a traffic sign, the traffic sign indicative of a traffic rule that is to be followed by the vehicle;

determining, based at least in part on a gesture performed by the person, a state associated with the traffic sign, wherein the state associated with the traffic sign is at least partially indicated by an orientation of the traffic sign; and

based at least in part on the state indicating that the traffic rule is in effect, causing the vehicle to perform an action associated with the traffic rule.

18. The one or more non-transitory computer-readable media of claim 17 , wherein analyzing the sensor data comprises analyzing a series of frames of the sensor data captured over a period of time, each frame of the series of frames representing the person at a different instance of time during the period of time.

19. The one or more non-transitory computer-readable media of claim 17 , the operations further comprising determining, based at least in part on analyzing the sensor data using the machine-learned model, an attribute associated with the person, wherein causing the vehicle to perform the action is further based at least in part on the attribute.

20. The one or more non-transitory computer-readable media of claim 17 , wherein causing the vehicle to perform the action comprises providing, to a planner component of the vehicle, an indication of a type associated with the traffic sign and the state associated with the traffic sign.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2023
From: ULUTAN, OYTUN; WANG, XIN; GOEL, KRATARTH; KARASEV, VASILIY; TARIQ, SARAH; XU, YI
To: ZOOX, INC.
Reel/Frame 063857/0159 →
Continuity (3)
Continuation 17320690 · May 14, 2021
Provisional Application 63117263 · Nov 23, 2020
Provisional Application 63028377 · May 21, 2020
Cited By (1)
US 12,198,071