IP Library › Granted Patent US 12,260,651
Granted Patent B2
US 12,260,651 · App. 18/621,922 · Granted Mar 25, 2025

Machine-learned architecture for efficient object attribute and/or intention classification

Inventors: Subhasis Das (San Mateo, CA); Oytun Ulutan (Buena Park, CA); Yi-Ting Lin (Foster City, CA); Derek Xiang Ma (San Carlos, CA)
Assignee: Zoox, Inc.
G06V20/58G05D1/0221G05D1/0246G05D1/249G06N3/04G06N3/08G06V40/23
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,651
App. No.
18/621,922
Granted
Mar 25, 2025
Kind
B2
Abstract

A system for faster object attribute and/or intent classification may include an machine-learned (ML) architecture that processes temporal sensor data (e.g., multiple instances of sensor data received at different times) and includes a cache in an intermediate layer of the ML architecture. The ML architecture may be capable of classifying an object's intent to enter a roadway, idling near a roadway, or active crossing of a roadway. The ML architecture may additionally or alternatively classify indicator states, such as indications to turn, stop, or the like. Other attributes and/or intentions are discussed herein.

Claims (98)

1. A method comprising:

receiving first sensor data associated with an object;

receiving second sensor data associated with the object;

determining, by a first machine-learned model and based at least in part on the first sensor data, a first output;

storing the first output in a cache;

retrieving, from the cache a second output previously generated by the first machine-learned model and based at least in part on the second sensor data;

aggregating, as an aggregated output, the first output and the second output;

determining, by a second machine-learned model and based at least in part on the aggregated output, an attribute associated with the object in an environment, the attribute indicating at least one of a motion of the object, a state of the object, a predicted intent of the object, or an association of the object with another object; and

controlling a vehicle based at least in part on the attribute.

2. The method of claim 1 , wherein the attribute indicates at least one of:

an indication of an object motion state,

an indication of an object indicator state,

an indication that the object is idling,

an indication that the object intends to enter a roadway, or

an indication that the object is not associated with the roadway.

3. The method of claim 1 , wherein the first machine-learned model comprises an object detection machine-learned model and a set of machine-learned layers that receives outputs from the object detection machine-learned model.

4. The method of claim 3 , wherein aggregating the first output and the second output comprises:

determining, by the object detection machine-learned model and based at least in part on the first sensor data, a first intermediate output;

determining, by the object detection machine-learned model and based at least in part on the second sensor data, a second intermediate output;

determining, based at least in part on a first machine-learned layer of the set of machine-learned layers and the first intermediate output, a third intermediate output;

determining, based at least in part on a second machine-learned layer of the set of machine-learned layers and the second intermediate output, a fourth intermediate output; and

determining the aggregated output by a third machine-learned layer.

5. The method of claim 1 , wherein:

the first sensor data is a first subset of a first set of sensor data;

the second sensor data is a second subset of a second set of sensor data; and

the first subset and the second subset are determined by a pre-processing machine-learned component based at least in part on the first set of sensor data and the second set of sensor data.

6. The method of claim 1 , wherein the cache stores n number of outputs of the first machine-learned model, wherein n is a positive integer associated with n previous time steps.

7. A system comprising:

one or more processors; and

a memory storing processor-executable instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

receiving first sensor data associated with an object;

receiving second sensor data associated with the object;

determining, by a first machine-learned model and based at least in part on the first sensor data, a first output, wherein the first machine-learned model comprises an object detection machine-learned model and a set of machine-learned layers that receives outputs from the object detection machine-learned model;

determining, by the first machine-learned model and based at least in part on the second sensor data, a second output;

aggregating, as an aggregated output, the first output and the second output;

determining, by a second machine-learned model and based at least in part on the aggregated output, an attribute associated with the object in an environment, wherein the attribute indicates at least one of:

an indication of an object motion state,

an indication of an object indicator state,

an indication that the object is idling,

an indication that the object intends to enter a roadway, or

an indication that the object is not associated with the roadway; and

controlling a vehicle based at least in part on the attribute.

8. The system of claim 7 , wherein aggregating the first output and the second output comprises:

determining, by the object detection machine-learned model and based at least in part on the first sensor data, a first intermediate output;

determining, by the object detection machine-learned model and based at least in part on the second sensor data, a second intermediate output;

determining, based at least in part on a first machine-learned layer of the set of machine-learned layers and the first intermediate output, a third intermediate output;

determining, based at least in part on a second machine-learned layer of the set of machine-learned layers and the second intermediate output, a fourth intermediate output; and

determining the aggregated output by a third machine-learned layer.

9. The system of claim 7 , wherein:

the first sensor data is a first subset of a first set of sensor data;

the second sensor data is a second subset of a second set of sensor data; and

the first subset and the second subset are determined by a pre-processing machine-learned component based at least in part on the first set of sensor data and the second set of sensor data.

10. The system of claim 7 , wherein the operations further comprise retrieving the second output from a memory, wherein the memory is a cache and the cache stores n number of outputs of the first machine-learned model, wherein n is a positive integer associated with n previous time steps.

11. The system of claim 7 , wherein the second sensor data is received before the first sensor data and the second output is determined before the first output.

12. One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving first sensor data associated with an object;

receiving second sensor data associated with the object;

determining, by a first machine-learned model and based at least in part on the first sensor data, a first output;

determining a second output by retrieving the second output from a memory, wherein the memory is a cache and the cache stores n number of outputs of the first machine-learned model, wherein n is a positive integer associated with n previous time steps;

aggregating, as an aggregated output, the first output and the second output;

determining, by a second machine-learned model and based at least in part on the aggregated output, an attribute associated with the object in an environment, wherein the attribute indicates at least one of:

an indication of an object motion state,

an indication of an object indicator state,

an indication that the object is idling,

an indication that the object intends to enter a roadway, or

an indication that the object is not associated with the roadway; and

controlling a vehicle based at least in part on the attribute.

13. The one or more non-transitory computer-readable media of claim 12 , wherein the first machine-learned model comprises an object detection machine-learned model and a set of machine-learned layers that receives outputs from the object detection machine-learned model.

14. The one or more non-transitory computer-readable media of claim 13 , wherein aggregating the first output and the second output comprises:

determining, by the object detection machine-learned model and based at least in part on the first sensor data, a first intermediate output;

determining, by the object detection machine-learned model and based at least in part on the second sensor data, a second intermediate output;

determining, based at least in part on a first machine-learned layer of the set of machine-learned layers and the first intermediate output, a third intermediate output;

determining, based at least in part on a second machine-learned layer of the set of machine-learned layers and the second intermediate output, a fourth intermediate output; and

determining the aggregated output by a third machine-learned layer.

15. The one or more non-transitory computer-readable media of claim 14 , wherein:

the first sensor data is a first subset of a first set of sensor data;

the second sensor data is a second subset of a second set of sensor data; and

the first subset and the second subset are determined by a pre-processing machine-learned component based at least in part on the first set of sensor data and the second set of sensor data.

16. The one or more non-transitory computer-readable media of claim 12 , wherein the second sensor data is received before the first sensor data.

17. The method of claim 1 , further comprising:

determining a confidence level associated with the attribute; and

outputting the attribute based on the confidence level at least one of:

meeting or exceeding a confidence threshold; or

being greater than other confidence scores associated with other attributes determined by the second machine-learned model.

18. The method of claim 1 , wherein the attribute comprises multiple attributes, and the method further comprises:

determining confidence levels associated with individual attributes of the multiple attributes;

outputting a subset of the multiple attributes based at least in part on the confidence levels; and

wherein controlling the vehicle is based at least in part on the subset.

19. The system of claim 7 , the operations further comprising:

determining a confidence level associated with the attribute; and

outputting the attribute based on the confidence level at least one of:

meeting or exceeding a confidence threshold; or

being greater than other confidence scores associated with other attributes determined by the second machine-learned model.

20. The one or more non-transitory computer-readable media of claim 12 , further comprising:

determining a confidence level associated with the attribute; and

outputting the attribute based on the confidence level at least one of:

meeting or exceeding a confidence threshold; or

being greater than other confidence scores associated with other attributes determined by the second machine-learned model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: DAS, SUBHASIS; LIN, YI-TING; MA, DEREK XIANG; ULUTAN, OYTUN
To: ZOOX, INC.
Reel/Frame 067067/0577 →
Continuity (2)
Continuation 17522832 · Nov 9, 2021
Related Publication 20240320985A1 · Sep 26, 2024
References Cited (47)
US 11087494B1 · Srinivasan · 2021 [cited by examiner]
US 11449065B2 · Dariush · 2022 [cited by examiner]
US 11460857B1 · Tan · 2022 [cited by examiner]
US 11521396B1 · Jain · 2022 [cited by examiner]
US 11537139B2 · Rankawat · 2022 [cited by examiner]
US 11682272B2 · Avadhanam · 2023 [cited by examiner]
US 11700356B2 · Raichelgauz · 2023 [cited by examiner]
US 11703566B2 · Cohen · 2023 [cited by examiner]
US 11710352B1 · Ulutan · 2023 [cited by examiner]
US 11731662B2 · Wyffels · 2023 [cited by examiner]
US 11748620B2 · Elluswamy · 2023 [cited by examiner]
US 11829449B2 · Parikh · 2023 [cited by examiner]
US 11851081B2 · Refaat · 2023 [cited by examiner]
US 11885886B2 · Alghanem · 2024 [cited by examiner]
US 12012127B2 · Das · 2024 [cited by examiner]
US 12065140B1 · Pronovost · 2024 [cited by examiner]
US 20180336692A1 · Wendel · 2018 [cited by examiner]
US 20190147372A1 · Luo · 2019 [cited by examiner]
US 20190333232A1 · Vallespi-Gonzalez · 2019 [cited by examiner]
US 20190354782A1 · Kee · 2019 [cited by examiner]
US 20200051252A1 · Brown · 2020 [cited by examiner]
US 20210049378A1 · Gautam · 2021 [cited by examiner]
US 20210049776A1 · Tan et al. · 2021 [cited by applicant]
US 20210053570A1 · Akella · 2021 [cited by examiner]
US 20210084451A1 · Williams · 2021 [cited by applicant]
US 20210103742A1 · Adeli-Mosabbeb · 2021 [cited by examiner]
US 20210157312A1 · Cella · 2021 [cited by examiner]
US 20210181837A1 · Jiang · 2021 [cited by examiner]
US 20210192748A1 · Morales Morales · 2021 [cited by examiner]
US 20210347377A1 · Siebert · 2021 [cited by examiner]
US 20220111873A1 · Wyffels · 2022 [cited by examiner]
US 20230144745A1 · Ulutan · 2023 [cited by applicant]
CA 2739989A1 · 2010 [cited by examiner]
CA 2739989C · 2016 [cited by applicant]
CA 2999498C · 2020 [cited by applicant]
CA 3128028A1 · 2020 [cited by examiner]
JP 2006065513A · 2006 [cited by applicant]
JP 2020204804A · 2020 [cited by applicant]
KR 20190115542A · 2019 [cited by applicant]
KR 20210084451A · 2021 [cited by applicant]
WO WO2018158642A1 · 2018 [cited by examiner]
WO WO2018218155A1 · 2018 [cited by examiner]
WO 2020198121A1 · 2020 [cited by applicant]
WO 2021046103A1 · 2021 [cited by applicant]
WO WO2021118697A1 · 2021 [cited by examiner]
WO WO2021137849A1 · 2021 [cited by examiner]
The PCT Search Report and Written Opinion mailed Mar. 14, 2023 for PCT application No. PCT/US2022/049434, 12 pages.; Application Number. [cited by applicant]