IP Library › Granted Patent US 12,572,522
Granted Patent B2
US 12,572,522 · App. 18/189,962 · Granted Mar 10, 2026

Identifying quality of labeled data

Inventors: Russell Brennan (Kirkland, WA); Boxin Li (Sammamish, WA); Wenjie Zhou (San Jose, CA)
Assignee: GM Cruise Holdings LLC
G06F16/2272G06F16/285G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,522
App. No.
18/189,962
Granted
Mar 10, 2026
Kind
B2
Abstract

Aspects of the subject technology relate to systems, methods, and computer-readable media for identifying a quality of labeled data. Labeled data of a data set that exists at a specific granularity level can be accessed. The labeled data can be sampled on a lower granularity level relative to the specific granularity level of the data set to generate sampled data of the data set. The sampled data can be labeled to generate ground truth labeled data of the data set. The labeled data can be compared to the ground truth labeled data to identify a labeling quality metric of the labeled data. Relabeling of the data set can be performed based on the labeling quality metric.

Claims (71)

1 . A computer-implemented method, comprising:

accessing labeled data of a data set that exists at a capture granularity level, wherein the capture granularity level includes data captured by a plurality of sensors within a temporal time frame;

sampling the labeled data at a lower granularity level relative to the capture granularity level of the data set to generate sampled data of the data set;

labeling the sampled data to generate ground truth labeled data;

comparing the labeled data to the ground truth labeled data to determine a labeling quality metric of labels of the labeled data representing an accuracy of the labels of the labeled data based on a number of labels within a predefined error tolerance;

determining that the labeling quality metric satisfies a criterion;

based on determining that the labeling quality metric satisfies the criterion:

determining a sampling percentage for relabeling based on a difference between the labeling quality metric and a target quality metric;

relabeling a portion of the labeled data based on the labeling quality metric and the sampling percentage; and

training a machine learning model based on the relabeled portion of the labeled data; and

deploying the trained machine learning model to an autonomous vehicle, the trained machine learning model, when deployed to the autonomous vehicle, is configured to:

execute the trained machine learning model in a control stack for making predictions based on input data;

generate commands for actuators of the autonomous vehicle, the actuators including at least one of a steering system, a braking system, or a propulsion system; and

navigate the autonomous vehicle based on the commands.

2 . The computer-implemented method of claim 1 , wherein relabeling the portion of the labeled data based on the labeling quality metric and the sampling percentage comprises relabeling the labeled data based on the labeling quality metric in relation to one or more thresholds.

3 . The computer-implemented method of claim 1 , wherein relabeling the portion of the labeled data based on the labeling quality metric and the sampling percentage comprises relabeling the labeled data based on a labeling quality metric defined based on a specific application of the labeled data to one or more models.

4 . The computer-implemented method of claim 1 , wherein the lower granularity level is an object level of granularity in the data set.

5 . The computer-implemented method of claim 1 , wherein the sampled data is part of a first subset of the labeled data that is used in verifying the labels of the labeled data, and wherein relabeling the labeled data comprises:

sampling the labeled data to generate a second subset of the data set; and

labeling the second subset of the data set to generate refined labeled data of the data set.

6 . The computer-implemented method of claim 5 , wherein the second subset of the data set is generated independently from the first subset of the data set.

7 . The computer-implemented method of claim 5 , further comprising applying one or more machine learning techniques to select an amount of the data set to sample to generate a specific size of the second subset of the data set.

8 . The computer-implemented method of claim 5 , further comprising applying one or more machine learning techniques to identify specific portions of the data set to include in the second subset of the data based on a likelihood that the specific portions of the data set are labeled incorrectly.

9 . The computer-implemented method of claim 5 , further comprising:

comparing the refined labeled data to the ground truth labeled data to identify a labeling quality metric of the refined labeled data; and

relabeling the refined labeled data based on the labeling quality metric of the refined labeled data.

10 . A system, comprising:

one or more processors; and

at least one computer-readable storage medium storing instructions which, when executed by the one or more processors, cause the one or more processors to:

access labeled data of a data set that exists at a capture granularity level, wherein the capture granularity level includes data captured by a plurality of sensors within a temporal time frame;

sample the labeled data at a lower granularity level relative to the capture granularity level of the data set to generate sampled data of the data set;

label the sampled data to generate ground truth labeled data;

compare the labeled data to the ground truth labeled data to determine a labeling quality metric of labels of the labeled data representing an accuracy of the labels of the labeled data based on a number of labels within a predefined error tolerance;

determine that the labeling quality metric satisfies a criterion;

based on determining that the labeling quality metric satisfies the criterion:

determine a sampling percentage for relabeling based on a difference between the labeling quality metric and a target quality metric;

relabel a portion of the labeled data based on the labeling quality metric and the sampling percentage; and

train a machine learning model based on the relabeled portion of the labeled data; and

deploy the trained machine learning model to an autonomous vehicle, the trained machine learning model, when deployed to the autonomous vehicle, is configured to:

execute the trained machine learning model in a control stack for making predictions based on input data;

generate commands for actuators of the autonomous vehicle, the actuators including at least one of a steering system, a braking system, or a propulsion system; and

navigate the autonomous vehicle based on the commands.

11 . The system of claim 10 , wherein the instructions further cause the one or more processors to relabel the portion of the labeled data set based on the labeling quality metric and the sampling percentage comprises relabeling the labeled data based on the labeling quality metric in relation to one or more thresholds.

12 . The system of claim 10 , wherein the labeling quality metric is defined based on a specific application of the labeled data to one or more models.

13 . The system of claim 10 , wherein the lower granularity level is an object level of granularity in the data set.

14 . The system of claim 10 , wherein the sampled data is part of a first subset of the labeled data that is used in verifying the labels of the labeled data, and wherein the instructions further cause the one or more processors to:

sample the labeled data to generate a second subset of the data set; and

label the second subset of the data set to generate refined labeled data of the data set.

15 . The system of claim 14 , wherein the second subset of the data set is generated independently from the first subset of the data set.

16 . The system of claim 14 , wherein the instructions further cause the one or more processors to apply one or more machine learning techniques to select an amount to sample the data set to generate a specific size of the second subset of the data set.

17 . The system of claim 14 , wherein the instructions further cause the one or more processors to apply one or more machine learning techniques to identify specific portions of the data set to include in the second subset of the data based on a likelihood that the specific portions of the data set are labeled incorrectly.

18 . The system of claim 14 , wherein the instructions further cause the one or more processors to:

compare the refined labeled data to the ground truth labeled data to identify a labeling quality metric of the refined labeled data; and

relabel the refined labeled data based on the labeling quality metric of the refined labeled data.

19 . A non-transitory computer-readable storage medium storing instructions for causing one or more processors to:

access labeled data of a data set that exists at a capture granularity level, wherein the capture granularity level includes data captured by a plurality of sensors within a temporal time frame;

sample the labeled data at a lower granularity level relative to the capture granularity level of the data set to generate sampled data of the data set;

label the sampled data to generate ground truth labeled data;

compare the labeled data to the ground truth labeled data to determine a labeling quality metric of labels of the labeled data representing an accuracy of the labels of the labeled data based on a number of labels within a predefined error tolerance;

determine that the labeling quality metric satisfies a criterion;

based on determining that the labeling quality metric satisfies the criterion:

determine a sampling percentage for relabeling based on a difference between the labeling quality metric and a target quality metric;

relabel a portion of the labeled data based on the labeling quality metric and the sampling percentage; and

train a machine learning model based on the relabeled portion of the labeled data; and

deploy the trained machine learning model to an autonomous vehicle, the trained machine learning model, when deployed to the autonomous vehicle, is configured to:

execute the trained machine learning model in a control stack for making predictions based on input data;

generate commands for actuators of the autonomous vehicle, the actuators including at least one of a steering system, a braking system, or a propulsion system; and

navigate the autonomous vehicle based on the commands.

20 . The non-transitory computer-readable storage medium of claim 19 , wherein the sampled data is part of a first subset of the labeled data that is used in verifying the labels of the labeled data, and wherein the instructions further cause the one or more processors to:

sample the labeled data set to generate a second subset of the data set; and

label the second subset of the data set to generate refined labeled data of the data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2023
From: BRENNAN, RUSSELL; LI, BOXIN; ZHOU, WENJIE
To: GM CRUISE HOLDINGS LLC
Reel/Frame 063098/0596 →
Continuity (1)
Related Publication 20240320206A1 · Sep 26, 2024
References Cited (39)
US 6298351B1 · Castelli · 2001 [cited by examiner]
US 10997461B2 · Elluswamy · 2021 [cited by examiner]
US 11308364B1 · Herman · 2022 [cited by examiner]
US 11379718B2 · Desmond · 2022 [cited by examiner]
US 20080103996A1 · Forman · 2008 [cited by examiner]
US 20160371601A1 · Grove · 2016 [cited by examiner]
US 20180373980A1 · Huval · 2018 [cited by examiner]
US 20190065989A1 · Kida · 2019 [cited by examiner]
US 20190114546A1 · Anil · 2019 [cited by examiner]
US 20200012963A1 · Johnston · 2020 [cited by examiner]
US 20200065712A1 · Wang · 2020 [cited by examiner]
US 20200151578A1 · Chen · 2020 [cited by examiner]
US 20200158516A1 · Gale · 2020 [cited by examiner]
US 20210089964A1 · Zhang · 2021 [cited by examiner]
US 20210174196A1 · Desmond · 2021 [cited by examiner]
US 20210175553A1 · Van Tassell · 2021 [cited by examiner]
US 20210241040A1 · Tong · 2021 [cited by examiner]
US 20210319333A1 · Lee · 2021 [cited by examiner]
US 20210350181A1 · Navratil · 2021 [cited by examiner]
US 20210365793A1 · Surya · 2021 [cited by examiner]
US 20220036128A1 · Levanony · 2022 [cited by examiner]
US 20220067588A1 · Büttner · 2022 [cited by examiner]
US 20220076077A1 · Reddy · 2022 [cited by examiner]
US 20220138561A1 · Prendki · 2022 [cited by examiner]
US 20220300557A1 · Basu · 2022 [cited by examiner]
US 20220335311A1 · Lahlou · 2022 [cited by examiner]
US 20230087292A1 · Wang · 2023 [cited by examiner]
US 20230244987A1 · Truong · 2023 [cited by examiner]
US 20240054390A1 · Wendt · 2024 [cited by examiner]
US 20240062051A1 · Baran Pouyan · 2024 [cited by examiner]
CN 104166706A · 2014 [cited by examiner]
CN 110826494A · 2020 [cited by examiner]
CN 114372532A · 2022 [cited by examiner]
CN 116635866A · 2023 [cited by examiner]
EP 4379606A1 · 2024 [cited by examiner]
WO WO2018224879A1 · 2018 [cited by examiner]
WO WO2023114514A1 · 2023 [cited by examiner]
Kang, Daniel, et al. “Finding label and model errors in perception data with learned observation assertions.” Proceedings of the 2022 international conference on management of data. 2022. (Year: 2022). [cited by examiner]
“Identifying and Determining Trustworthiness of a MachineLearned Model” https://priorart.ip.com/IPCOM/000252359; Jan. 5, 2018 (Year: 2018). [cited by examiner]