IP Library Granted Patent US 12,430,899
Granted Patent B2
US 12,430,899 · App. 17/817,235 · Granted Sep 30, 2025

De-biasing datasets for machine learning

Inventors: Nikita Jaipuria (Pittsburgh, PA); Xianling Zhang (San Jose, CA); Katherine Stevo (Wellesley, MA); Jinesh Jain (San Francisco, CA); Vidya Nariyambut Murali (Sunnyvale, CA); Meghana Laxmidhar Gaopande (Sunnyvale, CA)
Assignee: Ford Global Technologies, LLC
G06V10/778G06V10/7715
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,899
App. No.
17/817,235
Granted
Sep 30, 2025
Kind
B2
Abstract

A computer includes a processor and a memory storing instructions executable by the processor to receive a dataset of images; extract feature data from the images; optimize a number of clusters into which the images are classified based on the feature data; for each cluster, optimize a number of subclusters into which the images in that cluster are classified; determine a metric indicating a bias of the dataset toward at least one of the clusters or subclusters based on the number of clusters, the numbers of subclusters, distances between the respective clusters, and distances between the respective subclusters; and after determining the metric, train a machine-learning program using a training set constructed from the clusters and the subclusters.

Claims (31)

1. A computer comprising a processor and a memory storing instructions executable by the processor to:

receive a dataset of images;

extract feature data from the images;

optimize a number of clusters into which the images are classified based on the feature data, wherein optimizing the number of clusters includes performing k-means clustering for a plurality of values for the number of clusters, determining the silhouette score for each of the values, and selecting one of the values based on the silhouette scores;

for each cluster, determine a perceptual-similarity score between each pair of the images in that cluster;

for each cluster, reduce a dimensionality of the perceptual-similarity scores for that cluster, wherein, for each cluster, reducing the dimensionality of the perceptual-similarity scores includes performing principal component analysis;

for each cluster, optimize a number of subclusters into which the images in that cluster are classified, wherein, for each cluster, optimizing the number of subclusters in that cluster is based on the perceptual-similarity scores for that cluster after reducing the dimensionality;

in response to optimizing the number of the clusters and the numbers of the subclusters, determine a metric indicating a bias of the dataset toward at least one of the clusters or subclusters based on the number of clusters, the numbers of subclusters, distances between the respective clusters, and distances between the respective subclusters, wherein the bias toward the at least one of the clusters or subclusters indicates overrepresentation of the at least one of the clusters or subclusters in the dataset of the images; and

after determining the metric, train a machine-learning program using a training set constructed from the clusters and the subclusters.

2. The computer of claim 1 , wherein the instructions further include instructions to reduce a dimensionality of the feature data.

3. The computer of claim 2 , wherein optimizing the number of clusters is based on the feature data after reducing the dimensionality.

4. The computer of claim 2 , wherein the instructions further include instructions to, after reducing the dimensionality of the feature data, map the feature data to a limited number of dimensions.

5. The computer of claim 4 , wherein the limited number of dimensions is two.

6. The computer of claim 2 , wherein reducing the dimensionality of the feature data includes performing principal component analysis.

7. The computer of claim 6 , wherein the principal component analysis results in a matrix of principal components, and the instructions further include instructions to map the matrix of principal components to a limited number of dimensions.

8. The computer of claim 1 , wherein the machine-learning program is a second machine-learning program, and extracting the feature data includes executing a first machine-learning program.

9. The computer of claim 8 , wherein the first machine-learning program is trained to identify objects in images, and the feature data is output from an intermediate layer of the first machine-learning program.

10. The computer of claim 1 , wherein the distances between the respective clusters are distances between centroids of the respective clusters, and the distances between the respective subclusters are distances between centroids of the respective subclusters.

11. The computer of claim 1 , wherein, for each cluster, the principal component analysis results in a matrix of principal components, and the instructions further include instructions to map the matrix of principal components to a limited number of dimensions.

12. The computer of claim 11 , wherein the limited number of dimensions is two.

13. The computer of claim 1 , wherein the clusters and subclusters serve as ground truth when training the machine-learning program.

14. The computer of claim 1 , wherein the instructions further include instructions to construct the training set from the clusters and subclusters by performing at least one of automatic metadata annotation correction, complex edge/corner cases detection, training and test dataset curation for artificial-intelligence and machine-learning models, data augmentation guidance, or visualization and quantification of gaps between simulated and real datasets, based on the metric.

15. A method comprising:

receiving a dataset of images;

extracting feature data from the images;

optimizing a number of clusters into which the images are classified based on the feature data, wherein optimizing the number of clusters includes performing k-means clustering for a plurality of values for the number of clusters, determining the silhouette score for each of the values, and selecting one of the values based on the silhouette scores;

for each cluster, determining a perceptual-similarity score between each pair of the images in that cluster;

for each cluster, reduce a dimensionality of the perceptual-similarity scores for that cluster, wherein, for each cluster, reducing the dimensionality of the perceptual-similarity scores includes performing principal component analysis;

for each cluster, optimizing a number of subclusters into which the images in that cluster are classified, wherein, for each cluster, optimizing the number of subclusters in that cluster is based on the perceptual-similarity scores for that cluster after reducing the dimensionality;

in response to optimizing the number of the clusters and the numbers of the subclusters, determining a metric indicating a bias of the dataset toward at least one of the clusters or subclusters based on the number of clusters, the numbers of subclusters, distances between the respective clusters, and distances between the respective subclusters, wherein the bias toward the at least one of the clusters or subclusters indicates overrepresentation of the at least one of the clusters or subclusters in the dataset of the images; and

after determining the metric, training a machine-learning program using a training set constructed from the clusters and the subclusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2022
From: JAIPURIA, NIKITA; ZHANG, XIANLING; STEVO, KATHERINE; JAIN, JINESH; NARIYAMBUT MURALI, VIDYA; GAOPANDE, MEGHANA LAXMIDHAR
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 060711/0797 →
Continuity (2)
Provisional Application 63234763 · Aug 19, 2021
Related Publication 20240046625A1 · Feb 8, 2024
References Cited (21)
US 10977571B2 · Miserendino et al. · 2021 [cited by applicant]
US 11620471B2 · Brock · 2023 [cited by examiner]
US 20100329529A1 · Feldman · 2010 [cited by examiner]
US 20160041958A1 · Zhuang et al. · 2016 [cited by applicant]
US 20170006135A1 · Siebel · 2017 [cited by examiner]
US 20170249547A1 · Shrikumar et al. · 2017 [cited by applicant]
US 20170255840A1 · Jean et al. · 2017 [cited by applicant]
US 20180150724A1 · Brock · 2018 [cited by examiner]
US 20180189635A1 · Olarig · 2018 [cited by examiner]
US 20190244253A1 · Vij · 2019 [cited by examiner]
US 20200401851A1 · Mau · 2020 [cited by examiner]
US 20220167928A1 · Baek · 2022 [cited by examiner]
US 20220261628A1 · Surya · 2022 [cited by examiner]
US 20220296930A1 · Chen · 2022 [cited by examiner]
US 20240046625A1 · Jaipuria · 2024 [cited by examiner]
Van der Maaten, L., et al. “Visualizing Data using t-SNE,” Journal of Machine Learning Research 9, 2008, 27 pages. [cited by applicant]
Johnson, J., et al., Perceptual Losses for Real-Time Style Transfer and Super-Resolution, Department of Computer Science, Stanford University, 2016, 17 pages. [cited by applicant]
“7 Types of Data Bias in Machine Learning,” TELUS International, Jan. 1, 2021, 9 pages. [cited by applicant]
Wang, Z, et al., “Multi-Scale Structural Similarity for Image Quality Assessment,” Proceedings of the 37th IEEE Asilomar Conference on Signals, Systems and Computers, Nov. 2003, 5 pages. [cited by applicant]
McInnes, L, et al., “UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.” ArXiv, 2018, 5 pages. [cited by applicant]
Zhang, R., et al. “Perceptual Similarity Metric and Dataset,” CVPR, 2018, 8 pages. [cited by applicant]