IP Library › Granted Patent US 12,277,753
Granted Patent B1
US 12,277,753 · App. 18/751,166 · Granted Apr 15, 2025

Common-sense bias discovery and mitigation for machine-learning tasks

Inventors: Gaurav Bharaj (San Francisco, CA); Miao Zhang (Jersey City, NJ); Zee Fryer (Oakland, CA); Ben Colman (New York, NY); Ali Shahriyari (Las Vegas, NV)
Assignee: Reality Defender, Inc.
G06V10/774G06V10/751G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,753
App. No.
18/751,166
Granted
Apr 15, 2025
Kind
B1
Abstract

An exemplary method for reducing bias in a training image dataset for training a machine-learning model comprises: receiving a plurality of text strings comprising at least one text string describing each image in the training image dataset; generating a plurality of embeddings based on the plurality of text strings; identifying, based on the plurality of embeddings, a plurality of visual features in the training image dataset; identifying one or more correlations between the plurality of visual features in the training image dataset; receiving a user input identifying at least one biased correlation from the one or more correlations; and training the machine-learning model at least partially by adjusting one or more data sampling weights associated with one or more training images in the training image dataset based on the user input.

Claims (46)

1. A method for reducing bias in a training image dataset for training a machine-learning model, comprising:

receiving a plurality of text strings comprising at least one text string describing each image in the training image dataset;

generating a plurality of embeddings based on the plurality of text strings;

identifying, based on the plurality of embeddings, a plurality of visual features in the training image dataset;

identifying one or more correlations between the plurality of visual features in the training image dataset;

receiving a user input identifying at least one biased correlation from the one or more correlations; and

training the machine-learning model at least partially by adjusting one or more data sampling weights associated with one or more training images in the training image dataset based on the user input.

2. The method of claim 1 , wherein the machine-learning model is an image classification model, an object recognition model, an object segmentation model, or any combination thereof.

3. The method of claim 1 , wherein the one or more data sampling weights are adjusted to reduce the at least one biased correlation.

4. The method of claim 1 , further comprising:

determining, based on the user input, if one or more machine-learning models trained using the training image dataset is biased.

5. The method of claim 1 , wherein generating the plurality of embeddings comprises:

for each image in the training image dataset:

extracting one or more tokens from the at least one text string describing the respective image;

generating, based on each extracted token, an embedding using an encoder model.

6. The method of claim 5 , wherein the encoder model is a Universal Sentence Encoder (USE) or a Contrastive Language-Image Pretraining (CLIP) model.

7. The method of claim 1 , wherein identifying the plurality of visual features in the training image dataset comprises:

clustering the plurality of embeddings into a plurality of categories; and

clustering embeddings in each category of the plurality of categories to identify the plurality of visual features.

8. The method of claim 7 , wherein clustering the plurality of embeddings is performed using a K-means clustering algorithm, a hierarchical clustering algorithm, a mean shift algorithm, a Gaussian mixture model, or an agglomerative clustering algorithm.

9. The method of claim 7 , wherein clustering the embeddings in each category is performed using a K-means clustering algorithm, a hierarchical clustering algorithm, a mean shift algorithm, a Gaussian mixture model, or an agglomerative clustering algorithm.

10. The method of claim 1 , wherein identifying the one or more correlations between the plurality of visual features in the training image dataset comprises:

calculating a measurement indicative of an association between two visual features of the plurality of visual features; and

comparing the measurement with a predefined threshold.

11. The method of claim 10 , wherein the measurement comprises a Matthews correlation coefficient or a Pearson coefficient.

12. The method of claim 1 , further comprising: displaying the one or more correlations between the plurality of visual features in the training image dataset.

13. The method of claim 1 , further comprising: receiving a user input identifying at least one benign correlation from the one or more correlations.

14. The method of claim 1 , wherein the plurality of text strings are generated by one or more human annotators.

15. The method of claim 1 , wherein each text string of the plurality of text strings describes a person, an object, an attribute of the person, an attribute of the object, or any combination thereof.

16. A non-transitory computer-readable storage medium storing one or more programs for reducing bias in a training image dataset for training a machine-learning model, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a plurality of text strings comprising at least one text string describing each image in the training image dataset;

generate a plurality of embeddings based on the plurality of text strings;

identify, based on the plurality of embeddings, a plurality of visual features in the training image dataset;

identify one or more correlations between the plurality of visual features in the training image dataset;

receive a user input identifying at least one biased correlation from the one or more correlations; and

train the machine-learning model at least partially by adjusting one or more data sampling weights associated with one or more training images in the training image dataset based on the user input.

17. A system for reducing bias in a training image dataset for training a machine-learning model, comprising:

one or more processors;

one or more memories; and

one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a plurality of text strings comprising at least one text string describing each image in the training image dataset;

generating a plurality of embeddings based on the plurality of text strings;

identifying, based on the plurality of embeddings, a plurality of visual features in the training image dataset;

identifying one or more correlations between the plurality of visual features in the training image dataset;

receiving a user input identifying at least one biased correlation from the one or more correlations; and

training the machine-learning model at least partially by adjusting one or more data sampling weights associated with one or more training images in the training image dataset based on the user input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2025
From: BHARAJ, GAURAV; ZHANG, MIAO; FRYER, ZEE; COLMAN, BEN; SHAHRIYARI, ALI
To: REALITY DEFENDER, INC.
Reel/Frame 070473/0338 →
Continuity (1)
Provisional Application 63600577 · Nov 17, 2023
References Cited (66)
US 20230153687A1 · Vu · 2023 [cited by examiner]
US 20240013522A1 · Gubbi Lakshminarasimha · 2024 [cited by examiner]
Agrawal et al. “VQA: Visual Question Answering,” IEEE international conference on computer vision, Dec. 7-13, 2015, Santiago, Chile; pp. 1-25. [cited by applicant]
Ahn et al. “Mitigating dataset bias by using per-sample gradient,” ICCLR 2023, May 1-5, 2023, Kigali, Rwanda; pp. 1-36. [cited by applicant]
Alayrac et al. “Flamingo: a Visual Language Model for Few-Shot Learning,” NeurIPS 2022: 36th Conference on Neural Information Processing Systems, Nov. 28-Dec. 9, 2022, New Orleans, Louisiana; pp. 1-54. [cited by applicant]
Amini et al. “Uncovering and Mitigating Algorithmic Bias through Learned Latent Structure,” 2019 AAAI/ACM Conference on AI, Ethics, and Society, Jan. 27-28, 2019, Honolulu, Hawaii; pp. 289-295. [cited by applicant]
Amini et al. “Variational Autoencoder for End-to-End Control of Autonomous Driving with Novelty Detection and Training De-biasing,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems, Oct. 1-5, 201… [cited by applicant]
Bahng et al. “Learning De-biased Representations with Biased Representations,” 37th International Conference on Machine Learning, Jul. 13-18, 2020, Vienna, Austria; 17 pages. [cited by applicant]
Basu et al. (May 2023). “Inspecting the Geographical Representativeness of Images from Text-to-Image Models,” Indian Institute of Science; 15 pages. [cited by applicant]
Bommasani et al. (Aug. 2021). “On the Opportunities and Risks of Foundation Models,” Center for Research on Foundation Models (CRFM) and Stanford Institute for Human-Centered Artificial Intelligence (HAI); pp. 1-214. [cited by applicant]
Brown et al. (Jun. 2023). “Detecting shortcut learning for fair medicalAI using shortcut testing,” Nature Communications 14(4314); pp. 1-10. [cited by applicant]
Cer et al. (Mar. 2018). “Universal Sentence Encoder,” Google Research; 7 pages. [cited by applicant]
Zhang et al. (Mar. 2022). “Correct-n-Contrast: A Contrastive Approach for Improving Robustness to Spurious Correlations,” located at https://arxiv.org/abs/2203.01517; pp. 1-38. [cited by applicant]
Chen et al. (Apr. 2015). “Microsoft COCO Captions: Data Collection and Evaluation Server,” located at https://arxiv.org/abs/1504.00325; pp. 1-7. [cited by applicant]
Diomataris et al. “Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, Montreal, … [cited by applicant]
Dunlap et al. “Using language to extend to unseen domains,” The Eleventh International Conference on Learning Representations, Apr. 25-29, 2022, Virtual Event, Austria; pp. 1-19. [cited by applicant]
Hendrycks et al. “Benchmarking neural network robustness to common corruptions and perturbations,” International Conference on Learning Representations, May 6-9, 2019, New Orleans, Louisiana; pp. 1-16. [cited by applicant]
Jia et al. “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision,” 38th International Conference on Machine Learning, Jul. 18-24, 2021, Virtual Event; 14 pages. [cited by applicant]
Jiang et al. “Talk-to-Edit: Fine-Grained Facial Editing via Dialog,” IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, Montreal, Canada; pp. 1-22. [cited by applicant]
Karpathy et al. “Deep Visual-Semantic Alignments for Generating Image Descriptions,” IEEE conference on computer vision and pattern recognition, Jun. 7-12, 2015, Los Alamitos, California; pp. 3128-3137. [cited by applicant]
Khattak et al. “MaPLe: Multi-modal Prompt Learning,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2023, Vancouver, Canada; 13 pages. [cited by applicant]
Kim et al. “BiaSwap: Removing Dataset Bias with Bias-Tailored Swapping Augmentation,” IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, Montreal, Canada; 14 pages. [cited by applicant]
Krishnakumar et al. (2021) “UDIS: Unsupervised Discovery of Bias in Deep Visual Recognition Models,” British Machine Vision Conference (BMVC) 1; pp. 1-40. [cited by applicant]
Lake et al. (2017). “Building Machines That Learn and Think Like People,” Behavioral and Brain Sciences 40(e253); pp. 1-58. [cited by applicant]
Lang et al. “Explaining in Style: Training a GAN to explain a classifier in StyleSpace,” IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, Montreal, Canada; 10 pages. [cited by applicant]
LeCun et al. (2010) “The MNIST Database of handwritten digits,” The Courant Institute of Mathematical Sciences and Google Labs; 8 pages. [cited by applicant]
Li et al. “Discover and Mitigate Unknown Biases with Debiasing Alternate Networks,” European Conference on Computer Vision, Oct. 23-27, 2022, Tel Aviv, Israel; pp. 1-39. [cited by applicant]
Li et al. “Discover the Unknown Biased Attribute of an Image Classifier,” IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, Montreal, Canada; 22 pages. [cited by applicant]
Li et al. “REPAIR: Removing Representation Bias by Dataset Resampling,” IEEE/CVF conference on computer vision and pattern recognition, Jun. 15-20, 2019, Long Beach, California; 10 pages. [cited by applicant]
Li et al. (2020). “A Deeper Look at Facial Expression Dataset Bias,” IEEE Transactions on Affective Computing 13(2); pp. 1-15. [cited by applicant]
Lin et al. “Microsoft COCO: Common Objects in Context,” Computer Vision—ECCV 2014: 13th European Conference, Sep. 6-12, 2014, Zurich, Switzerland; pp. 1-15. [cited by applicant]
Liu et al. “Deep Learning Face Attributes in the Wild,” IEEE international conference on computer vision, Dec. 7-13, 2015, Santiago, Chile; 11 pages. [cited by applicant]
Liu et al. “Just Train Twice: Improving Group Robustness without Training Group Information,” International Conference on Machine Learning, Jul. 18-24, 2021, Virtual Conference; 16 pages. [cited by applicant]
Lloyd, S. P. (Mar. 1982). “Least Squares Quantization in PCM,” IEEE Transactions on Information Theory IT-28(2); pp. 129-137. [cited by applicant]
Manjunatha et al. “Explicit Bias Discovery in Visual Question Answering Models,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15-20, 2019, Long Beach, California; pp. 1-10. [cited by applicant]
Matthews, B. W. (May 1975). “Comparison of the Predicted and Observed Secondary Structure of T4 Phage Lysozyme,” Biochimica et Biophysica Acta; pp. 443-451. [cited by applicant]
McInnes et al. (Sep. 21, 2020). “UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction,” located at https://arxiv.org/abs/1802.03426; pp. 1-63. [cited by applicant]
Misra et al. “Seeing through the Human Reporting Bias: Visual Classifiers from Noisy Human-Centric Labels,” IEEE conference on computer vision and pattern recognition, Jun. 27-30, 2016, Las Vegas, Nevada; pp. 1-10. [cited by applicant]
Nam et al. (2020). “Learning from Failure: Training Debiased Classifier from Biased Classifier,” Advances in Neural Information Processing Systems; pp. 1-19. [cited by applicant]
Park et al. “Fair Contrastive Learning for Facial Attribute Classification,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, New Orleans, Louisiana; pp. 1-19. [cited by applicant]
Qraitem et al. “Bias Mimicking: A Simple Sampling Approach for Bias Mitigation,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2023, Vancouver, Canada; pp. 1-13. [cited by applicant]
Radford et al. “Learning Transferable Visual Models From Natural Language Supervision,” International conference on machine learning, Jul. 18-24, 2021, Virtual Conference; 48 pages. [cited by applicant]
Radford et al. (2019). “Language Models are Unsupervised Multitask Learners,” located at https://www.semanticscholar.org/paper/Language-Models-are-Unsupervised-Multitask-Learners-Radford-Wu/9405cc0d6169988371b2755e573cc… [cited by applicant]
Ramaswamy et al. “Fair Attribute Classification through Latent Space De-biasing,” IEEE/CVF conference on computer vision and pattern recognition, Jun. 20-25, 2021, Nashville, Tennessee; pp. 1-15. [cited by applicant]
Rombach et al. “High-Resolution Image Synthesis with Latent Diffusion Models,” the IEEE/CVF conference on computer vision and pattern recognition, Jun. 18-24, 2022, New Orleans, Louisiana; pp. 1-45. [cited by applicant]
Schramowski et al. (2022). “Large Pre-trained Language Models Contain Human-like Biases of What is Right and Wrong to Do,” Nature Machine Intelligence 4(3); pp. 1-38. [cited by applicant]
Seo et al. “Unsupervised Learning of Debiased Representations with Pseudo-Attributes,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, New Orleans, Louisiana; 15 pages. [cited by applicant]
Sharma et al. “Data Augmentation for Discrimination Prevention and Bias Disambiguation,” AAAI/ACM Conference on AI, Ethics, and Society, Feb. 7-8, 2020, New York, New York; pp. 358-364. [cited by applicant]
Singh et al. “Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13-19, 2020, Seattle, Washington; 14 pages. [cited by applicant]
Sohn et al. “Visual Prompt Tuning for Generative Transfer Learning,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 17-24, 2023, Vancouver, Canada; pp. 1-36. [cited by applicant]
Sohoni et al. (Apr. 2022). “No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification Problems,” Advances in Neural Information Processing Systems; pp. 1-40. [cited by applicant]
Tian et al. (Feb. 2022). “Image fairness in deep learning: problems, models, and challenges,” Neural Computing and Applications 34; pp. 12875-12893. [cited by applicant]
Torralba et al. “Unbiased Look at Dataset Bias,” IEEE International Conference on Automation Science and Engineering, Aug. 24-27, 2011, Shanghai, China; pp. 1521-1528. [cited by applicant]
Truong et al. “FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 17-24, 2023, Vancouver, Canada; 10 pages. [cited by applicant]
Wang et al. “Designing Theory-Driven User-Centric Explainable AI,” 2019 CHI conference on human factors in computing systems, May 4-9, 2019, Glasgow, United Kingdom; 16 pages. [cited by applicant]
Wang et al. “Overwriting Pretrained Bias with Finetuning Data,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 1-6, 2023, Paris, France; 15 pages. [cited by applicant]
Wang et al. “Towards Fairness in Visual Recognition: Effective Strategies for Bias Mitigation,” IEEE/CVF conference on computer vision and pattern recognition, Jun. 13-19, 2020, Seattle, Washington; 10 pages. [cited by applicant]
Wu et al. “Discover and Cure: Concept-aware Mitigation of Spurious Correlation,” 40th International Conference on Machine Learning, Jul. 23-29, 2023, Honolulu, Hawaii; pp. 1-22. [cited by applicant]
Wu et al. “Unified Visual-Semantic Embeddings: Bridging Vision and Language with Structured Meaning Representations,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15-20, 2019, Long Beach, Califor… [cited by applicant]
Wu et al. (2022). “A Survey of Human-in-the-Loop for Machine Learning,” Future Generation Computer Systems; pp. 1-22. [cited by applicant]
Yuan et al. (2021). “Florence: A New Foundation Model for Computer Vision,” located at https://arxiv.org/abs/2111.11432; 17 pages. [cited by applicant]
Zhang et al. “CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2023, Vancouver, Canada; 15 pages. [cited by applicant]
Zhang et al. “Diagnosing and Rectifying Vision Models Using Language,” International Conference on Learning Representations, May 1-5, 2023, Kigali, Rwanda; pp. 1-27. [cited by applicant]
Zhang et al. “Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes,” The Eleventh International Conference on Learning Representations, Apr. 25-29, 2022, Virtual Conference; pp. 1-18. [cited by applicant]
Zhang et al. “GLIPv2: Unifying Localization and VL Understanding,” Advances in Neural Information Processing Systems, Nov. 28-Dec. 9, 2022, New Orleans, Louisiana; pp. 1-25. [cited by applicant]
Zhong et al. “RegionCLIP: Region-based Language-Image Pretraining,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, New Orleans, Louisiana; pp. 1-12. [cited by applicant]
Cited By (3)
US 12,400,434 US 12,548,316 US 12,555,365