IP Library Granted Patent US 11,610,079
Granted Patent B2
US 11,610,079 · App. 16/777,912 · Granted Mar 21, 2023

Test suite for different kinds of biases in data

Inventor: Michael Yang (Valley Stream, NY)
Assignee: salesforce.com, inc.
G06K9/6257G06K9/6262G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,610,079
App. No.
16/777,912
Granted
Mar 21, 2023
Kind
B2
Abstract

There is provided computer implemented method for detecting and reducing or removing bias for generating a machine learning model, comprising: prior to generating the machine learning model: receiving a training dataset, comprising target inputs, each comprising parameters and labelled with a corresponding target output, wherein at least one of the parameters of at least of the target inputs comprises a sensitive parameter indicative of the corresponding target input assigned to a sensitive group that is potentially biased against other target inputs that are excluded from the sensitive group, analyzing the training dataset to identify target inputs affected by label bias when a statistically significant difference is detected between target inputs assigned to the sensitive group and target inputs excluded from the sensitive group, correcting labels of the target inputs affected by label bias, and generating the machine learning model using the corrected labels.

Claims (47)

1. A computer implemented method for detecting and reducing or removing bias for generating a machine learning model, comprising:

(i) prior to generating the machine learning model:

receiving a training dataset, comprising a plurality of target inputs, each comprising a plurality of parameters and labelled with a corresponding target output,

wherein at least one of the plurality of parameters of at least one of the plurality of target inputs comprises a corresponding sensitive parameter indicative of a corresponding target input assigned to a sensitive group that is potentially biased against other target inputs that are excluded from the sensitive group;

analyzing the training dataset to identify target inputs affected by label bias when a statistically significant difference is detected between target inputs assigned to the sensitive group and the other target inputs excluded from the sensitive group;

correcting labels of the target inputs affected by label bias; and

(ii) generating the machine learning model using the corrected labels.

2. The method of claim 1 , further comprising:

computing by a score computing machine learning model, for each respective target input, a probability of the respective target input being assigned to the sensitive group according to a respective value of the corresponding sensitive parameter.

3. The method of claim 2 , further comprising:

clustering the plurality of target inputs into a plurality of clusters according to corresponding computed probabilities, wherein for each respective cluster;

wherein the target inputs are assigned to the respective cluster and associated with probabilities within a certain probability value range,

wherein each cluster includes the target inputs assigned to the sensitive group and the other target inputs excluded from the sensitive group,

for each respective cluster:

determining whether the statistically significant difference exists between the target inputs assigned to the sensitive group and the other target inputs excluded from the sensitive group; and

identifying label bias for the respective cluster when the statistically significant difference is detected,

wherein the target inputs affected by label bias comprise the target inputs of the respective cluster, including the target inputs assigned to the sensitive group and the other target inputs excluded from the sensitive group.

4. The method of claim 3 , wherein correcting labels of the target inputs affected by label bias comprises correcting labels for the target inputs assigned to the sensitive groups and labels of the other target inputs excluded from the sensitive group.

5. The method of claim 4 , wherein the correcting labels comprises assigning the same label to all of the target inputs assigned to the sensitive groups and to all of the other target inputs excluded from the sensitive group.

6. The method of claim 2 , wherein the probability of the respective target input being assigned to the sensitive group is computed by the score computing machine learning model performing a causal inference process, wherein a treatment of the causal inference process is the sensitive parameter.

7. The method of claim 6 , wherein the causal inference process comprises a propensity score matching (PSM) process, and the probability denotes the propensity score.

8. The method of claim 2 , further comprising:

computing accuracy of the score computing machine learning model for computing the probability of the respective target input being assigned to the sensitive group; and

identifying sampling bias between target inputs of the sensitive group and the other target inputs excluded from the sensitive group when the accuracy of the score computing machine learning model is above a threshold.

9. The method of claim 8 , wherein (ii) generating further comprises generating one respective machine learning model for the target inputs of the sensitive group and generating another respective machine learning model for the other target inputs excluded from the sensitive group.

10. The method of claim 1 , wherein each of the plurality of target inputs comprises the sensitive parameter, wherein the target inputs assigned to the sensitive group include a value of the sensitive parameter meeting a requirement, and the other target inputs excluded from the sensitive group include another value of the sensitive parameter that does not meet the requirement.

11. The method of claim 1 , wherein the sensitive parameter is selected from a group consisting of: gender, race, and age.

12. A computer implemented method for detecting and reducing or removing bias for generating a machine learning model, comprising:

(i) prior to generating the machine learning model:

receiving a training dataset, comprising a plurality of target inputs, each comprising a plurality of parameters and labelled with a corresponding target output,

wherein at least one of the plurality of parameters of at least one of the plurality of target inputs comprises a sensitive parameter indicative of a corresponding target input assigned to a sensitive group that is potentially biased against other target inputs that are excluded from the sensitive group;

analyzing the training dataset to detect sampling bias between target inputs of the sensitive group and the other target inputs excluded from the sensitive group; and

(ii) generating one respective machine learning model for the target inputs of the sensitive group and generating another respective machine learning model for the other target inputs excluded from the sensitive group.

13. The method of claim 12 , further comprising computing accuracy of a score computing machine learning model that computes a probability of a certain target input being assigned to the sensitive group; and

detecting the sampling bias when the accuracy of the score computing machine learning model is above a threshold.

14. A system for detecting and reducing or removing bias for generating a machine learning model, comprising:

at least one hardware processor executing a code for:

(i) prior to generating the machine learning model:

receiving a training dataset, comprising a plurality of target inputs, each comprising a plurality of parameters and labelled with a corresponding target output,

wherein at least one of the plurality of parameters of at least one of the plurality of target inputs comprises a sensitive parameter indicative of a corresponding target input assigned to a sensitive group that is potentially biased against target inputs that are excluded from the sensitive group;

analyzing the training dataset to identify target inputs affected by label bias when a statistically significant difference is detected between target inputs assigned to the sensitive group and the target inputs excluded from the sensitive group;

correcting labels of the target inputs affected by label bias; and

(ii) generating the machine learning model using the corrected labels.

15. The system of claim 14 , further comprising, code for:

computing accuracy of a score computing machine learning model for computing a probability of a respective target input being assigned to the sensitive group; and

identifying sampling bias between target inputs of the sensitive group and target inputs of the sensitive group when the accuracy of the score computing machine learning model is above a threshold.

16. The system of claim 15 , wherein generating the machine learning model further comprises generating one respective machine learning model for the target inputs of the sensitive group and generating another respective machine learning model for the target inputs excluded from the sensitive group.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0475 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2020
From: YANG, MICHAEL
To: SALESFORCE.COM, INC.
Reel/Frame 053144/0570 →
Continuity (1)
Related Publication 20210241033A1 · Aug 5, 2021
Cited By (1)
US 12,468,936