IP Library › Granted Patent US 12,456,061
Granted Patent B2
US 12,456,061 · App. 17/132,786 · Granted Oct 28, 2025

Self-monitoring cognitive bias mitigator in predictive systems

Inventor: Debasis Ganguly (Dublin, IE)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,061
App. No.
17/132,786
Granted
Oct 28, 2025
Kind
B2
Abstract

One or more embodiments described herein facilitate identification and mitigation of cognitive bias in data-driven models. In one embodiment, a deep-learning system can comprise a memory that stores computer executable components; and a processor that executes the computer executable components stored in the memory. The computer executable components can comprise: an input component that receives data comprising primary task labels, secondary-identity attributes and a number of potential categories for one or more of the secondary-identity attributes; a machine-learning model that generates one or more predictions based on the received data; and a multi-objective learning component that trains the machine-learning model to mitigate bias from the one or more predictions.

Claims (314)

1 . A system, comprising:

a memory that stores computer executable components; and

a processor that executes at least one of the computer executable components that:

receives a set of data comprising data items, a primary classification task associated with the data items, primary task labels associated with the primary classification task, secondary-identity attributes associated with the data items, and respective categories for the secondary-identity attributes, wherein a machine learning model is trained to perform the primary classification task with primary task variables, and wherein the machine learning model has a structure comprising one or more layers associated with the primary classification task;

trains, using the set of data, the machine learning model to generate predictions with respect to the primary classification task associated with labeling the data items respectively with the primary task labels, and mitigate bias from the predictions, wherein the training comprises:

generating, using the machine learning model and the set of data, one or more of the predictions for the primary classification task associated with labeling one or more of the data items respectively with the primary task labels,

clustering, using a defined clustering process, the data items into clusters, wherein each cluster is associated with a distinct combination of the secondary-identity attributes and the respective categories of the secondary-identity attributes,

identifying one or more secondary-identity attributes that have bias based on a function of non-uniformity in distribution of respective posteriors of the clusters and respective mappings of the one or more of the predictions to the clusters,

generating one or more pseudo-task variables respectively associated with the one or more secondary-identity attributes that have the bias,

generating a pseudo-bias classification task associated with the one or more pseudo-task variables for mitigating the bias from the predictions,

modifying the machine learning model to concurrently perform the primary classification task and the pseudo-bias classification task, wherein the modifying comprise changing the structure of the machine learning model to comprise a first portion of the machine learning model associated with the primary classification task and primary task variables of the primary classification task, and second portion of the machine learning model associated with the pseudo-bias classification task and the one or more pseudo-task variables, and

training, using a multi-objective loss function, the machine learning model to mitigate the bias from the predictions by concurrently generating, using the machine learning model, for each data item of the set of data:

a prediction with respect to the primary classification task associated with labeling the data item with one of the primary task labels at above a first defined performance threshold of the multi-objective loss function that increases classification effectiveness of the machine learning model for the primary classification task, and

one or more additional predictions for the data item with respect to the pseudo-bias classification task associated with the one or more pseudo-task variables at below a second defined performance threshold of the multi-objective loss function that reduces classification effectiveness of the machine learning model for the pseudo-bias classification task, wherein generating the one or more additional predictions at below the second defined performance threshold trains the machine learning model to mitigate the bias in the predictions associated with the one or more secondary-identity attributes determined to have the bias.

2 . The system of claim 1 , wherein the at least one of the computer executable components further receives values of the secondary-identity attributes for a subset of the data items.

3 . The system of claim 1 , wherein the identifying comprises mapping the one or more of the predictions to the clusters.

4 . The system of claim 1 , wherein the training further comprises selectively turning on or off at least one pseudo-task variable based on the function of the non-uniformity in the distribution of the respective posteriors of the clusters.

5 . The system of claim 1 , wherein the training further comprises employing extrinsic data to identify a subset of the clusters that should be considered for setting the one or more pseudo-task variables.

6 . The system of claim 1 , wherein the training further comprises setting the one or more pseudo-task variables by employing a threshold-based mechanism that is based on the function:

y

i

B

(

x

)

=

{

1

,

𝕀

⁡

(

y

⁡

(

x

)

∈

P

s

∧

𝓏

i

∈

U

i

)

𝕀

⁡

(

y

⁡

(

x

)

∈

P

s

)

>

τ

0

,

otherwise

wherein x is a data item that is assigned a primary task label, y (x) is the primary task label y assigned to the data item x, P s is a category for which the bias needs to be removed, z i is a particular category of an identity attribute, U i is a set of under-represented categories within C i , C i is the categories of an i th identity attribute, and τ∈[0, 1].

7 . The system of claim 1 , wherein the multi-objective loss function comprises:

ℒ

=

P

⁡

(

y

❘

x

;

Θ

p

,

Θ

s

)

-

∑

i

=

1

n

P

⁡

(

y

i

B

❘

x

;

Θ

i

B

,

Θ

s

)

,

wherein y i B is a pseudo-task variable, x is a data item that is assigned a primary task label, y is the primary task label, P is the primary task labels, k is a quantity of the primary task labels, θ p ∈ p×k ,θ i B ∈ r×1 , θ s ∈ d×p , and d is a quantity of dimensions in a set of d-dimension real-valued vectors.

8 . A computer-implemented method comprising:

receiving, via by a system operatively coupled to a processor, a set of data comprising data items, a primary classification task associated with the data items, primary task labels associated with the primary classification task, secondary-identity attributes associated with the data items, and respective categories for the secondary-identity attributes, wherein a machine learning model is trained to perform the primary classification task with primary task variables, and wherein the machine learning model has a structure comprising one or more layers associated with the primary classification task;

training, by the system, using the set of data, the machine learning model to generate predictions with respect to the primary classification task associated with the labeling the data items respectively with the primary task labels, and mitigate bias from the predictions, wherein the training comprises:

generating, using the machine learning model and the set of data, one or more of the predictions for the primary classification task associated with labeling one or more of the data items respectively with the primary task labels,

clustering, using a defined clustering process, the data items into clusters, wherein each cluster is associated with a distinct combination of the secondary-identity attributes and the respective categories of the secondary-identity attributes,

identifying one or more secondary-identity attributes that have bias based on a function of non-uniformity in distribution of respective posteriors of the clusters and respective mappings of the one or more of the predictions to the clusters,

generating one or more pseudo-task variables respectively associated with the one or more secondary-identity attributes that have the bias,

generating a pseudo-bias classification task associated with the one or more pseudo-task variables for mitigating the bias from the predictions,

modifying the machine learning model to concurrently perform the primary classification task and the pseudo-bias classification task, wherein the modifying comprise changing the structure of the machine learning model to comprise a first portion of the machine learning model associated with the primary classification task and primary task variables of the primary classification task, and second portion of the machine learning model associated with the pseudo-bias classification task and the one or more pseudo-task variables, and

training, using a multi-objective loss function, the machine learning model to mitigate the bias from the predictions by concurrently generating, using the machine learning model, for each data item of the set of data:

a prediction with respect to the primary classification task associated with labeling the data item with one of the primary task labels at above a first defined performance threshold of the multi-objective loss function that increases classification effectiveness of the machine learning model for the primary classification task, and

one or more additional predictions for the data item with respect to the pseudo-bias classification task associated with the one or more pseudo-task variables at below a second defined performance threshold of the multi-objective loss function that reduces classification effectiveness of the machine learning model for the pseudo-bias classification task, wherein generating the one or more additional predictions at below the second defined performance threshold trains the machine learning model to mitigate the bias in the predictions associated with the one or more secondary-identity attributes determined to have the bias.

9 . The computer-implemented method of claim 8 , wherein the receiving comprises receiving values of the secondary-identity attributes for a subset of the data items.

10 . The computer-implemented method of claim 8 , wherein the identifying comprises mapping the one or more of the predictions to the clusters.

11 . The computer-implemented method of claim 8 , wherein the training further comprises selectively turning on or off at least one pseudo-task variable based on the function of the non-uniformity in the distribution of the respective posteriors of the clusters.

12 . The computer-implemented method of claim 10 , wherein the training further comprises employing extrinsic data to identify a subset of the clusters that should be considered for setting the one or more pseudo-task variables.

13 . The computer-implemented method of claim 8 , wherein the training further comprises setting the one or more pseudo-task variables by employing a threshold-based mechanism that is based on the function:

y

i

B

(

x

)

=

{

1

,

𝕀

⁡

(

y

⁡

(

x

)

∈

P

s

∧

𝓏

i

∈

U

i

)

𝕀

⁡

(

y

⁡

(

x

)

∈

P

s

)

>

τ

0

,

otherwise

wherein x is a data item that is assigned a primary task label, y (x) is the primary task label y assigned to the data item x, P s is a category for which the bias needs to be removed, z i is a particular category of an identity attribute, Ui is a set of under-represented categories within C i , C i is the categories of an i th identity attribute, and τ∈[0, 1].

14 . The computer-implemented method of claim 8 , wherein the multi-objective loss function comprises:

ℒ

=

P

⁡

(

y

❘

x

;

Θ

p

,

Θ

s

)

-

∑

i

=

1

n

P

⁡

(

y

i

B

❘

x

;

Θ

i

B

,

Θ

s

)

,

wherein y i B is a pseudo-task variable, x is a data item that is assigned a primary task label, y is the primary task label, P is the primary task labels, k is a quantity of the primary task labels, θ p ∈ p×k ,θ i B ∈ r×1 , θ s ∈ d×p , and d is a quantity of dimensions in a set of d-dimension real-valued vectors.

15 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

receive a set of data comprising data items, a primary classification task associated with the data items, primary task labels associated with the primary classification task, secondary-identity attributes associated with the data items, and respective categories for the secondary-identity attributes, wherein a machine learning model is trained to perform the primary classification task with primary task variables, and wherein the machine learning model has a structure comprising one or more layers associated with the primary classification task;

train, using the set of data, the machine learning model to generate predictions with respect to the primary classification task associated with labeling the data items respectively with the primary task labels, and mitigate bias from the predictions, wherein the training comprises:

generating, using the machine learning model and the set of data, one or more of the predictions for the primary classification task associated with labeling one or more of the data items respectively with the primary task labels,

clustering, using a defined clustering process, the data items into clusters, wherein each cluster is associated with a distinct combination of the secondary-identity attributes and the respective categories of the secondary-identity attributes,

identifying one or more secondary-identity attributes that have bias based on a function of non-uniformity in distribution of respective posteriors of the clusters and respective mappings of the one or more of the predictions to the clusters,

generating one or more pseudo-task variables respectively associated with the one or more secondary-identity attributes identified have the bias,

generating a pseudo-bias classification task associated with the one or more pseudo-task variables for mitigating the bias from the predictions,

modifying the machine learning model to concurrently perform the primary classification task and the pseudo-bias classification task, wherein the modifying comprise changing the structure of the machine learning model to comprise a first portion of the machine learning model associated with the primary classification task and primary task variables of the primary classification task, and second portion of the machine learning model associated with the pseudo-bias classification task and the one or more pseudo-task variables, and

training, using a multi-objective loss function, the machine learning model to mitigate the bias from the predictions by concurrently generating, using the machine learning model, for each data item of the set of data:

a prediction with respect to the primary classification task associated with labeling the data item with one of the primary task labels at above a first defined performance threshold of the multi-objective loss function that increases classification effectiveness of the machine learning model for the primary classification task, and

one or more additional predictions for the data item with respect to the pseudo-bias classification task associated with the one or more pseudo-task variables at below a second defined performance threshold of the multi-objective loss function that reduces classification effectiveness of the machine learning model for the pseudo-bias classification task, wherein generating the one or more additional predictions at below the second defined performance threshold trains the machine learning model to mitigate the bias in the predictions associated with the one or more secondary-identity attributes determined to have the bias.

16 . The computer program product of claim 15 , wherein the identifying comprises mapping the one or more of the predictions to the clusters.

17 . The computer program product of claim 15 , wherein the training further comprises selectively turning on or off at least one pseudo-task variable based on the function of the non-uniformity in the distribution of the respective posteriors of the clusters.

18 . The computer program product of claim 15 , wherein the training further comprises employing extrinsic data to identify a subset of the clusters that should be considered for setting the one or more pseudo-task variables.

19 . The computer program product of claim 15 , wherein the training further comprises setting the one or more pseudo-task variables by employing a threshold-based mechanism that is based on the function:

y

i

B

(

x

)

=

{

1

,

𝕀

⁡

(

y

⁡

(

x

)

∈

P

s

⋀

z

i

∈

U

i

)

𝕀

⁢

(

y

⁡

(

x

)

∈

P

s

)

>

𝒯

0

,

otherwise

wherein x is a data item that is assigned a primary task label, y (x) is the primary task label y assigned to the data item x, P s is a category for which the bias needs to be removed, z i is a particular category of an identity attribute, U i is a set of under-represented categories within C i , C i is the categories of an i th identity attribute, and τ∈[0, 1].

20 . The computer program product of claim 15 , wherein the multi-objective loss function comprises:

ℒ

=

P

⁡

(

y

⁢

❘

"\[LeftBracketingBar]"

x

;

Θ

p

,

Θ

s

)

-

∑

i

=

1

n

P

⁡

(

y

i

B

⁢

❘

"\[LeftBracketingBar]"

x

;

Θ

i

B

,

Θ

s

)

,

wherein y i B is a pseudo-task variable, x is a data item that is assigned a primary task label, y is the primary task label, P is the primary task labels, k is a quantity of the primary task labels, θ p ∈ p×k ,θ i B ∈ r×1 , θ s ∈ d×p , and d is a quantity of dimensions in a set of d-dimension real-valued vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2020
From: GANGULY, DEBASIS
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054741/0596 →
Continuity (1)
Related Publication 20220198297A1 · Jun 23, 2022
References Cited (24)
US 10275877B2 · Reicher et al. · 2019 [cited by applicant]
US 10593220B2 · Joseph et al. · 2020 [cited by applicant]
US 20180357557A1 · Williams et al. · 2018 [cited by applicant]
US 20190180358A1 · Nandan · 2019 [cited by examiner]
US 20200285939A1 · Baker · 2020 [cited by examiner]
US 20210174222A1 · Dodwell · 2021 [cited by examiner]
US 20220012591A1 · Dalli · 2022 [cited by examiner]
US 20220108222A1 · Brannon · 2022 [cited by examiner]
Wang, M. et al., “Racial Faces in the Wild: Reducing Racial Bias by Information Maximization Adaptation Network”, https://openaccess.thecvf.com/content_ICCV_2019/html/Wang_Racial_Faces_in_the_Wild_Reducing_Racial_Bias_b… [cited by examiner]
Meyerson, E. et al., “Pseudo-task Augmentation: From Deep Multitask Learning to Intratask Sharing—and Back”, https://arxiv.org/abs/1803.04062 (Year: 2018). [cited by examiner]
Farooq et al., “Artificial Intelligence based Smart Diagnosis of Alzheimer's Disease and Mild Cognitive Impairment,” International Smart cities conference, 2017, 4 pages. [cited by applicant]
Trinh et al., “A pedestrian path-planning model in accordance with obstacle's danger with reinforcement learning,” arXiv:1912.02945, 2019, 6 pages. [cited by applicant]
Das et al., “Knowledge-augmented col. Networks: Guiding Deep Learning with Advice,” arXiv:1906.01432v1 [cs.LG], 2019, 7 pages. [cited by applicant]
Patankar et al., “Bias Discovery in News Articles Using Word Vectors,” 16th IEEE International Conference on Machine Learning and Applications, 2017, 4 pages. [cited by applicant]
Wenskovitch et al., “The Effect of Semantic Interaction on Foraging in Text Analysis,” IEEE Conference on Visual Analytics Science and Technology, 2018, 12 pages. [cited by applicant]
Das, “Human-Allied Efficient and Effective Learning in Noisy Domains,” Dissertation, 2019, 208 pages. [cited by applicant]
Bolukbasi et al., “Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings,” In Proc. of NIPS, arXiv:1607.06520v1 [cs.CL], 2016, 25 pages. [cited by applicant]
Zafar et al., “Fairness beyond disparate treatment : Learning classification without disparate mistreatment,” In Proc. of WWW, arXiv:1610.08452v2 [stat.ML], 2017, 10 pages. [cited by applicant]
Beutel et al., “Data Decisions and Theoretical Implications when Adversarially Learning Fair Representations,” FAT/ML, arXiv:1707.00075v2 [cs.LG], 2017, 5 pages. [cited by applicant]
Kamiran et al., “Classifying without Discriminating,” International Conference on Computer, Control and Communication, 2009, 7 pages. [cited by applicant]
Yenala et al., “Deep learning for detecting inappropriate content in text,” International Journal of Data Science and Analytics, 2018, 14 pages. [cited by applicant]
Robinson, “Filtering inappropriate content with the Cloud Vision API,” https://cloud.google.com/blog/products/gcp/filtering-inappropriate-content-with-the-cloud-vision-api, Aug. 17, 2016, 8 pages. [cited by applicant]
Anonymous, “Learning to be Fair: A Self-Retrospective Multi-Objective Approach,” 8 pages. [cited by applicant]
Sen et al., “Towards Socially Responsible Al: Cognitive Bias-Aware Multi-Objective Learning,” Association for the Advancement of Artificial Intelligence, 2020, 8 pages. [cited by applicant]