IP Library › Granted Patent US 11,734,612
Granted Patent B2
US 11,734,612 · App. 17/855,323 · Granted Aug 22, 2023

Obtaining a generated dataset with a predetermined bias for evaluating algorithmic fairness of a machine learning model

Inventors: Sérgio Gabriel Pontes Jesus (Gondomar, PT); Duarte Miguel Rodrigues dos Santos Marques Alves (Lisbon, PT); José Maria Pereira Rosa Correia Pombal (Lisbon, PT); André Miguel Ferreira Da Cruz (Vila do Conde, PT); Joäo António Sobral Leite Veiga (Lisbon, PT); Joäo Guilherme Simöes Bravo Ferreira (Lisbon, PT); Catarina Garcia Belém (Seixal, PT); Marco Oliveira Pena Sampaio (Vila Nova de Gaia, PT); Pedro Dos Santos Saleiro (Lisbon, PT); Pedro Gustavo Santos Rodrigues Bizarro (Lisbon, PT)
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,734,612
App. No.
17/855,323
Granted
Aug 22, 2023
Kind
B2
Abstract

In various embodiments, a process for obtaining a generated dataset with a predetermined bias for evaluating algorithmic fairness of a machine learning model includes receiving an input dataset and generating an anonymized reconstructed dataset based at least on the input dataset. The process includes introducing a predetermined bias into the generated dataset, forming an evaluation dataset based at least on the generated dataset with the predetermined bias, and outputting the evaluation dataset. In various embodiments, a process for training a generative model includes configuring a generative model and receiving training data, where the training data includes a tabular dataset. The process includes using computer processor(s) and the received training data to train the generative model, where the generative model is sampled to generate a dataset with a predetermined bias.

Claims (49)

1. A method, comprising:

receiving an input dataset;

generating an anonymized reconstructed dataset based at least on the input dataset wherein the generated dataset includes tabular data;

introducing a predetermined bias into the generated dataset while training a generative adversarial network (GAN) model, wherein:

the generative adversarial network model is configured to append one or more columns to the generated dataset, the one or more columns including at least one dataset attribute or attribute of interest for fairness evaluation; and

a generative adversarial network sampler is configured to randomly sample the generated dataset appended with the one or more columns;

forming an evaluation dataset based at least on the generated dataset with the predetermined bias; and

outputting the evaluation dataset for evaluating algorithmic fairness.

2. The method of claim 1 , wherein a column corresponding to an attribute of interest for fairness evaluation includes an attribute including a first label for a majority group and a second label for a minority group.

3. The method of claim 2 , wherein the evaluation dataset is applied to evaluate testing group size disparity by being formed such that the majority group has a larger number of records than the minority group.

4. The method of claim 2 , wherein the evaluation dataset is applied to evaluate prevalence disparity by being formed such that prevalence with respect to a binary classification task of the majority group and prevalence with respect to a binary classification task of the minority group are disparate.

5. The method of claim 2 , wherein the evaluation dataset is applied to evaluate conditional class separability disparity by being formed such that predictive performance, including true positive rate, with respect to a binary classification task is disparate between the majority group and the minority group.

6. The method of claim 5 , wherein:

the conditional class separability disparity is introduced by selecting or adding at least one reference column sampled from a plurality of multivariate normal distributions, each distribution in the plurality of multivariate normal distributions being for a combination of group label and classification task label; and

the classification task is linearly separable with adjustable true-positive-rate and false-positive-rate for the majority and minority groups determined by the attribute of interest for fairness evaluation.

7. The method of claim 6 , wherein the predetermined bias is introduced while training a generative model including by adapting a value function of the generative model during training.

8. The method of claim 7 , wherein the generative model includes the generative adversarial network (GAN) model.

9. The method of claim 8 , wherein the GAN includes a tabular-data modeling conditional generative adversarial network (CTGAN).

10. The method of claim 1 , wherein the generated dataset is used to test a machine learning model for algorithmic fairness.

11. A method, comprising:

receiving an input dataset;

generating an anonymized reconstructed dataset based at least on the input dataset, wherein the generated dataset includes tabular data;

introducing a predetermined bias into the generated dataset while training a generative adversarial network (GAN) model, wherein:

the generative adversarial network model is configured to select one or more columns from the generated dataset as a dataset attribute or attribute of interest for fairness evaluation; and

a generative adversarial network sampler is configured to sample the generated dataset according to a predetermined distribution of the selected one or more columns;

forming an evaluation dataset based at least on the generated dataset with the predetermined bias; and

outputting the evaluation dataset for evaluating algorithmic fairness.

12. The method of claim 11 , wherein a column corresponding to an attribute of interest for fairness evaluation includes an attribute including a first label for a majority group and a second label for a minority group.

13. The method of claim 12 , wherein the evaluation dataset is applied to evaluate testing group size disparity by being formed such that the majority group has a larger number of records than the minority group.

14. The method of claim 12 , wherein the evaluation dataset is applied to evaluate prevalence disparity by being formed such that prevalence with respect to a binary classification task of the majority group and prevalence with respect to a binary classification task of the minority group are disparate.

15. The method of claim 12 , wherein the evaluation dataset is applied to evaluate conditional class separability disparity by being formed such that predictive performance, including true positive rate, with respect to a binary classification task is disparate between the majority group and the minority group.

16. The method of claim 15 , wherein:

the conditional class separability disparity is introduced by selecting or adding at least one reference column sampled from a plurality of multivariate normal distributions, each distribution in the plurality of multivariate normal distributions being for a combination of group label and classification task label; and

the classification task is linearly separable with adjustable true-positive-rate and false-positive-rate for the majority and minority groups determined by the attribute of interest for fairness evaluation.

17. The method of claim 16 , wherein:

the predetermined bias is introduced while training a generative model including by adapting a value function of the generative model during training; and

the generative model includes the generative adversarial network (GAN) model.

18. The method of claim 11 , wherein the GAN includes a tabular-data modeling conditional generative adversarial network (CTGAN).

19. The method of claim 11 , wherein the generated dataset is used to test a machine learning model for algorithmic fairness.

20. A system, comprising:

a processor configured to:

receive an input dataset;

generate an anonymized reconstructed dataset based at least on the input dataset, wherein the generated dataset includes tabular data;

introduce a predetermined bias into the generated dataset while training a generative adversarial network (GAN) model, including by at least one of:

appending one or more columns to the generated dataset, wherein the one or more columns includes at least one dataset attribute or attribute of interest for fairness evaluation; or

selecting one or more columns from the generated dataset as a dataset attribute or attribute of interest for fairness evaluation;

form an evaluation dataset based at least on the generated dataset with the predetermined bias; and

output the evaluation dataset for evaluating algorithmic fairness; and

a memory coupled to the processor and configured to provide the processor with instructions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2022
From: JESUS, SÉRGIO GABRIEL PONTES; ALVES, DUARTE MIGUEL RODRIGUES DOS SANTOS MARQUES; POMBAL, JOSÉ MARIA PEREIRA ROSA CORREIA; CRUZ, ANDRÉ MIGUEL FERREIRA DA; VEIGA, JOÃO ANTÓNIO SOBRAL LEITE; FERREIRA, JOÃO GUILHERME SIMÕES BRAVO; BELÉM, CATARINA GARCIA; SAMPAIO, MARCO OLIVEIRA PENA; SALEIRO, PEDRO DOS SANTOS; BIZARRO, PEDRO GUSTAVO SANTOS RODRIGUES
To: FEEDZAI - CONSULTADORIA E INOVAÇÃO TECNOLÓGICA, S.A.
Reel/Frame 061120/0962 →
Priority Claims (1)
EP 22175664 · May 26, 2022 · regional
Continuity (2)
Provisional Application 63237961 · Aug 27, 2021
Related Publication 20230074606A1 · Mar 9, 2023