IP Library Granted Patent US 12,675,738
Granted Patent B2
US 12,675,738 · App. 18/349,403 · Granted Jul 7, 2026

Systems and methods for providing fairness measures for regression machine learning models based on estimating conditional densities using gaussian mixtures

Inventors: Wei Sun (Fairfax, VA); Joshua Scott Andrews (Raleigh, NC); Xuning Tang (Mclean, VA)
Assignee: Verizon Patent and Licensing Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,738
App. No.
18/349,403
Filed
Jul 10, 2023
Granted
Jul 7, 2026
Kind
B2
Art Unit
3657
USPC
706/12
Abstract

A device may receive sensitive attribute data, model prediction data, and true target data associated with a regression machine learning model, and may determine quantities of Gaussian components for Gaussian mixtures. The device may generate Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data, and may determine parameters of estimates of conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data. The device may calculate an independence measure, a separation measure, and a sufficiency measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities. The device may perform actions based on one or more of the independence measure, the separation measure, or the sufficiency measure.

Claims (70)

1 . A method, comprising:

receiving, by a device, sensitive attribute data, model prediction data, and true target data associated with a regression machine learning model;

determining, by the device, quantities of Gaussian components for Gaussian mixtures associated with the regression machine learning model;

generating, by the device, Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data;

determining, by the device, parameters of estimates of conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data;

calculating, by the device, an independence measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities;

calculating, by the device, a separation measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities;

calculating, by the device, a sufficiency measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities; and

performing, by the device, one or more actions based on one or more of the independence measure, the separation measure, or the sufficiency measure.

2 . The method of claim 1 , wherein generating the Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data comprises:

utilizing a function to generate the Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data.

3 . The method of claim 1 , wherein determining the parameters of the estimates of the conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data comprises:

utilizing a function to determine the parameters of the estimates of the conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data.

4 . The method of claim 1 , wherein the independence measure, the separation measure, and the sufficiency measure represent a fairness associated with the regression machine learning model.

5 . The method of claim 1 , wherein performing the one or more actions comprises:

determining whether the regression machine learning model is biased based on comparing the separation measure and a separation threshold; and

providing an indication of whether the regression machine learning model is biased.

6 . The method of claim 1 , wherein performing the one or more actions comprises:

determining whether the regression machine learning model is biased based on comparing the sufficiency measure and a sufficiency threshold; and

providing an indication of whether the regression machine learning model is biased.

7 . The method of claim 1 , wherein performing the one or more actions comprises one or more of:

determining whether the regression machine learning model is biased based on one or more of the independence measure, the separation measure, or the sufficiency measure; or

retraining the regression machine learning model based on one or more of the independence measure, the separation measure, or the sufficiency measure.

8 . A device, comprising:

one or more processors configured to:

receive sensitive attribute data, model prediction data, and true target data associated with a regression machine learning model;

determine quantities of Gaussian components for Gaussian mixtures associated with the regression machine learning model;

generate Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data;

determine parameters of estimates of conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data;

calculate an independence measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities;

calculate a separation measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities;

calculate a sufficiency measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities; and

perform one or more actions based on one or more of the independence measure, the separation measure, or the sufficiency measure.

9 . The device of claim 8 , wherein the one or more processors, to generate the Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data, are configured to:

utilize a function to generate the Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data.

10 . The device of claim 8 , wherein the one or more processors, to determine the parameters of the estimates of the conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data, are configured to:

utilize a function to determine the parameters of the estimates of the conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data.

11 . The device of claim 8 , wherein the independence measure, the separation measure, and the sufficiency measure represent a fairness associated with the regression machine learning model.

12 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to:

determine whether the regression machine learning model is biased based on comparing the separation measure and a separation threshold; and

provide an indication of whether the regression machine learning model is biased.

13 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to:

determine whether the regression machine learning model is biased based on comparing the sufficiency measure and a sufficiency threshold; and

provide an indication of whether the regression machine learning model is biased.

14 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to one or more of:

determine whether the regression machine learning model is biased based on one or more of the independence measure, the separation measure, or the sufficiency measure; or

retrain the regression machine learning model based on one or more of the independence measure, the separation measure, or the sufficiency measure.

15 . A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

receive sensitive attribute data, model prediction data, and true target data associated with a regression machine learning model;

determine quantities of Gaussian components for Gaussian mixtures associated with the regression machine learning model;

generate Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data;

determine parameters of estimates of conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data;

calculate an independence measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities;

calculate a separation measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities;

calculate a sufficiency measure of the regression machine learning model based on the Gaussian mixtures and the parameters of the estimates of the conditional densities; and

perform one or more actions based on one or more of the independence measure, the separation measure, or the sufficiency measure.

16 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to generate the Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data, cause the device to:

utilize a function to generate the Gaussian mixtures of the quantities of Gaussian components based on the sensitive attribute data, the model prediction data, and the true target data.

17 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to determine the parameters of the estimates of the conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data, cause the device to:

utilize a function to determine the parameters of the estimates of the conditional densities by the Gaussian mixtures based on the sensitive attribute data, the model prediction data, and the true target data.

18 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to perform the one or more actions, cause the device to:

determine whether the regression machine learning model is biased based on comparing the separation measure and a separation threshold; and

provide an indication of whether the regression machine learning model is biased.

19 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to perform the one or more actions, cause the device to:

determine whether the regression machine learning model is biased based on comparing the sufficiency measure and a sufficiency threshold; and

provide an indication of whether the regression machine learning model is biased.

20 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to perform the one or more actions, cause the device to one or more of:

determine whether the regression machine learning model is biased based on one or more of the independence measure, the separation measure, or the sufficiency measure; or

retrain the regression machine learning model based on one or more of the independence measure, the separation measure, or the sufficiency measure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2023
From: SUN, WEI; ANDREWS, JOSHUA SCOTT; TANG, XUNING
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 064211/0350 →
Continuity (1)
Related Publication 20250021867A1 · Jan 16, 2025
References Cited (40)
US 11410050B2 · Baker · 2022 [cited by examiner]
US 11531900B2 · Baker · 2022 [cited by examiner]
US 11687788B2 · Baker · 2023 [cited by examiner]
US 12248882B2 · Baker · 2025 [cited by examiner]
US 12423586B2 · Baker · 2025 [cited by examiner]
US 20110055121A1 · Datta · 2011 [cited by examiner]
US 20200225655A1 · Cella · 2020 [cited by examiner]
US 20200285939A1 · Baker · 2020 [cited by examiner]
US 20200302524A1 · Kamkar · 2020 [cited by examiner]
US 20200320371A1 · Baker · 2020 [cited by examiner]
US 20200348662A1 · Cella · 2020 [cited by examiner]
US 20210157312A1 · Cella · 2021 [cited by examiner]
US 20220108262A1 · Cella · 2022 [cited by examiner]
US 20220164877A1 · Kamkar · 2022 [cited by examiner]
US 20220335305A1 · Baker · 2022 [cited by examiner]
US 20220383131A1 · Baker · 2022 [cited by examiner]
US 20230105547A1 · Kamkar · 2023 [cited by examiner]
US 20230289611A1 · Baker · 2023 [cited by examiner]
US 20250209333A1 · Baker · 2025 [cited by examiner]
US 20250217508A1 · Cherubin · 2025 [cited by examiner]
US 20250384341A1 · Cella · 2025 [cited by examiner]
US 20250390753A1 · Baker · 2025 [cited by examiner]
Kamiran et al., “Classification with No. Discrimination by Preferential Sampling,” Informal Proceedings of the 19th Annual Machine Learning Conference of Belgium and The Netherlands, 2010, 6 Pages. [cited by applicant]
Zhao et al., “Maximum Relevance and Minimum Redundancy Feature Selection Methods for a Marketing Machine Learning Platform,” IEEE International Conference on Data Science and Advanced Analytics (DSAA), Aug. 15, 2019, 11… [cited by applicant]
Steinberg et al., “Fairness Measures for Regression via Probabilistic Classification,” 2nd Ethics of Data Science Conference (EDSC 2020), Sydney, Australia, 9 Pages. [cited by applicant]
Agarwal et al., “Fair Regression: Quantitative Definitions and Reduction-based Algorithms,” International Conference on Machine Learning, May 30, 2019, 18 Pages. [cited by applicant]
Barocas et al., “Fairness in Machine Learning: Limitations and Opportunities,” Website: http://www.fairmlbook.org, 2019, 181 Pages. [cited by applicant]
Berk et al., “A Convex Framework for Fair Regression,” Proceedings of the Conference on Fairness, Accountability, and Transparency—FAT (2017), 5 Pages. [cited by applicant]
Bickel et al., “Discriminative Learning Under Covariate Shift,” Journal of Machine Learning Research, vol. 10, 2009, 19 Pages. [cited by applicant]
Christopher M. Bishop, “Pattern Recognition and Machine Learning,” Springer Publishing, 2006, 758 Pages. [cited by applicant]
Caton et al., “Fairness in Machine Learning: A Survey,” Website: https://arxiv.org/abs/2010.04053, Oct. 4, 2020, 33 Pages. [cited by applicant]
Chang et al., “Scalable Fusion with Mixture Distributions in Sensor Networks,” 11th Int. Conf. Control, Automation, Robotics and Vision, Dec. 7-10, 2010, 6 Pages. [cited by applicant]
Cover et al., “Elements of Information Theory,” Wiley-Interscience Publisher, 1991, 563 Pages. [cited by applicant]
Dua et al., “UCI Machine Learning Repository: Communities and Crime,” Website: http://archive.ics.uci.edu/ml, 2009, 17 Pages. [cited by applicant]
Fitzsimons et al., “A General Framework for Fair Regression,” Entropy vol. 21, No. 8, 2019, 22 Pages. [cited by applicant]
Jing Qin, “Inferences for Case-Control and Semiparametric Two-Sample Density Ratio Models,” Biometrika, vol. 85, No. 3, Sep. 1998, 13 Pages. [cited by applicant]
Rasmussen et al., “Gaussian Processes for Machine Learning,” MIT Press, 2006, 266 Pages. [cited by applicant]
Sugiyama et al., “Density Ratio Estimation in Machine Learning,” RIMS Kokyuroku, 2010, 12 Pages. [cited by applicant]
Sun et al., “Scalable Inference for Hybrid Bayesian Networks with Full Density Estimations,” Proceedings of the 13th International Conference on Information Fusion, 2010, 8 Pages. [cited by applicant]
Hsi Guang Sung, “Gaussian Mixture Regression and Classification,” Rice University PhD Thesis, May 2004, 117 Pages. [cited by applicant]