IP Library Granted Patent US 12,450,388
Granted Patent B2
US 12,450,388 · App. 18/202,440 · Granted Oct 21, 2025

Identifying and mitigating disparate group impact in differential-privacy machine-learned models

Inventors: Jesse Cole Cresswell (Toronto, CA); Atiyeh Ashari Ghomi (Toronto, CA); Yaqiao Luo (Toronto, CA); Maria Esipova (Toronto, CA)
Assignee: The Toronto-Dominion Bank
G06F21/6245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,388
App. No.
18/202,440
Granted
Oct 21, 2025
Kind
B2
Abstract

A model evaluation system evaluates the extent to which privacy-aware training processes affect the direction of training gradients for groups. A modified differential-privacy (“DP”) training process provides per-sample gradient adjustments with parameters that may be adaptively modified for different data batches. Per-sample gradients are modified with respect to a reference bound and a clipping bound. A scaling factor may be determined for each per-sample gradient based on the higher of the reference bound or a magnitude of the per-sample gradient. Per-sample gradients may then be adjusted based on a ratio of the clipping bound to the scaling factor. A relative privacy cost between groups may be determined as excess training risk based on a difference in group gradient direction relative to an unadjusted batch gradient and the adjusted batch gradient according to the privacy-aware training.

Claims (656)

1. A system for evaluating differential privacy for groups in training data, comprising:

one or more processors; and

a non-transitory computer-readable medium having instructions executable by the one or more processors for:

identifying a batch of training data samples, the batch of training data samples including at least one training data sample from a plurality of data groups;

determining a set of per-sample gradients for training a computer model by applying the computer model with a set of current model parameter values to the batch of training data samples;

determining an unadjusted batch gradient by combining the set of per-sample gradients;

determining an adjusted batch gradient by applying a differential-privacy algorithm to the set of per-sample gradients, the differential-privacy algorithm modifying per-sample gradients;

determining an excess risk for one or more data groups of the differential-privacy algorithm based on a direction error of data samples associated with the one or more groups, the direction error describing a change in direction between the adjusted batch gradient and the unadjusted batch gradient; and

training the model based on the excess risk of the data group.

2. The system of claim 1 , wherein the excess risk for a group is determined based on an orthogonal matrix describing a direction difference between the adjusted batch gradient and the unadjusted batch gradient.

3. The system of claim 2 , wherein the direction error is determined based on:

η

t

g

D

a

,

𝔼

[

g

¯

B

g

B

(

g

B

-

M

B

g

B

)

]

+

η

t

2

2

𝔼

[

g

¯

B

2

g

B

2

(

(

M

B

g

B

)

T

H

a

(

M

B

g

B

)

-

g

B

T

H

a

g

B

)

]

in which:

η t is a learning rate,

g D a is an unadjusted group gradient,

g B is the unadjusted batch gradient,

g B is the adjusted batch gradient,

M B is the orthogonal matrix,

H l a is a Hessian of a loss function l evaluated over data samples of the group a, and

is an expectation taken over the data samples of the group.

4. The system of claim 1 , wherein the excess risk is a disparate group-group excess risk.

5. The system of claim 1 , wherein the excess risk is determined for a first group relative to a second group is based on a difference between a first angle measured for an unadjusted group gradient and the unadjusted batch gradient and a second angle measured for the unadjusted group gradient and the adjusted batch gradient.

6. The system of claim 1 , wherein the excess risk is determined for a first group relative to the second group according to:

𝔼

[

g

¯

B

(

cos

θ

B

a

-

cos

θ

¯

B

a

)

]

>

g

D

b

g

D

a

𝔼

[

g

¯

B

(

cos

θ

B

b

-

cos

θ

¯

B

b

)

]

+

𝔼

[

g

¯

B

2

]

g

D

a

Where

:

θ

B

k

=

(

g

D

k

,

g

B

)

and

θ

¯

B

k

=

(

g

D

k

,

g

¯

B

)

g D a is an unadjusted group gradient for the first group,

g D b is an unadjusted group gradient for the second group,

g B is an unadjusted batch gradient,

g B is an adjusted batch gradient, and

is an expectation taken over the data samples of the respective groups and batch.

7. A method for evaluating differential privacy for groups in training data, the method comprising:

identifying a batch of training data samples, the batch of training data samples including at least one training data sample from a plurality of data groups;

determining a set of per-sample gradients for training a computer model by applying the computer model with a set of current model parameter values to the batch of training data samples;

determining an unadjusted batch gradient by combining the set of per-sample gradients;

determining an adjusted batch gradient by applying a differential-privacy algorithm to the set of per-sample gradients, the differential-privacy algorithm modifying per-sample gradients;

determining an excess risk for one or more data groups of the differential-privacy algorithm based on a direction error of data samples associated with the one or more groups, the direction error describing a change in direction between the adjusted batch gradient and the unadjusted batch gradient; and

training the model based on the excess risk of the data group.

8. The method of claim 7 , wherein the excess risk for a group is determined based on an orthogonal matrix describing a direction difference between the adjusted batch gradient and the unadjusted batch gradient.

9. The method of claim 8 , wherein the direction error is determined based on:

η

t

g

D

a

,

𝔼

[

g

¯

B

g

B

(

g

B

-

M

B

g

B

)

]

+

η

t

2

2

𝔼

[

g

¯

B

2

g

B

2

(

(

M

B

g

B

)

T

H

a

(

M

B

g

B

)

-

g

B

T

H

a

g

B

)

]

in which:

η t is a learning rate,

g D a is an unadjusted group gradient,

g B is the unadjusted batch gradient,

g B is the adjusted batch gradient,

M B is the orthogonal matrix,

H l a is a Hessian of a loss function l evaluated over data samples of the group a, and

is an expectation taken over the data samples of the group.

10. The method of claim 7 , wherein the excess risk is a disparate group-group excess risk.

11. The method of claim 7 , wherein the excess risk is determined for a first group relative to a second group is based on a difference between a first angle measured for an unadjusted group gradient and the unadjusted batch gradient and a second angle measured for the unadjusted group gradient and the adjusted batch gradient.

12. The method of claim 7 , wherein the excess risk is determined for a first group relative to the second group according to:

𝔼

[

g

¯

B

(

cos

θ

B

a

-

cos

θ

¯

B

a

)

]

>

g

D

b

g

D

a

𝔼

[

g

¯

B

(

cos

θ

B

b

-

cos

θ

¯

B

b

)

]

+

𝔼

[

g

¯

B

2

]

g

D

a

Where

:

θ

B

k

=

(

g

D

k

,

g

B

)

and

θ

¯

B

k

=

(

g

D

k

,

g

¯

B

)

g D a is an unadjusted group gradient for the first group,

g D b is an unadjusted group gradient the second group,

g B is an unadjusted batch gradient,

g B is an adjusted batch gradient, and

is an expectation taken over the data samples of the respective groups and batch.

13. A non-transitory computer-readable medium for training a computer model with differential privacy and reduced group-group privacy disparity, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to:

identify a batch of training data samples, the batch of training data samples including at least one training data sample from a plurality of data groups;

determine a set of per-sample gradients for training a computer model by applying the computer model with a set of current model parameter values to the batch of training data samples;

determine an unadjusted batch gradient by combining the set of per-sample gradients;

determine an adjusted batch gradient by applying a differential-privacy algorithm to the set of per-sample gradients, the differential-privacy algorithm modifying per-sample gradients;

determine an excess risk for one or more data groups of the differential-privacy algorithm based on a direction error of data samples associated with the one or more groups, the direction error describing a change in direction between the adjusted batch gradient and the unadjusted batch gradient; and

train the model based on the excess risk of the data group.

14. The non-transitory computer-readable medium of claim 13 , wherein the excess risk for a group is determined based on an orthogonal matrix describing a direction difference between the adjusted batch gradient and the unadjusted batch gradient.

15. The non-transitory computer-readable medium of claim 13 , wherein the direction error is determined based on:

η

t

g

D

a

,

𝔼

[

g

¯

B

g

B

(

g

B

-

M

B

g

B

)

]

+

η

t

2

2

𝔼

[

g

¯

B

2

g

B

2

(

(

M

B

g

B

)

T

H

a

(

M

B

g

B

)

-

g

B

T

H

a

g

B

)

]

in which:

η t is a learning rate,

g D a is an unadjusted group gradient,

g B is the unadjusted batch gradient,

g B is the adjusted batch gradient,

M B is the orthogonal matrix,

H l a is a Hessian of a loss function l evaluated over data samples of the group a, and

is an expectation taken over the data samples of the group.

16. The non-transitory computer-readable medium of claim 13 , wherein the excess risk is a disparate group-group excess risk.

17. The non-transitory computer-readable medium of claim 13 , wherein the excess risk is determined for a first group relative to a second group is based on a difference between a first angle measured for an unadjusted group gradient and the unadjusted batch gradient and a second angle measured for the unadjusted group gradient and the adjusted batch gradient.

18. The non-transitory computer-readable medium of claim 13 , wherein the excess risk is determined for a first group relative to the second group according to:

𝔼

[

g

¯

B

(

cos

θ

B

a

-

cos

θ

¯

B

a

)

]

>

g

D

b

g

D

a

𝔼

[

g

¯

B

(

cos

θ

B

b

-

cos

θ

¯

B

b

)

]

+

𝔼

[

g

¯

B

2

]

g

D

a

Where

:

θ

B

k

=

(

g

D

k

,

g

B

)

and

θ

¯

B

k

=

(

g

D

k

,

g

¯

B

)

g D a is an unadjusted group gradient for the first group,

g D b is an unadjusted group gradient for the second group,

g B is an unadjusted batch gradient,

g B is an adjusted batch gradient, and

is an expectation taken over the data samples of the respective groups and batch.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2025
From: CRESSWELL, JESSE COLE; GHOMI, ATIYEH ASHARI; LUO, YAQIAO; ESIPOVA, MARIA
To: THE TORONTO-DOMINION BANK
Reel/Frame 072242/0037 →
Continuity (3)
Provisional Application 63350333 · Jun 8, 2022
Provisional Application 63346812 · May 27, 2022
Related Publication 20230385444A1 · Nov 30, 2023
References Cited (53)
US 10402469B2 · McMahan · 2019 [cited by examiner]
US 10489605B2 · Nerurkar · 2019 [cited by examiner]
US 11120102B2 · McMahan · 2021 [cited by examiner]
US 11893133B2 · Hockenbrocht · 2024 [cited by examiner]
US 11914674B2 · Zadeh · 2024 [cited by examiner]
US 12001509B2 · Kim et al. · 2024 [cited by applicant]
US 12072998B2 · Nerurkar · 2024 [cited by examiner]
US 12136038B2 · Bhalgat et al. · 2024 [cited by applicant]
US 20190227980A1 · McMahan et al. · 2019 [cited by applicant]
US 20210049298A1 · Suresh et al. · 2021 [cited by applicant]
US 20210089887A1 · Johnson et al. · 2021 [cited by applicant]
US 20210158211A1 · Talwar et al. · 2021 [cited by applicant]
US 20210295201A1 · Kim et al. · 2021 [cited by applicant]
US 20210374605A1 · Qian et al. · 2021 [cited by applicant]
US 20220231648A1 · Li et al. · 2022 [cited by applicant]
US 20220318412A1 · Guo et al. · 2022 [cited by applicant]
US 20230351042A1 · De et al. · 2023 [cited by applicant]
Hu, et al., “Adaptive clipping bound of deep learning with differential privacy,” 2021 IEEE 20th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), Mar. 9, 2022, 8 pages;… [cited by applicant]
Palanisamy, et al., “Group privacy-aware disclosure of association graph data,” 2017 IEEE International Conference on Big Data (Big Data), Jan. 15, 2018, 10 pages; https://ieeexplore.ieee.org/abstract/document/8258028. [cited by applicant]
WIPO International Search Report with Written Opinion of the ISA mailed in PCT/CA2023/050724 dated Aug. 4, 2023, 10 pages. [cited by applicant]
Abadi, et al., “Deep Learning with Differential Privacy,” ACM SIGSAC Conference on Computer and Communications Security, arXiv:1607.00133, Oct. 25, 2016, 14 pages; https://systems.cs.columbia.edu/private-systems-class/p… [cited by applicant]
Abowd, et al., “The U.S. Census of Bureau Adopts Differential Privacy,” 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, ISBN 9781450355520, Aug. 2018, 3 pages; https://core.ac.uk/download/… [cited by applicant]
Adnan, et al., “Federated Learning and Differential Privacy for Medical Image Analysis,” Scientific Reports Feb. 4, 2022, 10 pages; https://www.nature.com/articles/s41598-022-05539-7.pdf. [cited by applicant]
Andrew, et al., “Differentially Private Learning with Adaptive Clipping,” Advances of Neural Information Processing Systems arXiv:1905.03871, May 9, 2022, 12 pages; https://arxiv.org/pdf/1905.03871.pdf. [cited by applicant]
Bagdasaryan, et al., “Differential Privacy has Disparate Impact on Model Accuracy,” Advances in Neural Information Processing Systems, 2019, 10 pages; https://proceedings.neurips.cc/paper/2019/file/fc0de4e0396fff257ea36… [cited by applicant]
Bu, et al., “On the Convergence and Calibration of Deep Learning with Differential Privacy,” arXiv preprint arXiv:2106.07830, Jun. 15, 2021, 26 pages; https://arxiv.org/pdf/2106.07830.pdf. [cited by applicant]
Buolamwini, et al., “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” 1st Conference on Fairness, Accountability, and Transparency, Jan. 21, 2018, 15 pages; https://proceedings.ml… [cited by applicant]
Carlini, et al., “The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks,” 28th USENIX Security Symposium (USENIX Security 19), arXiv:1802.08232, Jul. 16, 2019, 19 pages; https://arxiv.org/… [cited by applicant]
Chang, et al., “On the Privacy Risks of Algorithmic Fairness,” 2021 IEEE European Symposium on Security and Privacy (EuroS P), arXiv:2011.03731, Apr. 7, 2021, 12 pages; https://arxiv.org/pdf/2011.03731.pdf. [cited by applicant]
Chaudhuri, et al., “Differentially Private Empirical Risk Minimization,” Journal of Machine Learning Research, arXiv:0912.0071, Feb. 16, 2011, 40 pages; https://arxiv.org/pdf/0912.0071.pdf. [cited by applicant]
Chen, et al., “Understanding Gradient Clipping in Private SGD: A Geometric Perspective,” Advances in Neural Information Processing Systems, arXiv:2006.15429, Mar. 18, 2021, 10 pages; https://proceedings.neurips.cc/paper… [cited by applicant]
Choquette-Cho, et al., “CaPC Learning: Confidential and Private Collaborative Learning,” International Conference on Learning Representations, arXiv:2102.05188, Mar. 19, 2021, 23 pages; https://arxiv.org/pdf/2102.05188.… [cited by applicant]
Chouldechova, et al., “A Snapshot of the Frontiers of Fairness in Machine Learning,” Communications of the ACM, Apr. 2020, 10 pages; https://cacm.acm.org/magazines/2020/5/244336-a-snapshot-of-the-frontiers-of-fairness-i… [cited by applicant]
Dwork, et al., “Calibrating Noise to Sensitivity in Private Data Analysis,” Theory of Cryptography Conference TCC 2006, 20 pages; https://www.researchgate.net/publication/225124717_Calibrating_Noise_to_Sensitivity_in_Pr… [cited by applicant]
Ekstrand, et al., “Privacy for All: Ensuring Fair and Equitable Privacy Protections,” 1st Conference on Fairness, Accountability, and Transparency, vol. 81 of Proceedings of Machine Learning Research, 2018, 13 pages; ht… [cited by applicant]
Farrand, et al., Neither Private Nor Fair: Impact of Data Imbalance on Utility and Fairness in Differential Privacy, in Proceedings of the 2020 Workshop on Privacy-Preserving Machine Learning in Practice, arXiv:2009.063… [cited by applicant]
Gentry, “Fully Homomorphic Encryption Using Ideal Lattices,” Proceedings of the Forty-First Annual ACM Symposium on Theory of Computing, May 31, 2009, 10 pages; https://www.cs.cmu.edu/˜odonnell/hits09/gentry-homomorphic… [cited by applicant]
Hutchison, “A Stochastic Estimator of the Trace of the Influence Matrix for Laplacian Smoothing Splines,” Communications in Statistics—Simulation and Computation, 1990, 18 pages; https://www.researchgate.net/publication… [cited by applicant]
Jagielski., et al., “Differentially Private Fair Learning,” Proceedings of the 36th International Conference on Machine Learning, arXiv:1812.02696, May 31, 2019, 31 pages; https://arxiv.org/pdf/1812.02696.pdf. [cited by applicant]
Jaiswal, et al., “Privacy Enhanced Multimodal Neural Representations for Emotion Recognition,” Proceedings of the AAAI Conference on Artificial Intelligence, arXiv:1910.13212, Oct. 29, 2019, 8 pages; https://arxiv.org/p… [cited by applicant]
Kalra, et al., “ProxyFL: Decentralized Federated Learning through Proxy Model Sharing,” arXiv preprint arXiv:2111.11343, Nov. 22, 2021, 15 pages; https://arxiv.org/pdf/2111.11343.pdf. [cited by applicant]
Le Quy, et al., “A Survey on Datasets for Fairness-Aware Machine Learning,” WIREs Data Minig and Knowledge Discovery, arXiv:2110.00530, Jan. 21, 2022, 56 pages; https://arxiv.org/pdf/2110.00530.pdf. [cited by applicant]
McMahan, et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data,” in Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), arXiv:1602.05629, Fe… [cited by applicant]
Mehrabi, et al., “A Survey on Bias and Fairness in Machine Learning,” ACM Computing Surveys, arXiv:1908.09635, Jan. 25, 2022, 34 pages; https://arxiv.org/pdf/1908.09635.pdf. [cited by applicant]
Mironov, “Renyi Differential Privacy,” 2017 IEEE 30th Computer Security Foundations Symposium (CSF), arXiv:1702.07476, Aug. 25, 2017, 13 pages; https://arxiv.org/pdf/1702.07476.pdf. [cited by applicant]
Mironov, et al., “Renyi Differential Privacy of the Sampled Gaussian Mechanism,” arXiv preprint arXiv:1908:10530, Aug. 28, 2019, 14 pages; https://arxiv.org/pdf/1908.10530.pdf. [cited by applicant]
Mozannar, et al., “Fair Learning with Private Demographic Data,” International Conference on Machine Learning, arXiv:2002.11651, Jul. 13, 2020, 37 pages; https://arxiv.org/pdf/2002.11651.pdf. [cited by applicant]
Pichapati, et al., “AdaCliP: Adaptive Clipping for Private SGD,” CoRR, arXiv:1908.07643, Oct. 23, 2019, 19 pages; https://arxiv.org/pdf/1908.07643.pdf. [cited by applicant]
Pujol, et al., “Fair Decision Making Using Privacy-Protected Data,” Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, arXiv:1905.12744, Jan. 24, 2020, 12 pages; https://arxiv.org/pdf/1905… [cited by applicant]
Tran, et al., “Differentially Private and Fair Deep Learning: A Lagrangian Dual Approach,” Proceedings of the AAAI Conference on Artificial Intelligence, arXiv:2009.12562, Sep. 26, 2020, 20 pages; https://arxiv.org/pdf/… [cited by applicant]
Tran, et al., “Differentially Private Empirical Risk Minimization Under the Fairness Lens,” Advances in Neural Information Processing Systems, vol. 34, May 21, 2021, 11 pages, https://proceedings.neurips.cc/paper/2021/f… [cited by applicant]
Xu, et al., “Removing Disparate Impact on Model Accuracy in Differentially Private Stochastic Gradient Descent on Model Accuracy,” Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, arXi… [cited by applicant]
Yu, et al., “Differentially Private Model Publishing for Deep Learning,” IEEE Symposium on Security and Privacy (SP), arXiv:1904.02200, Dec. 19, 2019, 20 pages; https://arxiv.org/pdf/1904.02200.pdf. [cited by applicant]