IP Library Granted Patent US 12,373,598
Granted Patent B2
US 12,373,598 · App. 18/202,435 · Granted Jul 29, 2025

Identifying and mitigating disparate group impact in differential-privacy machine-learned models

Inventors: Jesse Cole Cresswell (Toronto, CA); Atiyeh Ashari Ghomi (Toronto, CA); Yaqiao Luo (Toronto, CA); Maria Esipova (Toronto, CA)
Assignee: The Toronto-Dominion Bank
G06F21/6245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,598
App. No.
18/202,435
Granted
Jul 29, 2025
Kind
B2
Abstract

A model evaluation system evaluates the extent to which privacy-aware training processes affect the direction of training gradients for groups. A modified differential-privacy (“DP”) training process provides per-sample gradient adjustments with parameters that may be adaptively modified for different data batches. Per-sample gradients are modified with respect to a reference bound and a clipping bound. A scaling factor may be determined for each per-sample gradient based on the higher of the reference bound or a magnitude of the per-sample gradient. Per-sample gradients may then be adjusted based on a ratio of the clipping bound to the scaling factor. A relative privacy cost between groups may be determined as excess training risk based on a difference in group gradient direction relative to an unadjusted batch gradient and the adjusted batch gradient according to the privacy-aware training.

Claims (58)

1. A system for training a computer model with differential privacy and reduced group-group privacy disparity, comprising:

one or more processors; and

a non-transitory computer-readable medium having instructions executable by the one or more processors for:

identifying a batch of training data samples;

determining a set of per-sample gradients for training a computer model by applying the computer model with a set of current model parameter values to the batch of training data samples;

determining a set of adjusted per-sample gradients by, for each per-sample gradient in the set of per-sample gradients:

setting a scaling factor to the higher of: a reference bound or a magnitude of the per-sample gradient;

determining an adjusted per-sample gradient by adjusting the per-sample gradient based on a a ratio of a clipping bound to the scaling factor;

determining a model update gradient based on the set of adjusted per-sample gradients,

wherein determining the model update gradient comprises averaging the set of adjusted per-sample gradients and adding noise; and

updating the current model parameter values based on the model update gradient.

2. The system of claim 1 , wherein determining the adjusted per-sample gradient for at least one of the per-sample gradients having the magnitude of the per-sample gradient higher than the reference bound comprises adjusting the per-sample gradient to a magnitude substantially equal to the clipping bound.

3. The system of claim 1 , wherein the batch of training data samples is one batch of a plurality of batches used to train the computer model; and wherein the instructions are further executable for:

modifying the reference bound based on the set of per-sample gradients for use of the modified reference bound with another batch of training data samples.

4. The system of claim 3 , wherein modifying the reference bound includes increasing or decreasing the reference bound based on a number of per-sample gradients having a magnitude above the reference bound.

5. The system of claim 3 , wherein modifying the reference bound includes modifying the reference bound with randomized noise.

6. The system of claim 3 , wherein modifying the reference bound comprises applying an exponential function based on:

a number of per-sample gradients having a magnitude higher than the reference bound by a threshold value;

a randomized noise;

a number of training data samples in the batch; and

a clipping learning rate.

7. A computer-implemented method for training a computer model with differential privacy and reduced group-group privacy disparity, comprising:

identifying a batch of training data samples;

determining, by one or more hardware processors, a set of per-sample gradients for training a computer model by applying the computer model with a set of current model parameter values to the batch of training data samples;

determining a set of adjusted per-sample gradients by, for each per-sample gradient in the set of per-sample gradients:

setting a scaling factor to the higher of: a reference bound or a magnitude of the per-sample gradient;

determining an adjusted per-sample gradient by adjusting the per-sample gradient based on a ratio of a clipping bound to the scaling factor;

determining a model update gradient based on the set of adjusted per-sample gradients,

wherein determining the model update gradient comprises averaging the set of adjusted per-sample gradients and adding noise; and

updating, by the one or more hardware processors, the current model parameter values based on the model update gradient.

8. The computer-implemented method of claim 7 , wherein determining the adjusted per-sample gradient for at least one of the per-sample gradients having the magnitude of the per-sample gradient higher than the reference bound comprises adjusting the per-sample gradient to a magnitude substantially equal to the clipping bound.

9. The computer-implemented method of claim 7 , wherein the batch of training data samples is one batch of a plurality of batches used to train the computer model; the method further comprising:

modifying the reference bound based on the set of per-sample gradients for use of the modified reference bound with another batch of training data samples.

10. The computer-implemented method of claim 9 , wherein modifying the reference bound includes increasing or decreasing the reference bound based on a number of per-sample gradients having a magnitude above the reference bound.

11. The computer-implemented method of claim 9 , wherein modifying the reference bound includes modifying the reference bound with randomized noise.

12. The computer-implemented method of claim 9 , wherein modifying the reference bound comprises applying an exponential function based on:

a number of per-sample gradients having a magnitude higher than the reference bound by a threshold value;

a randomized noise;

a number of training data samples in the batch; and

a clipping learning rate.

13. A non-transitory computer-readable medium for training a computer model with differential privacy and reduced group-group privacy disparity, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to:

identify a batch of training data samples;

determine a set of per-sample gradients for training a computer model by applying the computer model with a set of current model parameter values to the batch of training data samples;

determine a set of adjusted per-sample gradients by, for each per-sample gradient in the set of per-sample gradients:

set a scaling factor to the higher of: a reference bound or a magnitude of the per-sample gradient;

determine an adjusted per-sample gradient by adjusting the per-sample gradient based on a ratio of a clipping bound to the scaling factor;

determine a model update gradient based on the set of adjusted per-sample gradients, wherein determining the model update gradient comprises averaging the set of adjusted per-sample gradients and adding noise; and

update the current model parameter values based on the model update gradient.

14. The non-transitory computer-readable medium of claim 13 , wherein determining the adjusted per-sample gradient for at least one of the per-sample gradients having the magnitude of the per-sample gradient higher than the reference bound comprises adjusting the per-sample gradient to a magnitude substantially equal to the clipping bound.

15. The non-transitory computer-readable medium of claim 13 , wherein the batch of training data samples is one batch of a plurality of batches used to train the computer model; and wherein the instructions further cause the processor to:

modify the reference bound based on the set of per-sample gradients for use of the modified reference bound with another batch of training data samples.

16. The non-transitory computer-readable medium of claim 15 , wherein modifying the reference bound includes increasing or decreasing the reference bound based on a number of per-sample gradients having a magnitude above the reference bound.

17. The non-transitory computer-readable medium of claim 15 , wherein modifying the reference bound includes modifying the reference bound with randomized noise.

18. The non-transitory computer-readable medium of claim 15 , wherein modifying the reference bound comprises applying an exponential function based on:

a number of per-sample gradients having a magnitude higher than the reference bound by a threshold value;

a randomized noise;

a number of training data samples in the batch; and

a clipping learning rate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2025
From: CRESSWELL, JESSE COLE; GHOMI, ATIYEH ASHARI; LUO, YAQIAO; ESIPOVA, MARIA
To: THE TORONTO-DOMINION BANK
Reel/Frame 071409/0001 →
Continuity (3)
Provisional Application 63350333 · Jun 8, 2022
Provisional Application 63346812 · May 27, 2022
Related Publication 20230385443A1 · Nov 30, 2023
References Cited (54)
US 10402469B2 · McMahan et al. · 2019 [cited by applicant]
US 10489605B2 · Nerurkar et al. · 2019 [cited by applicant]
US 11120102B2 · McMahan et al. · 2021 [cited by applicant]
US 11893133B2 · Hockenbrocht et al. · 2024 [cited by applicant]
US 11914674B2 · Zadeh et al. · 2024 [cited by applicant]
US 12001509B2 · Kim · 2024 [cited by examiner]
US 12072998B2 · Nerurkar et al. · 2024 [cited by applicant]
US 12136038B2 · Bhalgat · 2024 [cited by examiner]
US 20190227980A1 · McMahan · 2019 [cited by examiner]
US 20210049298A1 · Suresh · 2021 [cited by examiner]
US 20210089887A1 · Johnson · 2021 [cited by examiner]
US 20210158211A1 · Talwar · 2021 [cited by examiner]
US 20210295201A1 · Kim · 2021 [cited by examiner]
US 20210374605A1 · Qian · 2021 [cited by examiner]
US 20220231648A1 · Li · 2022 [cited by examiner]
US 20220318412A1 · Guo · 2022 [cited by examiner]
US 20230351042A1 · De · 2023 [cited by examiner]
Non-Final Office Action mailed in U.S. Appl. No. 18/202,440 dated Feb. 5, 2025, 6 pages. [cited by applicant]
Abadi, et al., “Deep Learning with Differential Privacy,” ACM SIGSAC Conference on Computer and Communications Security, arXiv:1607.00133, Oct. 25, 2016, 14 pages; https://systems.cs.columbia.edu/private-systems-class/p… [cited by applicant]
Abowd, et al., “The U.S. Census of Bureau Adopts Differential Privacy,” 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, ISBN 9781450355520, Aug. 2018, 3 pages; https://core.ac.uk/download/… [cited by applicant]
Adnan, et al., “Federated Learning and Differential Privacy for Medical Image Analysis,” Scientific Reports Feb. 4, 2022, 10 pages; https://www.nature.com/articles/s41598-022-05539-7.pdf. [cited by applicant]
Andrew, et al., “Differentially Private Learning with Adaptive Clipping,” Advances of Neural Information Processing Systems arXiv:1905.03871, May 9, 2022, 12 pages; https://arxiv.org/pdf/1905.03871.pdf. [cited by applicant]
Bagdasaryan, et al., “Differential Privacy has Disparate Impact on Model Accuracy,” Advances in Neural Information Processing Systems, 2019, 10 pages; https://proceedings.neurips.cc/paper/2019/file/fc0de4e0396fff257ea36… [cited by applicant]
Bu, et al., “On the Convergence and Calibration of Deep Learning with Differential Privacy,” arXiv preprint arXiv:2106.07830, Jun. 15, 2021, 26 pages; https://arxiv.org/pdf/2106.07830.pdf. [cited by applicant]
Buolamwini, et al., “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” 1st Conference on Fairness, Accountability, and Transparency, Jan. 21, 2018, 15 pages; https://proceedings.ml… [cited by applicant]
Carlini, et al., “The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks,” 28th USENIX Security Symposium (USENIX Security 19), arXiv:1802.08232, Jul. 16, 2019, 19 pages; https://arxiv.org/… [cited by applicant]
Chang, et al., “On the Privacy Risks of Algorithmic Fairness,” 2021 IEEE European Symposium on Security and Privacy (EuroS P), arXiv:2011.03731, Apr. 7, 2021, 12 pages; https://arxiv.org/pdf/2011.03731.pdf. [cited by applicant]
Chaudhuri, et al., “Differentially Private Empirical Risk Minimization,” Journal of Machine Learning Research, arXiv:0912.0071, Feb. 16, 2011, 40 pages; https://arxiv.org/pdf/0912.0071.pdf. [cited by applicant]
Chen, et al., “Understanding Gradient Clipping in Private SGD: A Geometric Perspective,” Advances in Neural Information Processing Systems, arXiv:2006.15429, Mar. 18, 2021, 10 pages; https://proceedings.neurips.cc/paper… [cited by applicant]
Choquette-Cho, et al., “CaPC Learning: Confidential and Private Collaborative Learning,” International Conference on Learning Representations, arXiv:2102.05188, Mar. 19, 2021, 23 pages; https://arxiv.org/pdf/2102.05188.… [cited by applicant]
Chouldechova, et al., “A Snapshot of the Frontiers of Fairness in Machine Learning,” Communications of the ACM, Apr. 2020, 10 pages; https://cacm.acm.org/magazines/2020/5/244336-a-snapshot-of-the-frontiers-of-fairness-i… [cited by applicant]
Dwork, et al., “Calibrating Noise to Sensitivity in Private Data Analysis,” Theory of Cryptography Conference TCC 2006, 20 pages; https://www.researchgate.net/publication/225124717_Calibrating_Noise_to_Sensitivity_in_Pr… [cited by applicant]
Ekstrand, et al., “Privacy For All: Ensuring Fair and Equitable Privacy Protections,” 1st Conference on Fairness, Accountability, and Transparency, vol. 81 of Proceedings of Machine Learning Research, 2018, 13 pages; ht… [cited by applicant]
Farrand, et al., Neither Private Nor Fair: Impact of Data Imbalance on Utility and Fairness in Differential Privacy, in Proceedings of the 2020 Workshop on Privacy-Preserving Machine Learning in Practice, arXiv:2009.063… [cited by applicant]
Gentry, “Fully Homomorphic Encryption Using Ideal Lattices,” Proceedings of the Forty-First Annual ACM Symposium on Theory of Computing, May 31, 2009, 10 pages; https://www.cs.cmu.edu/˜odonnell/hits09/gentry-homomorphic… [cited by applicant]
Hutchison, “A Stochastic Estimator of the Trace of the Influence Matrix for Laplacian Smoothing Splines,” Communications in Statistics—Simulation and Computation, 1990, 18 pages; https://www.researchgate.net/publication… [cited by applicant]
Jagielski., et al., “Differentially Private Fair Learning,” Proceedings of the 36th International Conference on Machine Learning, arXiv:1812.02696, May 31, 2019, 31 pages; https://arxiv.org/pdf/1812.02696.pdf. [cited by applicant]
Jaiswal, et al., “Privacy Enhanced Multimodal Neural Representations for Emotion Recognition,” Proceedings of the AAAI Conference on Artificial Intelligence, arXiv:1910.13212, Oct. 29, 2019, 8 pages; https://arxiv.org/p… [cited by applicant]
Kalra, et al., “ProxyFL: Decentralized Federated Learning through Proxy Model Sharing,” arXiv preprint arXiv:2111.11343, Nov. 22, 2021, 15 pages; https://arxiv.org/pdf/2111.11343.pdf. [cited by applicant]
Le Quy, et al., “A Survey on Datasets for Fairness-Aware Machine Learning,” WIREs Data Minig and Knowledge Discovery, arXiv:2110.00530, Jan. 21, 2022, 56 pages; https://arxiv.org/pdf/2110.00530.pdf. [cited by applicant]
McMahan, et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data,” in Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), arXiv:1602.05629, Fe… [cited by applicant]
Mehrabi, et al., “A Survey on Bias and Fairness in Machine Learning,” ACM Computing Surveys, arXiv:1908.09635, Jan. 25, 2022, 34 pages; https://arxiv.org/pdf/1908.09635.pdf. [cited by applicant]
Mironov, “Renyi Differential Privacy,” 2017 IEEE 30th Computer Security Foundations Symposium (CSF), arXiv:1702.07476, Aug. 25, 2017, 13 pages; https://arxiv.org/pdf/1702.07476.pdf. [cited by applicant]
Mironov, et al., “Renyi Differential Privacy of the Sampled Gaussian Mechanism,” arXiv preprint arXiv:1908:10530. Aug. 28, 2019, 14 pages; https://arxiv.org/pdf/1908.10530.pdf. [cited by applicant]
Mozannar, et al., “Fair Learning with Private Demographic Data,” International Conference on Machine Learning, arXiv:2002.11651, Jul. 13, 2020, 37 pages; https://arxiv.org/pdf/2002.11651.pdf. [cited by applicant]
Pichapati, et al., “AdaCliP: Adaptive Clipping for Private SGD,” CoRR, arXiv:1908.07643, Oct. 23, 2019, 19 pages; https://arxiv.org/pdf/1908.07643.pdf. [cited by applicant]
Pujol, et al., “Fair Decision Making Using Privacy-Protected Data,” Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, arXiv:1905.12744, Jan. 24, 2020, 12 pages; https://arxiv.org/pdf/1905… [cited by applicant]
Tran, et al., “Differentially Private and Fair Deep Learning: A Lagrangian Dual Approach,” Proceedings of the AAAI Conference on Artificial Intelligence, arXiv:2009.12562, Sep. 26, 2020, 20 pages; https://arxiv.org/pdf/… [cited by applicant]
Tran, et al., “Differentially Private Empirical Risk Minimization Under the Fairness Lens,” Advances in Neural Information Processing Systems, vol. 34, May 21, 2021, 11 pages, https://proceedings.neurips.cc/paper/2021/f… [cited by applicant]
Xu, et al., “Removing Disparate Impact on Model Accuracy in Differentially Private Stochastic Gradient Descent on Model Accuracy,” Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, arXi… [cited by applicant]
Yu, et al., “Differentially Private Model Publishing for Deep Learning,” IEEE Symposium on Security and Privacy (SP), arXiv:1904.02200, Dec. 19, 2019, 20 pages; https://arxiv.org/pdf/1904.02200.pdf. [cited by applicant]
Hu, et al., “Adaptive clipping bound of deep learning with differential privacy,” 2021 IEEE 20th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), Mar. 9, 2022, 8 pages;… [cited by applicant]
Palanisamy, et al., “Group privacy-aware disclosure of association graph data,” 2017 IEEE International Conference on Big Data (Big Data), Jan. 15, 2018, 10 pages; https://ieeexplore.ieee.org/abstract/document/8258028. [cited by applicant]
WIPO International Search Report with Written Opinion of the ISA mailed in PCT/CA2023/050724 dated Aug. 4, 2023, 10 pages. [cited by applicant]