IP Library › Granted Patent US 12,547,759
Granted Patent B2
US 12,547,759 · App. 16/994,396 · Granted Feb 10, 2026

Privacy preserving machine learning model training

Inventors: Ananda Theertha Suresh (New York, NY); Xinnan Yu (New York, NY); Sanjiv Kumar (Jericho, NY); Sashank Jakkam Reddi (Jersey City, NJ); Venkatadheeraj Pichapati (La Jolla, CA)
Assignee: Google LLC
G06F21/6245G06F17/17G06F18/2148G06F18/2193G06N3/04G06N3/084G06V10/7747G06V10/7796
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,759
App. No.
16/994,396
Granted
Feb 10, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for privacy preserving training of a machine learning model.

Claims (98)

1 . A method of training of a machine learning model having a plurality of model parameters on user data from a plurality of users while preserving privacy of the user data, the method comprising:

initializing a respective mean noisy gradient estimate for each of the plurality of model parameters;

initializing a respective standard deviation noisy gradient estimate for each of the plurality of model parameters; and

at each of a plurality of training steps:

obtaining, over a data communication network, a plurality of training examples for the training step comprising data from a plurality of users; and

training the machine learning model using the plurality of training examples for the training step, the training comprising:

(i) updating the model parameters while preserving privacy of user data within the plurality of training examples using the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters, and

(ii) updating the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters using the plurality of training examples, wherein updating the model parameters using the plurality of training examples comprises:

for each of the plurality of training examples:

generating, for the training example and for each of the plurality of model parameters, a respective gradient of an objective function for the training example with respect to the model parameter; and

generating, for each of the plurality of model parameters, a respective privacy preserving noisy gradient for the training example that preserves privacy of user data contained within the training example by modifying the respective gradient for the model parameter for the training example based on (i) the mean noisy gradient estimate for the model parameter after being updated using privacy preserving noisy gradients with respect to the model parameter determined at a preceding training step and (ii) the standard deviation noisy gradient estimate for the model parameter after being updated using the mean noisy gradient estimate and the privacy preserving noisy gradients determined with respect to model parameter at the preceding training step, and

 wherein generating the respective privacy preserving noisy gradient comprises, for each model parameter:

 transforming the gradient for the model parameter based on (i) the mean noisy gradient estimate for the model parameter and (ii) the standard deviation noisy gradient estimates for the plurality of model parameters to generate a transformed gradient; and

 generating an adaptively clipped gradient, comprising clipping, to have a fixed norm, the transformed gradient for the model parameter that has been transformed using the respective mean and standard deviation noisy gradient estimates for the model parameter;

generating, for each of the model parameters, a respective privacy preserving update for the model parameter for the training examples for the training step that preserves privacy of user data contained within the training examples for the training step from the privacy preserving noisy gradients for the model parameter for the plurality of training examples; and

updating the model parameters with the privacy preserving updates;

and wherein updating the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters comprises, for each of the plurality of model parameters:

updating, for use at a subsequent training step, the standard deviation noisy gradient estimate for the model parameter based on the mean estimate for the model parameter and the privacy preserving noisy gradients for the model parameter; and

updating, for use at the subsequent training step, the mean noisy gradient estimate for the model parameter based on the privacy preserving noisy gradients for the model parameter.

2 . The method of claim 1 , further comprising:

initializing the mean noisy gradient estimate for each of the model parameters to a fixed value.

3 . The method of claim 1 , further comprising:

initializing the standard deviation noisy gradient estimate for each of the model parameters based on predetermined maximum and minimum values for a standard deviation of the noisy gradients for the model parameters.

4 . The method of claim 1 , further comprising:

initializing the model parameters using a machine learning parameter initialization technique.

5 . The method of claim 1 , wherein updating the mean noisy gradient estimate for the model parameter based on the privacy preserving noisy gradients for the model parameter comprises:

interpolating between the mean noisy gradient estimate and an average privacy preserving noisy gradient for the model parameter for the plurality of training examples in accordance with a mean interpolation weight hyperparameter.

6 . The method of claim 1 , wherein deriving, for each of the model parameters, a respective privacy preserving update for the model parameter for the training examples for the training step that preserves privacy of user data contained within the training examples for the training step from the privacy preserving noisy gradients for the model parameter for the plurality of training examples comprises:

computing an average noisy gradient for the model parameter from the privacy preserving noisy gradients for the model parameter for the plurality of training examples; and

applying a learning rate to the average noisy gradient for the model parameter to generate the privacy preserving update for the model parameter.

7 . The method of claim 1 , wherein updating the model parameters with the privacy preserving updates comprises:

for each model parameter, subtracting the privacy preserving update for the model parameter from the model parameter.

8 . The method of claim 1 , wherein deriving, for each of the plurality of model parameters, a respective privacy preserving noisy gradient for the training example that preserves privacy of user data contained within the training example further comprises:

adding noise to the adaptively clipped gradient.

9 . The method of claim 1 , wherein generating the transformed gradient for each of the plurality of model parameters comprises:

computing an auxiliary transformation parameter for the model parameter from the standard deviation noisy gradient estimate for the model parameter; and

transforming the gradient for the model parameter based on the mean noisy gradient estimate for the model parameter and the auxiliary transformation parameter for the model parameter to generate the transformed gradient.

10 . The method of claim 9 , wherein generating an adaptively clipped gradient further comprises:

after clipping the transformed gradient to have a fixed norm, adding noise to the transformed gradient to generate a noisy clipped transformed gradient.

11 . The method of claim 10 , wherein generating an adaptively clipped gradient further comprises:

rescaling the noisy clipped transformed gradient based on the mean noisy gradient estimate for the model parameter and the auxiliary transformation parameter for the model parameter to generate a transformed gradient.

12 . The method of claim 9 , wherein generating a transformed gradient further comprises:

subtracting the mean noisy gradient estimate for the model parameter from the gradient for the model parameter to generate a difference; and

dividing the difference by the auxiliary transformation parameter to generate the transformed gradient.

13 . The method of claim 11 , wherein rescaling the noisy clipped transformed gradient based on the mean noisy gradient estimate for the model parameter and the auxiliary transformation parameter for the model parameter comprises:

multiplying the noisy clipped transformed gradient by the auxiliary transformation parameter to generate a product; and

adding the mean noisy gradient estimate for the model parameter to the product.

14 . The method of claim 9 , wherein the auxiliary transformation parameter for the model parameter i is directly proportional to:

√{square root over ( s i t )}×√{square root over (Σ i=1 d s i t )},

where s i t is the standard deviation noisy gradient estimate for the model parameter i and d is the total number of model parameters.

15 . The method of claim 9 , wherein updating the standard deviation noisy gradient estimate for the model parameter based on the mean estimate for the model parameter and the privacy preserving noisy gradients for the model parameter comprises:

determining an initial updated estimate using the mean estimate for the model parameter and the privacy preserving noisy gradients for the model parameter; and

determining an interpolation between the standard deviation noisy gradient estimate and the initial updated estimate.

16 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training of a machine learning model having a plurality of model parameters on user data from a plurality of users while preserving privacy of the user data, the operations comprising:

initializing a respective mean noisy gradient estimate for each of the plurality of model parameters;

initializing a respective standard deviation noisy gradient estimate for each of the plurality of model parameters; and

at each of a plurality of training steps:

obtaining, over a data communication network, a plurality of training examples for the training step comprising data from a plurality of users; and

training the machine learning model using the plurality of training examples for the training step, the training comprising:

(i) updating the model parameters while preserving privacy of user data within the plurality of training examples using the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters, and

(ii) updating the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters using the plurality of training examples, wherein updating the model parameters using the plurality of training examples comprises:

for each of the plurality of training examples:

generating, for the training example and for each of the plurality of model parameters, a respective gradient of an objective function for the training example with respect to the model parameter; and

generating, for each of the plurality of model parameters, a respective privacy preserving noisy gradient for the training example that preserves privacy of user data contained within the training example by modifying the respective gradient for the model parameter for the training example based on (i) the mean noisy gradient estimate for the model parameter after being updated using privacy preserving noisy gradients with respect to the model parameter determined at a preceding training step and (ii) the standard deviation noisy gradient estimate for the model parameter after being updated using the mean noisy gradient estimate and the privacy preserving noisy gradients determined with respect to model parameter at the preceding training step, and

 wherein generating the respective privacy preserving noisy gradient comprises, for each model parameter:

 transforming the gradient for the model parameter based on (i) the mean noisy gradient estimate for the model parameter and (ii) the standard deviation noisy gradient estimates for the plurality of model parameters to generate a transformed gradient; and

 generating an adaptively clipped gradient, comprising clipping, to have a fixed norm, the transformed gradient for the model parameter that has been transformed using the respective mean and standard deviation noisy gradient estimates for the model parameter;

generating, for each of the model parameters, a respective privacy preserving update for the model parameter for the training examples for the training step that preserves privacy of user data contained within the training examples for the training step from the privacy preserving noisy gradients for the model parameter for the plurality of training examples; and

updating the model parameters with the privacy preserving updates;

and wherein updating the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters comprises, for each of the plurality of model parameters:

updating, for use at a subsequent training step, the standard deviation noisy gradient estimate for the model parameter based on the mean estimate for the model parameter and the privacy preserving noisy gradients for the model parameter; and

updating, for use at the subsequent training step, the mean noisy gradient estimate for the model parameter based on the privacy preserving noisy gradients for the model parameter.

17 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for training of a machine learning model having a plurality of model parameters on user data from a plurality of users while preserving privacy of the user data, the operations comprising:

initializing a respective mean noisy gradient estimate for each of the plurality of model parameters;

initializing a respective standard deviation noisy gradient estimate for each of the plurality of model parameters; and

at each of a plurality of training steps:

obtaining, over a data communication network, a plurality of training examples for the training step comprising data from a plurality of users; and

training the machine learning model using the plurality of training examples for the training step, the training comprising:

(i) updating the model parameters while preserving privacy of user data within the plurality of training examples using the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters, and

(ii) updating the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters using the plurality of training examples, wherein updating the model parameters using the plurality of training examples comprises:

for each of the plurality of training examples:

generating, for the training example and for each of the plurality of model parameters, a respective gradient of an objective function for the training example with respect to the model parameter; and

generating, for each of the plurality of model parameters, a respective privacy preserving noisy gradient for the training example that preserves privacy of user data contained within the training example by modifying the respective gradient for the model parameter for the training example based on (i) the mean noisy gradient estimate for the model parameter after being updated using privacy preserving noisy gradients with respect to the model parameter determined at a preceding training step and (ii) the standard deviation noisy gradient estimate for the model parameter after being updated using the mean noisy gradient estimate and the privacy preserving noisy gradients determined with respect to model parameter at the preceding training step, and

 wherein generating the respective privacy preserving noisy gradient comprises, for each model parameter:

 transforming the gradient for the model parameter based on (i) the mean noisy gradient estimate for the model parameter and (ii) the standard deviation noisy gradient estimates for the plurality of model parameters to generate a transformed gradient; and

 generating an adaptively clipped gradient, comprising clipping, to have a fixed norm, the transformed gradient for the model parameter that has been transformed using the respective mean and standard deviation noisy gradient estimates for the model parameter;

generating, for each of the model parameters, a respective privacy preserving update for the model parameter for the training examples for the training step that preserves privacy of user data contained within the training examples for the training step from the privacy preserving noisy gradients for the model parameter for the plurality of training examples; and

updating the model parameters with the privacy preserving updates;

and wherein updating the respective mean noisy gradient estimate and the respective standard deviation noisy gradient estimate for each of the model parameters comprises, for each of the plurality of model parameters:

updating, for use at a subsequent training step, the standard deviation noisy gradient estimate for the model parameter based on the mean estimate for the model parameter and the privacy preserving noisy gradients for the model parameter; and

updating, for use at the subsequent training step, the mean noisy gradient estimate for the model parameter based on the privacy preserving noisy gradients for the model parameter.

18 . The system of claim 17 , wherein generating the transformed gradient for each of the plurality of model parameters comprises:

computing an auxiliary transformation parameter for the model parameter from the standard deviation noisy gradient estimate for the model parameter; and

transforming the gradient for the model parameter based on the mean noisy gradient estimate for the model parameter and the auxiliary transformation parameter for the model parameter to generate a transformed gradient.

19 . The system of claim 18 , wherein generating an adaptively clipped gradient further comprises:

after clipping the transformed gradient to have a fixed norm, adding noise to the transformed gradient to generate a noisy clipped transformed gradient.

20 . The system of claim 19 , wherein generating an adaptively clipped gradient further comprises:

rescaling the noisy clipped transformed gradient based on the mean noisy gradient estimate for the model parameter and the auxiliary transformation parameter for the model parameter to generate a transformed gradient.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2020
From: SURESH, ANANDA THEERTHA; YU, XINNAN; KUMAR, SANJIV; JAKKAM REDDI, SASHANK; PICHAPATI, VENKATADHEERAJ
To: GOOGLE LLC
Reel/Frame 053856/0331 →
Continuity (2)
Provisional Application 62886889 · Aug 14, 2019
Related Publication 20210049298A1 · Feb 18, 2021
References Cited (55)
US 11449639B2 · Bernau · 2022 [cited by examiner]
US 20170140264A1 · Sutskever · 2017 [cited by examiner]
US 20180268283A1 · Gilad-Bachrach · 2018 [cited by examiner]
US 20190147298A1 · Rabinovich · 2019 [cited by examiner]
US 20190188402A1 · Wang · 2019 [cited by examiner]
US 20200184106A1 · Santana de Oliveira · 2020 [cited by examiner]
US 20200311520A1 · Zhao · 2020 [cited by examiner]
US 20210209247A1 · Mohassel · 2021 [cited by examiner]
Dwork et al., “The Algorithmic Foundations of Differential Privacy,” Foundations and Trends® in Theoretical Computer Science, vol. 9, Nos. 3-4, pp. 1-277 (2013) (Year: 2013). [cited by examiner]
Garcia et al., “Semantic Noise: Privacy-Protection of Nominal Microdata through Uncorrelated Noise Addition,” IEEE (2015) (Year: 2015). [cited by examiner]
Abadi et al., “Deep Learning with Differential Privacy,” arXiv (2016) (Year: 2016). [cited by examiner]
McMahan et al., “Learning Differentially Private Recurrent Language Models,” arXiv (Feb. 2018) (Year: 2018). [cited by examiner]
Xiping Liu, “Distributed Variational Inference and Privacy,” University of Cambridge (Aug. 12, 2019) (Year: 2019). [cited by examiner]
Zhang et al., “Analysis of Gradient Clipping and Adaptive Scaling with a Relaxed Smoothness Condition,” arXiv (May 2019) (Year: 2019). [cited by examiner]
Veen et al., “Three Tools for Practical Differential Privacy,” arXiv (2018) (Year: 2018). [cited by examiner]
Xiping Liu, “Distributed Variational Inference and Privacy,” University of Cambridge (Aug. 12, 2019, Thesis) (Year: 2019). [cited by examiner]
Truex et al., “A Hybrid Approach to Privacy-Preserving Federated Learning,” arXiv (2018) (Year: 2018). [cited by examiner]
Abadi et al., “Deep learning with differential privacy”, arXiv:1607.00133v2, Oct. 2016, 14 pages. [cited by applicant]
Agarwal et al., “cpSGD: Communication-efficient and differentially-private distributed SGD”, Advances in Neural Information Processing Systems, 2018, pp. 7564-7575. [cited by applicant]
Barak et al., “Privacy, accuracy, and consistency too: A holistic solution to contingency table release”, Proceedings of the 26th ACM SIDMOD-SIDACT-SIGART Symposium on Principles of Database Systems, 2007, pp. 273-282. [cited by applicant]
Bassily et al., “Differentially private empirical risk minimization: Efficient algorithms and tight error bounds”, arXiv:1405.7085v2, Oct. 2014, 39 pages. [cited by applicant]
Bonawitz et al., “Practical secure aggregation for privacy preserving machine learning”, Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 2017, pp. 1175-1191. [cited by applicant]
Bun et al., “Concentrated differential privacy: Simplifications extensions, and lower bounds”, Theory of Cryptography Conference, 2016, pp. 635-658. [cited by applicant]
Chaudhuri et al., “Differentially private empirical risk minimization”, Journal of Machine Learning Research, 2011, 12(Mar.): 1069-1109. [cited by applicant]
Chaudhuri et al., “Privacy-preserving logistic regression”, Advances in Neural Information Processing Systems, 2009, pp. 289-296. [cited by applicant]
Chaudhuri et al., “When random sampling preserves privacy”, Annual International Cryptologu Conference, 2006, pp. 198-213. [cited by applicant]
Duchi et al., “Adaptive subgradient methods for online learning and stochastic optimization”, Journal of Machine Learning Research, 2011, 12(Jul.):2121-2159. [cited by applicant]
Duchi et al., “Local privacy and statistical minimax rates”, 2013 IEEE 54th Annual Symposium on Foundations of Computer Science, pp. 429-438. [cited by applicant]
Dwork et al., “Boosting and differential privacy”, 2010 IEEE 51st Annual Symposium on Foundations of Computer Science, 2010, pp. 51-60. [cited by applicant]
Dwork et al., “Calibrating noise to sensitivity in private data analysis”, Theory of Cryptography Conference, 2006, pp. 265-284. [cited by applicant]
Dwork et al., “Differential privacy and robust statistics”, Proceedings of the 41st Annual ACM Symposium on Theory of Computing, 2009, pp. 371-380. [cited by applicant]
Dwork et al., “Our data, ourselves: Privacy via distributed noise generation”, Annual International Conference on the Theory and Applications of Cryptographic Techniques, 2006, pp. 486-503. [cited by applicant]
Dwork et al., “The algorithmic foundations of differential privacy”, Foundations and Trends in Theoretical Computer Science, 2014, 9(3-4):211-407. [cited by applicant]
Ellenberg, “The Netflix Challenge”, Wired-San Francisco, 2008, 16(3):114. [cited by applicant]
Geng et al., “The optimal noise-adding mechanism in additive differential privacy”, arXiv:1809.10224, 10 pages. [cited by applicant]
Geng et al., “The optimal noise-adding mechanism in differential privacy”, IEEE Transactions on Information Theory, 2016, 62(2):925-951. [cited by applicant]
Hard et al., “Federated learning for mobile keyboard prediction”, arXiv:1811.03604v2, Feb. 2009, 7 pages. [cited by applicant]
He et al., “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification”, Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1026-1034. [cited by applicant]
Kairouz et al., “The composition theorem for differential privacy”, IEEE Transactions on Information Theory, 2017, 63(6):4037-4049. [cited by applicant]
Kingma et al., “ADAM: A method for stochastic optimization”, arXiv: 1412.6980v9, Jan. 2017, 15 pages. [cited by applicant]
Krizhevsky et al., “Imagenet classification with deep convolutional neural networks”, Advances in Neural Information Processing Systems, 2012, pp. 1097-1105. [cited by applicant]
Kumar et al., “Lattice rescoring strategies for long short term memory language models in speech recognition”, arXiv:1711.05448v1, Nov. 2017, 8 pages. [cited by applicant]
LeCun et al., “Gradient based learning applied to document recognition”, Proceedings of the IEEE, Nov. 1998, 86(11):2278-2324. [cited by applicant]
McMahan et al., “Learning differentially private recurrent language models”, arXiv:1710.06963v3, Feb. 2018, 14 pages. [cited by applicant]
McSherry et al., “Mechanism design via differential privacy”, 48th Annual IEEE Symposium on Foundations of Computer Science, 2007, pp. 94-103. [cited by applicant]
McSherry et al., “Privacy integrated queries: an extensible platform for privacy-preserving data analysis”, Proceedings of the 2009 ACM SIGMOD International Conference on Management of Data, pp. 19-30. [cited by applicant]
Mikolov et al., “Recurrent neural network based language modeling in meeting recognition”, Eleventh Annual Conference of the International Speech Communication Association, 2010, 4 pages. [cited by applicant]
Polyak et al., “Some methods of speeding up the convergence of iteration methods”, USSR Computational mathematics and Mathematical Physics, 1964, 4(5):1-17. [cited by applicant]
Reddi et al., “Stochastic variance reduction for nonconvex optimization”, International Conference on Machine Learning, 2016, pp. 314-323. [cited by applicant]
Rubinstein et al., “Learning in a large function space: Privacy-preserving mechanisms for SVM learning”, arXiv: 0911.52708v1, Nov. 2009, 21 pages. [cited by applicant]
Shokri et al., “Privacy-preserving deep learning”, Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, 2015, pp. 1310-1321. [cited by applicant]
Vinyals et al., “Grammar as a foreign language”, Advances in Neural Information Processing Systems, 2015, pp. 2773-2781. [cited by applicant]
Wu et al., “Bolt-on differential privacy for scalable stochastic gradient descent-based analytics”, arXiv:1606.04722v3, Mar. 2017, 29 pages. [cited by applicant]
Zhang et al., “Differentially private releasing via deep generative model (technical report)”, arXiv:1801.01594v2, Mar. 2018, 12 pages. [cited by applicant]
Zhang et al., “Functional Mechanism: Regression analysis under differential privacy”, arXiv:1208.0219v1, Aug. 2012, 12 pages. [cited by applicant]