IP Library › Granted Patent US 12,340,810
Granted Patent B2
US 12,340,810 · App. 18/007,656 · Granted Jun 24, 2025

Server efficient enhancement of privacy in federated learning

Inventors: Om Thakkar (San Jose, CA); Abhradeep Guha Thakurta (Santa Clara, CA); Peter Kairouz (Seattle, WA); Borja de Balle Pigem (London, GB); Brendan McMahan (Seattle, WA)
Assignee: GOOGLE LLC
G10L15/30G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,810
App. No.
18/007,656
Granted
Jun 24, 2025
Kind
B2
Abstract

Techniques are disclosed that enable training a global model using gradients provided to a remote system by a set of client devices during a reporting window, where each client device randomly determines a reporting time in the reporting window to provide the gradient to the remote system. Various implementations include each client device determining a corresponding gradient by processing data using a local model stored locally at the client device, where the local model corresponds to the global model.

Claims (63)

1. A method implemented by one or more processors, the method comprising:

selecting, at a remote system, a set of client devices, from a plurality of client devices;

determining, at the remote system, a reporting window indicating a time frame for the set of client devices to provide one or more gradients, to update a global model;

transmitting, by the remote system, to each client device in the set of client devices, the reporting window, wherein transmitting the reporting window causes each of the client devices to at least selectively determine a corresponding reporting time, within the reporting window, for transmitting a corresponding locally generated gradient to the remote system;

receiving, in the reporting window, the corresponding locally generated gradients at the corresponding reporting times, wherein each of the corresponding locally generated gradients is generated by a corresponding one of the client devices based on processing, using a local model stored locally at the client device, data generated locally at the client device to generate a predicted output of the local model;

updating one or more portions of the global model, based on the received gradients;

selecting, at the remote system, an additional set of additional client devices, from the plurality of client devices;

determining, at the remote system, an additional reporting window indicating an additional time frame for the additional set of additional client devices to provide one or more additional gradients, to update the global model;

transmitting, by the remote system, to each additional client device in the additional set of additional client devices, the additional reporting window, wherein transmitting the additional reporting window causes each of the additional client devices to at least selectively determine a corresponding additional reporting time, within the additional reporting window, for transmitting a corresponding additional locally generated gradient to the remote system;

receiving, in the additional reporting window, the corresponding additional locally generated gradients at the corresponding additional reporting times, wherein each of the corresponding additional locally generated gradients is generated by a corresponding one of the additional client devices based on processing, using a local model stored locally at the additional client device, additional data generated locally at the additional client device to generate an additional predicted output of the local model; and

updating one or more additional portions of the global model, based on the received additional gradients.

2. The method of claim 1 , wherein processing, using the local model stored locally at the client device, data generated locally at the client device to generate the predicted output of the local model further comprises:

generating the gradient based on the predicted output of the local model and ground truth data generated by the client device.

3. The method of claim 2 , wherein the global model is a global automatic speech recognition (“ASR”) model, the local model is a local ASR model, and wherein generating the gradient based on the predicted output of the local model comprises:

processing audio data capturing a spoken utterance using the local ASR model to generate a predicted text representation of the spoken utterance; and

generating the gradient based on the predicted text representation of the spoken utterance and a ground truth representation of the spoken utterance generated by the client device.

4. The method of claim 3 , wherein each of the client devices at least selectively determining the corresponding reporting time, within the reporting window, for transmitting the corresponding locally generated gradient to the remote system comprises:

for each of the client devices, randomly determining the corresponding reporting time, within the reporting window, for transmitting the corresponding locally generated gradient to the remote system.

5. The method of claim 4 , wherein each of the client devices at least selectively determining the corresponding reporting time, within the reporting window, for transmitting the corresponding locally generated gradient to the remote system comprises:

for each of the client devices:

determining whether to transmit the corresponding locally generated gradient to the remote system; and

in response to determining to transmit the corresponding locally generated gradient, transmitting the corresponding locally generated gradient to the remote system.

6. The method of claim 5 , wherein determining whether to transmit the corresponding locally generated gradient to the remote system comprises:

randomly determining whether to transmit the corresponding locally generated gradient to the remote system.

7. The method of claim 1 , wherein at least one client device in the set of client devices, is in the additional set of additional client devices.

8. The method of claim 1 , wherein receiving, in the reporting window, the corresponding locally generated gradients at the corresponding reporting times comprises receiving a plurality of corresponding locally generated gradients at the same reporting time.

9. The method of claim 8 , wherein updating one or more portions of the global model, based on the received gradients comprises:

determining an update gradient based on the plurality of corresponding locally generated gradients received at the same reporting time; and

updating the one or more portions of the global model, based on the update gradient.

10. The method of claim 9 , wherein determining the update gradient based on the plurality of corresponding locally generated gradients, received at the same reporting time, comprises:

selecting the update gradient from the plurality of corresponding locally generated gradients received at the same reporting time.

11. The method of claim 10 , wherein selecting the update gradient from the plurality of corresponding locally generated gradients, received at the same reporting time, comprises:

randomly selecting the update gradient from the plurality of corresponding locally generated gradients, received at the same reporting time.

12. The method of claim 9 , wherein determining the update gradient based on the plurality of corresponding locally generated gradients, received at the same reporting time, comprises:

determining the update gradient based on an average of the plurality of corresponding locally generated gradients.

13. A method implemented by one or more processors, the method comprising:

selecting, at a remote system, a set of client devices, from a plurality of client devices;

determining, at the remote system, a reporting window indicating a time frame for the set of client devices to provide one or more gradients, to update a global model;

transmitting, by the remote system, to each client device in the set of client devices, the reporting window, wherein transmitting the reporting window causes each of the client devices to at least selectively determine a corresponding reporting time, within the reporting window, for transmitting a corresponding locally generated gradient to the remote system;

receiving, in the reporting window, the corresponding locally generated gradients at the corresponding reporting times, wherein each of the corresponding locally generated gradients is generated by a corresponding one of the client devices based on processing, using a local model stored locally at the client device, data generated locally at the client device to generate a predicted output of the local model,

wherein receiving, in the reporting window, the corresponding locally generated gradients at the corresponding reporting times comprises receiving a plurality of corresponding locally generated gradients at the same reporting time; and

updating one or more portions of the global model, based on the received gradients, wherein updating one or more portions of the global model, based on the received gradients comprises:

determining an update gradient based on the plurality of corresponding locally generated gradients received at the same reporting time; and

updating the one or more portions of the global model, based on the update gradient.

14. The method of claim 13 , wherein determining the update gradient based on the plurality of corresponding locally generated gradients, received at the same reporting time, comprises:

selecting the update gradient from the plurality of corresponding locally generated gradients received at the same reporting time.

15. The method of claim 14 , wherein selecting the update gradient from the plurality of corresponding locally generated gradients, received at the same reporting time, comprises:

randomly selecting the update gradient from the plurality of corresponding locally generated gradients, received at the same reporting time.

16. The method of claim 13 , wherein determining the update gradient based on the plurality of corresponding locally generated gradients, received at the same reporting time, comprises:

determining the update gradient based on an average of the plurality of corresponding locally generated gradients.

17. A remote system comprising:

memory storing instructions; and

one or more processors operable to execute the instructions to:

select a set of client devices from a plurality of client devices;

determine a reporting window indicating a time frame for the set of client devices to provide one or more gradients, to update a global model;

transmit, to each client device in the set of client devices, the reporting window, wherein transmitting the reporting window causes each of the client devices to at least selectively determine a corresponding reporting time, within the reporting window, for transmitting a corresponding locally generated gradient to the remote system;

receive, in the reporting window, the corresponding locally generated gradients at the corresponding reporting times, wherein each of the corresponding locally generated gradients is generated by a corresponding one of the client devices based on processing, using a local model stored locally at the client device, data generated locally at the client device to generate a predicted output of the local model;

update one or more portions of the global model, based on the received gradients;

select an additional set of additional client devices, from the plurality of client devices;

determine an additional reporting window indicating an additional time frame for the additional set of additional client devices to provide one or more additional gradients, to update the global model;

transmit to each additional client device in the additional set of additional client devices, the additional reporting window, wherein transmitting the additional reporting window causes each of the additional client devices to at least selectively determine a corresponding additional reporting time, within the additional reporting window, for transmitting a corresponding additional locally generated gradient to the remote system;

receive, in the additional reporting window, the corresponding additional locally generated gradients at the corresponding additional reporting times, wherein each of the corresponding additional locally generated gradients is generated by a corresponding one of the additional client devices based on processing, using a local model stored locally at the additional client device, additional data generated locally at the additional client device to generate an additional predicted output of the local model; and

update one or more additional portions of the global model, based on the received additional gradients.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2022
From: THAKKAR, OM; THAKURTA, ABHRADEEP GUHA; KAIROUZ, PETER; PIGEM, BORJA DE BALLE; MCMAHAN, BRENDAN
To: GOOGLE LLC
Reel/Frame 061957/0524 →
Continuity (2)
Provisional Application 63035559 · Jun 5, 2020
Related Publication 20230223028A1 · Jul 13, 2023
References Cited (60)
US 7596475B2 · Thiesson · 2009 [cited by examiner]
US 11062229B1 · Mnih · 2021 [cited by examiner]
US 11853391B1 · Ladkat · 2023 [cited by examiner]
US 20160267380A1 · Gemello · 2016 [cited by examiner]
US 20190213470A1 · Schmidt · 2019 [cited by examiner]
US 20200042362A1 · Cui · 2020 [cited by examiner]
US 20210174243A1 · Angel · 2021 [cited by examiner]
US 20210327410A1 · Beaufays · 2021 [cited by examiner]
US 20210365358A1 · Ide · 2021 [cited by examiner]
US 20210383280A1 · Shaloudegi · 2021 [cited by examiner]
US 20230188319A1 · Froelicher · 2023 [cited by examiner]
CN 107871160 · 2018 [cited by applicant]
CN 110572253 · 2019 [cited by applicant]
CN 111027715 · 2020 [cited by applicant]
Intellectual Property India; Examination Report issued in Application No. 202227058583; 7 pages; dated Feb. 3, 2023. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 202080100899.8; 20 pages; dated Jun. 29, 2024. [cited by applicant]
Erlingsson U. et al., Amplification by Shuffling: From Local to Central Differential Privacy via Anonymity; 12 pages; dated Jan. 6, 2019. [cited by applicant]
Erlingsson, U. et al., “Amplification by Shuffling: From Local to Central Differential Privacy via Anonymity;” Proceedings of the 2019 Annual ACM-SIAM Symposium on Discrete Algorithms; 12 pages; Jan. 6, 2019 Jan. 6, 201… [cited by applicant]
Balle, B. et al., “Privacy Amplification via Random Check-Ins;” arXiv.org; arXiv:2007.06605v1; 26 pages; Jul. 13, 2020 Jul. 13, 2020. [cited by applicant]
Shokri, R. et al., “Privacy-Preserving Deep Learning;” Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security; 12 pages; Jan. 1, 2015 Jan. 1, 2015. [cited by applicant]
Bittau, A. et al., “PROCHLO: Strong Privacy for Analytics in the Crowd;” Cornell University; arXiv.org; arXiv:1710.00901v1; 19 pages; Oct. 2, 2017 Oct. 2, 2017. [cited by applicant]
Chen, L. et al., “Crowdlearning: Crowded Deep Learning with Data Privacy;” 15th Annual IEEE International Conference on Sensing, Communication, and Networking (SECON); 9 pages; Jun. 11, 2018 Jun. 11, 2018. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT/US2020/055906; 15 pages; dated Mar. 4, 2021 Mar. 4, 2021. [cited by applicant]
Abadi, M. et al., “Deep learning with differential privacy;” In Proceedings of the 2016 ACM Conference on Computer and Communications Security (CCS); 14 pages; dated Oct. 2016. [cited by applicant]
Apple, Differential Privacy Team; “Learning with Privacy at Scale;” retrieved from Internet: https://docs-assets. developer.apple.com/ml-research/papers/learning-with-privacy-at-scale.pdf; 25 pages; dated Dec. 2017. [cited by applicant]
Augenstein, S et al., “Generative Models for Effective ML on Private, Decentralized Datasets;” Cornell University, arXiv.org; arXiv:1911.06679; 27 pages; dated Nov. 2019. [cited by applicant]
Balcer, V. et al., “Separating Local & Shuffled Differential Privacy via Histograms”; Cornell University, arXiv.org, arXiv:1911.06879; 13 pages; dated 2019. [cited by applicant]
Balle, B. et al., “Privacy Amplification by Subsampling: Tight Analyses via Couplings and Divergences”; Advances in Neural Information Processing Systems; Annual Conference on Neural Information Processing Systems (Neur… [cited by applicant]
Balle, B. et al., “The Privacy Blanket of the Shuffle Model”; In Advances in Cryptology-CRYPTO; 30 pages; dated 2019. [cited by applicant]
Bassily, R. et al., “Private Stochastic Convex Optimization with Optimal Rates;” Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems (NeurIPS 2019); 10 pages;… [cited by applicant]
Bassily, R. et al., “Private Empirical Risk Minimization: Efficient Algorithms and Tight Error Bounds;” IEEE 55th Annual Symposium on Foundations of Computer Science (FOCS); pp. 464-473; dated 2014. [cited by applicant]
Bonawitz, K. et al., “Towards Federated Learning at Scale: System Design;” Cornell University, arXiv.org, arXiv:1902.01046; 15 pages; dated 2019. [cited by applicant]
Cheu, A. et al., “Distributed Differential Privacy via Mixnets;” Cornell University, axRiv.org, arXiv:1808.01394v1; 38 pages; dated Aug. 4, 2018. [cited by applicant]
Ding, B. et al. “Collecting Telemetry Data Privately;” In Advances in Neural Information Processing Systems, Annual Conference on Neural Information Processing Systems; 10 pages; dated 2017. [cited by applicant]
Duchi, J.C. et al., “Local Privacy and Statistical Minimax Rates;” In 54th Annual IEEE Symposium on Foundations of Computer Science (FOCS); 10 pages; dated Oct. 2013. [cited by applicant]
Dwork, C. et al., “Our Data, Ourselves: Privacy Via Distributed Noise Generation;” In Advances in Cryptology-EUROCRYPT; pp. 486-503; dated 2006. [cited by applicant]
Dwork, C. et al., “Calibrating Noise to Sensitivity in Private Data Analysis;” Third Conference on Theory of Cryptography (TCC); pp. 265-284; dated 2006. [cited by applicant]
Dwork, C. et al., “The Algorithmic Foundations of Differential Privacy;” Foundations and Trends in Theoretical Computer Science, vol. 9: No. 3-4; pp. 211-407; dated 2014. [cited by applicant]
Erlingsson, U. et al., “Encode, Shuffle, Analyze Privacy Revisited: Formalizations and Empirical Evaluation;” Cornell University, arXiv.org, arXiv:2001.03618v1; dated Jan. 2020. [cited by applicant]
Erlingsson, U. et al., “RAPPOR: Randomized Aggregatable Privacy-Preserving Ordinal Response;” In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security; pp. 1054-1067; dated 2014. [cited by applicant]
Feldman, V. et al., “Privacy Amplification by Iteration;” In 59th Annual IEEE Symposium on Foundations of Computer Science (FOCS); pp. 521-532; dated 2018. [cited by applicant]
Ghazi, B. et al., “Scalable and Differentially Private Distributed Aggregation in the Shuffled Model;” Cornell University, arXiv.org, arXiv:1906.08320; 17 pages; dated 2019. [cited by applicant]
Kairouz, P. et al., “Advances and Open Problems in Federated Learning;” Cornell University, arXiv.org, arXiv:1912.04977v1; 105 pages; dated 2019. [cited by applicant]
Kairouz, P. et al., “The Composition Theorem for Differential Privacy;” IEEE Transactions on Information Theory, vol. 63, Issue 6; 14 pages; dated 2017. [cited by applicant]
Kasiviswanathan, S.P. et al., “What Can We Learn Privately?”; 49th Annual IEEE Symposium on Foundations of Computer Science (FOCS); pp. 531-540; dated 2008. [cited by applicant]
Kuo, Y. et al., “Differentially Private Hierarchical Count-Of-Counts Histograms;” Proceedings of the VLDB Endowment, vol. 11, No. 11; pp. 1509-1521; dated 2018. [cited by applicant]
McMahan, H.B. et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data;” In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS); 10 pages; dated… [cited by applicant]
McMahan, H.B. et al., “Learning Differentially Private Language Models Without Losing Accuracy;” Cornell University, arXiv.org, arXiv:1710.06963v1; 15 pages; dated Oct. 18, 2017. [cited by applicant]
Mitzenmacher, M et al., “Probability and Computing: Randomized Algorithms and Probabilistic Analysis;” Cambridge University Press; 487 pages, published Jan. 2005. [cited by applicant]
Pichapati, V. et al., “AdaCliP: Adaptive Clipping for Private SGD;” Cornell University, arXiv.org, arXiv:1908.07643; 19 pages; dated 2019. [cited by applicant]
Rogers, R. et al., LinkedIn's Audience Engagements API: A Privacy Preserving Data Analytics System at Scale; Cornell University, arXiv.org, arXiv:2002.05839; 28 pages; dated 2020. [cited by applicant]
Shamir, O. et al., “Stochastic Gradient Descent for Non-Smooth Optimization: Convergence Results and Optimal Averaging Schemes;” In International Conference on Machine Learning; 9 pages, dated 2013. [cited by applicant]
Smith, A. et al., “Is Interaction Necessary for Distributed Private Learning?”; IEEE Symposium on Security and Privacy; 20 pages; dated 2017. [cited by applicant]
Song, S. et al., “Stochastic Gradient Descent with Differentially Private Updates”; In IEEE Global Conference on Signal and Information Processing; pp. 245-248; dated 2013. [cited by applicant]
Thakkar, O. et al., “Differentially Private Learning with Adaptive Clipping”; Cornell University, arXiv.org, arXiv:1905.03871v1; 9 pages; dated May 2019. [cited by applicant]
Vitter; U.S. “Random Sampling with a Reservoir”; ACM Transactions on Mathematical Software, vol. 11, No. 1; pp. 37-57; dated Mar. 1985. [cited by applicant]
Wang, Y-X. et al., “Subsampled Renyi Differential Privacy and Analytical Moments Accountant”; Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS); 10 pages; dated 2019. [cited by applicant]
Wang, Y-X et al., “Privacy for Free: Posterior Sampling and Stochastic Gradient Monte Carlo”; In Proceedings of the 32nd International Conference on Machine Learning, vol. 37; 10 pages; dated 2015. [cited by applicant]
Wu, X. et al., “Bolt-on Differential Privacy for Scalable Stochastic Gradient Descent-based Analytics”; In Proceedings of the 2017 ACM International Conference on Management of Data, SIGMOD; pp. 1307-1322; dated 2017. [cited by applicant]
China National Intellectual Property Administration; Notification of Second Office Action issued in Application No. 202080100899.8; 14 pages; dated Jan. 22, 2025. [cited by applicant]