IP Library Granted Patent US 12,639,628
Granted Patent B2
US 12,639,628 · App. 18/202,459 · Granted May 26, 2026

Distributed model training with collaboration weights for private data sets

Inventors: Jesse Cole Cresswell (Toronto, CA); Brendan Leigh Ross (Toronto, CA); Ka Ho Yenson Lau (Toronto, CA); Junfeng Wen (Waterloo, CA); Yi Sui (Newmarket, CA)
Assignee: The Toronto-Dominion Bank
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,628
App. No.
18/202,459
Granted
May 26, 2026
Kind
B2
Abstract

Model training systems collaborate on model training without revealing respective private data sets. Each private data set learns a set of client weights for a set of computer models that are also learned during training. Inference for a particular private data set is determined as a mixture of the computer model parameters according to the client weights. During training, at each iteration, the client weights are updated in one step based on how well sampled models represent the private data set. In another step, gradients are determined for each sampled model and may be weighed according to the client weight for that model, relatively increasing the gradient contribution of a private data set for model parameters that correspond more highly to that private data set.

Claims (58)

1 . A system for distributed learning of computer model parameters for private client data, comprising:

one or more processors;

one or more non-transitory computer-readable media containing instructions executable by the one or more processors for:

selecting one or more sampled computer models to sample in a training iteration, each of the one or more sampled computer models being selected from a plurality of computer models having the same model architecture and each computer model having an associated client weight and a set of training parameters describing values for applying the respective computer model, wherein one of the plurality of computer models is a local model;

for each sampled computer model:

updating the client weight for the computer model based on the computer model applied to the private data set according to the associated set of training parameters;

determining an update gradient for the sampled model as applied to the private data set and weighed based on the updated client weight; and

sending the update gradient for application to the sampled computer model; and

sending the local model to another computer model training system having another private data set;

receiving an update gradient for the local model with respect to the local model applied to the other private data set weighed by a client weight of the other model for the local model;

updating the local model based on the update gradient; and

determining parameters for an inference model for performing inference of a private local data sample by combining training parameters of the plurality of computer models in proportion to the associated client weights.

2 . The system of claim 1 , wherein the sampled models from the plurality of computer models is selected based, in part, on the client weight of the sampled models.

3 . The system of claim 1 , wherein updating the client weight for a sampled computer model comprises:

determining a loss function with respect to the sampled computer model applied to the private data set.

4 . The system of claim 3 , wherein updating the client weight further comprises:

updating a plurality of loss functions based on the loss function with respect to the sampled computer model; and

wherein the client weight is updated based on the updated plurality of loss functions.

5 . The system of claim 3 , wherein updating the client weight for a sampled computer model comprises:

determining a moving average of the loss function and updating the client weight based on the moving average.

6 . The system of claim 1 , wherein the one or more sampled computer models include a local model and at least one computer model at another computer model training system.

7 . A method for distributed learning of computer model parameters for private client data, comprising:

selecting one or more sampled computer models to sample in a training iteration, each of the one or more sampled computer models being selected from a plurality of computer models having the same model architecture and each computer model having an associated client weight and a set of training parameters describing values for applying the respective computer model, wherein one of the plurality of computer models is a local model;

for each sampled computer model:

updating the client weight for the computer model based on the computer model applied to the private data set according to the associated set of training parameters;

determining an update gradient for the sampled model as applied to the private data set and weighed based on the updated client weight; and

sending the update gradient for application to the sampled computer model; and

sending the local model to another computer model training system having another private data set;

receiving an update gradient for the local model with respect to the local model applied to the other private data set weighed by a client weight of the other model for the local model;

updating the local model based on the update gradient; and

determining parameters for an inference model for performing inference of a private local data sample by combining training parameters of the plurality of computer models in proportion to the associated client weights.

8 . The method of claim 7 , wherein the sampled models from the plurality of computer models is selected based, in part, on the client weight of the sampled models.

9 . The method of claim 7 , wherein updating the client weight for a sampled computer model comprises:

determining a loss function with respect to the sampled computer model applied to the private data set.

10 . The method of claim 9 , wherein updating the client weight further comprises:

updating a plurality of loss functions based on the loss function with respect to the sampled computer model; and

wherein the client weight is updated based on the updated plurality of loss functions.

11 . The method of claim 9 , wherein updating the client weight for a sampled computer model comprises:

determining a moving average of the loss function and updating the client weight based on the moving average.

12 . The method of claim 7 , wherein the one or more sampled computer models include a local model and at least one computer model at another computer model training system.

13 . A non-transitory computer-readable medium for distributed learning of computer model parameters for private client data, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to:

select one or more sampled computer models to sample in a training iteration, each of the one or more sampled computer models being selected from a plurality of computer models having the same model architecture and each computer model having an associated client weight and a set of training parameters describing values for applying the respective computer model, wherein one of the plurality of computer models is a local model;

for each sampled computer model:

update the client weight for the computer model based on the computer model applied to the private data set according to the associated set of training parameters;

determine an update gradient for the sampled model as applied to the private data set and weighed based on the updated client weight; and

send the update gradient for application to the sampled computer model; and

send the local model to another computer model training system having another private data set;

receive an update gradient for the local model with respect to the local model applied to the other private data set weighed by a client weight of the other model for the local model;

update the local model based on the update gradient; and

determine parameters for an inference model for performing inference of a private local data sample by combining training parameters of the plurality of computer models in proportion to the associated client weights.

14 . The non-transitory computer-readable medium of claim 13 , wherein the sampled models from the plurality of computer models is selected based, in part, on the client weight of the sampled models.

15 . The non-transitory computer-readable medium of claim 13 , wherein updating the client weight for a sampled computer model comprises:

determining a loss function with respect to the sampled computer model applied to the private data set.

16 . The non-transitory computer-readable medium of claim 15 , wherein updating the client weight further comprises:

updating a plurality of loss functions based on the loss function with respect to the sampled computer model; and

wherein the client weight is updated based on the updated plurality of loss functions.

17 . The non-transitory computer-readable medium of claim 15 , wherein updating the client weight for a sampled computer model comprises:

determining a moving average of the loss function and updating the client weight based on the moving average.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2026
From: CRESSWELL, JESSE COLE; ROSS, BRENDAN LEIGH; LAU, KA HO YENSON; WEN, JUNFENG; SUI, YI
To: TORONTO-DOMINION BANK, THE
Reel/Frame 074226/0595 →
Continuity (3)
Provisional Application 63350342 · Jun 8, 2022
Provisional Application 63346820 · May 27, 2022
Related Publication 20230385694A1 · Nov 30, 2023
References Cited (35)
US 20190294864A1 · Chabanne et al. · 2019 [cited by applicant]
US 20210073677A1 · Peterson et al. · 2021 [cited by applicant]
US 20210073678A1 · Chu et al. · 2021 [cited by applicant]
US 20220237898A1 · Uehara · 2022 [cited by applicant]
International Search Report and Written Opinion issued in PCT/CA2023/050728 on Aug. 9, 2023; 8 pages. [cited by applicant]
Acar, et al., “Debiasing Model Updates for Improving Personalized Federated Training,” Proceedings of the 38th International Conference on Machine Learning, 2021, 11 pages; http://proceedings.mlr.press/v139/acar21a/acar… [cited by applicant]
Arivazhagan, et al., “Federated Learning with Personalization Layers,” arXiv:1912.00818v1 [cs.LG], Dec. 2, 2019, 13 pages; https://arxiv.org/pdf/1912.00818.pdf. [cited by applicant]
Corinzia, et al., “Variational Federated Multi-Task Learning,” arXiv:1906.06268v2 [cs.LG], Feb. 4, 2021, 12 pages; https://arxiv.org/pdf/1906.06268.pdf. [cited by applicant]
Dempster, et al., “Maximum Likelihood from Incomplete Data via the EM Algorithm,” Journal of the Royal Statistical Society, Series B (Methodological), vol. 39, No. 1, pp. 1-38, 1977,39 pages; https://cs.brown.edu/course… [cited by applicant]
Deng, et al., “Adaptive Personalized Federated Learning,” arXiv:2003.13461v3 [cs.LG], Nov. 6, 2020, 50 pages; https://arxiv.org/pdf/2003.13461.pdf. [cited by applicant]
Dinh, et al., “Personalized Federated Learning with Moreau Envelopes,” 34th Conference on Neural Information Processing Systems, arXiv:2006.08848v3 [cs.LG], Jan. 26, 2022, 23 pages; https://arxiv.org/pdf/2006.08848.pdf. [cited by applicant]
Fallah, et al., “Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach,” 34th Conference on Neural Information Processing Systems, 2020, 12 pages; https://proceedings.neuri… [cited by applicant]
Ghosh, et al., “An Efficient Framework for Clustered Federated Learning,” arXiv:2006.04088v2 [stat.ML], Jun. 8, 2021, 28 pages; https://arxiv.org/pdf/2006.04088.pdf. [cited by applicant]
Hanzely, et al., “Federated Learning of a Mixture of Global and Local Models,” arXiv:2002.05516v3 [cs.LG], Feb. 12, 2021, 40 pages; https://arxiv.org/pdf/2002.05516.pdf. [cited by applicant]
Huang, et al., “Personalized Cross-Silo Federated Learning on Non-IID Data,” 35th AAAI Conference on Artificial Intelligence, vol. 35, No. 9, pp. 7865-7873, arXiv:2007.03797v5 [cs.LG], Dec. 14, 2021, 9 pages; https://ar… [cited by applicant]
Jiang, et al., “Improving Federated Learning Personalization via Model Agnostic Meta Learning,” arXiv:1909.12488v2 [cs.LG], Jan. 18, 2023, 11 pages; https://arxiv.org/pdf/1909.12488.pdf. [cited by applicant]
Kairouz, et al., “Advances and Open Problems in Federated Learning,” Foundations and Trends® in Machine Learning, vol. 14, No. 1-2, arXiv:1912.04977v3 [cs.LG], Mar. 9, 2021, 121 pages; https://arxiv.org/pdf/1912.04977.p… [cited by applicant]
Kalra, et al., “ProxyFL: Decentralized Federated Learning Through Proxy Model Sharing,” arXiv:2111.11343v2 [cs.LG], Dec. 16, 2021, 16 pages; https://assets.researchsquare.com/files/rs-1168002/v1_covered.pdf?c=1639662082. [cited by applicant]
Khodak, et al., “Adaptive Gradient-Based Meta-Learning Methods,” 33rd Conference on Neural Information Processing Systems 32, arXiv:1906.02717v3 [cs.LG], Dec. 7, 2019, 42 pages; https://arxiv.org/pdf/1906.02717.pdf. [cited by applicant]
Krizhevsky, A., “Learning Multiple Layers of Features from Tiny Images,” University of Toronto, Ontario, Apr. 8, 2009, 60 pages; https://www.cs.toronto.edu/˜kriz/learning-features-2009-TR.pdf. [cited by applicant]
Kulkarni, et al., “Survey of Personalization Techniques for Federated Learning,” 2020 Fourth World Conference on Smart Trends in Systems, Security and Sustainability, pp. 794-797, arXiv:2003.08673v1 [cs.LG}, Mar. 19, 20… [cited by applicant]
Li, et al., “Decentralized Federated Learning via Mutual Knowledge Transfer,” IEEE Internet of Things Journal, vol. 9, No. 2, May 12, 2021, 12 pages; https://arxiv.org/ftp/arxiv/papers/2012/2012.13063.pdf. [cited by applicant]
Li, et al., “Ditto: Fair and Robust Federated Learning Through Personalization,” Proceedings of the 38th International Conference on Machine Learning, arXiv:2012.04221v3 [cs.LG], Jun. 15, 2021, 32 pages; https://arxiv.o… [cited by applicant]
Long, et al., “Federated Learning for Open Banking,” arXiv:2108.10749v1 [cs.DC], Aug. 24, 2021, 15 pages; https://arxiv.org/pdf/2108.10749.pdf. [cited by applicant]
Mansour, et al., “Three Approaches for Personalization with Applications to Federated Learning,” arXiv:2002.10619v2 [cs.LG], Jul. 19, 2020, 26 pages; https://arxiv.org/pdf/2002.10619.pdf. [cited by applicant]
Marfoq, et al., “Federated Multi-Task Learning Under a Mixture of Distributions,” 35th Conference on Neural Information Processing Systems, arXiv:2108.10252v4 [cs.LG], Nov. 7, 2022, 77 pages; https://arxiv.org/pdf/2108.… [cited by applicant]
Mcmahan, et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data,” 20th International Conference on Artificial Intelligence and Statistics, arXiv:1602.05629v4 [cs.LG], Jan. 26, 2023, 11 pages;… [cited by applicant]
Sadilek, et al., “Privacy-First Health Research with Federated Learning,” NPJ Digital Medicine, vol. 4, No. 1, pp. 132, Sep. 7, 2021, 8 pages; https://www.nature.com/articles/s41746-021-00489-2. [cited by applicant]
Sattler, et al., “Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization Under Privacy Constraints,” IEEE Transactions on Neural Networks and Learning Systems, vol. 32, No. 8, arXiv:1910.01991v1… [cited by applicant]
Shen, et al., “Federated Mutual Learning,” arXiv:2006.16765v3 [cs.LG], Sep. 17, 2020, 12 pages; https://arxiv.org/pdf/2006.16765.pdf. [cited by applicant]
Smith, et al., “Federated Multi-Task Learning,” 31st Conference on Neural Information Processing Systems, arXiv:1705.10467v2 [cs.LG], Feb. 27, 2018, 19 pages; https://arxiv.org/pdf/1705.10467.pdf. [cited by applicant]
Tan, et al., “Towards Personalized Federated Learning,” IEEE Transactions on Neural Networks and Learning Systems, arXiv:2103.00710v3 [cs.LG], Mar. 17, 2022, 17 pages; https://arxiv.org/pdf/2103.00710.pdf. [cited by applicant]
Venkateswara, et al., “Deep Hashing Network for Unsupervised Domain Adaptation,” IEEE conference on computer vision and pattern recognition, pp. 5018-5027, 2017, 10 pages; https://openaccess.thecvf.com/content_cvpr_2017… [cited by applicant]
Zhang, et al., “Personalized Federated Learning with First Order Model Optimization,” International Conference on Learning Representations 2021, arXiv:2012.08565v4 [cs.LG], Mar. 26, 2021, 17 pages; https://arxiv.org/pdf… [cited by applicant]
Zhao, et al., “Federated Learning with Non-IID Data,” arXiv preprint arXiv:1806.00582v2 [cs.LG], Jul. 21, 2022, 12 pages; https://arxiv.org/pdf/1806.00582.pdf. [cited by applicant]