IP Library › Granted Patent US 12,530,427
Granted Patent B2
US 12,530,427 · App. 17/406,377 · Granted Jan 20, 2026

Systems and methods of distributed optimization

Inventors: Hugh Brendan McMahan (Seattle, WA); Jakub Konecny (Edinburgh, GB); Eider Brantly Moore (Seattle, WA); Daniel Ramage (Seattle, WA); Blaise H. Aguera-Arcas (Seattle, WA)
Assignee: GOOGLE LLC
G06F17/17G06F17/11G06F30/00G06N20/00G06F2111/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,427
App. No.
17/406,377
Granted
Jan 20, 2026
Kind
B2
Abstract

Systems and methods of determining a global model are provided. In particular, one or more local updates can be received from a plurality of user devices. Each local update can be determined by the respective user device based at least in part on one or more data examples stored on the user device. The one or more data examples stored on the plurality of user devices are distributed on an uneven basis, such that no user device includes a representative sample of the overall distribution of data examples. The local updates can then be aggregated to determine a global model.

Claims (35)

1 . A computer-implemented method of performing computationally-distributed optimization of a differentially private global model based on unevenly distributed data that is distributed across a plurality of user devices, the method comprising:

providing, by a central computing system comprising a server or data center, a global model to the plurality of user devices;

receiving, by the central computing system, one or more local updates from the plurality of user devices, each local update being determined by the respective user device based at least in part on one or more data examples stored on the respective user device, wherein the one or more data examples stored on the plurality of user devices are distributed on an uneven basis, such that no user device includes a representative sample of an overall distribution of data examples, and wherein each respective user device has performed a differential privacy technique by adding random noise to its corresponding local update prior to transmission of the corresponding local update to the central computing system;

aggregating, by the central computing system, the received local updates to determine a global update for the global model; and

transmitting, by the central computing system, the global update to the plurality of user devices for respective use by the plurality of user devices in generating predictions at the plurality of user devices.

2 . The computer-implemented method of claim 1 , wherein each respective user device has added random noise to its corresponding local update prior to transmission of the corresponding local update to the central computing system.

3 . The computer-implemented method of claim 1 , wherein at least one of the local updates comprise a gradient vector associated with the data stored on the respective user device.

4 . The computer-implemented method of claim 1 , wherein the size of each local update is independent from the size of the data used to determine the local update.

5 . The computer-implemented method of claim 1 , wherein the one or more local updates are determined using one or more optimization algorithms.

6 . The computer-implemented method of claim 1 , wherein each local update is determined based at least in part on one or more stochastic iterations, each stochastic iteration having a stepsize that is inversely proportional to a number of data examples stored on the respective user device.

7 . The computer-implemented method of claim 6 , wherein the one or more stochastic iterations are determined at least in part by randomly sampling the data examples stored on the respective user device.

8 . The computer-implemented method of claim 1 , wherein the one or more local updates are determined using one or more diagonal scaling matrices.

9 . The computer-implemented method of claim 1 , wherein aggregating, by the central computing system, the received local updates to determine the global update for the global model comprises aggregating the received local updates based at least in part on the number of data examples stored on each user device, and wherein aggregating the received local updates includes determining a weighted average of the received local updates.

10 . The computer-implemented method of claim 1 , wherein the aggregating, by the central computing system, the received local updates to determine the global update for the global model comprises scaling the received local updates on a per-coordinate basis.

11 . The computer-implemented method of claim 1 , wherein aggregating, by the central computing system, the received local updates to determine the global update for the global model comprises aggregating the received local updates for at least one iteration.

12 . The computer-implemented method of claim 11 , wherein the at least one iteration is determined based at least in part on a threshold, and wherein the threshold is determined based at least in part on an amount of time required for communication of the one or more local updates.

13 . The computer-implemented method of claim 1 , wherein the number of data examples stored on each user device is smaller than the total number of user devices.

14 . The computer-implemented method of claim 1 , further comprising providing a gradient of a loss function to each of the one or more user devices.

15 . The computer-implemented method of claim 1 , wherein aggregating, by the central computing system, the received local updates to determine the global update for the global model comprises determining a weighted average of the received local updates.

16 . A computer-implemented method of updating a local machine learning model based on unevenly distributed data, the method comprising:

determining, by a user computing device, a gradient vector of a loss function associated with a global machine learning model;

determining, by the user computing device, a local model update based at least in part on the gradient vector of the loss function and one or more locally stored data examples, wherein the distribution of the one or more locally stored data examples is not representative of an overall distribution of the data examples used to train the global machine learning model;

performing, by the user computing device, a differential privacy technique by adding random noise to the local model update; and

after performing the differential privacy technique, providing, by the user computing device, the local model update to the central computing system for use in determination of a global update to the global machine learning model, the global update to the global machine learning model being determined based on one or more local model updates received from a plurality of user devices; and

receiving, from the central computing system, the global update, wherein the global update is used by the user computing device in generating predictions.

17 . The computer-implemented method of claim 16 , wherein determining, by the user computing device, a local model update comprises determining the local model update based at least in part on one or more stochastic iterations, each stochastic iteration having a stepsize that is inversely proportional to a number of data examples stored on the respective user device.

18 . The computer-implemented method of claim 17 , wherein the one or more stochastic iterations are determined at least in part by randomly sampling the data examples stored on the respective user device.

19 . The computer-implemented method of claim 16 , wherein determining, by the user computing device, a local model update comprises determining the local model update based at least in part on one or more diagonal scaling matrices.

20 . A computing system, comprising:

one or more processors; and

one or more memory devices, the one or more memory devices storing computer-readable instructions that when executed by the one or more processors cause the one or more processors to perform operations, the operations comprising:

determining a local model update associated with an objective function based at least in part on one or more local data examples stored by the computing system, the local model update being determined using one or more stochastic iterations, each stochastic iteration having a stepsize that is inversely proportional to the number of local data examples;

adding random noise to the local model update; and

after adding random noise to the local model update, providing the local model update to a central computing system for use in the determination of a global model based on a plurality of data examples stored on a plurality of computing devices, wherein the distribution of the one or more local data examples is not representative of an overall distribution of data examples stored on the plurality of computing devices; and

receiving, from the central computing system, a global update to the global model, wherein the global update is used in generating predictions.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2026
From: MOORE, EIDER BRANTLY; RAMAGE, DANIEL; AGUERA-ARCAS, BLAISE H.
To: GOOGLE INC.
Reel/Frame 073533/0822 →
CHANGE OF NAME Recorded Jan 21, 2026
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 074463/0632 →
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY NAME PREVIOUSLY RECORDED AT REEL: 57232 FRAME: 842. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 21, 2026
From: MCMAHAN, HUGH BRENDAN; KONECNY, JAKUB
To: GOOGLE INC.
Reel/Frame 074362/0847 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2021
From: MCMAHAN, HUGH BRENDAN; KONECNY, JAKUB
To: GOOGLE LLC
Reel/Frame 057232/0842 →
Continuity (5)
Continuation 17004324 · Aug 27, 2020
Continuation 16558945 · Sep 3, 2019
Continuation 15045707 · Feb 17, 2016
Provisional Application 62242771 · Oct 16, 2015
Related Publication 20210382962A1 · Dec 9, 2021
References Cited (140)
US 6687653B1 · Kurien · 2004 [cited by examiner]
US 6708163B1 · Kargupta · 2004 [cited by examiner]
US 6879944B1 · Tipping et al. · 2005 [cited by applicant]
US 7069256B1 · Campos · 2006 [cited by applicant]
US 7664249B2 · Horvitz et al. · 2010 [cited by applicant]
US 8018874B1 · Owechko · 2011 [cited by examiner]
US 8239396B2 · Byun · 2012 [cited by examiner]
US 8321412B2 · Yang et al. · 2012 [cited by applicant]
US 8429103B1 · Aradhye et al. · 2013 [cited by applicant]
US 8473451B1 · Hakkani-Tur · 2013 [cited by examiner]
US 8954357B2 · Faddoul et al. · 2015 [cited by applicant]
US 9190055B1 · Kiss et al. · 2015 [cited by applicant]
US 9275398B1 · Kumar et al. · 2016 [cited by applicant]
US 9336483B1 · Abeysooriya · 2016 [cited by examiner]
US 9390370B2 · Kingsbury · 2016 [cited by examiner]
US 9424836B2 · Lee · 2016 [cited by examiner]
US 10402469B2 · McMahan · 2019 [cited by examiner]
US 10769549B2 · Bonawitz · 2020 [cited by examiner]
US 11023561B2 · McMahan · 2021 [cited by examiner]
US 11120102B2 · McMahan · 2021 [cited by examiner]
US 11196800B2 · Suresh · 2021 [cited by examiner]
US 11488054B2 · Bonawitz · 2022 [cited by examiner]
US 12001536B1 · Nichols · 2024 [cited by examiner]
US 20020147579A1 · Kushner · 2002 [cited by examiner]
US 20050138571A1 · Keskar et al. · 2005 [cited by applicant]
US 20060224579A1 · Zheng · 2006 [cited by applicant]
US 20080209031A1 · Zhu et al. · 2008 [cited by applicant]
US 20100132044A1 · Kogan et al. · 2010 [cited by applicant]
US 20110085546A1 · Capello · 2011 [cited by examiner]
US 20120016816A1 · Yanase · 2012 [cited by examiner]
US 20120226639A1 · Burdick et al. · 2012 [cited by applicant]
US 20120272200A1 · Lai · 2012 [cited by examiner]
US 20120310870A1 · Caves et al. · 2012 [cited by applicant]
US 20130198372A1 · Wei · 2013 [cited by examiner]
US 20130290223A1 · Chapelle et al. · 2013 [cited by applicant]
US 20140129226A1 · Lee · 2014 [cited by examiner]
US 20140214735A1 · Harik · 2014 [cited by examiner]
US 20150170053A1 · Miao · 2015 [cited by applicant]
US 20150186798A1 · Vasseur · 2015 [cited by examiner]
US 20150193695A1 · Cruz Mota · 2015 [cited by examiner]
US 20150195144A1 · Vasseur · 2015 [cited by examiner]
US 20150242760A1 · Miao et al. · 2015 [cited by applicant]
US 20150324690A1 · Chilimbi et al. · 2015 [cited by applicant]
US 20170279849A1 · Weibel · 2017 [cited by examiner]
US 20180039905A1 · Anghel · 2018 [cited by examiner]
US 20180089587A1 · Suresh · 2018 [cited by examiner]
US 20180089590A1 · Suresh · 2018 [cited by examiner]
US 20190318268A1 · Wang · 2019 [cited by examiner]
US 20200034197A1 · Nagpal · 2020 [cited by examiner]
US 20200090031A1 · Jakkam Reddi · 2020 [cited by examiner]
US 20210067339A1 · Schiatti · 2021 [cited by examiner]
US 20210326757A1 · Rawat · 2021 [cited by examiner]
US 20210342749A1 · Wang · 2021 [cited by examiner]
US 20210382962A1 · McMahan · 2021 [cited by examiner]
US 20210406782A1 · Nakayama · 2021 [cited by examiner]
US 20220076169A1 · Wang · 2022 [cited by examiner]
US 20220101130A1 · Taherzadeh Boroujeni · 2022 [cited by examiner]
CN 103327094 · 2013 [cited by applicant]
CN 103534687 · 2014 [cited by applicant]
WO WO2015126858 · 2015 [cited by applicant]
Balcan et al. (Distributed Learning, Communication Complexity and Privacy, JMLR: Workshop and Conference Proceedings vol. 23 (2012) 26.1-26.22) (Year: 2012). [cited by examiner]
Cormode et al. (“Differentially Private Spatial Decompositions”, arXiv, 2012, pp. 1-18) (Year: 2012). [cited by examiner]
Ailon et al., “Approximate Nearest Neighbors and the Fast Johnson-Lindenstrauss Transform”, 38 [cited by applicant]
Alistarh et al., “QSGB: Randomized Quantization for Communication-Optimal Stochastic Gradient Descent”, arXiv:1610.02132v1, , Oct. 7, 2016, 22 pages. [cited by applicant]
Al-Rfou et al., “Conversational Contextual Cues: The Case of Personalization and History for Response Ranking”, arXiv:1606.00372v1, Jun. 1, 2016, 10 pages. [cited by applicant]
Arjevani et al., “Communication Complexity of Distributed Convex Learning and Optimization”, Neural Information Processing Systems, Montreal, Canada, Dec. 7-12, 2015, 9 pages. [cited by applicant]
Balcan et al., “Distributed Learning, Communication Complexity and Privacy”, Conference on Learning Theory, Edinburgh, Scotland, Jun. 25-27, 2012, 19 pages. [cited by applicant]
Bonawitz et al., “Practical Secure Aggregation for Federated Learning on User-Held Data”, arXiv1611.04482v1, Nov. 14, 2016, 5 pages. [cited by applicant]
Braverman et al., “Communication Lower Bounds for Statistical Estimation Problems via a Distributed Data Processing Inequality”, 48 [cited by applicant]
Chaudhuri et al., “Differentially Private Empirical Risk Minimization”, Journal of Machine Learning Research, vol. 12, Jul. 12, 2011, pp. 1069-1109. [cited by applicant]
Chen et al., “Revisiting Distributed Synchronous SGD”, arXiv:1604.00981v3, Mar. 21, 2017, 10 pages. [cited by applicant]
Chen et al., “Communication-Optimal Distributed Clustering”, Neural Information Processing Systems, Barcelona, Spain, Dec. 5-10, 2016, 9 pages. [cited by applicant]
Chilimbi et al., “Project Adam: Building an Efficient and Scalable Deep Learning Training System”, 11th USENIX Symposium on Operating Systems Design and Implementation, Broomfield, Colorado, Oct. 6-8, 2014, pp. 571-582. [cited by applicant]
European Search Report for Application No. 16856233.8, mailed on Nov. 20, 2018, 9 pages. [cited by applicant]
Dasgupta et al., “An Elementary Proof of a Theorem of Johnson and Lindenstrauss”, Random Structures & Algorithms, vol. 22, Issue 1, 2003, pp. 60-65. [cited by applicant]
Dean et al., “Large Scale Distributed Deep Networks”, Neural Information Processing Systems, Dec. 3-8, 2012, Lake Tahoe, 9 pages. [cited by applicant]
Denil et al., “Predicting Parameters in Deep Learning”,26th International Conference on Neural Information Processing Systems, Lake Tahoe, Nevada, Dec. 5-10, 2013, pp. 2148-2156. [cited by applicant]
Duchi et al., “Privacy Aware Learning”, arXiv:1210.2085v2, Oct. 10, 2013, 60 pages. [cited by applicant]
Dwork et al., “The Algorithmic Foundations of Differential Privacy”, Foundations and Trends in Theoretical Computer Science, vol. 9, Nos. 3-4, 2014, pp. 211-407. [cited by applicant]
Efron et al., “The Jackknife Estimate of Variance”, The Annals of Statistics, vol. 9, Issue 3, May 1981, pp. 586-596. [cited by applicant]
Elias, “Universal Codeword Sets and Representations of the Integers”, IEEE Transactions on Information Theory, vol. 21, Issue 2, Mar. 1975, pp. 194-203. [cited by applicant]
Falahatgar et al., “Universal Compression of Power-Law Distributions”, arXiv:1504.08070v2, May 1, 2015, 20 pages. [cited by applicant]
Fercoq et al., “Fast Distributed Coordinate Descent for Non-Strongly Convex Losses”, arXiv:1405.5300v1, May 21, 2014, 6 pages. [cited by applicant]
Gamal et al., “On Randomized Distributed Coordinate Descent with Quantized Updates”, arXiv:1609.05539v1, Sep. 18, 2016, 5 pages. [cited by applicant]
Garg et al., “On Communication Cost of Distributed Statistical Estimation and Dimensionality”, Neural Information Processing Systems, Montreal, Canada, Dec. 8-13, 2014, 9 pages. [cited by applicant]
Golovin et al., “Large-Scale Learning with Less Ram via Randomization”, arXiv:1303.4664v1, Mar. 19, 2013, 10 pages. [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding”, arXiv:1510.00149v5, Nov. 20, 2015, 13 pages. [cited by applicant]
Ho et al., “More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server”, Neural Information Processing Systems, Dec. 5-10, 2013, Lake Tahoe, 9 pages. [cited by applicant]
Horadam, “Hadamard Matrices and Their Applications”, Princeton University Press, 2007. [cited by applicant]
International Search Report from PCT/US2016/056954 mailed Jan. 25, 2017, 15 pages. [cited by applicant]
Jaggi et al., “Communication-Efficient Distributed Dual Coordinate Ascent”, arXiv:1409.1458v2, Sep. 29, 2014, 15 pages. [cited by applicant]
Johnson et al., “Accelerating Stochastic Gradient Descent Using Predictive Variance Reduction” Advances in Neural Information Processing Systems, Lake Tahoe, Nevada, Dec. 5-10, 2013, pp. 315-323. [cited by applicant]
Konecny et al., “Federated Optimization: Distributed Machine Learning for On-Device Intelligence”, arXiv:1610.02527v1, Oct. 8, 2016, 38 pages. [cited by applicant]
Konecny et al., “Federated Optimization: Distributed Optimization Beyond the Datacenter”, arXiv:1511.03575v1, Nov. 11, 2015, 5 pages. [cited by applicant]
Konecny et al., “Semi-Stochastic Gradient Descent Methods”, arXiv:1312.1666v1, Dec. 5, 2013, 19 pages. [cited by applicant]
Konecny et al., “Federated Learning: Strategies for Improving Communication Efficiency”, arXiv:610.05492v1, Oct. 18, 2016, 5 pages. [cited by applicant]
Konecny et al., “Randomized Distributed Mean Estimation: Accuracy vs. Communication”, arXiv:1611.07555v1, Nov. 22, 2016, 19 pages. [cited by applicant]
Krichevsky et al., “The Performance of Universal Encoding”, IEEE Transactions on Information Theory, vol. 27, Issue 2, Mar. 1981, pp. 199-207. [cited by applicant]
Krizhevsky, “Learning Multiple Layers of Features from Tiny Images”, Technical Report, Apr. 8, 2009, 60 pages. [cited by applicant]
Krizhevsky, “One Weird Trick for Parallelizing Convolutional Neural Networks”, arXiv:1404.59997v2, Apr. 26, 2014, 7 pages. [cited by applicant]
Kumar et al., “Fugue: Slow-Worker-Agnostic Distributed Learning for Big Models on Big Data”, Journal of Machine Learning Research: Workshop and Conference Proceedings, Apr. 2014, 9 pages. [cited by applicant]
Livni et al., “An Algorithm for Training Polynomial Networks”, arXiv:1304.7045v1, Apr. 26, 2013, 22 pages. [cited by applicant]
Livni et al., “On the Computational Efficiency of Training Neural Networks” arXiv:1410.1141v2, Oct. 28, 2014, 15 pages. [cited by applicant]
Lloyd, “Least Squares Quantization in PCM”, IEEE Transactions on Information Theory, vol. 28, Issue 2, Mar. 1982, pp. 129-137. [cited by applicant]
Ma et al., “Adding vs. Averaging in Distributed Primal-Dual Optimization”, arXiv:1502.03508v2, Jul. 3, 2015, 19 pages. [cited by applicant]
Ma et al., “Distributed Optimization with Arbitrary Local Solvers”, arXiv:1512.04039v2, Aug. 3, 2016, 38 pages. [cited by applicant]
MacKay, “Information Theory, Inference and Learning Algorithms”, Cambridge University Press, 2003. [cited by applicant]
Mahajan, et al., “A Functional Approximation Based Distributed Learning Algorithm”, Oct. 31, 2013, https://arXiv.org/pdf/1310.8418v1, retrieved on Nov. 15, 2018. [cited by applicant]
Mahajan, et al., “An Efficient Distributed Learning Algorithm Based On Approximations”, Journal of Machine Learning Research, Mar. 16, 2015, pp. 1-32, https://arXiv.org/pdf/1310.8418.pdf, retrieved on Nov. 15, 2018. [cited by applicant]
McDonald et al., “Distributed Training Strategies for the Structures Perceptron”, Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, L… [cited by applicant]
McMahan et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data”, arXiv:1602.05629v3, Feb. 28, 2017, 11 pages. [cited by applicant]
McMahan et al., “Federated Learning: Collaborative Machine Learning without Centralized Training Data”, Apr. 6, 2017, https://research.googleblog.com/2017/04/federated-learning-collaborative.html, retrieved on Oct. 3, 2… [cited by applicant]
McMahan et al., “Federated Learning of Deep Networks using Model Averaging”, arXiv:1602.05629, Feb. 17, 2016, 11 pages. [cited by applicant]
Nedic et al., “Distributed Subgradient Methods for Multi-Agent Optimization”, Transactions on Automatic Control, vol. 54, No. 1, Jan. 2009. [cited by applicant]
Povey et al., “Parallel Training of Deep Neural Networks with Natural Gradient and Parameter Averaging”, arVix:1410.7455v1, Oct. 27, 2014, 21 pages. [cited by applicant]
Qu et al., “Coordinate Descent with Arbitrary Sampling I: Algorithms and Complexity”, arXiv:1412.8060v2, Jun. 15, 2015, 32 pages. [cited by applicant]
Qu et al., “Quartz: Randomized Dual Coordinate Ascent with Arbitrary Sampling”, arXiv:1411.5873v1, Nov. 21, 2014, 34 pages. [cited by applicant]
Rabbat et al., “Quantized Incremental Algorithms for Distributed Optimization”, Journal on Selected Areas in Communications, vol. 23, No. 4, 2005, pp. 798-808. [cited by applicant]
Reddi et al., “AIDE: Fast and Communication Efficient Distributed Optimization”, arXiv:1608.06879v1, Aug. 24, 2016, 23 pages. [cited by applicant]
Richtarik et al., “Distributed Coordinate Descent Method for Learning with Big Data”, arXiv:1310.2059v1, Oct. 8, 2013, 11 pages. [cited by applicant]
Seide et al., “1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs”, 15th Annual Conference of the International Speech Communication Association, Singapore, Sep. 14-18… [cited by applicant]
Shamir et al., “Communication-Efficient Distributed Optimization Using an Approximate Newton-Type Method”, arXiv1312.7853v4, May 13, 2013, 22 pages. [cited by applicant]
Shamir et al., “Distributed Stochastic Optimization and Learning”, 52nd Annual Allerton Conference on Communication, Control, and Computing, Monticello, Illinois, Oct. 1-3, 2014, pp. 850-857. [cited by applicant]
Shokri et al., “Privacy-Preserving Deep Learning” Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, Denver, Colorado, Oct. 12-16, 2015, 12 pages. [cited by applicant]
Springenberg et al., “Striving for Simplicity: The All Convolutional Net”, arXiv:1412.6806v3, Apr. 13, 2015, 14 pages. [cited by applicant]
Suresh et al., “Distributed Mean Estimation with Limited Communication”, arXiv.1611.00429v3, Sep. 25, 2017, 17 pages. [cited by applicant]
Tsitsiklis et al., “Communication Complexity of Convex Optimization”, Journal of Complexity, vol. 3, Issue 3, Sep. 1, 1987, pp. 231-243. [cited by applicant]
Wikipedia, “Rounding”, https://en.wikipedia.org/wiki/Rounding, retrieved on Aug. 14, 2017, 13 pages. [cited by applicant]
Woodruff, “Sketching as a Tool for Numerical Linear Algebra”, arXiv:1411.4357v3, Feb. 10, 2015, 139 pages. [cited by applicant]
Xie et al., “Distributed Machine Learning via Sufficient Factor Broadcasting”, arXiv:1409.5705v2, Sep. 7, 2015, 15 pages. [cited by applicant]
Xing et al., “Petuum: A New Platform for Distributed Machine Learning on Big Data”, Conference on Knowledge Discovery and Data Mining, Aug. 10-13, 2015, Hilton, Sydney, 10 pages. [cited by applicant]
Yadan et al., “Multi-GPU Training of ConvNets”, International Conference on Learning Representations, Apr. 14-16, 2014, Banff, Canada, 4 pages. [cited by applicant]
Yang, “Trading Computation for Communication: Distributed Stochastic Dual Coordinate Ascent”, Advances in Neural Information Processing Systems, Lake Tahoe, Nevada, Dec. 5-10, 2013, pp. 629-637. [cited by applicant]
Yu et al., “Circulant Binary Embedding”, arXiv:1405.3162v1, May 13, 2014, 9 pages. [cited by applicant]
Yu et al., “Orthogonal Random Features”, Neural Information Processing Systems, Barcelona, Spain, Dec. 5-10, 2016, 9 pages. [cited by applicant]
Zhang et al., “Communication-Efficient Algorithms for Statistical Optimization”, arXiv.1209.4129v3, Oct. 11, 2013, 44 pages. [cited by applicant]
Zhang et al., “Communication-Efficient Distributed Optimization of Self Concordant Empirical Loss”, arXiv:1501.00263v1, Jan. 1, 2015, 46 pages. [cited by applicant]
Zhang et al., “DISCO: Distributed Optimization for Self-Concordant Empirical Loss”, 32nd International Conference on Machine Learning, vol. 37, 2015, pp. 362-370. [cited by applicant]
Zhang et al., “Information-Theoretic Lower Bounds for Distributed Statistical Estimation with Communication Constraints”, Neural Information Processing Systems, Lake Tahoe, Nevada, Dec. 5-10, 2013, 9 pages. [cited by applicant]
Zhang et al., “Poseidon: A System Architecture for Efficient GPU-based Deep Learning on Multiple Machines”, arXiv:1512.06216v1, Dec. 19, 2015, 14 pages. [cited by applicant]