IP Library Granted Patent US 12,450,484
Granted Patent B2
US 12,450,484 · App. 18/320,385 · Granted Oct 21, 2025

Communication optimizations for distributed machine learning

Inventors: Srinivas Sridharan (Bangalore, IN); Karthikeyan Vaidyanathan (Bangalore, IN); Dipankar Das (Pune, IN); Chandrasekaran Sakthivel (Sunnyvale, CA); Mikhail E. Smorkalov (Nizhniy Novgorod, RU)
Assignee: Intel Corporation
G06N3/08G06F9/50G06F9/5061G06F9/5077G06N3/04G06N3/044G06N3/045G06N3/063G06N3/084G06N3/088G06N3/048G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,484
App. No.
18/320,385
Granted
Oct 21, 2025
Kind
B2
Abstract

Embodiments described herein provide an apparatus comprising an interconnect switch configured to couple with a plurality of graphics processors via a plurality of point-to-point interconnects and one or more processors including a graphics processor coupled with the interconnect switch via a point-to-point interconnect of the plurality of point-to-point interconnects.

Claims (34)

1. An apparatus comprising:

an interconnect switch configured to couple with a plurality of graphics processors via a plurality of point-to-point interconnects; and

one or more processors including a graphics processor, the graphics processor coupled with the interconnect switch via a point-to-point interconnect of the plurality of point-to-point interconnects, the graphics processor comprising:

a cluster of graphics multiprocessors configured for single instruction multiple thread (SIMT) operation, the cluster of graphics multiprocessors interconnected via a data interconnect and configured to exchange data via the data interconnect, the cluster of graphics multiprocessors including a graphics multiprocessor configured to:

receive data associated with a first thread group to be executed via the graphics multiprocessor during execution of operations associated with a second thread group, the data to be received via a point-to-point interconnect coupled with the interconnect switch in association with a communication pattern for messages to be transmitted between worker nodes of a first group of worker nodes configured to perform distributed training of a neural network; and

transmit data processed by the second thread group via the point-to-point interconnect coupled with the interconnect switch during execution of operations the first thread group.

2. The apparatus of claim 1 , wherein the communication pattern is an allreduce pattern.

3. The apparatus of claim 1 , wherein worker nodes in the first group of worker nodes are grouped according to topological locality and configured to perform allreduce synchronization of data between the worker nodes.

4. The apparatus of claim 3 , wherein the data associated with the first thread group includes parameter data.

5. The apparatus of claim 4 , wherein the first thread group is to transmit updated parameter data to a parameter server via the interconnect switch via the point-to-point interconnect.

6. The apparatus of claim 5 , the parameter server to synchronize the updated parameter data with second group of worker nodes, the second group of worker nodes grouped according to topological locality.

7. The apparatus of claim 6 , further comprising a second graphics processor of the plurality of graphics processors coupled with the interconnect switch via a second point-to-point interconnect, the second graphics processor configured as a worker node in the first group of worker nodes.

8. A method comprising:

transmitting configuration data associated with a set of worker nodes of a distributed training system configured to perform distributed training of a neural network, each worker node including a graphics processor configured to perform compute operations associated with a machine learning framework workflow, the set of worker nodes interconnected via an interconnect switch configured to couple with a plurality of graphics processors via a plurality of point-to-point interconnects and enable communication between the plurality of graphics processors;

creating a group of worker nodes based on a communication pattern for messages to be transmitted between during the distributed training of the neural network, wherein worker nodes of the group of worker nodes are automatically grouped according to locality; and

facilitating transmission between worker nodes in the group of worker nodes according to the communication pattern.

9. The method of claim 8 , further comprising configuring the interconnect switch to accelerate transmission of data between the worker nodes according to the communication pattern.

10. The method of claim 8 , further comprising transparently adjusting communication paths between the worker nodes based on the communication pattern.

11. The method of claim 10 , at least a portion of the worker nodes additionally interconnected via point-to-point interconnects between interconnected worker nodes.

12. The method of claim 11 , further comprising transparently adjusting communication paths between the worker nodes based on communication congestion points between worker nodes of the group of worker nodes.

13. The method as in claim 8 , further comprising performing an all-reduce synchronization of parameter data between the worker nodes within the group of worker nodes via the interconnect switch.

14. A data processing system comprising:

a host processor;

an interconnect switch coupled with the host processor, the interconnect switch configured to couple with a plurality of graphics processors via a plurality of point-to-point interconnects; and

one or more processors including a graphics processor, the graphics processor coupled with the interconnect switch via a point-to-point interconnect of the plurality of point-to-point interconnects, the graphics processor comprising:

a cluster of graphics multiprocessors configured for single instruction multiple thread (SIMT) operation, the cluster of graphics multiprocessors interconnected via a data interconnect and configured to exchange data via the data interconnect, the cluster of graphics multiprocessors including a graphics multiprocessor configured to:

receive data associated with a first thread group to be executed via the graphics multiprocessor during execution of operations associated with a second thread group, the data to be received via a point-to-point interconnect coupled with the interconnect switch in association with a communication pattern for messages to be transmitted between worker nodes of a first group of worker nodes configured to perform distributed training of a neural network; and

transmit data processed by the second thread group via the point-to-point interconnect coupled with the interconnect switch during execution of operations the first thread group.

15. The data processing system of claim 14 , wherein the communication pattern is an allreduce pattern.

16. The data processing system of claim 14 , wherein worker nodes in the first group of worker nodes are grouped according to topological locality and configured to perform allreduce synchronization of data between the worker nodes.

17. The data processing system of claim 16 , wherein the data associated with the first thread group includes parameter data.

18. The data processing system of claim 17 , wherein the first thread group is to transmit updated parameter data to a parameter server via the interconnect switch via the point-to-point interconnect.

19. The data processing system of claim 18 , the parameter server to synchronize the updated parameter data with second group of worker nodes, the second group of worker nodes grouped according to topological locality.

20. The data processing system of claim 19 , further comprising a second graphics processor of the plurality of graphics processors coupled with the interconnect switch via a second point-to-point interconnect, the second graphics processor configured as a worker node in the first group of worker nodes.

Continuity (3)
Continuation 17685462 · Mar 3, 2022
Continuation 15859180 · Dec 29, 2017
Related Publication 20230376762A1 · Nov 23, 2023
References Cited (66)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 9164807B2 · Blanc et al. · 2015 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11373266B2 · Das et al. · 2022 [cited by applicant]
US 20080034130A1 · Perego · 2008 [cited by examiner]
US 20090031070A1 · Purcell et al. · 2009 [cited by applicant]
US 20090089078A1 · Bursey · 2009 [cited by applicant]
US 20090135766A1 · Vitebsky et al. · 2009 [cited by applicant]
US 20110029471A1 · Chakradhar et al. · 2011 [cited by applicant]
US 20110091127A1 · Kisilev · 2011 [cited by examiner]
US 20110093854A1 · Blanc · 2011 [cited by applicant]
US 20120082171A1 · Georgiou · 2012 [cited by examiner]
US 20140328172A1 · Kumar et al. · 2014 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160313984A1 · Meixner · 2016 [cited by applicant]
US 20160321777A1 · Jin · 2016 [cited by applicant]
US 20160352598A1 · Reinhardt et al. · 2016 [cited by applicant]
US 20170153914A1 · Rausch et al. · 2017 [cited by applicant]
US 20180005074A1 · Shacham et al. · 2018 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180218257A1 · Xu · 2018 [cited by examiner]
US 20180322606A1 · Das et al. · 2018 [cited by applicant]
US 20180336458A1 · Tomioka · 2018 [cited by examiner]
CN 102708404B · 2016 [cited by applicant]
CN 106297774A · 2017 [cited by applicant]
CN 107463990A · 2017 [cited by applicant]
DE 10353532A1 · 2005 [cited by applicant]
EP 3506095A2 · 2019 [cited by applicant]
WO 2015192812A1 · 2015 [cited by applicant]
Afsahi, A. (2016, May 27). Topology-aware rank reordering for MPI collectives. IEEE Xplore. https://ieeexplore.ieee.org/abstract/ document/7530080 (Year: 2016). [cited by applicant]
Alex et al., “Meet Horovod: Uber's Open Source Distributed Deep Learning Framework for TensorFlow”, Oct. 17, 2017, 12 pages. [cited by applicant]
Awan, A. (2017, Jul. 28). Optimized broadcast for deep learning workloads on Dense-GPU Infini Band clusters: MPI or NCCL? arXiv.org. https://arxiv.org/abs/1707 .09414 (Year: 2017). [cited by applicant]
Chun et al., “Dolphin: Runtime Optimization for Distributed Machine Leaming”, 2016, 6 pages. [cited by applicant]
Chun, B. (Jun. 22, 2016). Dolphin: Runtime optimization for distributed machine learning. Markus Weimer. https://www .markusweimer.com/publication/2016/06/22/Dophin/ (Year: 2016). [cited by applicant]
Das, R. (Feb. 27, 2013). Application-to-core mapping policies to reduce memory system interference in multi-core systems. IEEE Xplore. https://ieeexplore.ieee.org/document/6522311 (Year: 2013). [cited by applicant]
Delimitrou, C. (Aug. 2015). Christina Delimitrou. Computer Systems Laboratory—Cornell University. https://www.csl.cornell.edu/-delimitrou/Publications.html (Year: 2015). [cited by applicant]
Extended European Search Report for EP Application No. 18209320.3, Aug. 22, 2019, 14 pages. [cited by applicant]
Fairhurst, G. (Oct. 1, 2001). Non-Return to Zero (NRZ) Encoding. https://erg.abdn.ac.uk/users/gorry/eg3567/phy-pages/nrz.html (Year: 2001). [cited by applicant]
Gibiansky, A. (Feb. 21, 2017). Bringing HPC techniques to deep learning—Andrew Gibiansky. Andrew Gibiansky's Blog. https:// andrew.gibiansky.com/blog/machine-learning/baidu-allreduce/ (Year: 2017). [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Intel. (Apr. 2015). Improving Real-Time Performance by Utilizing Cache Allocation Technology. Intel I Data Center Solutions, IoT, and PC Innovation. https://www.intel.com/content/dam/www/public/us/en/documents/white-pap… [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/859,180 mailed May 28, 2021, 25 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/869,551 mailed Jun. 4, 2021, 12 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/859,180 mailed Oct. 25, 2021, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/859,180 mailed Oct. 5, 2021, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/869,551 mailed Nov. 10, 2021, 8 pages. [cited by applicant]
Notification of CN Publication for CN Application No. 201811549383.2, Pub No. CN110135575A, 5 pages, Aug. 29, 2019. [cited by applicant]
Parashar et al., “SCNN: Accelerator for Compressed Convolution Neural Networks”, ⋅12 pages, May 23, 2017. [cited by applicant]
Partial European Search Report for EP Application No. 18209320.3, May 22, 2019, 15 pages. [cited by applicant]
Rajchl Martin et al., DeepCut: Object Segmentation From Bounding Box Annotations Using Convolution Neural Networks, Nov. 2016, IEEE, 36(2): 674-683. (Year: 2016). [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Sergeev, A. (Oct. 17, 2017). Meet Horovod: Ube r's open source distributed deep learning framework for TensorFlow. Uber Engineering Blog. https://eng.uber.com/horovod/ (Year: 2017). [cited by applicant]
Shafik, R. (Jun. 2016). Learning transfer-based adaptive energy minimization in embedded systems. IEEE Xplore. https:// ieeexplore.ieee.org/document/7308001 (Year: 2016). [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
STMicroelectronics. (2016). ST25 NFC guide. https://www.st.com/resource/en/technical_note/dm00190233-st25- nfc-guide-stmicroelectronics.pdf (Year: 2016). [cited by applicant]
TN1216 Technical Note, ST25 Nfc Guide, 38 pages, Oct. 2016. [cited by applicant]
Zhang, H. (Jul. 12, 2017). Poseidon I Proceedings of the 2017 USENIX conference on Usenix annual technical conference. ACM Digital Library. https://dl.acm.org/doi/10.5555/3154690.3154708 (Year: 2017). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/685,462 mailed Mar. 3, 2023, 9 pages. [cited by applicant]
Cabezas et al, Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications, IEEE Transactions on Parallel and Distributed Systems 26(5): 1405-1418. (Year: 2015). [cited by applicant]
Intention to Grant for EP Application No. 18209320.3 Oct. 8, 2024, 5 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/849,968 Sep. 30, 2024, 9 pages. [cited by applicant]
Decision to Grant EP Application No. 18209320.3 mailed Feb. 20, 2025. [cited by applicant]