IP Library › Granted Patent US 12,277,444
Granted Patent B2
US 12,277,444 · App. 17/993,564 · Granted Apr 15, 2025

Software-defined tensor streaming multiprocessor for large-scale machine learning

Inventors: Dennis Charles Abts (Eau Claire, WI); Jonathan Ross (Palo Alto, CA); Garrin Kimmell (Mountain View, CA); Michael Bye (Chippewa Falls, WI); Matthew Boyd (Gresham, OR); Andrew Ling (Toronto, CA)
Assignee: Groq, Inc.
G06F9/4881G06F9/5072G06F15/163G06F15/7867
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,444
App. No.
17/993,564
Granted
Apr 15, 2025
Kind
B2
Abstract

A system contains a network of processors arranged in a plurality of nodes. Each node comprises a respective plurality of processors connected via local links, and different nodes are connected via global links. The processors of the network communicate with each other to establish a global counter for the network, enabling deterministic communication between the processors of the network. A compiler is configured to explicitly schedule communication traffic across the global and local links of the network of processors based upon the deterministic links between the processors, which enable software-scheduled networking with explicit send or receive instructions executed by functional units of the processors at specific times, to establish a specific ordering of operations performed by the network of processors. In some embodiments, the processors of the network of processors are tensor streaming processors (TSPs).

Claims (33)

1. A system comprising:

a network of processors, comprising:

a plurality of processors arranged in a plurality of nodes, comprising at least a first node having a first plurality of deterministic processors connected by local links;

wherein the plurality of nodes are connected by global links; and wherein the plurality of processors communicate with each other to establish a global counter for the network, enabling deterministic communication between the plurality of processors of the network; and

a compiler to explicitly schedule communication traffic across the global and local links with explicit send or receive instructions executed at specific times to establish a specific ordering of operations performed by the network of processors.

2. The system of claim 1 , wherein each processor of the network of processors is a deterministic processor.

3. The system of claim 2 , wherein each deterministic processor comprises one or more deterministic data paths connecting rows of deterministic functional units, each functional unit configured to process data received from, or output processed data onto, at least one data path, in accordance with a plurality of instructions which are executed at a known clock cycle.

4. The system of claim 1 , wherein each of the plurality of processors comprises n local links and m global links, and wherein the first plurality of processors comprises n+1 processors, wherein each of the n+1 processors is connected via a respective local link to each other processor of the first node, and wherein each local link comprises a synchronous link.

5. The system of claim 4 , wherein the plurality of nodes comprises up to ((n+1)*m)+1 nodes, and wherein each of the up to ((n+1)*m)+1 nodes is connected to each remaining node of the up to ((n+1)*m)+1 nodes via a respective global link and wherein each global link comprises a synchronous link.

6. The system of claim 1 , wherein the plurality of nodes are grouped into racks each having a set of global links, each rack having a set of nodes connected to each other via a first portion of the set of global links, and is connected to each remaining rack of the plurality of racks via a second portion of the set of global links.

7. The system of claim 6 , wherein the set of nodes within each rack are doubly connected to each other via the first portion of the set of global links, and wherein the racks are singly connected via the second portion of the set of global links.

8. The system of claim 6 , wherein the set of nodes of each rack includes a hot spare node, and wherein, responsive to detection of a critical error, the system is configured to perform a runtime reply of an inference during which the critical error was detected, where computation of a node of a rack associated with the detected critical error is switched to the hot spare node.

9. The system of claim 1 , wherein each processor of the plurality of processors implements a first counter, wherein the first counter of each of the plurality of processors is synchronized with those of other processors of the plurality of processors to establish the global counter for the network.

10. The system of claim 9 , wherein synchronizing the first counter of a first processor with the first counter of a second processor of the plurality of processors comprises:

transmitting, at a first time, a first value of the first counter of the first processor corresponding to the first time to the second processor, wherein the second processor is configured to, upon receiving the first value of the first counter from the first processor, reflect the first value back to the first processor;

determining a latency value of a link connecting the first processor and the second processor, based upon a difference between the first value of the first counter with a second value of the first counter of the first processor corresponding to a second time at which the reflected first value is received by the first processor;

transmitting, at a third time, a third value of the first counter of the first processor corresponding to the third time; and

responsive to receiving the third value at the second processor, adjusting a value of the first counter of the second processor based upon the third value and the determined latency value.

11. The system of claim 10 , wherein the first processor is configured to transmit a value of its first counter to the second processor periodically.

12. The system of claim 9 , wherein a first processor and a second processor of the plurality of processors are configured to perform an initial program alignment operation, comprising:

placing the second processor into a synchronization loop, wherein the second processor checks for an instruction transmitted from the first processor at each overflow boundary of the first counter of the second processor;

at a first time corresponding to an overflow boundary of the first counter of the first processor, transmitting an instruction from the first processor to the second processor;

causing the second processor to exit the synchronization loop and begin synchronized computation at an overflow boundary of the first counter of the second processor occurring after the instruction from the first processor is received; and

causing the first processor to begin synchronized computation at an overflow boundary of the first counter of the first processor aligned with the overflow boundary of the first counter of the second processor.

13. The system of claim 9 , wherein each processor of the plurality of processors further implements a second counter, and is configured to periodically delay for a target number of cycles based upon a difference between the first counter and the second counter.

14. The system of claim 1 , wherein the plurality of processors are configured to operate based upon a compiled program, wherein the compiled program includes an explicit schedule of instructions causing the plurality of processors to communicate with each other with a predetermined timing.

15. The system of claim 14 , wherein the compiled program schedules a transmission of data between a first processor and a second processor of the plurality of processors by scheduling a first instruction at the first processor configured to cause a functional unit of the first processor to read data onto a data path of the first processor at a first time, and a second instruction at the second processor configured to cause a functional unit of the second processor to consume data on a data path of the second processor at a second time.

16. The system of claim 15 , wherein the transmission occurs without the second processor sending a communication to the first processor to request the transmission of data.

17. The system of claim 14 , wherein the compiled program is generated by a compiler based on a model, and wherein the compiler schedules each instruction of the schedule of instructions to prevent routing deadlock on the local and global links due to transmission of data between the processors of the network, based on a topology of the network and determined latency values of the local and global links.

18. The system of claim 17 , where the compiler is configured to divide the compiled program into a plurality of sub-programs, and assigns each sub-program to a respective processor of the plurality of processors.

19. The system of claim 18 , wherein the compiler is configured to select the respective processors of the plurality of processors by identifying a pinch point within the model corresponding to a portion of the model with a reduced volume of data communication, and assigning sub-programs corresponding to different sides of pinch point to respective processors of different nodes of the plurality of nodes.

20. The system of claim 17 , wherein the compiler is configured to schedule a transmission of data between a first processor and a second processor of the network by dividing the data into a plurality of data portions, and scheduling instructions to cause the first processor to transmit each of the plurality of data portions via a different non-minimal path to the second processor.

21. The system of claim 20 , wherein each non-minimal path comprises at least a first link connecting the first processor to one or more intermediate processors, and a second link connecting the one or more intermediate processors to the second processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: ABTS, DENNIS CHARLES; ROSS, JONATHAN; KIMMELL, GARRIN; BYE, MICHAEL; BOYD, MATTHEW; LING, ANDREW
To: GROQ, INC.
Reel/Frame 062091/0315 →
Continuity (2)
Provisional Application 63283094 · Nov 24, 2021
Related Publication 20230161621A1 · May 25, 2023
References Cited (43)
US 11474557B2 · Thorson et al. · 2022 [cited by applicant]
US 20190121779A1 · Osborne · 2019 [cited by examiner]
US 20210326173A1 · Shah · 2021 [cited by examiner]
US 20210342673A1 · Shah · 2021 [cited by examiner]
US 20220075633A1 · Desai · 2022 [cited by examiner]
US 20220357984A1 · Kotler · 2022 [cited by examiner]
US 20230116614A1 · Shi · 2023 [cited by examiner]
Hadi Jooybar, Wilson W. L. Fung, Mike O'Connor, Joseph Devietti, Tor M. Aamodt, “GPUDet: a deterministic GPU architecture”, AC , pp. 1-12 (Year: 2013). [cited by examiner]
Abts, D. et al. “Age-based packet arbitration in large-radix k-ary n-cubes,” ACM/IEEE conference on Supercomputing, Association for Computing Machinery, Article No. 5, Nov. 2007, pp. 1-11. [cited by applicant]
Abts, D. et al. “Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads,” ACM/IEEE 47th Annual International Symposium on Computer Architecture, May 30-Jun. 3, 2020, pp. 145-158. [cited by applicant]
Ahn, J.H. et al. “HyperX: topology, routing, and packaging of efficient large-scale networks,” Conference on High Performance Computing Networking, Article No. 41, Nov. 2009, pp. 1-11. [cited by applicant]
Alverson, B. et al. “Cray XC Series Network,” Cray Inc., White Paper WP-Aries01-1112, 2012, pp. 1-28. [cited by applicant]
Brown, T. et al. “Language Models are Few-Shot Learners,” Advances in Neural Information Processing Systems, No. 33, 2020, pp. 1-24. [cited by applicant]
Cerebras.net “Cerebras CS-1,” Apr. 12, 2021, six pages, Retrieved from the Internet Archive URL: <https://web.archive.org/web/20210412190256/https://cerebras.net/>. [cited by applicant]
Dally, W.J. “Virtual-Channel Flow Control,” IEEE Transactions on Parallel and Distributed Systems, vol. 3, No. 2, Mar. 1992, pp. 194-205. [cited by applicant]
Dally, W.J. et al. “Principles and Practices of Interconnection Networks,” Morgan Kaufmann Inc, 2004. [cited by applicant]
Flajslik, M. et al. “Megafly: A Topology for Exascale Systems,” International Conference on High Performance Computing, Jun. 24, 2018, pp. 289-310. [cited by applicant]
Glass, C. et al. “The Turn Model for Adaptive Routing,” 25 years of the International Symposia on Computer architecture, Aug. 1, 1998, pp. 441-450. [cited by applicant]
Haidar, A. et al. “High-performance Cholesky factorization for GPU-only execution,” General Purpose GPUs, Feb. 4-5, 2017, pp. 42-52. [cited by applicant]
Herault, T. et al. “Generic Matrix Multiplication for Multi-GPU Accelerated Distributed-Memory Platforms over PaRSEC,” IEEE/ACM 10th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems, Nov. 18, 2… [cited by applicant]
Huang, Y. et al. “Pipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism,” Advances in Neural Information Processing Systems, No. 32, 2019, pp. 1-10. [cited by applicant]
Jiang, N. et al. “Indirect Adaptive Routing on Large Scale Interconnection Networks,” 36th Annual International Symposium on Computer Architecture, Jun. 20, 2009, pp. 220-231. [cited by applicant]
Karpathy, A. “Keynote,” Conference on Computer Vision and Pattern Recognition, Workshop on Autonomous Driving, Jun. 20, 2021, Retrieved from Youtube <URL: https://www.youtube.com/watch?v=g6bOwQdCJrc>. [cited by applicant]
Kermani, P. et al. “Virtual cut-through: A new computer communication switching technique,” Computer Networks (1976), vol. 3, No. 4, Sep. 1979, pp. 267-286. [cited by applicant]
Kim, J. et al. “Flattened Butterfly Topology for On-Chip Networks,” 40th Annual IEEE/ACM International Symposium on Microarchitecture, Dec. 1-7, 2007, pp. 172-182. [cited by applicant]
Kim, J. et al. “Microarchitecture of a High-Radix Router,” 32nd International Symposium on Computer Architecture, 2005, pp. 420-431. [cited by applicant]
Kim, J. et al. “Technology-Driven, Highly-Scalable Dragonfly Topology,” 2008 International Symposium on Computer Architecture, Jun. 21-25, 2008, pp. 77-88. [cited by applicant]
Lee, M. et al. “Probabilistic Distance-Based Arbitration: Providing Equality of Service for Many-Core CMPs,” 43rd Annual IEEE/ACM International Symposium on Microarchitecture, Dec. 4-8, 2010, pp. 509-519. [cited by applicant]
McKeown, N. et al. “OpenFlow: Enabling Innovation in Campus Networks,” ACM SIGCOMM Computer Communication Review, vol. 38, No. 2, Mar. 31, 2008, pp. 69-74. [cited by applicant]
Medina, E. “Habana Labs Approach to Scaling AI Training,” Hot Chips 31 Symposium, Aug. 18-20, 2019, pp. 1-29. [cited by applicant]
Narayanan, D. et al. “Efficient Large-Scale Language Model Training on GPU Clusters,” International Conference for High Performance Computing, Networking, Storage and Analysis, Article No. 58, Nov. 13, 2021, pp. 1-15. [cited by applicant]
Narayanan, D. et al. “PipeDream: Generalized Pipeline Parallelism for DNN Training,” 27th ACM Symposium on Operating Systems Principles, Oct. 27, 2019, pp. 1-15. [cited by applicant]
NVIDIA “Matrix Multiplication Background,” User's Guide, Nov. 2021, pp. 1-17. [cited by applicant]
NVIDIA “NVIDIA Dgx BasePOD,” Date Unknown, eight pages, [Online] [Retrieved on Dec. 2, 2022] Retrieved from the Internet URL: < https://www.nvidia.com/en-us/data-center/dgx-basepod/>. [cited by applicant]
Prabhakar, R. et al. “Plasticine: A Reconfigurable Architecture for Parallel Patterns,” ACM/IEEE 44th Annual International Symposium on Computer Architecture, Jun. 24-28, 2017, pp. 389-402. [cited by applicant]
Scott, S. “Rosetta: A 64-port Switch for Cray's Slingshot Interconnect,” IEEE Symposium on High-Performance Interconnects, Aug. 16, 2019, [Abstract Only]. [cited by applicant]
Scott, S. et al. “The Blackwidow High-Radix Clos Network,” ACM SIGARCH Computer Architecture News, vol. 34, No. 2, May 1, 2006, pp. 16-28. [cited by applicant]
Shoeybi, M. et al. “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism,” Sep. 17, 2019, pp. 1-15. [cited by applicant]
Shpiner, A. et al. “Dragonfly+: Low Cost Topology for Scaling Datacenters,” IEEE 3rd International Workshop on High-Performance Interconnection Networks in the Exascale and Big-Data Era, Feb. 5, 2017, pp. 1-8. [cited by applicant]
U.S. Appl. No. 17/203,214, filed Mar. 16, 2021, Inventor Dennis Charles Abts and Jonathan Alexander Ross (copy not enclosed). [cited by applicant]
Williams, S. et al. “Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures,” Communications of the ACM, vol. 52, No. 4, Apr. 1, 2009, pp. 65-76. [cited by applicant]
Won, J. et al. “Overcoming Far-End Congestion in Large-Scale Networks,” IEEE 21st International Symposium on High Performance Computer Architecture, Feb. 7-11, 2015, pp. 415-427. [cited by applicant]
Young, C. “Evaluation of the Tensor Processing Unit: A Deep Neural Network Accelerator for the Datacenter,” Hot Chips 29 Symposium, May 3, 2017, Retrieved from Youtube <URL: https://www.youtube.com/watch?v=fhHAArxwzvQ>. [cited by applicant]