IP Library › Granted Patent US 12,314,218
Granted Patent B2
US 12,314,218 · App. 17/367,879 · Granted May 27, 2025

Task synchronization for accelerated deep learning

Inventors: Sean Lie (Los Gatos, CA); Michael Morrison (Sunnyvale, CA); Srikanth Arekapudi (Los Altos Hills, CA); Michael Edwin James (San Carlos, CA); Gary R. Lauterbach (Los Altos, CA)
Assignee: Cerebras Systems Inc.
G06F15/825G06F9/30036G06F9/3005G06F9/3016G06F9/30192G06F9/323G06F9/324G06F9/3836G06F9/3851G06F9/3887G06F9/45533G06F9/4881G06F9/52G06F13/00G06F17/16G06N3/04G06N3/045G06N3/063G06N3/08G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,218
App. No.
17/367,879
Granted
May 27, 2025
Kind
B2
Abstract

Techniques in advanced deep learning provide improvements in one or more of accuracy, performance, and energy efficiency. An array of processing elements performs flow-based computations on wavelets of data. Each processing element has a compute element and a routing element. Each compute element has memory. Each router enables communication via wavelets with at least nearest neighbors in a 2D mesh. Routing is controlled by respective virtual channel specifiers in each wavelet and routing configuration information in each router. A compute element conditionally selects for task initiation a previously received wavelet specifying a particular one of the virtual channels. The conditional selecting excludes the previously received wavelet for selection until at least block/unblock state maintained for the particular virtual channel is in an unblock state. The compute element executes block/unblock instructions to modify the block/unblock state.

Claims (85)

1. A method comprising:

sending a first fabric packet by a sending processing element, the first fabric packet comprising a virtual channel specifier specifying one of a plurality of virtual channels;

routing the first fabric packet from the sending processing element to a receiving processing element via one or more of routing elements in accordance with the virtual channel specifier;

in the receiving processing element, picking for processing the first fabric packet, the picking in accordance with a respective block/unblock state comprised in the receiving processing element and maintained for each of the plurality of virtual channels, the picking excluding the first fabric packet from being picked until at least the respective block/unblock state maintained for the specified virtual channel is in the unblock state;

wherein the sending processing element implements at least a portion of a first node of a plurality of nodes of a dataflow graph and the receiving processing element implements at least a portion of a second node of the dataflow graph, and the sending processing element and the receiving processing element are respective instances of a plurality of processing elements interconnected as a fabric; and

wherein the first fabric packet is an instance of a plurality of fabric packets and the routing of the first fabric packet is an instance of routing between the processing elements of the fabric, and wherein the routing between the processing elements of the fabric comprises each processing element transmitting each of one or more of the plurality of fabric packets over a selected group of preconfigured groups of one or more physical couplings between connected neighbors of the processing element, and the virtual channel specifier of each transmitted fabric packet is used to select the group.

2. The method of claim 1 , further comprising:

receiving in the receiving processing element a second fabric packet, the second fabric packet enabling execution of a selected one of a block instruction and an unblock instruction, the selected instruction having been stored in the receiving processing element prior to the receiving and comprising an immediate source operand specifying one or more of the virtual channels;

decoding the selected instruction in the receiving processing element; and

in the receiving processing element, setting the respective block/unblock state maintained for the specified one or more of the virtual channels in accordance with the decoding.

3. The method of claim 1 , further comprising:

receiving in the receiving processing element a second fabric packet of the fabric packets, the second fabric packet enabling execution of a selected one of a block instruction and an unblock instruction, the selected instruction having been stored in the receiving processing element prior to the receiving and comprising other than an immediate source operand;

decoding the selected instruction in the receiving processing element; and

in the receiving processing element, setting the respective block/unblock state maintained for all the virtual channels in accordance with the decoding.

4. The method of claim 1 , further comprising, in the receiving processing element, setting the respective block/unblock state maintained for a particular virtual channel of the virtual channels to a blocked state in response to a block instruction specifying the particular virtual channel of the virtual channels.

5. The method of claim 1 , further comprising, in the receiving processing element, setting the respective block/unblock state maintained for a particular virtual channel of the virtual channels to an unblocked state in response to an unblock instruction specifying the particular virtual channel of the virtual channels.

6. The method of claim 1 , wherein the sending processing element and the receiving processing element are fabricated via wafer-scale integration on separate die of a single wafer.

7. The method of claim 1 , wherein the sending processing element implements at least a portion of a first neuron of a neural network and the receiving processing element implements at least a portion of a second neuron of the neural network.

8. The method of claim 1 , wherein the sending processing element implements at least a portion of a first layer of a neural network and the receiving processing element implements at least a portion of a second layer of the neural network.

9. The method of claim 1 , wherein the sending processing element and the receiving processing element implement respective portions of at least a partitioned neuron of a neural network.

10. The method of claim 1 , further comprising, with respect to the receiving processing element, managing fabric packet input queues to have generally equal average rates of production and consumption by stalling/resuming task activities via manipulation of the respective block/unblock state.

11. The method of claim 1 , further comprising, with respect to the receiving processing element, managing one or more priorities within and between tasks by stalling/resuming task activities via manipulation of the respective block/unblock state.

12. The method of claim 1 , further comprising, with respect to the receiving processing element, managing one or more dependencies within and between tasks by stalling/resuming task activities via manipulation of the respective block/unblock state.

13. The method of claim 1 , further comprising, with respect to the receiving processing element, synchronizing one or more of computations and communications of one or more tasks, via manipulation of the respective block/unblock state.

14. The method of claim 1 , further comprising, with respect to the receiving processing element, implementing task software interlocks via manipulation of the respective block/unblock state.

15. The method of claim 1 , further comprising synchronizing data sourced via unequal delay paths by manipulation of the respective block/unblock state of one or more of the processing elements along the delay paths.

16. The method of claim 1 , further comprising shaping at least some dataflow in at least part of the fabric by manipulation of the respective block/unblock state of one or more of the processing elements.

17. The method of claim 1 , further comprising:

wherein the one of the plurality of virtual channels is used for communicating at least one of control and data associated with one or more of: computing an activation of a neural network, computing a partial sum of activations of the neural network, computing an error of the neural network, computing a gradient estimate of the neural network, and updating a weight of the neural network; and

wherein the first fabric packet comprises the at least one of control and data associated with the one or more of: computing the activation of the neural network, computing the partial sum of activations of the neural network, computing the error of the neural network, computing the gradient estimate of the neural network, and updating the weight of the neural network.

18. The method of claim 1 , wherein the sending processing element, the one or more of the routing elements, and the receiving processing element are fabricated via wafer-scale integration.

19. The method of claim 1 , further comprising:

initializing the fabric with all parameters and task software required for concurrent execution of communications and computations respectively corresponding to the dataflow graph; and

concurrently executing all layers of the dataflow graph using all-digital techniques for one or more of inference and training.

20. The method of claim 19 , wherein except for defects, the fabric is homogeneous, the plurality of processing elements numbers in the hundreds of thousands, each processing element comprises in the tens of kB of private local storage for instructions and data, and the parameters and task software comprises in the tens of GB of instruction and data storage for the one or more of inference and training.

21. The method of claim 19 , wherein the concurrently executing does not require any access to storage external to the fabric for any intermediate state or additional parameters of the dataflow graph.

22. The method of claim 19 , wherein the dataflow graph is a neural network, the plurality of nodes corresponds to neurons, and at least some of parameters of the dataflow graph correspond to a plurality of weights of the neural network.

23. The method of claim 1 , further comprising:

wherein the processing comprises selective dataflow-based processing and selective instruction-based processing;

wherein the selective dataflow-based processing comprises the routing and the picking, each in accordance with the virtual channel specifier; and

wherein the selective instruction-based processing is in accordance with a starting address for fetching instructions executable by the processing elements, the starting address identified at least in part by a selected one of the virtual channel specifier and a task specifier of the first fabric packet.

24. An apparatus comprising:

means for sending a first fabric packet by a sending processing element, the first fabric packet comprising a virtual channel specifier specifying one of a plurality of virtual channels;

means for routing the first fabric packet from the sending processing element to a receiving processing element via one or more of routing elements in accordance with the virtual channel specifier;

in the receiving processing element, means for picking for processing the first fabric packet, the picking in accordance with a respective block/unblock state comprised in the receiving processing element and maintained for each of the plurality of virtual channels, the picking excluding the first fabric packet from being picked until at least the respective block/unblock state maintained for the specified virtual channel is in the unblock state;

wherein the sending processing element implements at least a portion of a first node of a plurality of nodes of a dataflow graph and the receiving processing element implements at least a portion of a second node of the dataflow graph, and the sending processing element and the receiving processing element are respective instances of a plurality of processing elements interconnected as a fabric; and

wherein the first fabric packet is an instance of a plurality of fabric packets and the means for routing of the first fabric packet is a part of means for routing between the processing elements of the fabric, and wherein the means for routing between the processing elements of the fabric comprises means for each processing element transmitting each of one or more of the plurality of fabric packets over a selected group of preconfigured groups of one or more physical couplings between connected neighbors of the processing element, and the virtual channel specifier of each transmitted fabric packet is used to select the group.

25. The apparatus of claim 24 , further comprising:

means for receiving in the receiving processing element a second fabric packet, the second fabric packet enabling execution of a selected one of a block instruction and an unblock instruction, the selected instruction having been stored in the receiving processing element prior to the receiving and comprising an immediate source operand specifying one or more of the virtual channels;

means for decoding the selected instruction in the receiving processing element; and

in the receiving processing element, means for setting the respective block/unblock state maintained for the specified one or more of the virtual channels in accordance with the decoding.

26. The apparatus of claim 24 , further comprising:

means for receiving in the receiving processing element a second fabric packet of the fabric packets, the second fabric packet enabling execution of a selected one of a block instruction and an unblock instruction, the selected instruction having been stored in the receiving processing element prior to the receiving and comprising other than an immediate source operand;

means for decoding the selected instruction in the receiving processing element; and

in the receiving processing element, means for setting the respective block/unblock state maintained for all the virtual channels in accordance with the decoding.

27. The apparatus of claim 24 , further comprising, in the receiving processing element, means for setting the respective block/unblock state maintained for a particular virtual channel of the virtual channels to a blocked state in response to a block instruction specifying the particular virtual channel of the virtual channels.

28. The apparatus of claim 24 , further comprising, in the receiving processing element, means for setting the respective block/unblock state maintained for a particular virtual channel to an unblocked state in response to an unblock instruction specifying the particular virtual channel of the virtual channels.

29. The apparatus of claim 24 , wherein the sending processing element and the receiving processing element are fabricated via wafer-scale integration on separate die of a single wafer.

30. The apparatus of claim 24 , wherein the sending processing element implements at least a portion of a first neuron of a neural network and the receiving processing element implements at least a portion of a second neuron of a neural network.

31. The apparatus of claim 24 , wherein the sending processing element implements at least a portion of a first layer of a neural network and the receiving processing element implements at least a portion of a second layer of a neural network.

32. The apparatus of claim 24 , wherein the sending processing element and the receiving processing element implement respective portions of at least a partitioned neuron of a neural network.

33. A method comprising:

specifying communications and computations respectively corresponding to a plurality of nodes of a dataflow graph, wherein a sending processing element implements at least a portion of a first node of the dataflow graph and a receiving processing element implements at least a portion of a second node of the dataflow graph, and the sending processing element and the receiving processing element are respective instances of a plurality of processing elements interconnected as a fabric;

initializing the fabric with all parameters and task software required for a concurrent execution of the communications and computations respectively corresponding to the dataflow graph;

sending a first fabric packet by the sending processing element, the first fabric packet comprising a virtual channel specifier specifying one of a plurality of virtual channels;

routing the first fabric packet from the sending processing element to the receiving processing element via one or more of routing elements in accordance with the virtual channel specifier;

in the receiving processing element, the virtual channel specifier of the first fabric packet identifying a respective one of a plurality of queues of the receiving processing element, the respective queue holding at least a task specifier of the first fabric packet; and

in the receiving processing element, picking for processing the first fabric packet, the picking in accordance with a respective block/unblock state maintained for each of the plurality of queues of the receiving processing element, the picking excluding the first fabric packet from being picked until at least the respective block/unblock state maintained for the respective queue is in the unblock state.

34. The method of claim 33 , further comprising, in the receiving processing element, setting the respective block/unblock state maintained for a particular virtual channel of the virtual channels to a blocked state in response to a block instruction specifying the particular virtual channel of the virtual channels.

35. The method of claim 34 , wherein the block instruction is an instance of the task software stored in the receiving processing element by the initializing.

36. The method of claim 34 , wherein the block instruction is associated with a previously stored fabric packet received via the fabric.

37. The method of claim 33 , further comprising, in the receiving processing element, setting the respective block/unblock state maintained for a particular virtual channel to an unblocked state in response to an unblock instruction specifying the particular virtual channel of the virtual channels.

38. The method of claim 37 , wherein the unblock instruction is an instance of the task software stored in the receiving processing element by the initializing.

39. The method of claim 37 , wherein the unblock instruction is associated with a previously stored fabric packet received via the fabric.

40. The method of claim 33 , further comprising, with respect to the receiving processing element, managing one or more dependencies within and between tasks by stalling/resuming task activities via manipulation of the respective block/unblock state.

41. The method of claim 33 , further comprising, with respect to the receiving processing element, synchronizing one or more of computations and communications of one or more tasks, via manipulation of the respective block/unblock state.

42. The method of claim 33 , further comprising, with respect to the receiving processing element, implementing task software interlocks via manipulation of the respective block/unblock state.

43. The method of claim 33 , further comprising synchronizing data sourced via unequal delay paths by manipulation of the respective block/unblock state of one or more of the processing elements along the delay paths.

44. The method of claim 33 , further comprising shaping at least some dataflow in at least part of the fabric by manipulation of the respective block/unblock state of one or more of the processing elements.

45. The method of claim 33 , further comprising:

wherein the dataflow graph is a neural network for one or more of inference and training, the plurality of nodes corresponds to neurons, and at least some of parameters of the dataflow graph correspond to a plurality of weights of the neural network; and

concurrently executing all layers of the dataflow graph using all-digital techniques, the concurrently executing comprising the sending, routing, identifying, and picking.

46. The method of claim 33 , wherein the sending processing element implements at least a portion of a first neuron of a neural network and the receiving processing element implements at least a portion of a second neuron of the neural network.

47. The method of claim 33 , wherein the sending processing element implements at least a portion of a first layer of a neural network and the receiving processing element implements at least a portion of a second layer of the neural network.

48. The method of claim 33 , wherein the sending processing element and the receiving processing element implement respective portions of at least a partitioned neuron of a neural network.

Assignments (2)
SECURITY INTEREST Recorded Jun 18, 2026
From: CEREBRAS SYSTEMS INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS THE COLLATERAL AGENT
Reel/Frame 075845/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2024
From: LIE, SEAN; MORRISON, MICHAEL; AREKAPUDI, SRIKANTH; JAMES, MICHAEL EDWIN; LAUTERBACH, GARY R.
To: CEREBRAS SYSTEMS INC.
Reel/Frame 067935/0388 →
Continuity (14)
Continuation 16603647
Provisional Application 62486372 · Apr 17, 2017
Provisional Application 62517949 · Jun 11, 2017
Provisional Application 62520433 · Jun 15, 2017
Provisional Application 62522065 · Jun 19, 2017
Provisional Application 62522081 · Jun 19, 2017
Provisional Application 62542645 · Aug 8, 2017
Provisional Application 62542657 · Aug 8, 2017
Provisional Application 62580207 · Nov 1, 2017
Provisional Application 62628773 · Feb 9, 2018
Provisional Application 62628784 · Feb 9, 2018
Provisional Application 62652933 · Apr 5, 2018
Provisional Application 62655210 · Apr 9, 2018
Related Publication 20220172031A1 · Jun 2, 2022
References Cited (84)
US 5361334A · Cawley · 1994 [cited by applicant]
US 5761103A · Oakland et al. · 1998 [cited by applicant]
US 6167051A · Nagami · 2000 [cited by examiner]
US 6185213B1 · Katsube · 2001 [cited by examiner]
US 6199089B1 · Mansingh · 2001 [cited by applicant]
US 6212627B1 · Dulong et al. · 2001 [cited by applicant]
US 6542986B1 · White · 2003 [cited by applicant]
US 7075937B1 · El Kolli · 2006 [cited by examiner]
US 7293002B2 · Starzyk · 2007 [cited by applicant]
US 7359274B2 · Noguchi et al. · 2008 [cited by applicant]
US 8667439B1 · Kumar · 2014 [cited by examiner]
US 9529400B1 · Kumar · 2016 [cited by examiner]
US 9627095B1 · Xu et al. · 2017 [cited by applicant]
US 10268679B2 · Li et al. · 2019 [cited by applicant]
US 10289816B1 · Malassenet et al. · 2019 [cited by applicant]
US 10614357B2 · Lie · 2020 [cited by examiner]
US 10762420B2 · Teig et al. · 2020 [cited by applicant]
US 11062200B2 · Lie · 2021 [cited by examiner]
US 20070019552A1 · Senarath · 2007 [cited by examiner]
US 20070140240A1 · Dally et al. · 2007 [cited by applicant]
US 20080031157A1 · Tamboise et al. · 2008 [cited by applicant]
US 20080107105A1 · Reilly · 2008 [cited by examiner]
US 20080222646A1 · Sigal et al. · 2008 [cited by applicant]
US 20090135739A1 · Hoover et al. · 2009 [cited by applicant]
US 20090313195A1 · Mcdaid · 2009 [cited by examiner]
US 20100272110A1 · Allan · 2010 [cited by examiner]
US 20110282926A1 · Amemiya · 2011 [cited by examiner]
US 20120084533A1 · Sperber et al. · 2012 [cited by applicant]
US 20130073497A1 · Akopyan et al. · 2013 [cited by applicant]
US 20130073498A1 · Izhikevich et al. · 2013 [cited by applicant]
US 20130322459A1 · Xu · 2013 [cited by applicant]
US 20150242741A1 · Campos et al. · 2015 [cited by applicant]
US 20150302295A1 · Rivera et al. · 2015 [cited by applicant]
US 20150324684A1 · Alvarez-Icaza Rivera et al. · 2015 [cited by applicant]
US 20150324690A1 · Chilimbi et al. · 2015 [cited by applicant]
US 20160019061A1 · Chatha et al. · 2016 [cited by applicant]
US 20160182398A1 · Davis et al. · 2016 [cited by applicant]
US 20160239647A1 · Johnson et al. · 2016 [cited by applicant]
US 20160246337A1 · Colgan · 2016 [cited by examiner]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170295061A1 · Wittenschlaeger · 2017 [cited by applicant]
US 20170316116A1 · Elliott · 2017 [cited by applicant]
US 20180046894A1 · Yao · 2018 [cited by applicant]
US 20180132055A1 · Leedy · 2018 [cited by applicant]
US 20180174041A1 · Imam et al. · 2018 [cited by applicant]
US 20180189652A1 · Korthikanti et al. · 2018 [cited by applicant]
US 20190042377A1 · Teig et al. · 2019 [cited by applicant]
US 20190130250A1 · Park et al. · 2019 [cited by applicant]
US 20190132928A1 · Rodinger et al. · 2019 [cited by applicant]
US 20190138423A1 · Agerstam et al. · 2019 [cited by applicant]
US 20190244058A1 · Franca-Neto et al. · 2019 [cited by applicant]
US 20190244933A1 · Or-Bach et al. · 2019 [cited by applicant]
US 20190260504A1 · Philip · 2019 [cited by examiner]
US 20190303743A1 · Venkataramani et al. · 2019 [cited by applicant]
US 20190324759A1 · Yang et al. · 2019 [cited by applicant]
US 20190340064A1 · Sity et al. · 2019 [cited by applicant]
US 20190341091A1 · Sity et al. · 2019 [cited by applicant]
US 20190347555A1 · Park et al. · 2019 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210072894A1 · Chawla et al. · 2021 [cited by applicant]
US 20210092069A1 · Musleh · 2021 [cited by examiner]
US 20210142155A1 · Lie et al. · 2021 [cited by applicant]
US 20210142167A1 · Lie et al. · 2021 [cited by applicant]
US 20210166109A1 · Lie et al. · 2021 [cited by applicant]
US 20210224639A1 · Lie et al. · 2021 [cited by applicant]
US 20210248453A1 · Lauterbach et al. · 2021 [cited by applicant]
US 20210255860A1 · Morrison et al. · 2021 [cited by applicant]
JP H025173A · 2009 [cited by applicant]
JP 2009129447A · 2009 [cited by applicant]
WO 9716792A1 · 1997 [cited by applicant]
WO 2016186813A1 · 2016 [cited by applicant]
WO 2022034542A1 · 2022 [cited by applicant]
Feb. 24, 2022 , List of References Used in Art Rejections in Cases Related to Docket No. CS-17-06US-1, Feb. 24, 2022, 3 pages. [cited by applicant]
European search report, European Application No. 18756971.0, Aug. 19, 2021, 8 pages. [cited by applicant]
Japanese Notice of Reasons for Refusal Application No. 2019-556711, Oct. 5, 2021, 7 pages. [cited by applicant]
International Search Report in the related case PCT/IB2021/057456, Nov. 19, 2021, 4 pages. [cited by applicant]
Written Opinion of the International Searching Authority in the related case PCT/IB2021/057456, Nov. 19, 2021, 4 pages. [cited by applicant]
International Preliminary Report On Patentability (Ch II) in PCT/IB2020/060231, Jan. 27, 2022, 5 pages. [cited by applicant]
International Preliminary Report on Patentability (Ch II) in PCT/IB2020/060232 Jan. 28, 2022, 5 pages. [cited by applicant]
European search report, European Application No. 18756971.0, Feb. 18, 2022, 7 pages. [cited by applicant]
Carlos Zamarreno-Ramos et al: “Multicasting Mesh AER: A Scalable Assembly Approach for Reconfigurable Neuromorphic Structured AER Systems. Application to ConvNets”, IEEE Transactions on Biomedical Circuits and Systems, … [cited by applicant]
International Preliminary Report on Patentability (Ch II) in PCT/IB2020/060188, Jan. 26, 2022, 5 pages. [cited by applicant]
Narayanamurthy N et al.: “Evolving bio plausible design with heterogeneous Noc”, The 15th International Conference on Advanced Communications Technology-ICACT2013, Jan. 27, 2013, (pp. 451-456), 6 pages. [cited by applicant]
Shen, X-W., et al., “An Efficient Network-on-Chip Router for Dataflow Architecture”, Journal of Computer Science and Technology, vol. 32, Jan. 2017, pp. 11-25. [cited by applicant]