IP Library Granted Patent US 12,218,666
Granted Patent B1
US 12,218,666 · App. 18/133,413 · Granted Feb 4, 2025

Application specific integrated circuit accelerators

Inventors: Michial Allen Gunter (San Francisco, CA); Charles Henry Leichner, IV (Palo Alto, CA); Tammo Spalink (Mountain View, CA)
Assignee: Google LLC
H03K19/17744G06F15/8046G06F15/8053G06N3/04G06N3/063G06N3/082H03K19/017509H03K19/017545H03K19/017581H03K19/1774G06F2015/763
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,218,666
App. No.
18/133,413
Granted
Feb 4, 2025
Kind
B1
Abstract

An application specific integrated circuit (ASIC) chip includes: a systolic array of cells; and multiple controllable bus lines configured to convey data among the systolic array of cells, in which the systolic array of cells is arranged in multiple tiles, each tile of the multiple tiles including 1) a corresponding subarray of cells of the systolic array of cells, 2) a corresponding subset of controllable bus lines of the multiple controllable bus lines, and 3) memory coupled to the subarray of cells.

Claims (37)

1. An integrated circuit, comprising:

a systolic array of cells arranged in a plurality of subarrays of cells, wherein:

each cell comprises circuitry for performing an arithmetic computation;

the plurality of subarrays of cells is arranged as a grid extending along a first direction and a second direction;

a plurality of memory circuits, wherein each memory circuit of the plurality of memory circuits is arranged adjacent to a different respective subarray of cells of the plurality of subarrays of cells; and

a vector processing unit, wherein a first subset of the plurality of subarrays of cells is arranged in a first section on a first side of the vector processing unit, and a second subset of the plurality of subarrays of cells is arranged in a second section on a second side of the vector processing unit that is opposite to the first side of the vector processing unit,

the integrated circuit operable to:

transfer a first data set along the first direction to a first subarray of cells of the plurality of subarrays of cells;

transfer a first set of control instructions along the second direction to the first subarray of cells;

perform, by the first subarray of cells and in accordance with the first set of control instructions, the arithmetic computation based on the first data set to obtain a second data set; and

transfer the second data set along the first direction from the first subarray of cells.

2. The integrated circuit of claim 1 , wherein transferring the second data set along the first direction from the first subarray of cells comprises transferring the second data set to a second subarray of cells of the plurality of subarrays of cells.

3. The integrated circuit of claim 1 , wherein the first data set is transferred to the first subarray of cells from a second subarray of cells of the plurality of subarrays of cells.

4. The integrated circuit of claim 3 , wherein the first data set comprises a first partial-sum data set obtained from the second subarray of cells, and wherein the second data set comprises a second partial-sum data set obtained from the first subarray of cells.

5. The integrated circuit of claim 1 , wherein transferring the second data set along the first direction from the first subarray of cells comprises transferring the second data set to a second subarray of cells of the plurality of subarrays of cells, the integrated circuit further operable to:

transfer a second set of control instructions along the second direction to the second subarray of cells;

perform, by the second subarray of cells and in accordance with the second set of control instructions, the arithmetic computation on the second data set to obtain a third data set;

transfer the third data set to the vector processing unit.

6. The integrated circuit of claim 1 , wherein transferring the first data set along the first direction to the first subarray of cells comprises transferring the first data set on at least one controllable bus line.

7. The integrated circuit of claim 6 , wherein the at least one controllable bus line comprises a switch.

8. The integrated circuit of claim 7 , wherein the switch comprises a flip-flop circuit element or a multiplexer.

9. The integrated circuit of claim 6 , wherein transferring the first data set on at least one controllable bus line comprises transferring the first data set according to a predetermined schedule.

10. The integrated circuit of claim 1 , wherein transferring the first set of control instructions along the second direction to the first subarray of cells comprises transferring the first data set on at least one controllable bus line.

11. The integrated circuit of claim 10 , wherein the at least one controllable bus line comprises a switch.

12. The integrated circuit of claim 11 , wherein the switch comprises a flip-flop circuit element or a multiplexer.

13. The integrated circuit of claim 12 , wherein transferring the first set of control instructions on at least one controllable bus line comprises transferring the first set of control instructions according to a predetermined schedule.

14. The integrated circuit of claim 1 , comprising transferring a second data set to the first subarray of cells from a first memory circuit arranged adjacent to the first subarray of cells.

15. The integrated circuit of claim 14 , wherein performing the arithmetic computation is further based on the second data set and on a third data set.

16. The integrated circuit of claim 15 , wherein the first data set comprises a partial-sum data set from a second subarray of cells of the plurality of subarrays of cells, the second data set comprises a plurality of weight inputs, and the third data set comprises a plurality of activation inputs, and wherein the arithmetic computation comprises a multiply and accumulate operation.

17. The integrated circuit of claim 1 , wherein transferring the second data set along the first direction from the first subarray of cells comprises transferring the second data set to a second subarray of cells of the plurality of subarrays of cells, the integrated circuit further operable to:

transfer a second set of control instructions along the second direction to the second subarray of cells;

perform, by the second subarray of cells and in accordance with the second set of control instructions, the arithmetic computation on the second data set to obtain a third data set;

transfer the third data set to the vector processing unit; and

perform a non-linear computation on the third data set in the vector processing unit to provide a vector computation output.

18. The integrated circuit of claim 17 , further comprising transferring the vector computation output to a communication interface coupled to the systolic array of cells.

19. The integrated circuit of claim 1 , comprising storing the first set of control instructions in a first memory circuit arranged adjacent to the first subarray of cells.

20. The integrated circuit of claim 19 , wherein transferring the first set of control instructions along the second direction to the first subarray of cells comprises transferring the first set of control instructions from the first memory circuit to the first subarray of cells.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 069880/0046 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2023
From: GUNTER, MICHIAL ALLEN; LEICHNER IV, CHARLES HENRY; SPALINK, TAMMO
To: X DEVELOPMENT LLC
Reel/Frame 063330/0872 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 063331/0083 →
Continuity (4)
Continuation 17397465 · Aug 9, 2021
Continuation 16827409 · Mar 23, 2020
Continuation 16042839 · Jul 23, 2018
Provisional Application 62535652 · Jul 21, 2017
References Cited (36)
US 5260881A · Agrawal et al. · 1993 [cited by applicant]
US 5455525A · Ho et al. · 1995 [cited by applicant]
US 5469003A · Kean · 1995 [cited by applicant]
US 5581199A · Pierce et al. · 1996 [cited by applicant]
US 5883525A · Tavana et al. · 1999 [cited by applicant]
US 5892962A · Cloutier · 1999 [cited by examiner]
US 6092174A · Roussakov · 2000 [cited by applicant]
US 6255849B1 · Mohan · 2001 [cited by applicant]
US 6362650B1 · New et al. · 2002 [cited by applicant]
US 6396303B1 · Young · 2002 [cited by applicant]
US 6943581B1 · Cruz et al. · 2005 [cited by applicant]
US 7613900B2 · Gonzalez et al. · 2009 [cited by applicant]
US 7765382B2 · Chester · 2010 [cited by applicant]
US 8018849B1 · Wentzlaff · 2011 [cited by applicant]
US 8860460B1 · Cashman · 2014 [cited by applicant]
US 8924455B1 · Barman et al. · 2014 [cited by applicant]
US 9323525B2 · Kim et al. · 2016 [cited by applicant]
US 9490811B2 · Ngai · 2016 [cited by applicant]
US 10175980B2 · Temam et al. · 2019 [cited by applicant]
US 10879904B1 · Gunter et al. · 2020 [cited by applicant]
US 11451229B1 · Gunter et al. · 2022 [cited by applicant]
US 11652484B1 · Gunter · 2023 [cited by examiner]
US 20080074142A1 · Henderson · 2008 [cited by applicant]
US 20090128571A1 · Smith et al. · 2009 [cited by applicant]
US 20100077374A1 · Qiu · 2010 [cited by applicant]
US 20160154717A1 · Alvarez-Icaza Rivera et al. · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170206939A1 · Verma · 2017 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180121377A1 · Woo et al. · 2018 [cited by applicant]
US 20180157465A1 · Bittner et al. · 2018 [cited by applicant]
US 20180189231A1 · Fleming, Jr. · 2018 [cited by examiner]
US 20180336165A1 · Phelps · 2018 [cited by examiner]
US 20190303743A1 · Venkataramani · 2019 [cited by applicant]
US 20230010315A1 · Gunter et al. · 2023 [cited by applicant]
Chen et al. “DaDianNao: A Machine-Learning Supercomputer,” Proceedings of the 47th Annual IEEE/ACM International Symposium on MicroArchitecture, IEEE Computer Society, Dec. 2014, 14 pages. [cited by applicant]