IP Library Granted Patent US 12,530,579
Granted Patent B2
US 12,530,579 · App. 17/874,573 · Granted Jan 20, 2026

Systolic array processor for neural network computation

Inventors: Jonathan Ross (Mountain View, CA); Norman Paul Jouppi (Palo Alto, CA); Andrew Everett Phelps (Middleton, WI); Reginald Clifford Young (Palo Alto, CA); Thomas Norrie (San Jose, CA); Gregory Michael Thorson (Waunakee, WI); Dan Luu (Madison, WI)
Assignee: Google LLC
G06N3/08G06F15/8046G06N3/063G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,579
App. No.
17/874,573
Granted
Jan 20, 2026
Kind
B2
Abstract

A circuit for performing neural network computations for a neural network comprising a plurality of neural network layers, the circuit comprising: a matrix computation unit configured to, for each of the plurality of neural network layers: receive a plurality of weight inputs and a plurality of activation inputs for the neural network layer, and generate a plurality of accumulated values based on the plurality of weight inputs and the plurality of activation inputs; and a vector computation unit communicatively coupled to the matrix computation unit and configured to, for each of the plurality of neural network layers: apply an activation function to each accumulated value generated by the matrix computation unit to generate a plurality of activated values for the neural network layer.

Claims (45)

1 . A system comprising one or more processors, wherein the one or more processors comprise a matrix computation unit, the matrix computation unit comprising a multi-dimensional systolic array comprising a first group of cells and a second group of cells, the second group of cells being adjacent to the first group of cells, the matrix computation unit configured to:

receive, by the first group of cells and along a first dimension of the matrix computation unit, weight inputs and activation inputs;

perform, by the first group of cells, a portion of neural network computations for a neural network layer of a neural network, using at least the weight inputs and the activation inputs;

process, by the first group of cells, an accumulated output using the received weight inputs and activation inputs;

provide the accumulated output across a second dimension to the second group of cells; and

generate a final accumulated output using at least the accumulated output across the first and second group of cells.

2 . The system of claim 1 , wherein the neural network computations comprise multiplication of a respective weight input of the weight inputs with a respective activation input of the activation inputs.

3 . The system of claim 1 , wherein the accumulated output is stored in a respective accumulator unit of a plurality of accumulator units, the respective accumulator unit being connected to the first group of cells.

4 . The system of claim 1 , wherein:

the first group of cells receives the weight inputs along the first dimension and the activation inputs along the second dimension; and

either the first dimension is a row dimension and the second dimension is a column dimension, or the first dimension is a column dimension and the second dimension is a row dimension.

5 . The system of claim 1 , wherein the matrix computation unit is further configured to generate, using at least the first group of cells, the activation inputs.

6 . The system of claim 5 , wherein:

the neural network layer is a first neural network layer; and

to generate the activation inputs, the matrix computation unit is configured to at least partially perform neural network computations of a second neural network layer preceding the first neural network in the neural network.

7 . The system of claim 1 , wherein each cell further comprises a control register configured to store a control signal for determining whether to shift either a weight input or an activation input to an adjacent cell.

8 . The system of claim 7 , wherein the matrix computation unit is further configured to pass the control signal from a first cell to a second cell adjacent to the first cell in the systolic array.

9 . A method comprising:

receiving, by a first group of cells along a first dimension of a matrix computation unit, weight inputs and activation inputs;

performing, by the first group of cells, a portion of neural network computations for a neural network layer of a neural network, using at least the weight inputs and the activation inputs;

processing, by the first group of cells, an accumulated output using the received weight inputs and activation inputs;

providing the accumulated output across a second dimension to a second group of cells adjacent to the first group of cells; and

generating a final accumulated output using at least the accumulated output across the first and second group of cells.

10 . The method of claim 9 , wherein the neural network computations comprise multiplication of a respective weight input of the weight inputs with a respective activation input of the activation inputs.

11 . The method of claim 9 , wherein the accumulated output is stored in a respective accumulator unit of a plurality of accumulator units, the respective accumulator unit being connected to the first group of cells.

12 . The method of claim 9 , wherein:

the first group of cells receives the weight inputs along the first dimension and the activation inputs along the second dimension; and

either the first dimension is a row dimension and the second dimension is a column dimension, or the first dimension is a column dimension and the second dimension is a row dimension.

13 . The method of claim 9 , further comprising generating, by the matrix computation unit using at least the first group of cells, the activation inputs.

14 . The method of claim 13 , wherein:

the neural network layer is a first neural network layer; and

to generate the activation inputs, the matrix computation unit is configured to at least partially perform neural network computations of a second neural network layer preceding the first neural network in the neural network.

15 . The method of claim 9 , wherein each cell further comprises a control register configured to store a control signal for determining whether to shift either a weight input or an activation input to an adjacent cell.

16 . The method of claim 15 , further comprising passing, by the matrix computation unit, the control signal from a first cell to a second cell adjacent to the first cell in a systolic array of the matrix computation unit.

17 . One or more non-transitory computer readable storage media encoding instructions that when executed by one or more processors including a matrix computation unit, the matrix computation unit including a multi-dimensional systolic array including a first group of cells and a second group of cells, the second group of cells being adjacent to the first group of cells, cause the matrix computation unit to perform operations comprising:

receiving, by the first group of cells and along a first dimension of the matrix computation unit, weight inputs and activation inputs;

performing, by the first group of cells, a portion of neural network computations for a neural network layer of a neural network, using at least the weight inputs and the activation inputs;

processing, by the first group of cells, an accumulated output using the received weight inputs and activation inputs;

providing the accumulated output across a second dimension to the second group of cells; and

generating a final accumulated output using at least the accumulated output across the first and second group of cells.

18 . The non-transitory computer readable storage media of claim 17 , wherein the neural network computations comprise multiplication of a respective weight input of the weight inputs with a respective activation input of the activation inputs.

19 . The non-transitory computer readable storage media of claim 17 , wherein the accumulated output is stored in a respective accumulator unit of a plurality of accumulator units, the respective accumulator unit being connected to the first group of cells.

20 . The non-transitory computer readable storage media of claim 17 , wherein:

the first group of cells receives the weight inputs along the first dimension and the activation inputs along the second dimension; and

either the first dimension is a row dimension and the second dimension is a column dimension, or the first dimension is a column dimension and the second dimension is a row dimension.

Assignments (2)
CHANGE OF NAME Recorded Jul 28, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 060991/0502 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2022
From: ROSS, JONATHAN; JOUPPI, NORMAN PAUL; PHELPS, ANDREW EVERETT; YOUNG, REGINALD CLIFFORD; NORRIE, THOMAS; THORSON, GREGORY MICHAEL; LUU, DAN
To: GOOGLE INC.
Reel/Frame 060642/0156 →
Continuity (5)
Continuation 16915161 · Jun 29, 2020
Continuation 15686615 · Aug 25, 2017
Continuation 14844524 · Sep 3, 2015
Provisional Application 62164931 · May 21, 2015
Related Publication 20220366255A1 · Nov 17, 2022
References Cited (103)
US 5014235A · Morton · 1991 [cited by applicant]
US 5136717A · Morley et al. · 1992 [cited by applicant]
US 5138695A · Means · 1992 [cited by examiner]
US 5146543A · Vassiliadis et al. · 1992 [cited by applicant]
US 5212766A · Arima · 1993 [cited by examiner]
US 5337395A · Vassiliadis et al. · 1994 [cited by applicant]
US 5471627A · Means et al. · 1995 [cited by applicant]
US 5509106A · Pechanek et al. · 1996 [cited by applicant]
US 5544336A · Kato · 1996 [cited by applicant]
US 5799134A · Chiueh et al. · 1998 [cited by applicant]
US 5812993A · Ginosar et al. · 1998 [cited by applicant]
US 5892962A · Cloutier · 1999 [cited by applicant]
US 6038337A · Lawrence · 2000 [cited by applicant]
US 6184753B1 · Ishimi et al. · 2001 [cited by applicant]
US 7082419B1 · Lightowler · 2006 [cited by applicant]
US 7136710B1 · Hoffberg · 2006 [cited by applicant]
US 8184696B1 · Chirila-Rus · 2012 [cited by applicant]
US 8417758B1 · Rao et al. · 2013 [cited by applicant]
US 8468109B2 · Moussa et al. · 2013 [cited by applicant]
US 8924455B1 · Barman et al. · 2014 [cited by applicant]
US 10699188B2 · Ross et al. · 2020 [cited by applicant]
US 20020168100A1 · Woodall · 2002 [cited by applicant]
US 20050044053A1 · Moreno · 2005 [cited by applicant]
US 20070022063A1 · Lightowler · 2007 [cited by applicant]
US 20070086655A1 · Simard et al. · 2007 [cited by applicant]
US 20080319933A1 · Moussa · 2008 [cited by applicant]
US 20110029471A1 · Chakradhar et al. · 2011 [cited by applicant]
US 20140142929A1 · Seide et al. · 2014 [cited by applicant]
US 20140180984A1 · Arthur et al. · 2014 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337262A1 · Kato et al. · 2014 [cited by applicant]
US 20160267111A1 · Shoaib · 2016 [cited by applicant]
CN 1333518A · 2002 [cited by applicant]
CN 101681450A · 2010 [cited by applicant]
CN 104035751A · 2014 [cited by applicant]
CN 104238993A · 2014 [cited by applicant]
EP 0422348A2 · 1991 [cited by applicant]
EP 3064130A1 · 2016 [cited by applicant]
TW 201128542A · 2011 [cited by applicant]
TW 201232429A · 2012 [cited by applicant]
TW 201331855A · 2013 [cited by applicant]
TW 201421381A · 2014 [cited by applicant]
TW 201421382A · 2014 [cited by applicant]
TW 201435757A · 2014 [cited by applicant]
Kung et al., Systolic Arrays for (VLSI); Department of Computer Science Carnegie-Mellon University; Apr. 1978; Total pp. 31 ( Year: 1978). [cited by examiner]
Sudha et al., “Systolic array realization of a neural network-based face recognition system,” Industrial Electronics and Applications, 2008, ICIEA 2008, 3rd IEEE Conference on, pp. 1864-1869, Jun. 2009. [cited by applicant]
TW Office Action in Taiwanese Application No. 107113688, dated Jan. 14, 2020, 20 pages (with English translation). [cited by applicant]
Wong et al., “A New Scalable Systolic Array Processor Architecture for Discrete Convolution,” College of Engineering at the University of Kentucky, Master Thesis, 2003, 175 pages. [cited by applicant]
Wu et al., “Flip-Rotate-Pooling Convolution and Split Dropout on Convolution Neural Networks for Image Classification,” Jul. 31, 2015, arXiv:1507.08754v1 pp. 1-9, XP055296122. [cited by applicant]
Yiping et al. “A High Performance Digital Neural Processor Design by Network on Chip Architecture” IEEE VLSI Design, Automation and Test, Apr. 25, 2011, 4 pages. [cited by applicant]
Office Action for Taiwanese Patent Application No. 111106513 dated Nov. 11, 2022. 8 pages. [cited by applicant]
Office Action for Taiwanese Patent Application No. 111106513 dated May 5, 2023. 4 pages. [cited by applicant]
Ahm Shapri and N.A.Z Rahman. “Performance Analysis of Two-Dimensional Systolic Array Matrix Multiplication with Orthogonal Interconnections,” International Journal on New Computer Architectures and Their Applications, 1… [cited by applicant]
Beamer et al., “Ivy Bridge Server Graph Processing Bottlenecks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 56 pages. [cited by applicant]
Bermak and Bouzerdoum. VLSI Implementation of a Neural Network Classifier Based on the Saturating Linear Activation Function. 2002. Proceedings of the 9th International Conference on Neural Information Processing. ICONI… [cited by applicant]
Bo et al., “String Kernel Testing Acceleration Using Micron's Automata Processor,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 21 pages. [cited by applicant]
Carlo et al., “An Area-Efficient 2-D Convolution Implementation on FPGA for Space Applications,” IEEE Computer Society, Dec. 11, 2011, pp. 1-7. [cited by applicant]
Chen and Li, “Hardware Acceleration for Neuromorphic Computing—An Evolving View,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 38 pages. [cited by applicant]
Chillet et al., “A Neural Network Model for Real-Time Scheduling on Heterogeneous SoC Architectures, ” Proceedings of International Joint Conference on Neural Networks, Aug. 2007, pp. 102-107. [cited by applicant]
Comu et al., “Design, Implementation, and Test of a Multi-Model Systolic Neural-Network Accelerator,” Scientific Programming—Parallel Computing Projects of the Swiss Priority Programme, vol. 5, No. 1, Jan. 1, 1996, DD. … [cited by applicant]
Dawwd, “The multi 2D systolic design and implementation of Convolutional Neural Networks,” 2013 IEEE 20th International Conference on Electronics, Circuits, and Systems (ICECS), IEEE, Dec. 8, 2013, DD. 221-224, XP032595… [cited by applicant]
Dielman, Sander, Kyle W. Willett, and Joni Dambre. “Rotation-invariant convolutional neural networks for galaxy morphology prediction,” Monthly notices of the royal astronomical society, 450.2, 2015, Dages 1441-1459. [cited by applicant]
Farabet et al., “Hardware Accelerated Convolutional Neural Networks for Synthetic Vision Systems,” Circuits and Systems (ISCAS), Proceedings of 2010 IEEE International Symposium on, May-Jun. 2010, pp. 257-260. [cited by applicant]
Francesco Pazienti, “A Systolic Array for Neural Network Implementation”, 1991, IEEE, pp. 981-984. [cited by applicant]
Ginosar, “Accelerators for Machine Learning of Big Data,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 13 pages. [cited by applicant]
Gokhale, “Enabling Machines to Understand our World,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 18 pages. [cited by applicant]
Graf et al., “A Massively Parallel Digital Learning Processor,” Proceedings of the 22nd annual conference on Neural Information Processing Systems (NIPS), Dec. 2008, 8 pages, XP055016863. [cited by applicant]
Hecht et al., “An advanced programmable 2D-convolution chip for, real time image processing,” Signal Image and Video Processing, Jun. 1991; [Proceedings of the International Symposium on Circuits and Systems], vol. SYMP… [cited by applicant]
Indiveri, “Neuromoiphic circuits for building autonomous cognitive systems,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 3 7 pages. [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2016/029294, issued on Nov. 21, 2017, 7 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029294, mailed Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029965, mailed Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029968, mailed Sep. 1, 2016, 14 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029986, mailed Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030515, mailed Aug. 25, 2016, 19 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030536, mailed Aug. 31, 2016, 17 pages. [cited by applicant]
Kane. An Instruction Systolic Array Architecture for Multiple Neural Network Types. Sep. 1998. Loughborough University. Thesis. 317 pages. <https://www.semanticscholar.org/paper/An-instruction-systolicarray-architecture… [cited by applicant]
Khan and Ling, “Systolic architectures for artificial neural nets,” Neural Networks, 1991. 1991 IEEE International Joint Conference on, vol. 1, Nov. 1991, pp. 620-627. [cited by applicant]
Khan et al., Two-Dimensional Multirate Systolic Array Design for Artificial Neural Networks, 1991, IEEE, pp. 186-193. [cited by applicant]
Kim et al. “Efficient Hardware Architecture for Sparse Coding,” IEEE Transactions on Signal Processing 62.16, Augustst 15, 2014, 14 pages. [cited by applicant]
Kim et al., “A Large-Scale Architecture for Restricted Boltzmann Machines,” Field-Programmable Custom Computing Machines (FCCM), 2010 18th IEEE Annual International Symposium on, IEEE, May 2, 2010, pp. 201-208, XP031681… [cited by applicant]
Krizhevsky et al., “ImageNet classification with deep convolutional neural networks,” The 26th annual conference on Neural Information Processing Systems (NIPS '25), Dec. 2012, pp. 1-9, XP55113686. [cited by applicant]
Kung et al., “Two-level pipelined systolic array for multidimensional convolution,” Image and Vision Computing, Elsevier, vol. 1, No. 1, Feb. 2, 1983, DD. 30-36, XP024237511. [cited by applicant]
Kung, “VLSI Array Processors,” IEEE Assp Magazine, IEEE, vol. 2, No. 3, Jul. 1, 1985, pp. 4-22, XP011370547. [cited by applicant]
Lee and Song, “Implementation of the Super-Systolic Array for Convolution,” Design Automation Conference, 2003. Proceedings of the ASP-DAC 2003. Asia and South Pacific, Jan. 2003, pp. 491-494. [cited by applicant]
Lee, Yim-Kul, and William T. Rhodes. “Nonlinear image processing by a rotating kernel transformation,” Optics letters 15.23, 1990, pp. 1383-1385. [cited by applicant]
Lehmann et al., “A generic systolic array building block for neural networks with on-chip learning,” Neural Networks, IEEE Transactions on, 4(3):400-407, May 1993. [cited by applicant]
Lipasti et al., Mimicking the Self-Organizing Properties of the Visual Cortex, The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 23 pages. [cited by applicant]
Lo, Shih-Chung B., et al. “Artificial convolutional neural network for medical image pattern recognition,” Neural networks 8.7, 1995, pp. 1201-1214. [cited by applicant]
Mahapatra et al., “Mapping of Neural Network Models onto Systolic Arrays,” Journal of Parallel and Distributed Computing 60, 677-689, Jan. 2000. [cited by applicant]
Merolla et al. “A digital Neurosynaptic Core Using Embedded Crossbar Memory with 45pJ per Spike in 45mm,” IEEE CICC, Sep. 19, 2011, 4 pages. [cited by applicant]
Office Action for Taiwanese Patent Application No. 109143265 dated Jul. 8, 2021. 17 pages. [cited by applicant]
Office Action for Taiwanese Patent Application No. 111106513 dated Jun. 22, 2022. 9 pages. [cited by applicant]
Office Action in Taiwanese Application No. 105115859, mailed on Nov. 16, 2016, 10 pages. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware in the Datacenter,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 33 pages. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware,” Microsoft Research, [online] [retrieved Mar. 25, 2021]. Retrieved from the Internet: <URL:http:/www.microsoft.com/en-us/res… [cited by applicant]
Patil et al., “Hardware Architecture for Large Parallel Array of Random Feature Extractors applied to Image Recognition,” Dec. 24, 2015, arXiv:1512.07783vl, 18 pages, XP055296121. [cited by applicant]
Pearce, “You Have No (Predictive) Power Here, SPEC!” The FirstInternational Workshop Computer Architecture for Machine Learning, Jun. 2015, 15 pages. [cited by applicant]
Rojas, “Hardware for Neural Networks,” Neural Networks, Springer-Verlag, Berlin, 1996, pp. 451-478. [cited by applicant]
Shaaban, “Systolic Architectures,” PowerPoint Presentation, Mar. 2003, 9 pages. [cited by applicant]
Shapri and Rahman, “Performance Analysis of Two-Dimensional Systolic Array Matrix Multiplication with Orthogonal Interconnections,” International Journal on New Computer Architectures and Their Applications (IJNCAA) 1 (… [cited by applicant]
Smith, “Biologically Plausible Spiking Neural Networks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 77 pages. [cited by applicant]