IP Library Granted Patent US 12,333,417
Granted Patent B2
US 12,333,417 · App. 18/464,935 · Granted Jun 17, 2025

Rotating data for neural network computations

Inventors: Jonathan Ross (Mountain View, CA); Gregory Michael Thorson (Waunakee, WI)
Assignee: Google LLC
G06N3/063G06F15/8046G06N3/045G06N3/08G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,417
App. No.
18/464,935
Granted
Jun 17, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for computing a layer output for a convolutional neural network layer, the method comprising: receiving a plurality of activation inputs; forming a plurality of vector inputs from the plurality of activation inputs, each vector input comprising values from a distinct region within the multi-dimensional matrix; sending the plurality of vector inputs to one or more cells along a first dimension of the systolic array; generating a plurality of rotated kernel structures from each of the plurality of kernel; sending each kernel structure and each rotated kernel structure to one or more cells along a second dimension of the systolic array; causing the systolic array to generate an accumulated output based on the plurality of value inputs and the plurality of kernels; and generating the layer output from the accumulated output.

Claims (62)

1. A method for computing neural network inferences using a special-purpose processor comprising a hardware integrated circuit configured to implement the neural network, the method comprising:

receiving, at the special-purpose processor:

i) a kernel weight matrix comprising weights for a layer of the neural network and

ii) activations provided as inputs to the layer;

loading the kernel weight matrix into a matrix unit of the special-purpose processor by shifting the weights for the layer along a first dimension of the matrix unit;

loading the activations into the matrix unit by shifting activations representing inputs to the layer along a second dimension of the matrix unit;

computing, at the matrix unit, accumulated values from convolutions executed on the weights and activations shifted along dimensions of the matrix unit; and

generating an output for the layer based on the accumulated values.

2. The method of claim 1 , wherein:

i) the matrix unit comprises a two-dimensional (“2D”) array of compute cells; and

ii) the first and second dimensions of the matrix unit are respective dimensions of the 2D array of compute cells.

3. The method of claim 2 , wherein loading the kernel weight matrix into the matrix unit comprises:

rotating, in hardware at the matrix unit, the kernel weight matrix across compute cells along the first dimension of the 2D array of compute cells.

4. The method of claim 3 , wherein:

i) the first dimension is a column dimension of the 2D array of compute cells; and

ii) the second dimension is a row dimension of the 2D array of compute cells.

5. The method of claim 3 , further comprising:

controlling a directional dataflow of rotated data elements to the 2D array of compute cells; and

computing dot products over a plurality of clock cycles using the rotated data elements.

6. The method of claim 5 , further comprising:

controlling a directional dataflow of data elements that are shifted to the 2D array of compute cells along the first dimension of the matrix unit; and

computing dot products over the plurality of clock cycles using the shifted data elements.

7. The method of claim 6 , further comprising:

generating a set of accumulated values from the dot products computed over the plurality of clock cycles using the rotated data elements or the shifted data elements.

8. The method of claim 1 , wherein the kernel weight matrix is an n×n weight matrix representing an input data structure processed at the special-purpose processor, and the method further comprises:

rotating the n×n weight matrix by shifting weights corresponding to elements of the matrix along a particular dimension of the n×n weight matrix, wherein n is an integer greater than one.

9. The method of claim 1 , wherein:

i) the 2D array of compute cells is an N×N array of multiply-accumulate units; and

ii) N is an integer greater than one.

10. The method of claim 2 , wherein the weights for the layer are shifted along the first dimension of the matrix unit by a shift module coupled to the matrix unit.

11. A system for computing neural network inferences, the system comprising:

a special-purpose processor comprising a hardware integrated circuit configured to implement the neural network; and

a non-transitory machine-readable storage device for storing instructions that are executable by the special-purpose processor to cause performance of operations comprising:

receiving, at the special-purpose processor:

i) a kernel weight matrix comprising weights for a layer of the neural network and

ii) activations provided as inputs to the layer;

loading the kernel weight matrix into a matrix unit of the special-purpose processor by shifting the weights for the layer along a first dimension of the matrix unit;

loading the activations into the matrix unit by shifting activations representing inputs to the layer along a second dimension of the matrix unit;

computing, at the matrix unit, accumulated values from convolutions executed on the weights and activations shifted along dimensions of the matrix unit; and

generating an output for the layer based on the accumulated values.

12. The system of claim 11 , wherein:

i) the matrix unit comprises a two-dimensional (“2D”) array of compute cells; and

ii) the first and second dimensions of the matrix unit are respective dimensions of the 2D array of compute cells.

13. The system of claim 12 , wherein loading the kernel weight matrix into the matrix unit comprises:

rotating, in hardware at the matrix unit, the kernel weight matrix across compute cells along the first dimension of the 2D array of compute cells.

14. The system of claim 13 , wherein:

i) the first dimension is a column dimension of the 2D array of compute cells; and

ii) the second dimension is a row dimension of the 2D array of compute cells.

15. The system of claim 13 , wherein the operations further comprise:

controlling a directional dataflow of rotated data elements to the 2D array of compute cells; and

computing dot products over a plurality of clock cycles using the rotated data elements.

16. The system of claim 15 , wherein the operations further comprise:

controlling a directional dataflow of data elements that are shifted to the 2D array of compute cells along the first dimension of the matrix unit; and

computing dot products over the plurality of clock cycles using the shifted data elements.

17. The system of claim 16 , wherein the operations further comprise:

generating a set of accumulated values from the dot products computed over the plurality of clock cycles using the rotated data elements or the shifted data elements.

18. The system of claim 11 , wherein the kernel weight matrix is an n×n weight matrix representing an input data structure processed at the special-purpose processor, and the operations further comprise:

rotating the n×n weight matrix by shifting weights corresponding to elements of the matrix along a particular dimension of the n×n weight matrix, wherein n is an integer greater than one.

19. The system of claim 12 , wherein:

i) the 2D array of compute cells is an N×N array of multiply-accumulate units; and

ii) N is an integer greater than one.

20. The system of claim 11 , wherein the weights for the layer are shifted along the first dimension of the matrix unit by a shift module coupled to the matrix unit.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: ROSS, JONATHAN; THORSON, GREGORY MICHAEL
To: GOOGLE INC.
Reel/Frame 064926/0893 →
CHANGE OF NAME Recorded Sep 15, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 064927/0117 →
Continuity (6)
Continuation 17520919 · Nov 8, 2021
Continuation 16857808 · Apr 24, 2020
Continuation 15792872 · Oct 25, 2017
Continuation 14845022 · Sep 3, 2015
Provisional Application 62164998 · May 21, 2015
Related Publication 20240185047A1 · Jun 6, 2024
References Cited (82)
US 5014235A · Morton · 1991 [cited by applicant]
US 5136717A · Morley et al. · 1992 [cited by applicant]
US 5138695A · Means et al. · 1992 [cited by applicant]
US 5146543A · Vassiliadis et al. · 1992 [cited by applicant]
US 5337395A · Vassiliadis et al. · 1994 [cited by applicant]
US 5471627A · Means et al. · 1995 [cited by applicant]
US 5544336A · Kato et al. · 1996 [cited by applicant]
US 5799134A · Chiueh et al. · 1998 [cited by applicant]
US 5812993A · Ginosar et al. · 1998 [cited by applicant]
US 6038337A · Lawrence et al. · 2000 [cited by applicant]
US 6184753B1 · Ishimi et al. · 2001 [cited by applicant]
US 7136710B1 · Hoffberg et al. · 2006 [cited by applicant]
US 8184696B1 · Chirila-Rus et al. · 2012 [cited by applicant]
US 8468109B2 · Moussa et al. · 2013 [cited by applicant]
US 8924455B1 · Barman et al. · 2014 [cited by applicant]
US 20050044053A1 · Moreno et al. · 2005 [cited by applicant]
US 20070022063A1 · Lightowler · 2007 [cited by applicant]
US 20070086655A1 · Simard et al. · 2007 [cited by applicant]
US 20080319933A1 · Moussa et al. · 2008 [cited by applicant]
US 20110029471A1 · Chakradhar et al. · 2011 [cited by applicant]
US 20140142929A1 · Seide et al. · 2014 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337262A1 · Kato et al. · 2014 [cited by applicant]
US 20160267111A1 · Shoaib et al. · 2016 [cited by applicant]
US 20170103310A1 · Henry et al. · 2017 [cited by applicant]
CN 104035751 · 2014 [cited by applicant]
EP 0422348 · 1991 [cited by applicant]
EP 3064130 · 2016 [cited by applicant]
TW 201232429 · 2012 [cited by applicant]
TW 201331855 · 2013 [cited by applicant]
Dhabi, “2013 IEEE 20th International Conference on Electronics, Circuits, and Systems (Icecs)”, IEEE, Dec. 8, 2013, pp. 221-224. [cited by applicant]
Beamer et al., “Ivy Bridge Server Graph Processing Bottlenecks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 56 pages. [cited by applicant]
Bo et al., “String Kernel Testing Acceleration Using Micron's Automata Processor,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 21 pages. [cited by applicant]
Carlo et al., “An Area-Efficient 2-D Convolution Implementation on FPGA for Space Applications,” IEEE Computer Society, Dec. 11, 2011, 7 pages. [cited by applicant]
Chen et al., “Hardware Acceleration for Neuromorphic Computing - An Evolving View,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 38 pages. [cited by applicant]
Chillet et al., “A Neural Network Model for Real-Time Scheduling on Heterogeneous SoC Architectures,” Proceedings of International Joint Conference on Neural Networks, Aug. 2007, pp. 102-107. [cited by applicant]
Cornu et al., “Design, Implementation, and Test of a Multi-Model Systolic Neural-Network Accelerator,” Scientific Programming—Parallel Computing Projects of the Swiss Priority Programme, Jan. 1, 1996, 5(1):47-61. [cited by applicant]
Dawwd, “The multi 2D systolic design and implementation of Convolutional Neural Networks,” IEEE, Dec. 2013, pp. 221-224. [cited by applicant]
Dielman et al., “Rotation-invariant convolutional neural networks for galaxy morphology prediction,” Monthly notices of the royal astronomical society, 2015, pp. 1441-1459. [cited by applicant]
Farabet et al., “Hardware Accelerated Convolutional Neural Networks for Synthetic Vision Systems,” Circuits and Systems (ISCAS), Proceedings of 2010 IEEE International Symposium on, May-Jun. 2010, pp. 257-260. [cited by applicant]
Ginosar, “Accelerators for Machine Learning of Big Data,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 13 pages. [cited by applicant]
Gokhale, “Enabling Machines to Understand our World,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 18 pages. [cited by applicant]
Graf et al., “A Massively Parallel Digital Learning Processor,” Proceedings of the 22nd annual conference on Neural Information Processing Systems (NIPS), Dec. 2008, 8 pages. [cited by applicant]
Hecht et al., “An advanced programmable 2D-convolution chip for, real time image processing,” Signal Image and Video Processing, Jun. 11, 1991, pp. 1897-1900. [cited by applicant]
Henriques et al. “Exploiting the circulant structure of tracking-by-detection with kernels,” European conference on computer vision, 2012, 14 pages. [cited by applicant]
Indiveri, “Neuromorphic circuits for building autonomous cognitive systems,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 37 pages. [cited by applicant]
International Preliminary Report on Patentability in International Application No. PCT/US2016/030536, mailed on Nov. 30, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029294, mailed Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029965, mailed Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029968, mailed Sep. 1, 2016, 14 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029986, mailed Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030515, mailed Aug. 25, 2016, 19 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030536, mailed Aug. 31, 2016, 17 pages. [cited by applicant]
Kane, “An instruction systolic array architecture for multiple neural network types,” Loughborough University, Doctoral Thesis, Sep. 1998, 315 pages. [cited by applicant]
Khan et al., “Systolic architectures for artificial neural nets,” Neural Networks, Nov. 1991, pp. 620-627. [cited by applicant]
Kim et al. “Efficient Hardware Architecture for Sparse Coding,” IEEE Transactions on Signal Processing 62.16, Aug. 15, 2014, 14 pages. [cited by applicant]
Kim et al., “A Large-Scale Architecture for Restricted Boltzmann Machines,” Field-Programmable Custom Computing Machines (FCCM), 2010 18th IEEE Annual International Symposium on, IEEE, May 2, 2010, pp. 201-208. [cited by applicant]
Krizhevsky et al., “ImageNet classification with deep convolutional neural networks,” The 26th annual conference on Neural Information Processing Systems (NIPS'25), Dec. 2012, 9 pages. [cited by applicant]
Kung et al., “Two-level pipelined systolic array for multidimensional convolution,” Image and Vision Computing, Feb. 2, 1983, 1(1):30-36. [cited by applicant]
Kung, “VLSI Array Processors,” IEEE ASSP Magazine, Jul. 1, 1985, 2(3):4-22. [cited by applicant]
Lee et al., “Implementation of the Super-Systolic Array for Convolution,” Design Automation Conference, Jan. 2003, pp. 491-494. [cited by applicant]
Lee et al., “Nonlinear image processing by a rotating kernel transformation,” Optics letters 15.23, 1990, pp. 1383-1385. [cited by applicant]
Lehmann et al., “A generic systolic array building block for neural networks with on-chip learning,” Neural Networks, IEEE Transactions on, May 1993, 4(3):400-407. [cited by applicant]
Lipasti et al., Mimicking the Self-Organizing Properties of the Visual Cortex, The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 23 pages. [cited by applicant]
Lo et al. “Artificial convolutional neural network for medical image pattern recognition,” Neural networks 8.7, 1995, pp. 1201-1214. [cited by applicant]
Mahapatra et al., “Mapping of Neural Network Models onto Systolic Arrays,” Journal of Parallel and Distributed Computing 60, Jan. 2000, 677-689. [cited by applicant]
Merolla et al. “A digital Neurosynaptic Core Using Embedded Crossbar Memory with 45pJ per Spike in 45nm,” IEEE Cicc, Sep. 19, 2011, 4 pages. [cited by applicant]
Office Action in Taiwanese Application No. 105115859, mailed on Nov. 16, 2016, 10 pages. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware in the Datacenter,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 33 pages. [cited by applicant]
Patil et al., “Hardware Architecture for Large Parallel Array of Random Feature Extractors applied to Image Recognition,” arXiv, Dec. 24, 2015, 18 pages. [cited by applicant]
Pearce, “You Have No (Predictive) Power Here, SPEC!,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 15 pages. [cited by applicant]
Rojas, “Hardware for Neural Networks,” Neural Networks, 1996, pp. 451-478. [cited by applicant]
Shaaban, “Systolic Architectures,” PowerPoint Presentation, Mar. 2003, 9 pages. [cited by applicant]
Shapri et al., “Performance Analysis of Two-Dimensional Systolic Array Matrix Multiplication with Orthogonal Interconnections,” International Journal on New Computer Architectures and Their Applications (IJNCAA), Dec. 2… [cited by applicant]
Shapri et al., “Performance Analysis of Two-Dimensional Systolic Array Matrix Multiplication with Orthogonal Interconnections,” International Journal on New Computer Architectures and Their Applications, 2001, 1(3):1090… [cited by applicant]
Smith, “Biologically Plausible Spiking Neural Networks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 77 pages. [cited by applicant]
Sudha et al., “Systolic array realization of a neural network-based face recognition system,” Industrial Electronics and Applications, ICIEA 2008, 3rd IEEE Conference on, pp. 1864-1869. [cited by applicant]
TW Office Action in Taiwanese Application No. 107115823, dated Jan. 16, 2020, 4 pages (translation only). [cited by applicant]
Wong et al., “A New Scalable Systolic Array Processor Architecture for Discrete Convolution,” College of Engineering at the University of Kentucky, Master Thesis, 2003, 175 pages. [cited by applicant]
Wu et al., “Flip-Rotate-Pooling Convolution and Split Dropout on Convolution Neural Networks for Image Classification,” arXiv, Jul. 31, 2015, 9 pages. [cited by applicant]
Yiping et al. “A High Performance Digital Neural Processor Design by Network on Chip Architecture,” IEEE VLSI Design, Automation and Test, Apr. 25, 2011, 4 pages. [cited by applicant]