IP Library › Granted Patent US 12,705,472
Granted Patent B2
US 12,705,472 · App. 18/386,037 · Granted Aug 11, 2026

Prefetching weights for use in a neural network processor

Inventor: Jonathan Ross (Mountain View, CA)
Assignee: Google LLC
G06N3/063G06F15/8046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,472
App. No.
18/386,037
Filed
Nov 1, 2023
Granted
Aug 11, 2026
Kind
B2
Art Unit
2124
USPC
706/16
Abstract

A circuit for performing neural network computations for a neural network, the circuit comprising: a systolic array comprising a plurality of cells; a weight fetcher unit configured to, for each of the plurality of neural network layers: send, for the neural network layer, a plurality of weight inputs to cells along a first dimension of the systolic array; and a plurality of weight sequencer units, each weight sequencer unit coupled to a distinct cell along the first dimension of the systolic array, the plurality of weight sequencer units configured to, for each of the plurality of neural network layers: shift, for the neural network layer, the plurality of weight inputs to cells along the second dimension of the systolic array over a plurality of clock cycles and where each cell is configured to compute a product of an activation input and a respective weight input using multiplication circuitry.

Claims (45)

1 . A system for performing neural network computations for a neural network having a plurality of neural network layers, the system comprising:

a matrix computation unit comprising circuitry configured to:

obtain a weight input for a neural network layer of the plurality of neural network layers;

receive a control signal; and

determine, based on the control signal, whether to reuse the weight input for a different neural network layer of the plurality of neural network layers at a subsequent clock cycle.

2 . The system of claim 1 , wherein the circuitry is further configured to, in response to a determination to reuse the weight input, shift the weight input.

3 . The system of claim 1 , wherein the circuitry is further configured to:

obtain an activation input for the neural network layer; and

determine, based on the control signal, whether to reuse the activation input at a subsequent clock cycle.

4 . The system of claim 3 , wherein obtaining the activation input comprises obtaining the activation input from a value loader.

5 . The system of claim 3 , wherein:

obtaining the weight input comprises obtaining a shifted weight input; and

obtaining the respective activation input comprises obtaining a shifted activation input.

6 . The system of claim 3 , further comprising:

a first memory configured to provide activation inputs for the plurality of neural network layers; and

a second memory configured to provide weight inputs for the plurality of neural network layers.

7 . The system of claim 6 , further comprising a vector computation unit comprising circuitry configured to:

receive one or more accumulated values from the matrix computation unit;

determine a vector based on the one or more accumulated values; and

provide the vector to the first memory.

8 . The system of claim 6 , further comprising sequencer circuitry configured to provide one or more control signals to at least one of the first memory, the second memory, the vector computation unit, or the matrix computation unit.

9 . The system of claim 1 , wherein obtaining the weight input comprises obtaining the weight input from a weight fetcher interface.

10 . The system of claim 1 , wherein determining, based on the control signal, whether to reuse the weight input comprises determining that the control signal meets a predetermined value.

11 . The system of claim 1 , wherein the circuitry comprises:

one or more weight control registers configured to store the control signal; and

one or more weight registers configured to load the weight input.

12 . The system of claim 1 , further comprising a sequencer comprising circuitry configured to provide the control signal to the matrix computation unit.

13 . The system of claim 12 , wherein the circuitry of the sequencer comprises decrement circuitry configured to decrement, at each clock cycle, a value of the control signal.

14 . A method for performing neural network computations for a neural network having a plurality of neural network layers, the method comprising:

obtaining, by a matrix computation unit, a weight input for a neural network layer of the plurality of neural network layers;

receiving, by the matrix computation unit, a control signal; and

determining, based on the control signal, whether to reuse the weight input for a different neural network layer of the plurality of neural network layers at a subsequent clock cycle.

15 . The method of claim 14 , further comprising, in response to determining to reuse the weight input, shifting the weight input.

16 . The method of claim 14 , further comprising:

obtaining, by the matrix computation unit, an activation input for the neural network layer; and

determining, based on the control signal, whether to reuse the activation input at a subsequent clock cycle.

17 . The method of claim 14 , wherein determining, based on the control signal, whether to reuse the weight input comprises determining that the control signal meets a predetermined value.

18 . A matrix computation unit for performing neural network computations for a neural network having a plurality of neural network layers, the matrix computation unit comprising circuitry configured to:

obtain a weight input for a neural network layer of the plurality of neural network layers;

receive a control signal; and

determine, based on the control signal, whether to reuse the weight input for a different neural network layer of the plurality of neural network layers at a subsequent clock cycle.

19 . The matrix computation unit of claim 18 , wherein the circuitry is further configured to, in response to a determination to reuse the weight input at the subsequent clock cycle, shift the respective weight input.

20 . The matrix computation unit of claim 18 , wherein the circuitry is further configured to:

obtain an activation input for the neural network layer; and

determine, based on the control signal, whether to reuse the activation input at a subsequent clock cycle.

Assignments (2)
CHANGE OF NAME Recorded Nov 2, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 065434/0481 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2023
From: ROSS, JONATHAN
To: GOOGLE INC.
Reel/Frame 065422/0240 →
Continuity (6)
Continuation 17134936 · Dec 28, 2020
Continuation 16826466 · Mar 23, 2020
Continuation 16053305 · Aug 2, 2018
Continuation 14844670 · Sep 3, 2015
Provisional Application 62164981 · May 21, 2015
Related Publication 20240062055A1 · Feb 22, 2024
References Cited (119)
US 5014235A · Morton · 1991 [cited by applicant]
US 5136717A · Morley et al. · 1992 [cited by applicant]
US 5138695A · Means · 1992 [cited by applicant]
US 5146543A · Vassiliadis et al. · 1992 [cited by applicant]
US 5337395A · Vassiliadis et al. · 1994 [cited by applicant]
US 5471627A · Means et al. · 1995 [cited by applicant]
US 5544336A · Kato · 1996 [cited by applicant]
US 5799134A · Chiueh et al. · 1998 [cited by applicant]
US 5812993A · Ginosar et al. · 1998 [cited by applicant]
US 6038337A · Lawrence · 2000 [cited by applicant]
US 6184753B1 · Ishimi et al. · 2001 [cited by applicant]
US 7136710B1 · Hoffberg · 2006 [cited by applicant]
US 8184696B1 · Chirila-Rus · 2012 [cited by applicant]
US 8468109B2 · Moussa et al. · 2013 [cited by applicant]
US 8924455B1 · Barman et al. · 2014 [cited by applicant]
US 20040117710A1 · Patil et al. · 2004 [cited by applicant]
US 20050044053A1 · Moreno · 2005 [cited by applicant]
US 20070022063A1 · Lightowler · 2007 [cited by applicant]
US 20070086655A1 · Simard et al. · 2007 [cited by applicant]
US 20080319933A1 · Moussa · 2008 [cited by applicant]
US 20110029471A1 · Chakradhar et al. · 2011 [cited by applicant]
US 20140142929A1 · Seide et al. · 2014 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337262A1 · Kato et al. · 2014 [cited by applicant]
US 20160267111A1 · Shoaib · 2016 [cited by applicant]
CN 1126318A · 1996 [cited by applicant]
CN 1150847A · 1997 [cited by applicant]
CN 1781076A · 2006 [cited by applicant]
CN 103109262A · 2013 [cited by applicant]
CN 104035751A · 2014 [cited by applicant]
EP 0422348A2 · 1991 [cited by applicant]
EP 3064130A1 · 2016 [cited by applicant]
JP S58039106A · 1983 [cited by applicant]
JP S60028345A · 1985 [cited by applicant]
JP S63293668A · 1988 [cited by applicant]
JP H03131965A · 1991 [cited by applicant]
JP H04229362A · 1992 [cited by applicant]
JP H04287265A · 1992 [cited by applicant]
JP H05346914A · 1993 [cited by applicant]
JP H06131308A · 1994 [cited by applicant]
JP H07141454A · 1995 [cited by applicant]
JP 2004157756A · 2004 [cited by applicant]
KR 100189195B1 · 1999 [cited by applicant]
KR 20030082255A · 2003 [cited by applicant]
TW 200923803A · 2009 [cited by applicant]
TW 201232429A · 2012 [cited by applicant]
TW 201331855A · 2013 [cited by applicant]
TW I417798B · 2013 [cited by applicant]
Islam, Mobarakol, and K. Murase. “A new weight freezing method for reducing training time in designing artificial neural networks.” In 2001 IEEE International Conference on Systems, Man and Cybernetics. e-Systems and e-… [cited by examiner]
Notice of Allowance for Korean Patent Application No. 10-2022-7021145 dated Nov. 28, 2023. 3 pages. [cited by applicant]
Office Action for European Patent Application No. 21205423.3 dated Apr. 19, 2024. 5 pages. [cited by applicant]
Beamer et al., “Ivy Bridge Server Graph Processing Bottlenecks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 56pages. [cited by applicant]
Bo et al., “String Kernel Testing Acceleration Using Micron's Automata Processor,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 21 pages. [cited by applicant]
Carlo et al., “An Area-Efficient 2-D Convolution Implementation on FPGA for Space Applications,” IEEE Computer Society, Dec. 11, 2011, pp. 1-7. [cited by applicant]
Chen and Li, “Hardware Acceleration for Neuromorphic Computing—An Evolving View,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 38 pages. [cited by applicant]
Chillet et al., “A Neural Network Model for Real-Time Scheduling on Heterogeneous SoC Architectures,” Proceedings of International Joint Conference on Neural Networks, Aug. 2007, pp. 102-107. [cited by applicant]
CN Office Action in Chinese Application No. 201680020202, dated Mar. 25, 2020, 23 pages (with English translation). [cited by applicant]
Combined Search and Examination Report for United Kingdom Patent Application No. 2112401.1 dated Nov. 30, 2021. 4 pages. [cited by applicant]
Cornu et al., “Design, Implementation, and Test of a Multi-Model Systolic Neural-Network Accelerator,” Scientific Programming—Parallel Computing Projects of the Swiss Priority Programme, vol. 5, No. 1, Jan. 1, 1996, pp.… [cited by applicant]
Dawwd, “The multi 2D systolic design and implementation of Convolutional Neural Networks,” 2013 IEEE 20.sup.th International Conference on Electronics, Circuits, and Systems (ICECS), IEEE, Dec. 8, 2013, pp. 221-224, XP0… [cited by applicant]
Dielman, Sander, Kyle W. Willett, and Joni Dambre. “Rotation-invariant convolutional neural networks for galaxy morphology prediction,” Monthly notices of the royal astronomical society, 450.2, 2015, pp. 1441-1459. [cited by applicant]
EP Office Action in European Application 16725355, dated Feb. 14, 2020, 4 pages. [cited by applicant]
EP Office Action in European Application No. 16725266.7, dated Nov. 2, 2020, 5 pages. [cited by applicant]
Examination Report for United Kingdom Patent Application No. 1715437.8 dated Apr. 15, 2021. 4 pages. [cited by applicant]
Extended European Search Report for European Patent Application No. 21205423.3 dated Nov. 15, 2021. 10 pages. [cited by applicant]
Farabet et al., “Hardware Accelerated Convolutional Neural Networks for Synthetic Vision Systems,” Circuits and Systems (ISCAS), Proceedings of 2010 IEEE International Symposium on, May-Jun. 2010, pp. 257-260. [cited by applicant]
First Examinaton Report for Indian Patent Application No. 202048055746 dated Dec. 21, 2021. 7 pages. [cited by applicant]
First Office Action for Japanese Patent Application No. 2020-069854 dated May 18, 2021. 3 pages. [cited by applicant]
Ginosar, “Accelerators for Machine Learning of Big Data,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 13 pages. [cited by applicant]
Gokhale, “Enabling Machines to Understand our World,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 18 pages. [cited by applicant]
Graf et al., “A Massively Parallel Digital Learning Processor,” Proceedings of the 22.sup.nd annual conference on Neural Information Processing Systems (NIPS), Dec. 2008, 8 pages, XP055016863. [cited by applicant]
Hecht et at., “An advanced programmable 2D-convolution chip for, real time image processing,” Signal Image and Video Processing, Jun. 1991; [Proceedings of the International Symposium on Circuits and Systems], vol. SYMP… [cited by applicant]
In Office Action in Indian Application No. 201747034437, dated Dec. 13, 2019, 6 pages (with English translation). [cited by applicant]
Indiveri, “Neuromorphic circuits for building autonomous cognitive systems,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 37 pages. [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2016/029965, dated Nov. 30, 2017, 7 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029294, dated Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029965, dated Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029968, dated Sep. 1, 2016, 14 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029986, dated Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030515, dated Aug. 25, 2016, 19 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030536, dated Aug. 31, 2016, 17 pages. [cited by applicant]
Kane, “An instruction systolic array architecture for multiple neural network types,” Loughborough University, Doctoral Thesis, Sep. 1998, 315 pages. [cited by applicant]
Khan and Ling, “Systolic architectures for artificial neural nets,” Neural Networks, 1991. 1991 IEEE International Joint Conference on, vol. 1, Nov. 1991, pp. 620-627. [cited by applicant]
Kim et al. “Efficient Hardware Architecture for Sparse Coding,” IEEE Transactions on Signal Processing 62.16, Aug. 15, 2014, 14 pages. [cited by applicant]
Kim et al., “A Large-Scale Architecture for Restricted Boltzmann Machines,” Field-Programmable Custom Computing Machines (FCCM), 2010 18th IEEE Annual International Symposium on, IEEE, May 2, 2010, pp. 201-208, XP031681… [cited by applicant]
KR Notice of Allowance in Korean Application No. 10-2017-7028188, dated Jan. 21, 2020, 3 pages (with English translation). [cited by applicant]
Krizhevsky et al., “ImageNet classification with deep convolutional neural networks,” The 26th annual conference on Neural Information Processing Systems (NIPS'25), Dec. 2012, pp. 1-9, XP55113686. [cited by applicant]
Kung et al., “Two-level pipelined systolic array for multidimensional convolution,” Image and Vision Computing, Elsevier, vol. 1, No. 1, Feb. 2, 1983, pp. 30-36, XP024237511. [cited by applicant]
Kung, “VLSI Array Processors,” IEEE ASSP Magazine, IEEE, vol. 2, No. 3, Jul. 1, 1985, pp. 4-22, XP011370547. [cited by applicant]
Lecun et al., “Efficient BackProp” Springer, 1998, 44 pages. [cited by applicant]
Lecun et al. Efficient BackProp. Sep. 13, 2013. ICIAP: International Conference on Image Analysis and Processing, 17th International Conference, Naples, Italy, Sep. 9-13, 2013, pp. 9-48. [cited by applicant]
Lee and Song, “Implementation of the Super-Systolic Array for Convolution,” Design Automation Conference, 2003. Proceedings of the ASP-DAC 2003. Asia and South Pacific, Jan. 2003, pp. 491-494. [cited by applicant]
Lee, Yim-Kul, and William T. Rhodes. “Nonlinear image processing by a rotating kernel transformation,” Optics letters 15.23, 1990, pp. 1383-1385. [cited by applicant]
Lehmann et al., “A generic systolic array building block for neural networks with on-chip learning,” Neural Networks, IEEE Transactions on, 4(3):400-407, May 1993. [cited by applicant]
Lipasti et al., Mimicking the Self-Organizing Properties of the Visual Cortex, The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 23 pages. [cited by applicant]
Lo, Shih-Chung B., et al. “Artificial convolutional neural network for medical image pattern recognition,” Neural networks 8.7, 1995, pp. 1201-1214. [cited by applicant]
Mahapatra et al., “Mapping of Neural Network Models onto Systolic Arrays, ” Journal of Parallel and Distributed Computing 60, 677-689, Jan. 2000. [cited by applicant]
Merolla et al. “A digital Neurosynaptic Core Using Embedded Crossbar Memory with 45pJ per Spike in 45nm,” IEEE CICC, Sep. 19, 2011, 4 pages. [cited by applicant]
Office Action for German Patent Application No. 112016002298.0 dated Jan. 3, 2023. 12 pages. [cited by applicant]
Office Action for Japanese Patent Application No. 2021-159352 dated Dec. 21, 2021. 3 pages. [cited by applicant]
Office Action for Japanese Patent Application No. 2022-076581 dated Apr. 4, 2023. 6 pages. [cited by applicant]
First Office Action for Chinese Patent Application No. 202011278833.6 dated Apr. 28, 2024. 7 pages. [cited by applicant]
Office Action for Korean Patent Application No. 10-2020-7011552 dated Dec. 24, 2021. 4 pages. [cited by applicant]
Office Action for Korean Patent Application No. 10-2022-7021145 dated Feb. 3, 2023. 5 pages. [cited by applicant]
Office Action in Japanese Application No. 2017-550913, dated Jun. 4, 2019, 11 pages (with English translation). [cited by applicant]
Office Action in Taiwanese Application No. 105115859, dated Nov. 16, 2016, 10 pages. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware in the Datacenter,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 33 pages. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware,” Microsoft Research, [online] [retrieved Mar. 25, 2021]. Retrieved from the Internet: <URL:http:/www.microsoft.com/en-us/res… [cited by applicant]
Patil et al., “Hardware Architecture for Large Parallel Array of Random Feature Extractors applied to Image Recognition,” Dec. 24, 2015, arXiv:1512.07783v1, 18 pages, XP055296121. [cited by applicant]
Pearce, “You Have No (Predictive) Power Here, SPEC!” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 15 pages. [cited by applicant]
Rojas, “Hardware for Neural Networks,” Neural Networks, Springer-Verlag, Berlin, 1996, pp. 451-478. [cited by applicant]
Shaaban, “Systolic Architectures,” PowerPoint Presentation, Mar. 2003, 9 pages. [cited by applicant]
Shapri and Rahman, “Performance Analysis of Two-Dimensional Systolic Array Matrix Multiplication with Orthogonal Interconnections,” International Journal on New Computer Architectures and Their Applications (IJNCAA) 1 (… [cited by applicant]
Smith, “Biologically Plausible Spiking Neural Networks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 77 pages. [cited by applicant]
Sudha et al., “Systolic array realization of a neural network-based face recognition system,” Industrial Electronics and Applications, 2008, ICIEA 2008, 3rd IEEE Conference on, pp. 1864-1869, Jun. 2009. [cited by applicant]
Wong et al., “A New Scalable Systolic Array Processor Architecture for Discrete Convolution,” College of Engineering at the University of Kentucky, Master Thesis, 2003, 175 pages. [cited by applicant]
Wu et al., “Flip-Rotate-Pooling Convolution and Split Dropout on Convolution Neural Networks for Image Classification,” Jul. 31, 2015, arXiv:1507.08754v1, pp. 1-9, XP055296122. [cited by applicant]
Yiping et al (“A High Performance Digital Neural Processor Design by Network on Chip Architecture” IEEE 2011). [cited by applicant]