IP Library Granted Patent US 12,277,496
Granted Patent B2
US 12,277,496 · App. 17/575,799 · Granted Apr 15, 2025

Batch processing in a neural network processor

Inventor: Reginald Clifford Young (Palo Alto, CA)
Assignee: Google LLC
G06N3/08G06N3/063G06N3/06G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,496
App. No.
17/575,799
Granted
Apr 15, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating a respective neural network output for each of a plurality of inputs, the method comprising, for each of the neural network layers: receiving a plurality of inputs to be processed at the neural network layer; forming one or more batches of inputs from the plurality of inputs, each batch having a number of inputs up to the respective batch size for the neural network layer; selecting a number of the one or more batches of inputs to process, where a count of the inputs in the number of the one or more batches is greater than or equal to the respective associated batch size of a subsequent layer in the sequence; and processing the number of the one or more batches of inputs to generate the respective neural network layer output.

Claims (48)

1. A system comprising:

a processor comprising processing elements, configured to:

receive data representing at least a plurality of neural network layers arranged in a graph structure, the data comprising weights for one or more of the plurality of neural network layers;

receive one or more batches of input for the one or more neural network layers in accordance with a weight reuse value;

receive one or more control signals for loading the weights into respective processing elements;

load the weights into the respective processing elements in accordance with the one or more control signals; and

reuse the loaded weights in accordance with the weight reuse value to process the one or more batches of input.

2. The system of claim 1 , wherein the weight reuse value corresponds to the number of the one or more batches of input.

3. The system of claim 1 , wherein the graph structure is a directed graph structure of a neural network comprising the plurality of neural network layers.

4. The system of claim 1 , wherein:

the one or more neural network layers are convolutional neural network layers; and

processing the one or more batches of input comprises calculating convolutions.

5. The system of claim 1 , wherein:

the received data comprises one or more instructions; and

the processor is further configured to convert the one or more instructions into the one or more control signals.

6. The system of claim 5 , wherein the one or more instructions comprise configuration information for the plurality of neural network layers.

7. The system of claim 6 , wherein the configuration information comprises respective input and output sizes for each of the plurality of neural network layers.

8. The system of claim 1 , wherein the processor comprises one or more accumulators for accumulating values calculated from processing the one or more batches of input.

9. A method, comprising:

receiving, by a processor comprising processing elements, data representing at least a plurality of neural network layers arranged in a graph structure, the data comprising weights for one or more of the plurality of neural network layers;

receiving, by the processor, one or more batches of input for the one or more neural network layers in accordance with a weight reuse value;

receiving, by the processor, one or more control signals for loading the weights into respective processing elements;

loading, by the processor, the weights into the respective processing elements in accordance with the one or more control signals; and

reusing, by the processor, the loaded weights in accordance with the weight reuse value to process the one or more batches of input.

10. The method of claim 9 , wherein the weight reuse value corresponds to the number of the one or more batches of input.

11. The method of claim 9 , wherein the graph structure is a directed graph structure of a neural network comprising the plurality of neural network layers.

12. The method of claim 9 , wherein:

the one or more neural network layers are convolutional neural network layers; and

processing the one or more batches of input comprises calculating convolutions.

13. The method of claim 9 , wherein:

the received data comprises one or more instructions; and

the method further comprises converting, by the processor, the one or more instructions into the one or more control signals.

14. The method of claim 13 , wherein the one or more instructions comprise configuration information for the plurality of neural network layers.

15. The method of claim 9 , wherein the processor comprises one or more accumulators for accumulating values calculated from processing the one or more batches of input.

16. One or more non-transitory computer-readable storage media having instructions stored thereon that when executed by a processor comprising processing elements, causes the processor to perform operations comprising:

receiving data representing at least a plurality of neural network layers arranged in a graph structure, the data comprising weights for one or more of the plurality of neural network layers;

receiving one or more batches of input for the one or more neural network layers in accordance with a weight reuse value;

receiving one or more control signals for loading the weights into respective processing elements;

loading the weights into the respective processing elements in accordance with the one or more control signals; and

reusing the loaded weights in accordance with the weight reuse value to process the one or more batches of input.

17. The one or more non-transitory computer-readable storage media of claim 16 , wherein the weight reuse value corresponds to the number of the one or more batches of input.

18. The one or more non-transitory computer-readable storage media of claim 16 , wherein the graph structure is a directed graph structure of a neural network comprising the plurality of neural network layers.

19. The one or more non-transitory computer-readable storage media of claim 16 , wherein:

the one or more neural network layers are convolutional neural network layers; and

processing the one or more batches of input comprises calculating convolutions.

20. The one or more non-transitory computer-readable storage media of claim 16 , wherein:

the received data comprises one or more instructions; and

the operations further comprise converting the one or more instructions into the one or more control signals.

Assignments (2)
CHANGE OF NAME Recorded Jan 19, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 058768/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2022
From: YOUNG, REGINALD CLIFFORD
To: GOOGLE INC.
Reel/Frame 058659/0534 →
Continuity (5)
Continuation 17226256 · Apr 9, 2021
Continuation 16139258 · Sep 24, 2018
Continuation 14844431 · Sep 3, 2015
Provisional Application 62165020 · May 21, 2015
Related Publication 20220138577A1 · May 5, 2022
References Cited (118)
US 5014235A · Morton · 1991 [cited by applicant]
US 5136717A · Morley et al. · 1992 [cited by applicant]
US 5138695A · Means · 1992 [cited by applicant]
US 5146543A · Vassiliadis et al. · 1992 [cited by applicant]
US 5337395A · Vassiliadis et al. · 1994 [cited by applicant]
US 5471627A · Means et al. · 1995 [cited by applicant]
US 5544336A · Kato · 1996 [cited by applicant]
US 5574827A · Wang · 1996 [cited by applicant]
US 5586223A · Bryant et al. · 1996 [cited by applicant]
US 5799134A · Chiueh et al. · 1998 [cited by applicant]
US 5812993A · Ginosar et al. · 1998 [cited by applicant]
US 6038337A · Lawrence · 2000 [cited by applicant]
US 6184753B1 · Ishimi et al. · 2001 [cited by applicant]
US 6917703B1 · Steffens et al. · 2005 [cited by applicant]
US 7136710B1 · Hoffberg · 2006 [cited by applicant]
US 8184696B1 · Chirila-Rus · 2012 [cited by applicant]
US 8468109B2 · Moussa et al. · 2013 [cited by applicant]
US 8924455B1 · Barman et al. · 2014 [cited by applicant]
US 20020143720A1 · Anderson et al. · 2002 [cited by applicant]
US 20050044053A1 · Moreno · 2005 [cited by applicant]
US 20070022063A1 · Lightowler · 2007 [cited by applicant]
US 20070086655A1 · Simard et al. · 2007 [cited by applicant]
US 20080319933A1 · Moussa · 2008 [cited by applicant]
US 20110029471A1 · Chakradhar et al. · 2011 [cited by applicant]
US 20140142929A1 · Seide et al. · 2014 [cited by applicant]
US 20140164299A1 · Sainath et al. · 2014 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337262A1 · Kato et al. · 2014 [cited by applicant]
US 20150066499A1 · Wang et al. · 2015 [cited by applicant]
US 20150127337A1 · Heigold et al. · 2015 [cited by applicant]
US 20160267111A1 · Shoaib · 2016 [cited by applicant]
CN 1150847A · 1997 [cited by applicant]
CN 104035751A · 2014 [cited by applicant]
EP 0422348A2 · 1991 [cited by applicant]
EP 3064130A1 · 2016 [cited by applicant]
JP H06052132A · 1994 [cited by applicant]
JP H06203005A · 1994 [cited by applicant]
JP 2004157756A · 2004 [cited by applicant]
JP 2013105377A · 2013 [cited by applicant]
KR 19970049687A · 1997 [cited by applicant]
KR 20100100105A · 2010 [cited by applicant]
KR 20140054267A · 2014 [cited by applicant]
KR 20140092879A · 2014 [cited by applicant]
KR 20150016089A · 2015 [cited by applicant]
SG 182933A1 · 2012 [cited by applicant]
TW 201232429A · 2012 [cited by applicant]
TW 201331855A · 2013 [cited by applicant]
WO 2015003436A1 · 2015 [cited by applicant]
Zhang, Chen, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. “Optimizing FPGA-based accelerator design for deep convolutional neural networks.” In Proceedings of the 2015 ACM/SIGDA international symposiu… [cited by examiner]
McFall, Kevin S. An artificial neural network method for solving boundary value problems with arbitrary irregular boundaries. Georgia Institute of Technology, 2006. (Year: 2006). [cited by examiner]
Notice of Allowance for Korean Patent Application No. 10-2023-7041445 dated Jun. 26, 2024. 3 pages. [cited by applicant]
IN Office Action in Indian Application No. 201747034501, dated May 3, 2020, 5 pages (with English translation). [cited by applicant]
CN Office Action in Chinese Application No. 201680020154, dated Mar. 24, 2020, 10 pages (with English translation). [cited by applicant]
EP Office Action in European Application No. 16722484.9, dated Feb. 17, 2020, 6 pages. [cited by applicant]
KR Notice of Allowance in Korean Application No. 10-2017-7027872, dated Jan. 23, 2020, 3 pages (with English translation). [cited by applicant]
JP Office Action in Japan Application No. 2017550720, dated Jun. 11, 2019, 5 pages. [cited by applicant]
Office Action for European Patent Application No. 16722484.9 dated Nov. 4, 2021. 5 pages. [cited by applicant]
Notice of Allowance for Korean Patent Application No. 10-2020-7011889 dated Nov. 8, 2021. 4 pages. [cited by applicant]
Notice of Allowance for Korean Patent Application No. 10-2022-7004246 dated Oct. 21, 2022. 3 pages. [cited by applicant]
Extended European Search Report for European Patent Application No. 23169823.4 dated Aug. 14, 2023. 11 pages. [cited by applicant]
Moreau et al. SNNAP: Approximate computing on programmable SoCs via neural acceleration. 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA), IEEE, Feb. 7, 2015 (Feb. 7, 2015), pp. 60… [cited by applicant]
Rogers et al. Using the BSP Cost Model to Optimise Parallel Neural Network Training. Future Generation Computer Systems, Elsevier Science Publishers. Amsterdam, NL, vol. 14, Dec. 1, 1998 (Dec. 1, 1998), pp. 409-424. [cited by applicant]
Tang et al. EF-Train: Enable Efficient On-device CNN Training on FPGA Through Data Reshaping for Online Adaptation or Personalization. ACM Transactions on Design Automation of Electronic Systems, ACM, New York, NY, US, … [cited by applicant]
Hearing Notice for Indian Patent Application No. 201747034501 dated Nov. 29, 2023. 3 pages. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware,” Microsoft Research, [online] [retrieved Mar. 25, 2021]. Retrieved from the Internet: <URL:http:/www.microsoft.com/en-us/res… [cited by applicant]
Beamer et al., “Ivy Bridge Server Graph Processing Bottlenecks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 56pages. [cited by applicant]
Bo et al., “String Kernel Testing Acceleration Using Micron's Automata Processor,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 21 pages. [cited by applicant]
Chen and Li, “Hardware Acceleration for Neuromorphic Computing—An Evolving View,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 38 pages. [cited by applicant]
Chillet et al., “A Neural Network Model for Real-Time Scheduling on Heterogeneous SoC Architectures,” Proceedings of International Joint Conference on Neural Networks, Aug. 2007, pp. 102-107. [cited by applicant]
Farabet et al., “Hardware Accelerated Convolutional Neural Networks for Synthetic Vision Systems,” Circuits and Systems (ISCAS), Proceedings of 2010 IEEE International Symposium on, May-Jun. 2010, pp. 257-260. [cited by applicant]
Ginosar, “Accelerators for Machine Learning of Big Data,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 13 pages. [cited by applicant]
Gokhale, “Enabling Machines to Understand our World,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 18 pages. [cited by applicant]
Indiveri, “Neuromorphic circuits for building autonomous cognitive systems,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 37 pages. [cited by applicant]
Kane, “An instruction systolic array architecture for multiple neural network types,” Loughborough University, Doctoral Thesis, Sep. 1998, 315 pages. [cited by applicant]
Khan and Ling, “Systolic architectures for artificial neural nets,” Neural Networks, 1991. 1991 IEEE International Joint Conference on, vol. 1, Nov. 1991, pp. 620-627. [cited by applicant]
Lee and Song, “Implementation of the Super-Systolic Array for Convolution,” Design Automation Conference, 2003. Proceedings of the ASP-DAC 2003. Asia and South Pacific, Jan. 2003, pp. 491-494. [cited by applicant]
Lehmann et al., “A generic systolic array building block for neural networks with on-chip learning,” Neural Networks, IEEE Transactions on, 4(3):400-407, May 1993. [cited by applicant]
Lipasti et al., Mimicking the Self-Organizing Properties of the Visual Cortex, The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 23 pages. [cited by applicant]
Mahapatra et al., “Mapping of Neural Network Models onto Systolic Arrays,” Journal of Parallel and Distributed Computing 60, 677-689, Jan. 2000. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware in the Datacenter,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 33 pages. [cited by applicant]
Pearce, “You Have No (Predictive) Power Here, SPEC!” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 15 pages. [cited by applicant]
Rojas, “Hardware for Neural Networks,” Neural Networks, Springer-Verlag, Berlin, 1996, pp. 451-478. [cited by applicant]
Shaaban, “Systolic Architectures,” PowerPoint Presentation, Mar. 2003, 9 pages. [cited by applicant]
Shapri and Rahman, “Performance Analysis of Two-Dimensional Systolic Array Matrix Multiplication with Orthogonal Interconnections,” International Journal on New Computer Architectures and Their Applications (IJNCAA) 1(3… [cited by applicant]
Smith, “Biologically Plausible Spiking Neural Networks,” The First International Workshop Computer Architecture for Machine Learning, Jun. 2015, 77 pages. [cited by applicant]
Sudha et al., “Systolic array realization of a neural network-based face recognition system,” Industrial Electronics and Applications, 2008, ICIEA 2008, 3rd IEEE Conference on, pp. 1864-1869, Jun. 2009. [cited by applicant]
Wong et al., “A New Scalable Systolic Array Processor Architecture for Discrete Convolution,” College of Engineering at the University of Kentucky, Master Thesis, 2003, 175 pages. [cited by applicant]
Dielman, Sander, Kyle W. Willett, and Joni Dambre. “Rotation-invariant convolutional neural networks for galaxy morphology prediction,” Monthly notices of the royal astronomical society, 450.2, 2015, pp. 1441-1459. [cited by applicant]
Kim et al. “Efficient Hardware Architecture for Sparse Coding,” IEEE Transactions on Signal Processing 62.16, Aug. 15, 2014, 14 pages. [cited by applicant]
Lee, Yim-Kul, and William T. Rhodes. “Nonlinear image processing by a rotating kernel transformation,” Optics letters 15.23, 1990, pp. 1383-1385. [cited by applicant]
Lo, Shih-Chung B., et al. “Artificial convolutional neural network for medical image pattern recognition,” Neural networks 8.7, 1995, pp. 1201-1214. [cited by applicant]
Merolla et al. “A digital Neurosynaptic Core Using Embedded Crossbar Memory with 45pJ per Spike in 45nm,” IEEE CICC, Sep. 19, 2011, 4 pages. [cited by applicant]
Yiping et al (“A High Performance Digital Neural Processor Design by Network on Chip Architecture” IEEE 2011). [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029968, dated Sep. 1, 2016, 14 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029294, dated Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029986, dated Sep. 1, 2016, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/029965, dated Sep. 1, 2016, 13 pages. [cited by applicant]
Krizhevsky et al., “ImageNet classification with deep convolutional neural networks,” The 26th annual conference on Neural Information Processing Systems (NIPS'25), Dec. 2012, pp. 1-9, XP55113686. [cited by applicant]
Kung, “VLSI Array Processors,” IEEE ASSP Magazine, IEEE, vol. 2, No. 3, Jul. 1, 1985, pp. 4-22, XP011370547. [cited by applicant]
Carlo et al., “An Area-Efficient 2-D Convolution Implementation on FPGA for Space Applications,” IEEE Computer Society, Dec. 11, 2011, pp. 1-7. [cited by applicant]
Office Action in Taiwanese Application No. 105115859, dated Nov. 16, 2016, 10 pages. [cited by applicant]
Cornu et al., “Design, Implementation, and Test of a Multi-Model Systolic Neural-Network Accelerator,” Scientific Programming—Parallel Computing Projects of the Swiss Priority Programme, vol. 5, No. 1, Jan. 1, 1996, pp.… [cited by applicant]
Dawwd, “The multi 2D systolic design and implementation of Convolutional Neural Networks,” 2013 IEEE 20.sup.th International Conference on Electronics, Circuits, and Systems (ICECS), IEEE, Dec. 8, 2013, pp. 221-224, XP0… [cited by applicant]
Graf et al., “A Massively Parallel Digital Learning Processor,” Proceedings of the 22.sup.nd annual conference on Neural Information Processing Systems (NIPS), Dec. 2008, 8 pages, XP055016863. [cited by applicant]
Hecht et at., “An advanced programmable 2D-convolution chip for, real time image processing,” Signal Image and Video Processing, Jun. 1991; [Proceedings of the International Symposium on Circuits and Systems], vol. SYMP… [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030515, dated Aug. 25, 2016, 19 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/030536, dated Aug. 31, 2016, 17 pages. [cited by applicant]
Kim et al., “A Large-Scale Architecture for Restricted Boltzmann Machines,” Field-Programmable Custom Computing Machines (FCCM), 2010 18th IEEE Annual International Symposium on, IEEE, May 2, 2010, pp. 201-208, XP031681… [cited by applicant]
Kung et al., “Two-level pipelined systolic array for multidimensional convolution,” Image and Vision Computing, Elsevier, vol. 1, No. 1, Feb. 2, 1983, pp. 30-36, XP024237511. [cited by applicant]
Patil et al., “Hardware Architecture for Large Parallel Array of Random Feature Extractors applied to Image Recognition,” Dec. 24, 2015, arXiv:1512.07783v1, 18 pages, XP055296121. [cited by applicant]
Wu et al., “Flip-Rotate-Pooling Convolution and Split Dropout on Convolution Neural Networks for Image Classification,” Jul. 31, 2015, arXiv:1507.08754v1, pp. 1-9, XP055296122. [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2016/030515, dated Nov. 30, 2017, 12 pages. [cited by applicant]
IN Office Action in Indian Application No. 201747034501, dated Jun. 3, 2020, 5 pages (with English translation). [cited by applicant]
JP Office Action in Japanese Application No. 2019-234606, dated Jun. 23, 2020, 4 pages (with English translation). [cited by applicant]
Bengio. Practical Recommendations for Gradient-Based Traing of Deep Architectures. Sep. 16, 2012. Neural Networks: Tricks of the Trade, Lecture Notes in Computer Science, vol. 7700. Springer, 33 pages. [cited by applicant]
Notice of Allowance for Korean Patent Application No. 10-2023-7002538 dated Feb. 28, 2023. 2 pages. [cited by applicant]
Notice of Allowance for Korean Patent Application No. 10-2023-7018341 dated Aug. 31, 2023. 3 pages. [cited by applicant]