IP Library Granted Patent US 12,387,089
Granted Patent B2
US 12,387,089 · App. 17/530,852 · Granted Aug 12, 2025

Efficient neural network accelerator dataflows

Inventors: Yakun Shao (Santa Clara, CA); Rangharajan Venkatesan (Sunnyvale, CA); Miaorong Wang (Cambridge, MA); Daniel Smith (Los Gatos, CA); William James Dally (Incline Village, NV); Joel Emer (Acton, MA); Stephen W. Keckler (Austin, TX); Brucek Khailany (Austin, TX)
Assignee: Nvidia Corp.
G06N3/063G06F9/30036G06F9/3877G06F17/16G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,089
App. No.
17/530,852
Granted
Aug 12, 2025
Kind
B2
Abstract

A distributed deep neural net (DNN) utilizing a distributed, tile-based architecture includes multiple chips, each with a central processing element, a global memory buffer, and a plurality of additional processing elements. Each additional processing element includes a weight buffer, an activation buffer, and vector multiply-accumulate units to combine, in parallel, the weight values and the activation values using stationary data flows.

Claims (38)

1. A data processor comprising:

a plurality of processing elements, each comprising:

a weight buffer;

an activation buffer;

an accumulation memory buffer;

a plurality of vector multiply-accumulate units configured to compute in parallel a convolution of weights from the weight buffer and activations from the activation buffer, each of the vector multiply-accumulate units comprising:

a first collector disposed between the vector multiply-accumulate unit and the weight buffer; and

a second collector disposed between the vector multiply-accumulate unit and the accumulation memory buffer; and

configuration logic to configure a depth of the collectors to adjust a level of data-stationary computation of the convolution.

2. The data processor of claim 1 , further comprising:

the configuration logic to configure a depth of the first collector on one or more of the vector multiply-accumulate units to adjust a level of weight-stationary computation of the convolution.

3. The data processor of claim 1 , further comprising:

the configuration logic to configure a depth of the second collector on one or more of the vector multiply-accumulate units to adjust a level of output-stationary computation of the convolution.

4. The data processor of claim 1 , further comprising:

the configuration logic to configure one or both of a depth of the first collector and a depth of the second collector on one or more of the vector multiply-accumulate units to implement one or more of multi-level weight-stationary computation and output-stationary computation of the convolution.

5. The data processor of claim 1 , the processing elements further comprising:

a third collector disposed between the activation buffer and the vector multiply-accumulate units.

6. The data processor of claim 5 , further comprising:

the configuration logic to configure a depth of the third collector to adjust a level of input-stationary computation of the convolution.

7. The data processor of claim 5 , further comprising:

the configuration logic to configure one or more of a depth of the first collector, a depth of the second collector, and a depth of the third collector of one or more of the vector multiply-accumulate units to implement one or more of multi-level weight-stationary, output-stationary, and input stationary computation of the convolution.

8. The data processor of claim 1 , further comprising:

a global memory; buffer; and

the configuration logic to configure the processing elements to utilize the global memory buffer to store the activations and to apply the activations between layers of a neural network.

9. The data processor of claim 1 , further comprising:

the configuration logic to configure one or more of the vector multiply-accumulate units to compute a portion of the convolution as a partial result and to forward the partial result from the accumulation memory buffer to neighboring processing elements.

10. The data processor of claim 1 , further comprising:

the configuration logic to distribute the weights and the activations among the processing elements spatially by a depth of an input of a neural network, and temporally by a height and a width of the input to the neural network.

11. The data processor of claim 1 , further comprising:

the configuration logic to distribute the weights and the activations among the processing elements spatially and temporally by configurable combinations of input dimensions of a neural network and dimensions of the weights.

12. A neural network computation method comprising:

distributing weight values and activation values for a neural network computation among a plurality of processing elements of the neural network spatially by a depth of an input to the neural network, and temporally by a height and a width of the input to the neural network;

configuring a depth of at least one collector of a plurality of vector multiply-accumulate units of the processing elements to implement a stationary data flow by the vector multiply-accumulate units during the neural network computation.

13. The neural network computation method of claim 12 , wherein the stationary data flow is a weigh-stationary data flow.

14. The neural network computation method of claim 12 , wherein the data flow is an output-stationary data flow.

15. A neural network computation method comprising:

distributing weight values and activation values for a neural network computation among a plurality of processing elements of the neural network spatially and temporally by configurable combinations of different dimensions of inputs to the neural network and weights of the neural network; and

configuring a depth of at least one collector of a plurality of vector multiply-accumulate units of a plurality of processing elements to implement a stationary data flow by the vector multiply-accumulate units during the neural network computation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2023
From: SHAO, YAKUN; WANG, MIAORONG
To: NVIDIA CORP.
Reel/Frame 063938/0357 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2022
From: VENKATESAN, RANGHARAJAN; SMITH, DANIEL; DALLY, WILLIAM JAMES; EMER, JOEL; KECKLER, STEPHEN W.; KHAILANY, BRUCEK
To: NVIDIA CORP.
Reel/Frame 058875/0036 →
Continuity (3)
Division 16672918 · Nov 4, 2019
Provisional Application 62817413 · Mar 12, 2019
Related Publication 20220076110A1 · Mar 10, 2022
References Cited (42)
US 5008833A · Agranat et al. · 1991 [cited by applicant]
US 5091864A · Baji et al. · 1992 [cited by applicant]
US 5297232A · Murphy · 1994 [cited by applicant]
US 5319587A · White · 1994 [cited by applicant]
US 5704016A · Shigematsu et al. · 1997 [cited by applicant]
US 5768476A · Sugaya et al. · 1998 [cited by applicant]
US 6560582B1 · Woodall · 2003 [cited by applicant]
US 7082419B1 · Lightowler · 2006 [cited by applicant]
US 10817260B1 · Huang et al. · 2020 [cited by applicant]
US 10929746B2 · Litvak et al. · 2021 [cited by applicant]
US 20030061184A1 · Masgonty et al. · 2003 [cited by applicant]
US 20090210218A1 · Collobert et al. · 2009 [cited by applicant]
US 20180004689A1 · Marfatia et al. · 2018 [cited by applicant]
US 20180046897A1 · Kang et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally · 2018 [cited by examiner]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180173676A1 · Tsai et al. · 2018 [cited by applicant]
US 20190050717A1 · Temam et al. · 2019 [cited by applicant]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190179975A1 · Rassekh et al. · 2019 [cited by applicant]
US 20190180170A1 · Huang et al. · 2019 [cited by applicant]
US 20190340510A1 · Li et al. · 2019 [cited by applicant]
US 20190370645A1 · Lee et al. · 2019 [cited by applicant]
US 20190392287A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20210004668A1 · Moshovos et al. · 2021 [cited by applicant]
US 20210125046A1 · Moshovos et al. · 2021 [cited by applicant]
CN 107463932A · 2017 [cited by applicant]
Yavits, Leonid, Amir Morad, and Ran Ginosar. “Sparse matrix multiplication on an associative processor.” IEEE Transactions on Parallel and Distributed Systems 26.11 (2014): 3175-3183. (Year: 2014). [cited by applicant]
Angshuman Parashar et al. “SCN N: An Accelerator for Compressed-sparse Convolutional Neural Networks”, (Year: 2017). [cited by applicant]
Duckhwan Kim et al. “Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory”, (Year: 2016). [cited by applicant]
Eunhyeok Park et al. “Energy-efficient Neural Network Accelerator Based on Outlier-aware Low-precision Computation”, Jun. 1, 2018 (Year: 2018). [cited by applicant]
U.S. Appl. No. 16/517,431, filed Jul. 19, 2019, Yakun Shao. [cited by applicant]
Andri, YodaN N: An Ultra-Low Power Convolutional Neural Network Accelerator Based on Binary Weights, IEEE, pp. 238-241 (Year: 2016). [cited by applicant]
Han, EIE: Efficient Inference Engine on Compressed Deep Neural Network, IEEE, pp. 243-254 (Year: 2016). [cited by applicant]
Kartik Hegde et al. “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition” ,Apr. 18, 2018. (Year: 2018). [cited by applicant]
Streat, Lennard, Dhireesha Kudithipudi, and Kevin Gomez. “Non-volatile hierarchical temporal memory: Hardware for spatial pooling.” arXiv preprint arXiv:1611.02792 (2016). (Year: 2016). [cited by applicant]
F. Sijstermans. The NVIDIA deep learning accelerator. In Hot Chips, 2018. [cited by applicant]
K. Kwon et al. Co-design of deep neural nets and neural net accelerators for embedded vision applications. In Proc. DAC, 2018. [cited by applicant]
N. P. Jouppi et al. In-datacenter performance analysis of a tensor processing unit. In Proc. ISCA, 2017. [cited by applicant]
T. Chen et al. DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning. In Proc. ASPLOS, 2014. [cited by applicant]
W. Lu et al. FlexFlow: A flexible dataflow accelerator architecture for convolutional neural networks. In Proc. HPCA, 2017. [cited by applicant]
Y. Chen, J. Emer, and V. Sze. Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks. In Proc. ISCA, 2016. [cited by applicant]