IP Library Granted Patent US 11,270,197
Granted Patent B2
US 11,270,197 · App. 16/672,918 · Granted Mar 8, 2022

Efficient neural network accelerator dataflows

Inventors: Yakun Shao (Santa Clara, CA); Rangharajan Venkatesan (San Jose, CA); Miaorong Wang (San Jose, CA); Daniel Smith (Los Gatos, CA); William James Dally (Incline Village, NV); Joel Emer (Acton, MA); Stephen W. Keckler (Austin, TX); Brucek Khailany (Austin, TX)
Assignee: NVIDIA Corp.
G06N3/063G06F9/3877G06F17/16G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,270,197
App. No.
16/672,918
Granted
Mar 8, 2022
Kind
B2
Abstract

A distributed deep neural net (DNN) utilizing a distributed, tile-based architecture includes multiple chips, each with a central processing element, a global memory buffer, and a plurality of additional processing elements. Each additional processing element includes a weight buffer, an activation buffer, and vector multiply-accumulate units to combine, in parallel, the weight values and the activation values using stationary data flows.

Claims (20)

1. A semiconductor neural network accelerator comprising a plurality of chips, each chip comprising:

a controller;

a global memory buffer;

a plurality of processing elements, each comprising a weight buffer, an accumulation memory buffer, and a plurality of vector multiply-accumulate units;

at least one of the vector multiply-accumulate units comprising:

a first collector interposed to receive weights from the weight buffer and to provide the weights for a convolution computation; and

a second collector interposed to receive results of the convolution computation and to provide the results to the accumulation memory buffer; and

configuration logic to configure a depth of one or both of the first collector and the second collector to implement a stationary data flow by the neural network during the convolution computation.

2. The semiconductor neural network accelerator of claim 1 , further comprising:

the configuration logic to configure a depth of one or both of the first collector and the second collector to implement multi-level weight stationary/output stationary data flows for the convolution computation.

3. The semiconductor neural network accelerator of claim 1 , further comprising:

the configuration logic to configure a depth of one or both of the first collector and the second collector to implement one or more of multi-level weight stationary, output stationary, and input stationary data flows for the convolution.

4. The semiconductor neural network accelerator of claim 1 , wherein the configuration logic configures at least two or more of the processing elements to:

compute a portion of the convolution and forward the portion to neighboring processing elements for completion; and

communicate results of completion of the convolution to the global memory buffer.

5. The semiconductor neural network accelerator of claim 4 , configured such that results of completion of the convolution are staged in the global memory buffer for communication between layers of the neural network.

6. The semiconductor neural network accelerator of claim 1 , one or more of the processing elements further comprising:

a third collector disposed between an activation buffer and the vector multiply-accumulate units.

7. The semiconductor neural network accelerator of claim 6 , further comprising:

the configuration logic to configure a depth of the third collector to adjust a stationary level of input activations during the convolution computation by the vector multiply-accumulate units.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2023
From: VENKATESAN, RANGHARAJAN; SMITH, DANIEL; DALLY, WILLIAM JAMES; EMER, JOEL; KECKLER, STEPHEN W; KHAILANY, BRUCEK
To: NVIDIA CORP.
Reel/Frame 063938/0332 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2020
From: WANG, MIAORONG; SHAO, YAKUN
To: NVIDIA CORP.
Reel/Frame 051407/0788 →
Continuity (2)
Provisional Application 62817413 · Mar 12, 2019
Related Publication 20200293867A1 · Sep 17, 2020