IP Library Granted Patent US 11,610,100
Granted Patent B2
US 11,610,100 · App. 16/504,275 · Granted Mar 21, 2023

Accelerator for deep neural networks

Inventors: Patrick Judd (Toronto, CA); Jorge Albericio (San Jose, CA); Alberto Delmas Lascorz (Toronto, CA); Andreas Moshovos (Toronto, CA); Sayeh Sharifymoghaddam (Toronto, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/063G06N3/049
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,610,100
App. No.
16/504,275
Granted
Mar 21, 2023
Kind
B2
Abstract

A system for bit-serial computation in a neural network is described. The system may be embodied on an integrated circuit and include one or more bit-serial tiles for performing bit-serial computations in which each bit-serial tile receives input neurons and synapses, and communicates output neurons. Also included is an activation memory for storing the neurons and a dispatcher and a reducer. The dispatcher reads neurons and synapses from memory and communicates either the neurons or the synapses bit-serially to the one or more bit-serial tiles. The other of the neurons or the synapses are communicated bit-parallelly to the one or more bit-serial tiles, or according to a further embodiment, may also be communicated bit-serially to the one or more bit-serial tiles. The reducer receives the output neurons from the one or more tiles, and communicates the output neurons to the activation memory.

Claims (45)

1. A system for bit-serial computation in a neural network, comprising:

one or more bit-serial tiles for performing bit-serial computations in a neural network wherein each bit-serial tile processes two or more windows in parallel, each bit-serial tile receiving input neurons in two or more windows and synapses in two or more filters, and generating output neurons;

an activation memory for storing neurons and in communication with the one or more bit-serial tiles via a dispatcher and a reducer,

wherein the dispatcher reads neurons from the activation memory and communicates the neurons to the one or more bit-serial tiles via two or more window lanes for each bit-serial tile,

and wherein the dispatcher reads synapses from a synapse buffer and communicates the synapses to the one or more bit-serial tiles via two or more filter lanes for each bit-serial tile,

and wherein the reducer receives the output neurons from the one or more bit-serial tiles, and communicates the output neurons to the activation memory;

and wherein one of the neurons or the synapses are communicated to the one or more bit-serial tiles bit-serially and the other of the neurons or the synapses are communicated to the one or more bit-serial tiles bit-parallelly.

2. The system of claim 1 , wherein the dispatcher comprises a shuffler to collect the neurons in one or more bricks and a transposer to convert the bricks into serial bit streams and wherein the dispatcher collects the one or more bricks into one of more groups.

3. The system of claim 1 , wherein the activation memory is a dedicated memory to the one or more bit-serial tiles.

4. The system of claim 1 , wherein each window lane comprises one or more bit-serial neuron lanes.

5. The system of claim 1 , wherein the bit-serial tiles each further comprise an input neuron buffer holding input neurons from the dispatcher and a neuron output buffer holding output neurons pending communication to the reducer.

6. The system of claim 5 , wherein each filter lane comprises one or more synapse lanes.

7. The system of claim 6 , wherein the synapse buffer and the input neuron buffer are in communication with a 2-dimensional array of one or more serial inner product subunits.

8. The system of claim 7 , wherein each of the one or more serial inner product subunits produces one output neuron.

9. The system of claim 8 , wherein the filter lanes of the synapse buffer are in communication with the corresponding serial inner product subunits via an interconnect.

10. The system of claim 9 , wherein the window lanes of the input neuron buffer are in communication with the corresponding serial inner product subunits via an interconnect.

11. The system of claim 8 , further comprising a synapse register for providing one or more synapse groups to the serial inner product subunits.

12. The system of claim 8 , wherein each serial inner product subunit comprises a multiple input adder tree.

13. The system of claim 12 , wherein each serial inner product subunit further comprises one or more negation blocks.

14. The system of claim 12 , wherein each serial inner product subunit further comprises a comparator.

15. The system of claim 1 , wherein the dispatcher comprises a shuffler to collect the neurons in one or more bricks and a transposer to convert the bricks into serial bit streams and wherein the shuffler comprises one or more multiplexers.

16. The system of claim 1 , wherein the synapses are communicated via a bit-parallel interface.

17. A system for bit-serial computation in a neural network, comprising:

one or more bit-serial tiles for performing bit-serial computations in a neural network wherein the one or more bit-serial tiles process two or more windows in parallel, each bit-serial tile receiving input neurons in two or more windows and synapses in two or more filters, and communicating output neurons;

an activation memory for storing neurons and in communication with the one or more bit-serial tiles via a dispatcher and a reducer,

wherein the dispatcher reads neurons from the activation memory and communicates the neurons to the one or more bit-serial tiles via two or more window lanes for each bit-serial tile,

and wherein the dispatcher reads synapses from a memory and communicates the synapses to the one or more bit-serial tiles via two or more filter lanes for each bit-serial tile,

and wherein the reducer receives the output neurons from the one or more bit-serial tiles, and communicates the output neurons to the activation memory;

and wherein the neurons and the synapses are communicated to the one or more bit-serial tiles bit-serially.

18. The system of claim 17 , wherein the dispatcher reduces the precision of an input synapse, based on a most significant bit value or a least significant bit value of the input neuron.

19. The system of claim 17 , wherein the dispatcher reduces the precision of the input synapse based on the most significant bit value and the least significant bit value of the input neuron.

20. An integrated circuit comprising a bit-serial neural network accelerator, the integrated circuit comprising:

one or more bit-serial tiles for performing bit-serial computations in a neural network wherein the one or more bit-serial tiles process two or more windows in parallel, each bit-serial tile receiving input neurons in two or more windows and synapses in two or more filters, and generating output neurons;

an activation memory for storing neurons and in communication with the one or more bit-serial tiles via a dispatcher and a reducer,

wherein the dispatcher reads neurons from the activation memory and communicates the neurons to the one or more bit-serial tiles via two or more window lanes for each bit-serial tile,

and wherein the dispatcher reads synapses from a memory and communicates the synapses to the one or more bit-serial tiles via two or more filter lanes for each bit-serial tile,

and wherein the reducer receives the output neurons from the one or more bit-serial tiles, and communicates the output neurons to the activation memory;

and wherein one of the neurons or the synapses are communicated to the one or more bit-serial tiles bit-serially and the other of the neurons or the synapses are communicated to the one or more bit-serial tiles bit-parallelly.

21. An integrated circuit comprising a bit-serial neural network accelerator, the integrated circuit comprising:

one or more bit-serial tiles for performing bit-serial computations in a neural network wherein the one or more bit-serial tiles process two or more windows in parallel, each bit-serial tile receiving input neurons and synapses, and communicating output neurons;

an activation memory for storing neurons and in communication with the one or more bit-serial tiles via a dispatcher and a reducer,

wherein the dispatcher reads neurons from the activation memory and communicates the neurons to the one or more bit-serial tiles via two or more window lanes for each bit-serial tile,

and wherein the dispatcher reads synapses from a memory and communicates the synapses to the one or more bit-serial tiles via two or more filter lanes for each bit-serial tile,

and wherein the reducer receives the output neurons from the one or more bit-serial tiles, and communicates the output neurons to the activation memory;

and wherein the neurons and the synapses are communicated to the one or more bit-serial tiles bit-serially.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: TARTAN AI LTD.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059516/0525 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2021
From: JUDD, PATRICK; ALBERICIO, JORGE; MOSHOVOS, ANDREAS; SHARIFY, SAYEH; DELMAS LASCORZ, ALBERTO
To: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
Reel/Frame 055089/0114 →
NUNC PRO TUNC ASSIGNMENT Recorded Jan 31, 2021
From: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
To: TARTAN AI LTD.
Reel/Frame 055089/0185 →
Continuity (9)
Continuation 15606118 · May 26, 2017
Provisional Application 62490659 · Apr 27, 2017
Provisional Application 62454268 · Feb 3, 2017
Provisional Application 62448454 · Jan 20, 2017
Provisional Application 62416782 · Nov 3, 2016
Provisional Application 62395027 · Sep 15, 2016
Provisional Application 62381202 · Aug 30, 2016
Provisional Application 62341814 · May 26, 2016
Related Publication 20200125931A1 · Apr 23, 2020