IP Library Granted Patent US 12,725,018
Granted Patent B2
US 12,725,018 · App. 16/968,678 · Granted Sep 1, 2026

Neural network accelerator

Inventors: Andreas Moshovos (Toronto, CA); Alberto Delmas Lascorz (Toronto, CA); Zisis Poulos (Toronto, CA); Dylan Malone Stuart (Toronto, CA); Patrick Judd (Merlin, CA); Sayeh Sharifymoghaddam (Toronto, CA); Mostafa Mahmoud (Toronto, CA); Milos Nikolic (Toronto, CA); Kevin Chong Man Siu (Toronto, CA); Jorge Albericio (Los Altos Hills, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/063G06F13/4282
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,018
App. No.
16/968,678
Granted
Sep 1, 2026
Kind
B2
Abstract

Described is a neural network accelerator tile for exploiting input sparsity. The tile includes a weight memory to supply each weight lane with a weight and a weight selection metadata, an activation selection unit to receive a set of input activation values and rearrange the set of input activation values to supply each activation lane with a set of rearranged activation values, a set of multiplexers including at least one multiplexer per pair of activation and weight lanes, where each multiplexer is configured to select a combination activation value for the activation lane from the activation lane set of rearranged activation values based on the weight lane weight selection metadata, and a set of combination units including at least one combination unit per multiplexer, where each combination unit is configured to combine the activation lane combination value with the weight lane weight to output a weight lane product.

Claims (66)

1 . A neural network accelerator tile for exploiting input sparsity defining a set of weight lanes and a set of activation lanes, each weight lane corresponding to an activation lane, the neural network accelerator tile comprising:

a weight memory configured to supply each weight lane of the set of weight lanes with a weight and a weight selection metadata of the respective weight lane weight indicating a position of the weight lane weight in a weight schedule, the position of the weight lane weight indicating one or more cycles during which the weight lane weight is to be combined with one or more input activation values determined based on the activation lanes corresponding to the one or more input activation values;

an activation selection unit comprising buffers configured to receive a set of input activation values and activation multiplexers configured to rearrange the set of input activation values based on the position of the weight lane weight in the weight schedule, to supply each activation lane with a set of rearranged activation values;

a set of multiplexers, the set of multiplexers including at least one multiplexer per pair of activation and weight lanes, each multiplexer configured to select a combination activation value for the activation lane from the activation lane set of rearranged activation values based on the weight lane weight selection metadata of the respective weight lane weight indicating the position of the weight lane weight in the weight schedule; and

a set of combination units, the set of combination units including at least one combination unit per multiplexer, each combination unit comprising at least one of a multiplier, an adder, and a shifter configured to combine the activation lane combination value with the weight lane weight to output a weight lane product,

wherein a first multiplexer of the set of multiplexers selects a first rearranged activation value associated with a first weight lane weight from one of an activation lane for a lookahead operation of a first weight lane and an activation lane for a lookaside operation of the first weight lane based on a first weight selection metadata of the first weight lane,

wherein the lookahead operation is an operation to remove ineffectual weights of a first weight schedule of the first weight lane using first weight lane weights of the first weight lane,

wherein the lookaside operation is an operation to remove ineffectual weights of the first weight schedule of the first weight lane using second weight lane weights of a second weight lane, and

wherein each effective weight of the first weight schedule of the first weight lane has a corresponding piece of metadata to identify the position of the effective weight in an original dense weight schedule such that the effective weight is matched at runtime with a corresponding activation value.

2 . The neural network accelerator tile of claim 1 , further comprising an activation memory configured to supply the set of input activation values to the activation selection unit.

3 . The neural network accelerator tile of claim 1 , wherein each multiplexer of the set of multiplexers is configured to select the combination activation from the corresponding set of rearranged activation values and from a set of additional lane activation values, the set of additional lane activation values formed of at least one rearranged activation value of at least one additional activation lane.

4 . The neural network accelerator tile of claim 1 , further comprising an adder tree configured to receive at least two eight lane products.

5 . The neural network accelerator tile of claim 1 , wherein the weight lane weights of the set of weight lanes define at least one neural network filter.

6 . The neural network accelerator tile of claim 1 , wherein the combination unit is one of a multiplier, an adder, and a shifter.

7 . A neural network accelerator comprising at least two tiles of claim 1 .

8 . The neural network accelerator tile of claim 1 , wherein each set of rearranged activation values includes a standard weight activation value and at least one lookahead activation value.

9 . The neural network accelerator tile of claim 1 , implemented on an activation efficiency exploiting accelerator structure.

10 . The neural network accelerator tile of claim 1 , wherein the set of input activation values are activation bits.

11 . The neural network accelerator tile of claim 1 , wherein the set of input activation values are signed powers of two.

12 . The neural network accelerator tile of claim 3 , wherein the set of multiplexers is a set of multiplexers of a uniform size.

13 . The neural network accelerator tile of claim 12 , wherein the uniform size is a power of two.

14 . The neural network accelerator tile of claim 13 , wherein the size of the set of rearranged activation values is larger than the size of the set of additional lane activation values.

15 . The neural network accelerator tile of claim 12 , wherein the set of rearranged activation values and the set of additional lane activation values are combined to generate a combined set of activation values, and the combined set of activation values contains 8 activations.

16 . The neural network accelerator tile of claim 3 , wherein the set of additional lane activation values is formed of at least one rearranged activation value from each of at least two additional activation lanes.

17 . The neural network accelerator tile of claim 16 , wherein the at least two additional activation lanes are non-contiguous activation lanes.

18 . The neural network accelerator tile of claim 1 , wherein the neural network accelerator tile is configured to receive the set of input activation values as at least one set of packed activation values stored bitwise to a required precision defined by a precision value, the neural network accelerator tile configured to unpack the at least one set of packed activation values.

19 . The neural network accelerator tile of claim 18 , wherein the at least one set of packed activation values includes a first set of packed activation values and a second set of packed activation values, the first set of packed activation values stored bitwise to a first required precision defined by a first precision value and the second set of packed activation values stored bitwise to a second required precision defined by a second precision value, the first precision value independent of the second precision value.

20 . The neural network accelerator tile of claim 18 , wherein the neural network accelerator tile is configured to receive a set of bit vectors including a bit vector corresponding to each set of packed activation values of the set of input activation values, the neural network accelerator tile configured to unpack each set of packed activation values to insert zero values as indicated by the corresponding bit vector.

21 . The neural network accelerator tile of claim 1 , wherein the neural network accelerator tile is configured to receive the weight lane weights of the set of weight lanes as at least one set of packed weight lane weights stored bitwise to a required precision defined by a precision value, the neural network accelerator tile configured to unpack the at least one set of weight lane weights.

22 . The neural network accelerator tile of claim 1 , wherein the set of activation lanes is at least two sets of column activation lanes, each set of column activation lanes forming a column in which each activation lane corresponds to a weight lane, the neural network accelerator tile further including at least one connection between at least two columns to transfer at least one weight lane product between the columns.

23 . A system for bit-serial computation in a neural network, comprising:

one or more bit-serial tiles configured according to claim 1 for performing bit-serial computations in a neural network, each bit-serial tile receiving input neurons and synapses, the input neurons including at least one set of input activation values and the synapses including at least one set of weights and at least one set of weight selection metadata, the one or more bit-serial tiles generating output neurons, each output neuron formed using at least one weight lane product;

an activation memory for storing neurons and in communication with the one or more bit-serial tiles via a dispatcher and a reducer,

wherein the dispatcher reads neurons from the activation memory and communicates the neurons to the one or more bit-serial tiles via a first interface,

and wherein the dispatcher reads synapses from a memory and communicates the synapses to the one or more bit-serial tiles via a second interface;

and wherein the reducer receives the output neurons from the one or more bit-serial tiles, and communicates the output neurons to the activation memory via a third interface;

and wherein one of the first interface and the second interface communicates the neurons or the synapses to the one or more bit-serial tiles bit-serially and the other of the first interface and the second interface communicates the neurons or the synapses to the one or more bit-serial tiles bit-parallelly.

24 . A system for computation of layers in a neural network, comprising:

one or more tiles configured according to claim 1 for performing computations in a neural network, each tile receiving input neurons and synapses, the input neurons each including at least one offset, each offset including at least one activation value, and the synapses including at least one set of weights and at least one set of weight selection metadata, the one or more tiles generating output neurons, each output neuron formed using least one weight lane product;

an activation memory for storing neurons and in communication with the one or more tiles via a dispatcher and an encoder,

wherein the dispatcher reads neurons from the activation memory and communicates the neurons to the one or more tiles, and wherein the dispatcher reads synapses from a memory and communicates the synapses to the one or more tiles,

and wherein the encoder receives the output neurons from the one or more tiles, encodes them and communications the output neurons to the activation memory;

and wherein the offsets are processed by the tiles in order to perform computations on only non-zero neurons.

25 . Use of the neural network accelerator tile of claim 1 for training.

26 . The neural network accelerator tile of claim 1 , wherein the weight lane weight selection metadata indexes a table that specifies a multiplexer select signal.

27 . An accelerator tile, comprising:

an activation selection unit comprising activation buffers configured to receive a set of activation values and activation multiplexers configured to rearrange the set of activation values into at least one set of multiplexer input values based on at least one weight selection metadata of at least one weight indicating a position of the at least one weight in a weight schedule, the position of the at least one weight indicating one or more cycles during which the at least one weight is to be combined with one or more input activation values determined based on activation lanes of the one or more input activation values;

a set of weight value receptors comprising weight buffers configured to receive at least one weight and the at least one weight selection metadata;

at least one multiplexer configured to receive at least one of the at least one set of multiplexer input values and the at least one weight selection metadata, the at least one multiplexer configured to apply the at least one weight selection metadata to select at least one combination activation value from the at least one set of multiplexer input values;

at least one combinator comprising at least one of a multiplier, an adder, and a shifter configured to apply the at least one combination activation value to the at least one weight to produce at least one product, wherein the at least one product is output; and

a first multiplexer of a set of multiplexers configured to select a first rearranged activation value associated with a first weight lane weight from one of an activation lane for a lookahead operation of a first weight lane and an activation lane for a lookaside operation of the first weight lane based on a first weight selection metadata of the first weight lane,

wherein the lookahead operation is an operation to remove ineffectual weights of a first weight schedule of the first weight lane using first weight lane weights of the first weight lane, and

wherein the lookaside operation is an operation to remove ineffectual weights of the first weight schedule of the first weight lane using second weight lane weights of a second weight lane, and

wherein each effective weight of the first weight schedule of the first weight lane has a corresponding piece of metadata to identify the position of the effective weight in an original dense weight schedule such that the effective weight is matched at runtime with a corresponding activation value.

28 . A neural network accelerator comprising the accelerator tile of claim 27 .

29 . The accelerator tile of claim 27 , further including an activation memory configured to supply the set of activation values to the activation selection unit.

30 . The accelerator tile of claim 27 , wherein the at least one set of multiplexer input values is at least two sets of multiplexer input values and the at least one multiplexer is configured to receive at least one of the at least two sets of multiplexer input values and at least one activation value from at least one other set of multiplexer input values.

31 . The accelerator tile of claim 27 , wherein the combinator is at least one of a multiplier, an adder, and a shifter.

32 . The accelerator tile of claim 27 , wherein each set of multiplexer input values includes a standard activation value and at least one lookahead activation value.

33 . The accelerator tile of claim 27 , implemented on an activation efficiency exploiting accelerator structure.

34 . The accelerator tile of claim 27 , wherein the set of activation values are activation bits.

35 . The accelerator tile of claim 27 , wherein the set of activation values are signed powers of two.

36 . The accelerator tile of claim 27 , wherein the size of each multiplexer of the at least one multiplexer is a power of two.

37 . The accelerator tile of claim 36 , wherein the size of each multiplexer of the at least one multiplexer is 8.

38 . Use of the accelerator tile of claim 27 for training.

39 . The accelerator tile of claim 27 , wherein the weight selection metadata indexes a table that specifies a multiplexer select signal.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2026
From: ALBERICIO, JORGE
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 075263/0308 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: TARTAN AI LTD.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059516/0525 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2021
From: MOSHOVOS, ANDREAS; DELMAS LASCORZ, ALBERTO; POULOS, ZISIS PARASKEVAS; MALONE STUART, DYLAN; JUDD, PATRICK; SHARIFY, SAYEH; MAHMOUD, MOSTAFA; NIKOLIC, MILOS; SIU, KEVIN CHONG MAN
To: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
Reel/Frame 055089/0164 →
NUNC PRO TUNC ASSIGNMENT Recorded Jan 31, 2021
From: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
To: TARTAN AI LTD.
Reel/Frame 055089/0185 →
Continuity (3)
Provisional Application 62664190 · Apr 29, 2018
Provisional Application 62710488 · Feb 16, 2018
Related Publication 20210004668A1 · Jan 7, 2021
References Cited (37)
US 5751913A · Chiueh · 1998 [cited by applicant]
US 6199057B1 · Tawel · 2001 [cited by applicant]
US 9710265B1 · Temam · 2017 [cited by applicant]
US 9818059B1 · Woo · 2017 [cited by examiner]
US 10127494B1 · Cantin · 2018 [cited by examiner]
US 10417555B2 · Brothers · 2019 [cited by examiner]
US 10467795B2 · Sarel · 2019 [cited by examiner]
US 10521488B1 · Ross · 2019 [cited by examiner]
US 20150310311A1 · Shi · 2015 [cited by applicant]
US 20160342893A1 · Ross · 2016 [cited by examiner]
US 20160358069A1 · Brothers · 2016 [cited by examiner]
US 20160364644A1 · Brothers · 2016 [cited by examiner]
US 20170344876A1 · Brothers · 2017 [cited by examiner]
US 20180046900A1 · Dally · 2018 [cited by examiner]
US 20180089562A1 · Jin · 2018 [cited by examiner]
US 20180129935A1 · Kim · 2018 [cited by examiner]
US 20180173571A1 · Huang · 2018 [cited by examiner]
US 20180197067A1 · Mody · 2018 [cited by examiner]
US 20180218518A1 · Yan · 2018 [cited by examiner]
US 20190050734A1 · Li · 2019 [cited by examiner]
US 20190171634A1 · Nowakiewicz · 2019 [cited by examiner]
CN 107533667A · 2018 [cited by applicant]
WO 2017214728A1 · 2017 [cited by applicant]
Albericio et al. Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing, 2016, 13 pages (Year: 2016). [cited by examiner]
Judd et al. Proteus: Exploiting Numerical Precision Variability in Deep Neural Networks, Jun. 2016, 12 pages (Year: 2016). [cited by examiner]
Judd et al. Stripes: Bit-Serial Deep Neural Network Computing, 2016, 12 pages (Year: 2016). [cited by examiner]
Parash et al., “SCNN: An accelerator for compressed-sparse convolutional neural networks”, ISCA '17: Proceedings of the 44th Annual International Symposium on Computer Architecture, Association for Computing Machinery, … [cited by applicant]
Qin et al., “SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training”, 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), IEEE, San Diego, CA, USA, 20… [cited by applicant]
Gupta et al., “MASR: A Modular Accelerator for Sparse RNNs”, 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT), IEEE, Seattle, WA, USA, 2019, pp. 1-14. [cited by applicant]
Gondimalla et al., “SparTen: A Sparse Tensor Accelerator for Convolutional Neural Networks”, Micro '52: Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, Association for Computing Mac… [cited by applicant]
Kung et al., “Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization”, Proceedings of the Twenty-Fourth International Conference on Architect… [cited by applicant]
Zhang et al., “Cambricon-X: An accelerator for sparse neural networks”, 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (Micro), IEEE, Taipei, Taiwan, 2016, pp. 1-12. [cited by applicant]
Han et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), IEEE, Seoul, South Korea, 2016, pp. 243-254. [cited by applicant]
Kim et al., “ZeNA: Zero-Aware Neural Network Accelerator”, IEEE Design & Test, vol. 35, No. 1, IEEE, Feb. 2018, pp. 39-46. [cited by applicant]
Han et al., “ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA”, FPGA '17: Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Association for Computing Machinery… [cited by applicant]
Singapore Office Action issued on Jun. 29, 2022, in counterpart Singapore Patent Application No. 11202007532T (9 pages in English). [cited by applicant]
Chinese Office Action issued on Jan. 5, 2024, in counterpart Chinese Patent Application No. 201980014141.X (7 pages in English, 6 pages in Chinese). [cited by applicant]