IP Library Granted Patent US 12,299,559
Granted Patent B2
US 12,299,559 · App. 17/052,178 · Granted May 13, 2025

Neural network processing element of accelerator tile

Inventors: Andreas Moshovos (Toronto, CA); Mostafa Mahmoud (Toronto, CA); Sayeh Sharifymoghaddam (Toronto, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/063G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,559
App. No.
17/052,178
Granted
May 13, 2025
Kind
B2
Abstract

Described is a neural network accelerator tile. It includes an activation memory interface for interfacing with an activation memory to receive a set of activation representations and a weight memory interface for interfacing with a weight memory to receive a set of weight representations, and a processing element. The processing element is configured to implement a one-hot encoder, a histogrammer, an aligner, a reducer, and an accumulation sub-element which process the set of activation representations and the set of weight representations to produce a set of output representations.

Claims (33)

1. A neural network accelerator tile, comprising:

an activation memory interface for interfacing with an activation memory to receive a set of activation representations, wherein the set of activation representations is represented by an exponential expression based on activation exponent values;

a weight memory interface for interfacing with a weight memory to receive a set of weight representations, wherein the set of weight representations is represented by an exponential expression based on weight exponent values; and

a processing element configured to implement an exponent sub-element, a one-hot encoder, a histogrammer, an aligner, a reducer, and an accumulation sub-element to process the set of activation representations and the set of weight representations to produce a set of output representations,

wherein the exponent sub-element performs addition operations using the activation exponent values and the weight exponent values to generate exponent values.

2. The accelerator tile of claim 1 , wherein the activation memory interface is configured to provide the set of activation representations to the processing element as a set of activation one-offset pairs, and the weight memory interface is configured to provide the set of weight representations to the processing element as a set of weight one-offset pairs.

3. The accelerator tile of claim 2 , wherein the exponent sub-element is configured to combine the set of activation one-offset pairs and the set of weight one-offset pairs to produce a set of one-offset pair products.

4. The accelerator tile of claim 3 , wherein the exponent sub-element includes a set of magnitude adders and a corresponding set of sign gates, one pair of magnitude adder and sign gate to provide each one-offset pair product of the set of one-offset pair products.

5. The accelerator tile of claim 1 , wherein the one-hot encoder includes a set of decoders to perform one-hot encoding of a set of one-hot encoder inputs.

6. The accelerator tile of claim 2 , wherein the histogrammer is configured to sort a set of histogrammer inputs by value.

7. The accelerator tile of claim 2 , wherein the aligner is configured to shift a set of aligner inputs to provide a set of shifted inputs to be reduced.

8. The accelerator tile of claim 2 , wherein the reducer includes an adder tree to reduce a set of reducer inputs into a partial sum.

9. The accelerator tile of claim 2 , wherein the accumulation sub-element includes an accumulator and is configured to receive a partial sum and add the partial sum to the accumulator to accumulate a product over multiple cycles.

10. The accelerator tile of claim 1 , wherein the processing element is further configured to implement a concatenator.

11. The accelerator tile of claim 10 , wherein the concatenator is provided to concatenate a set of concatenator inputs to produce a set of grouped counts to be shifted and reduced by the aligner and reducer to produce a partial product.

12. A neural network comprising a set of neural network accelerator tiles according to claim 1 .

13. The neural network of claim 12 , wherein the neural network is implemented as a convolutional neural network.

14. The neural network of claim 12 , wherein the set of neural network accelerator tiles includes a column of accelerator tiles, the tiles of the column of accelerator tiles configured to each receive the same set of activation representations and to each receive a unique set of weight representations.

15. The accelerator tile of claim 1 , wherein the processing element is configured to, over multiple cycles, process an activation representation of the set of activation representations with a weight representation of the set of weight representations.

16. A method of producing a neural network partial product, implementing processing components including an exponent sub-element, a one-hot encoder, a histogrammer, an aligner and a reducer, the method comprising:

receiving a set of activation representations, wherein the set of activation representations is represented by an exponential expression based on activation exponent values;

receiving a set of weight representations, each weight representation corresponding to an activation representation of the set of activation representations, wherein the set of weight representations is represented by an exponential expression based on weight exponent values;

generating, by the exponent sub-element, a set of partial exponent value results by combining the set of weight representations with the set of activation representations by combining each weight representation with its corresponding activation representation;

generating, by the one-hot encoder, a set of one-hot representations by encoding the set of partial exponent value results;

accumulating, by the histogrammer, the set of one-hot representations into a set of histogram bucket counts;

aligning, by the aligner, the counts of the set of histogram bucket counts according to their size; and

generating, by the reducer, the neural network partial product by reducing the aligned counts of the set of histogram bucket counts,

wherein the combination of the set of weight representations with the set of activation representations is based on an addition operation using the activation exponent values and the weight exponent values, and

wherein the aligner and the reducer operate based on an output, of the histogrammer, that is generated by the histogrammer to not ha q non-zero bits.

17. The method of claim 16 , further comprising outputting the neural network partial product to an accumulator to accumulate a product.

18. The method of claim 17 , further comprising outputting the product to an activation memory.

19. The method of claim 16 , further comprising, prior to aligning the counts of the set of histogram bucket counts, recursively grouping the counts of the set of histogram bucket counts into a set of grouped counts and providing the set of grouped counts to be aligned and reduced in place of the set of histogram bucket counts.

20. The method of claim 16 , wherein combining a weight representation of the set of weight representations with an activation representation of the set of activation representations is performed over multiple cycles.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: TARTAN AI LTD.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059516/0525 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2021
From: MAHMOUD, MOSTAFA; MOSHOVOS, ANDREAS; SHARIFY, SAYEH
To: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
Reel/Frame 055089/0150 →
NUNC PRO TUNC ASSIGNMENT Recorded Jan 31, 2021
From: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
To: TARTAN AI LTD.
Reel/Frame 055089/0185 →
Continuity (2)
Provisional Application 62668363 · May 8, 2018
Related Publication 20210125046A1 · Apr 29, 2021
References Cited (38)
US 9928460B1 · Nowatzyk et al. · 2018 [cited by applicant]
US 20100146030A1 · Erle · 2010 [cited by examiner]
US 20130159367A1 · Dubrovin · 2013 [cited by examiner]
US 20180246855A1 · Redfern · 2018 [cited by examiner]
CA 2990709C · 2017 [cited by applicant]
CA 2990712C · 2017 [cited by applicant]
CN 107851214A · 2018 [cited by applicant]
WO WO2017201627A1 · 2017 [cited by applicant]
WO WO2017214728A1 · 2017 [cited by applicant]
Albericio et al. (Bit-Pragmatic Deep Neural Network Computing, Oct. 2016, pp. 1-12) (Year: 2016). [cited by examiner]
Liu et al. (Design of Approximate Radix-4 Booth Multipliers for Error-Tolerant Computing, Aug. 2017, pp. 1435-1441) (Year: 2017). [cited by examiner]
Zhao et al. (An energy-efficient coarse grained spatial architecture for convolutional neural networks AlexNet, Jul. 2017, pp. 1-12) (Year: 2017). [cited by examiner]
He et al. (A New Redundant Binary Booth Encoding for Fast 2n-Bit Multiplier Design, Jun. 2009, pp. 1192-1201) (Year: 2009). [cited by examiner]
Ganesh et al. (Constructing a low power multiplier using Modified Booth Encoding Algorithm in redundant binary number system, May 2012, pp. 2734-2740) (Year: 2012). [cited by examiner]
Efstathiou et al. (Efficient modulo 2n+1 multiply and multiply-add units based on modified booth encoding, Apr. 2013, pp. 140-147) (Year: 2013). [cited by examiner]
International Search Report and Written Opinion, PCT/CA2019/050525, Jul. 16, 2019. [cited by applicant]
Knagge, Geoff, “Booth Recoding”. GeotlKnagge.com, Jul. 27, 2010 (Jul. 27, 2010),, [online] [retrieved on Jun. 28, 2019 (Jun. 28, 2019)]. Retrieved from the Internet: <https://web.archive.org/web/20160303211337/http://ww… [cited by applicant]
Albericio, A. Delmás, P. Judd, S. Sharify, G. O'Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network computing,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture, M… [cited by applicant]
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. Enright Jerger,and A. Moshovos, “CNVLUTIN: Ineffectual-Neuron-Free DeepNeural Network Computing,” in Proceedings of the InternationalSymposium on Computer Architectu… [cited by applicant]
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan,B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, Scnn: An accelerator for compressed-sparse convolutional neural networks,in Proceedings of the 44th Annu… [cited by applicant]
P. Judd, J. Albericio, T. Hetherington, T. Aamodt, and A. Moshovos, Stripes: Bit-serial Deep Neural Network Computing ,in Proceedings of the 49th Annual IEEE/ACM International Symposium on Microarchitecture, MICRO-49, 2… [cited by applicant]
A. Delmas, P. Judd, S. Sharify, and A. Moshovos, Dynamic stripes: Exploiting the dynamic precision requirements of activation values in neural networks, CoRR, vol. abs/1706.00504, 2017. [cited by applicant]
S. Sharify, A. D. Lascorz, P. Judd, and A. Moshovos, Loom: Exploiting weight and activation precisions to accelerate convolutional neural networks,CoRR, vol. abs/1706.07853, 2017. [cited by applicant]
Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, Dadiannao: A machine-learning supercomputer,in Microarchitecture (MICRO), 2014 47th Annual IEEE/ACM International Symposium on… [cited by applicant]
Synopsys, Design Compiler Graphical. http://www.https://www.synopsys.com/implementation-and-signoff/rtl-synthesis-test/design-compiler-graphical.html., downloaded on Feb. 26, 2021. [cited by applicant]
Cadence, “Encounter rtl compiler. https://www.cadence.com/zh_CN/home/company/newsroom/press-releases/pr/2004/cadencedeliversencounterrtlcompilerultrawithsupportforvhdl.html”, downloaded on Feb. 26, 2021. [cited by applicant]
N. Muralimanohar and R. Balasubramonian, Cacti 6.0: A tool to understand large caches,2015. [cited by applicant]
M. Poremba, S. Mittal, D. Li, J. Vetter, and Y. Xie, Destiny: A tool for modeling emerging 3d nvm and edram caches,in Design, Automation Test in Europe Conference Exhibition (DATE), 2015, pp. 1543-1546, Mar. 2015. [cited by applicant]
Yang, Tien-Ju and Chen, Yu-Hsin and Sze, Vivienne, Designing Energy-Efi)cient Convolutional Neural Networks using Energy-Aware Pruning,n IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. [cited by applicant]
A. Delmas, S. Sharify, P. Judd, M. Nikolic, and A. Moshovos, Dpred: Making typical activation values matter in deep learning computing,CoRR, vol. abs/1804.06732, 2018. [cited by applicant]
Ghodrati et al., “Bit-Parallel Vector Composability for Neural Acceleration”, 2020 57th ACM/IEEE Design Automation Conference (DAC), IEEE, San Francisco, CA, USA, Jul. 2020, pp. 1-6. [cited by applicant]
“Cadence Delivers Encounter RTL Compiler Ultra with Support for VHDL Cadence Provides Superior Quality of Silicon throughout the Design Chain”, Cadence.com, Cadence Design Systems, Inc., San Jose, CA, USA, Feb. 17, 2004… [cited by applicant]
Awad et al., “FPRaker: A Processing Element for Accelerating Neural Network Training”, ArXiv.org, ArXiv, Oct. 15, 2020. [cited by applicant]
Delmas Lascorz et al., “Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks”, ASPLOS '19: Proceedings of the Twenty-Fourth International Conference on Architectural Support… [cited by applicant]
Albericio, Jorge, et al. “Bit-Pragmatic Deep Neural Network Computing.” Proceedings of the 50th annual IEEE/ACM international symposium on microarchitecture. 2017, (13 pages). [cited by applicant]
Garland, James, et al. “Low Complexity Multiply-Accumulate Units for Convolutional Neural Networks with Weight-Sharing.” ACM Transactions on Architecture and Code Optimization (TACO) 15.3 (2018): 1-24. [cited by applicant]
Delmas, Alberto, et al. “Bit-tactical: Exploiting Ineffectual Computations in Convolutional Neural Networks: Which, Why, and How.” arXiv preprint arXiv:1803.03688, Mar. 9, 2018, (14 pages). [cited by applicant]
Chinese Office Action issued on Jan. 5, 2024, in counterpart Chinese Patent Application No. 201980031107.3 (10 pages in English, 9 pages in Chinese). [cited by applicant]