IP Library › Granted Patent US 12,242,894
Granted Patent B2
US 12,242,894 · App. 18/193,635 · Granted Mar 4, 2025

Technique for hardware activation function computation in RNS artificial neural networks

Inventors: Athanasios Stouraitis (Abu Dhabi, AE); Sakellariou Vasileios (Abu Dhabi, AE); Vasileios Paliouras (Patras, GR); Ioannis Kouretas (Patras, GR); Hani Saleh (Abu Dhabi, AE)
Assignee: Khalifa University of Science and Technology
G06F9/5027G06F7/72
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,894
App. No.
18/193,635
Granted
Mar 4, 2025
Kind
B2
Abstract

A device can be used to implement a neural network in hardware. The device can include a processor, a memory, and a neural network accelerator. The neural network accelerator can be configured to implement, in hardware, a neural network by using a residue number system (RNS). At least one function of the neural network can have a corresponding approximation in the RNS system, and the at least one function can be provided by implementing the corresponding approximation in hardware.

Claims (43)

1. A device comprising:

a processor;

a non-transitory computer-readable memory comprising instructions executable by the processor to cause the processor to perform one or more operations associated with at least one of an input to or an output from a neural network; and

a neural network accelerator configured to implement, in hardware, at least a part of the neural network by using a residue number system (RNS), wherein:

at least one function of the neural network has a corresponding approximation in the RNS,

the at least one function is provided by implementing the corresponding approximation in hardware, and

the operations include:

receiving the input,

performing a base extension on the input,

generating a mapped value based on the base extension, and

determining an index by at least using the mapped value and a lookup table operation.

2. The device of claim 1 , wherein the corresponding approximation includes a piecewise linear approximation, and wherein the piecewise linear approximation is configured to minimize a maximum approximation error.

3. The device of claim 1 , wherein the at least one function comprises at least one of a tanh function or a sigmoid function, and wherein the corresponding approximation includes at least one of a scaling operation or a comparison operation.

4. The device of claim 1 , wherein the corresponding approximation is configured to partition a domain of the at least one function into a plurality of successive intervals by a sequence of points and to use the sequence of points for approximation.

5. The device of claim 4 , wherein the corresponding approximation is configured to use a first factor at and a second factor b i for an interval of the plurality of successive intervals, wherein the first factor and the second factor are constrained for the interval to “1” and “0” when the at least one function is a tanh function and to “¼” and “½” when the at least one function is a sigmoid function.

6. The device of claim 4 , wherein the plurality of successive intervals includes five successive intervals.

7. The device of claim 1 , wherein the corresponding approximation is configured to perform a base extension of adding one or more changes to an RNS base for a division function.

8. A method implemented by a device that includes a neural network accelerator, the method comprising:

receiving data to be processed by a neural network, the neural network accelerator configured to implement, in hardware, at least a part of the neural network by using a residue number system (RNS), wherein at least one function of the neural network has a corresponding approximation in the RNS, and wherein the at least one function is provided by implementing the corresponding approximation in hardware, wherein the data is received in an RNS domain and is represented by a modulus set that comprises one or more residue representations of the data, and wherein the one or more residue representations comprise a representation range of the data;

performing a base extension on the representation range of the data to determine a last-channel offset between the data and mapped input data;

generating an input to the neural network accelerator based on the data; and

receiving an output of the neural network accelerator.

9. The method of claim 8 , wherein the neural network comprises a long short-term memory (LSTM) layer, wherein the data comprises sinusoidal data, and wherein receiving the data to be processed by the neural network comprises receiving the sinusoidal data via the LSTM layer.

10. The method of claim 8 , further comprising using a lookup table operation to determine a particular interval based on the base extension, wherein the lookup table operation involves distinguishing between a first interval of a plurality of intervals and a second interval of the plurality of intervals using the mapped input data.

11. The method of claim 8 , further comprising determining a plurality of intervals by partitioning the representation range of the data into a plurality of sub-intervals, wherein a number of sub-intervals included in the plurality of sub-intervals is equal to one or more values included in the modulus set.

12. The method of claim 8 , further comprising determining a plurality of intervals without converting the data to a binary representation, and wherein determining the plurality of intervals involve one-channel-wide operations.

13. A system comprising:

a first computing device; and

a second computing device communicatively coupled to the first computing device and configured to receive input data from the first computing device and generate output data to transmit to the first computing device, the second computing device comprising:

a processor;

a non-transitory computer-readable memory comprising instructions executable by the processor to cause the processor to perform one or more operations associated with at least one of an input to or an output from a neural network; and

a neural network accelerator configured to implement, in hardware, at least a part of the neural network by using a residue number system (RNS), wherein:

at least one function of the neural network has a corresponding approximation in the RNS,

the at least one function is provided by implementing the corresponding approximation in hardware, and

the one or more operations include:

receiving the input,

performing a base extension on the input,

generating a mapped value based on the base extension, and

determining an index by at least using the mapped value and a lookup table operation.

14. The system of claim 13 , wherein the corresponding approximation includes a piecewise linear approximation, and wherein the piecewise linear approximation is configured to minimize a maximum approximation error.

15. The system of claim 13 , wherein the at least one function comprises at least one of a tanh function or a sigmoid function, and wherein the corresponding approximation includes at least one of a scaling operation or a comparison operation.

16. The system of claim 13 , wherein the corresponding approximation is configured to partition a domain of the at least one function into a plurality of successive intervals by a sequence of points and to use the sequence of points for approximation.

17. The system of claim 16 , wherein the corresponding approximation is configured to use a first factor a i and a second factor b i for an interval of the plurality of successive intervals, wherein the first factor and the second factor are constrained for the interval to “1” and “0” when the at least one function is a tanh function and to “¼” and “½” when the at least one function is a sigmoid function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: STOURAITIS, ATHANASIOS; VASILEIOS, SAKELLARIOU; PALIOURAS, VASILEIOS; KOURETAS, IOANNIS; SALEH, HANI
To: KHALIFA UNIVERSITY OF SCIENCE AND TECHNOLOGY
Reel/Frame 063181/0288 →
Priority Claims (1)
GR 20220100431 · May 24, 2022 · national
Continuity (1)
Related Publication 20230385115A1 · Nov 30, 2023
References Cited (28)
US 6898613B1 · Robinson · 2005 [cited by examiner]
US 10387122B1 · Olsen · 2019 [cited by examiner]
Valueva, M.V et al., Application of the residue number system to reduce hardware costs of convolutional neural network implementation 2020, Elsevier,pp. 232-243. (Year: 2020). [cited by examiner]
Wang, X. et al., ReTanh: An activation function with vanishing gradient resistance for SAE-based DNNs and its application to rotating machinery fault diagnosis, 2019, Elsevier, pp. 88-98. (Year: 2019). [cited by examiner]
Abdelouahab, K. et al., PhD Forum:Why TahH is a Hardware Friendly Activation Function for CNNs, 2017,ACM, 3 pages. (Year: 2017). [cited by examiner]
Bank-Tavakoli POLAR: A Pipelined/Overlapped FPGA-Based LSTM Accelerator,2020,IEEE,pp. 838-842. (Year: 2020). [cited by examiner]
Chervyakov ,N.I. et al.,Residue Number System-Based Solution for Reducing the Hardware Cost of a Convolutional Neural Network.2020, Elsevier ,439-453. (Year: 2020). [cited by examiner]
Lin , Liang Yu. et al., Residue Number System Design Automation for Neural Network Acceleration, 2020,IEEE,2 pages. (Year: 2020). [cited by examiner]
M. Carreras, G. Deriu, and P. Meloni, “Flexible acceleration of convolutions on FPGAs: NEURAghe 2.0,” in CPS Summer School, PhD Workshop, 2019. [cited by applicant]
M. Zainab, A. R. Usmani, S. Mehrban, and M. Hussain, “FPGA Based Implementations of RNN and CNN: A Brief Analysis,” in 2019 International Conference on Innovative Computing (ICIC), Nov. 2019, pp. 1-8. [cited by applicant]
J. Sim, S. Lee, and L. Kim, “An energy-efficient deep convolutional neural network inference processor with enhanced output stationary dataflow in 65-nm CMOS,” IEEE Transactions on Very Large Scale Integration (VLSI) Sy… [cited by applicant]
E. Bank-Tavakoli, S. A. Ghasemzadeh, M. Kamal, A. Afzali-Kusha, and M. Pedram, “POLAR: A Pipelined/Overlapped FPGA-Based LSTM Accelerator,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 28, No. … [cited by applicant]
E. Azari and S. Vrudhula, “An Energy-Efficient Reconfigurable LSTM Accelerator for Natural Language Processing,” in 2019 IEEE Interna-tional Conference on Big Data (Big Data), Dec. 2019, pp. 4450-4459. [cited by applicant]
A. Marchisio, M. A. Hanif, F. Khalid, G. Plastiras, C. Kyrkou, T. Theocharides, and M. Shafique, “Deep Learning for Edge Computing: Current Trends, Cross-Layer Optimizations, and Open Research Chal-lenges,” in 2019 IEEE… [cited by applicant]
A. Munir, E. Blasch, J. Kwon, J. Kong, and A. Aved, “Artificial Intel-ligence and Data Fusion at the Edge,” IEEE Aerospace and Electronic Systems Magazine, vol. 36, No. 7, pp. 62-78, 2021. [cited by applicant]
M. Valueva, N. Nagornov, P. Lyakhov, G. Valuev, and N. Chervyakov, “Application of the residue number system to reduce hardware costs of the convolutional neural network implementation,” Mathematics and Computers in Sim… [cited by applicant]
N. Samimi, M. Kamal, A. Afzali-Kusha, and M. Pedram, “Res-DNN: A Residue Number System-Based DNN Accelerator Unit,” IEEE Trans- actions on Circuits and Systems I: Regular Papers, vol. 67, No. 2, pp. 658-671, 2020. [cited by applicant]
M. Abdelhamid and S. Koppula, “Applying the residue number system to network inference,” arXiv preprint arXiv:1712.04614, 2017. [cited by applicant]
S. Salamat, M. Imani, S. Gupta, and T. Rosing, “RNSnet: In-Memory Neural Network Acceleration Using Residue Number System,” 11 2018, pp. 1-12. [cited by applicant]
K. Greff, R. K. Srivastava, J. Koutn'ik, B. R. Steunebrink, and J. Schmid-huber, “LSTM: A Search Space Odyssey,” IEEE Transactions on Neural Networks and Learning Systems, vol. 28, No. 10, pp. 2222-2232, Oct. 2017. [cited by applicant]
Y. Kong and B. Phillips, “Fast scaling in the Residue Number System,” IEEE Trans. VLSI Syst., vol. 17, pp. 443-447, Mar. 2009. [cited by applicant]
E. Olsen, “RNS hardware matrix multiplier for high precision neural network acceleration: “RNS TPU”,” May 2018, pp. 1-5. [cited by applicant]
H. Vergos, C. Efstathiou, and D. Nikolos, “Diminished-one modulo 2n+ 1 adder design,” IEEE Transactions on Computers—TC, vol. 51, pp. 1389-1399, Jan. 2002. [cited by applicant]
C. Efstathiou, H. Vergos, G. Dimitrakopoulos, and D. Nikolos, “Efficient diminished1 modulo 2n + 1 multipliers,” IEEE Transactions on Computers—TC, vol. 54, pp. 491-496, Apr. 2005. [cited by applicant]
B. Liao, J. Zhang, C. Wu, D. McIlwraith, T. Chen, S. Yang, Y. Guo, and F. Wu, “Deep sequence learning with auxiliary information for traffic prediction,” in Proceedings of the 24th ACM SIGKDD International Conference on… [cited by applicant]
E. Azari and S. Vrudhula, “ELSA: a throughput-optimized design of an LSTM accelerator for energy-constrained devices,” ACM Transactions on Embedded Computing Systems (TECS), vol. 19, No. 1, pp. 1-21, 2020. [cited by applicant]
I. Kouretas and V. Paliouras, “Hardware aspects of long short term memory,” in 2018 25th IEEE International Conference on Electronics, Circuits and Systems (ICECS), 2018, pp. 525-528. [cited by applicant]
D. Kadetotad, V. Berisha, C. Chakrabarti, and J.-S. Seo, “A 8.93—TOPS/W LSTM recurrent neural network accelerator featuring hierarchical coarse-grain sparsity with all parameters stored on-chip,” IEEE Solid-State Circui… [cited by applicant]