IP Library Granted Patent US 12,400,112
Granted Patent B2
US 12,400,112 · App. 17/115,285 · Granted Aug 26, 2025

Efficient method for VLSI implementation of useful neural network activation functions

Inventors: Jun Sawada (Austin, TX); Myron D. Flickner (San Jose, CA); Andrew Stephen Cassidy (Austin, TX); John Vernon Arthur (San Jose, CA); Pallab Datta (San Jose, CA); Dharmendra S. Modha (San Jose, CA); Steven Kyle Esser (San Jose, CA); Brian Seisho Taba (Cupertino, CA); Jennifer Klamo (San Jose, CA); Rathinakumar Appuswamy (San Jose, CA); Filipp Akopyan (New Windsor, NY); Carlos Ortega Otero (Los Angeles, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/08G06N3/04G06N3/063G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,112
App. No.
17/115,285
Granted
Aug 26, 2025
Kind
B2
Abstract

A neural inference chip is provided, including at least one neural inference core. The at least one neural inference core is adapted to apply a plurality of synaptic weights to a plurality of input activations to produce a plurality of intermediate outputs. The at least one neural inference core comprises a plurality of activation units configured to receive the plurality of intermediate outputs and produce a plurality of activations. Each of the plurality of activation units is configured to apply a configurable activation function to its input. The configurable activation function has at least a re-ranging term and a scaling term, the re-ranging term determining the range of the activations and the scaling term determining the scale of the activations. Each of the plurality of activations units is configured to obtain the re-ranging term and the scaling term from one or more look up tables.

Claims (30)

1. A neural inference chip comprising:

at least one neural inference core adapted to apply a plurality of synaptic weights to a plurality of input data tensors to produce a plurality of intermediate output data tensors,

the at least one neural inference core comprising a plurality of activation units configured to receive the plurality of intermediate output data tensors and produce a plurality of activations,

each of the plurality of activation units being configured to apply a configurable activation function to its input data tensor,

the configurable activation function having at least a re-ranging term and a scaling term, the re-ranging term applying a bias that determines a range of the plurality of activations and the scaling term applying a slope that determines a scale of the plurality of activations, wherein the re-ranging term and the scaling term are obtained separately from predetermined sets, each predetermined set being associated with one of the re-ranging term or the scaling term,

each of the plurality of activations units obtaining the re-ranging term and the scaling term from one or more look up tables.

2. The neural inference chip of claim 1 , wherein the plurality of activations have flexible precision.

3. The neural inference chip of claim 1 , wherein the plurality of activations have a floating point value.

4. The neural inference chip of claim 1 , wherein the plurality of input data tensors have 16 -bit precision.

5. The neural inference chip of claim 1 , wherein the plurality of input data tensors have 32 -bit precision.

6. The neural inference chip of claim 1 , wherein each predetermined set is a lookup table.

7. The neural inference chip of claim 1 , wherein each of the re-ranging term and the scaling term are learned.

8. The neural inference chip of claim 1 , wherein the configurable activation function is selected from a list consisting of: Boolean, trinary, linear, ReLU, shifted ReLU, ExpReLU, sigmoid, and tanh.

9. An integrated circuit, comprising:

at least one neural inference core, adapted to apply a plurality of synaptic weights to a plurality of input data tensors to produce a plurality of intermediate output data tensors,

the at least one neural inference core comprising a plurality of activation units configured to receive the plurality of intermediate output data tensors and produce a plurality of activations,

each of the plurality of activation units being configured to apply an activation function to its input data tensor, the activation function being Boolean, trinary, linear, ReLU, shifted ReLU, ExpReLU, sigmoid, or tanh,

the activation function having at least a re-ranging term and a scaling term, the re-ranging term applying a bias that determines a range of the plurality of activations and the scaling term applying a slope that determines a scale of the plurality of activations, each of the re-ranging term and the scaling term having an associated lookup table,

each of the plurality of activations units obtaining the re-ranging term and the scaling term separately from the associated lookup tables.

10. A computer-implemented method, comprising:

applying a plurality of synaptic weights to a plurality of input data tensors to produce a plurality of intermediate output data tensors;

receiving the plurality of intermediate output data tensors and produce therefrom a plurality of activations, wherein producing the plurality of activations comprises:

applying a configurable activation function to the plurality of intermediate output data tensors, the configurable activation function having at least a re-ranging term and a scaling term, the re-ranging term applying a bias that determines a range of the plurality of activations and the scaling term applying a slope that determines a scale of the plurality of activations, each of the re-ranging term and the scaling term having an associated lookup table, and obtaining the re-ranging term and the scaling term separately from the associated lookup tables.

11. The method of claim 10 , wherein the plurality of activations have flexible precision.

12. The method of claim 10 , wherein the plurality of activations have a floating point value.

13. The method of claim 10 , wherein the plurality of input data tensors have 16-bit precision.

14. The method of claim 10 , wherein the plurality of input data tensors have 32-bit precision.

15. The method of claim 10 , wherein each of the re-ranging term and the scaling term have one associated lookup table.

16. The method of claim 10 , wherein each of the re-ranging term and the scaling term are learned.

17. The method of claim 10 , wherein the configurable activation function is selected from a list consisting of: Boolean, trinary, linear, ReLU, shifted ReLU, ExpReLU, sigmoid, and tanh.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2020
From: SAWADA, JUN; FLICKNER, MYRON D.; CASSIDY, ANDREW STEPHEN; ARTHUR, JOHN VERNON; DATTA, PALLAB; MODHA, DHARMENDRA S.; ESSER, STEVE KYLE; TABA, BRIAN SEISHO; KLAMO, JENNIFER; APPUSWAMY, RATHINAKUMAR; AKOPYAN, FILIPP; OTERO, CARLOS ORTEGA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054598/0076 →
Continuity (1)
Related Publication 20220180177A1 · Jun 9, 2022
References Cited (31)
US 10037306B2 · Lin et al. · 2018 [cited by applicant]
US 10127494B1 · Cantin · 2018 [cited by examiner]
US 11270187B2 · Choi · 2022 [cited by examiner]
US 20170102921A1 · Henry · 2017 [cited by examiner]
US 20180060278A1 · Lin · 2018 [cited by examiner]
US 20190147323A1 · Li et al. · 2019 [cited by applicant]
US 20190385048A1 · Cassidy et al. · 2019 [cited by applicant]
US 20210157549A1 · Elmer · 2021 [cited by examiner]
US 20210303977A1 · Sun · 2021 [cited by examiner]
CN 109643327A · 2019 [cited by applicant]
CN 109754063A · 2019 [cited by applicant]
CN 109816105A · 2019 [cited by applicant]
CN 110770762A · 2020 [cited by applicant]
CN 111226233A · 2020 [cited by applicant]
JP 2020004398A · 2020 [cited by applicant]
KR 20190051755A · 2019 [cited by applicant]
Ortega-Zamorano et al., “High precision FPGA implementation of neural network activation functions,” 2014 IEEE symposium on intelligent embedded systems (IES), pp. 55-60. [cited by applicant]
Li et al., “An Efficient Hardware Architecture for Activation Function in Deep Learning Processor,” 2018 IEEE 3rd International Conference on Image, Vision and Computing (ICIVC), pp. 911-918. [cited by applicant]
Search and Examination Report for United Kingdom Application No. GB2116839.8 dated Aug. 31, 2022. [cited by applicant]
United Kingdom Search Report for Application No. GB2116839.8 dated Nov. 2, 2023. [cited by applicant]
The State Intellectual Property Office of People's Republic of China, “First Office Action”, Mar. 3, 2025, 21 Pages, CN Application No. 202111438045.3. [cited by applicant]
Abdelouahab et al., “PhD Forum: Why TanH can be a Hardware Friendly Activation Function for CNNs”, HAL Id: hal-01654697 https://hal.archives-ouvertes.fr/hal-01654697, Dec. 4, 2017, 04 Pages. [cited by applicant]
Abdelsalam et al., “Accurate and Efficient Hyperbolic Tangent Activation Function on FPGA using the DCT Interpolation Filter”, Computer and Software Engineering Department Polytechnique Montreal, Sep. 25, 2016, 08 Pages. [cited by applicant]
Armato et al., “Low-error digital hardware implementation of artificial neuron activation functions and their derivative”, Microprocessors and Microsystems, Aug. 2011, pp. 557-567, vol. 35, Issue 6. [cited by applicant]
German Patent and Trademark Office, “Office Action,” Jan. 7, 2025, 11 Pages, DE Application No. 102021128932.7. [cited by applicant]
Hao Yufeng, “A General Neural Network Hardware Architecture on FPGA”, Dept. of Electronic, Electrical and Systems Engineering, Nov. 6, 2017, 06 Pages. [cited by applicant]
Larkin et al., “An Efficient Hardware Architecture for a Neural Network Activation Function Generator”, Centre for Digital Video Processing, Year 2006, pp. 08. [cited by applicant]
Raut et al., “A Cordic based configurable activation function for ANN applications”, In: 2020 IEEE computer society annual symposium on VLSI (ISVLSI). IEEE, 2020, pp. 78-83. [cited by applicant]
Tiwari et al., “Hardware implementation of neural network with Sigmoidal activation functions using CORDIC”, Microprocessors and Microsystems, Year 2015, pp. 373-381. [cited by applicant]
Zhang Lei, “Implementation of Fixed-point Neuron Models with Threshold, Ramp and Sigmoid Activation Functions”, IOP Conf. Series: Materials Science and Engineering 224, Year 2017, 06 Pages, doi: 10.1088/1757-899X/224/1/… [cited by applicant]
Japan Patent Office, “Notice of Reasons for Refusal” Mar. 18, 2025, 08 Pages, JP Application No. 2021-188192. [cited by applicant]