IP Library Granted Patent US 12,361,271
Granted Patent B2
US 12,361,271 · App. 17/338,910 · Granted Jul 15, 2025

Optimizing deep neural network mapping for inference on analog resistive processing unit arrays

Inventor: Malte Johannes Rasch (Chappaqua, NY)
Assignee: International Business Machines Corporation
G06N3/065G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,271
App. No.
17/338,910
Granted
Jul 15, 2025
Kind
B2
Abstract

A computer-implemented method, computer program product, and/or computer system that performs the following operations: (i) receiving a deep neural network (DNN) having a set of inputs nodes, a set of output nodes, and a set of weight parameters; (ii) configuring the DNN for application to a set of analog resistive processing unit (RPU) arrays, the configuring including applying a set of modifiers to respective outputs of the set of output nodes, the set of modifiers corresponding to the set of analog RPU arrays; (iii) training the DNN using a training process, the training yielding an updated set of weight parameters and an updated set of modifiers; and (iv) transferring the updated set of weight parameters and the updated set of modifiers to the set of analog RPU arrays.

Claims (50)

1. A computer-implemented method comprising:

receiving, by one or more computer processors, a deep neural network (DNN) having a set of inputs nodes, a set of output nodes, and a set of weight parameters, wherein the DNN is digitally trained and stored;

configuring, by the one or more computer processors, the DNN for application to a set of analog resistive processing unit (RPU) arrays, the configuring including:

applying a set of modifiers to respective output nodes of the set of output nodes, wherein each modifier of the set of modifiers is applied to a respective corresponding output node of the set of output nodes, and wherein each modifier of the set of modifiers a corresponds and applies to a respective corresponding analog RPU array of the set of analog RPU arrays; and

digitally recreating an array configuration of the set of analog RPU arrays in one or more layers of the DNN;

training, by the one or more computer processors, the DNN using a training process, the training yielding an updated set of weight parameters and an updated set of modifiers;

converting, by the one or more computer processors, the updated set of weight parameters to a set of conductance values corresponding to the set of analog RPU arrays; and

transferring, by the one or more computer processors, the updated set of weight parameters and the updated set of modifiers to the set of analog RPU arrays.

2. The computer-implemented method of claim 1 , wherein the converting is completed prior to the transferring the set of conductance values corresponding to the set of analog RPU arrays.

3. The computer-implemented method of claim 1 , further comprising:

instructing, by the one or more computer processors, the set of analog RPU arrays to perform an analog inference task utilizing the updated set of weight parameters and the updated set of modifiers.

4. The computer-implemented method of claim 1 , wherein the configuring further includes mapping the set of weight parameters to the set of conductance values corresponding to the set of analog RPU arrays.

5. The computer-implemented method of claim 4 , wherein the training further includes, upon completion of a training epoch, remapping the set of weight parameters to an updated set of conductance values corresponding to the set of analog RPU arrays.

6. The computer-implemented method of claim 1 , wherein the configuring further includes clipping a range of the set of output nodes and a range of the set of weight parameters according to a conductance range of a set of output nodes of the set of analog RPU arrays and a conductance range of a set of weights of the set of analog RPU arrays, respectively.

7. The computer-implemented method of claim 1 , wherein the training process utilizes stochastic gradient descent.

8. The computer-implemented method of claim 1 , wherein modifiers of the set of modifiers correspond to respective modifiers applied to respective output nodes of the set of analog RPU arrays.

9. The computer-implemented method of claim 1 , wherein one or more modifiers of the set of modifiers are affine transformations.

10. The computer-implemented method of claim 9 , wherein the set of modifiers include, for each output node of the set of output nodes, a respective linear scaling factor and a respective bias term.

11. A computer program product comprising one or more computer readable storage media and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by one or more computer processors to cause the one or more computer processors to:

receive a deep neural network (DNN) having a set of inputs nodes, a set of output nodes, and a set of weight parameters, wherein the DNN is digitally trained and stored;

configure the DNN for application to a set of analog resistive processing unit (RPU) arrays, the configuring including:

applying a set of modifiers to respective output nodes of the set of output nodes, wherein each modifier of the set of modifiers is applied to a respective corresponding output node of the set of output nodes, and wherein each modifier of the set of modifiers a corresponds and applies to a respective corresponding analog RPU array of the set of analog RPU arrays; and

digitally recreating an array configuration of the set of analog RPU arrays in one or more layers of the DNN;

train the DNN using a training process, the training yielding an updated set of weight parameters and an updated set of modifiers;

convert the updated set of weight parameters to a set of conductance values corresponding to the set of analog RPU arrays; and

transfer the updated set of weight parameters and the updated set of modifiers to the set of analog RPU arrays.

12. The computer program product of claim 11 , wherein the converting is completed prior to the transferring the set of conductance values corresponding to the set of analog RPU arrays.

13. The computer program product of claim 11 , the program instructions executable by one or more computer processors further causing the one or more computer processors to:

instruct the set of analog RPU arrays to perform an analog inference task utilizing the updated set of weight parameters and the updated set of modifiers.

14. The computer program product of claim 11 , wherein one or more modifiers of the set of modifiers are affine transformations.

15. The computer program product of claim 14 , wherein the set of modifiers include, for each output node of the set of output nodes, a respective linear scaling factor and a respective bias term.

16. A computer system comprising:

one or more analog resistive processing unit (RPU) arrays;

one or more computer processors; and

one or more computer readable storage media;

wherein:

the one or more computer processors are structured, located, connected and/or programmed to execute program instructions collectively stored on the one or more computer readable storage media; and

the program instructions, when executed by the one or more computer processors, cause the one or more computer processors to:

receive a deep neural network (DNN) having a set of inputs nodes, a set of output nodes, and a set of weight parameters, wherein the DNN is digitally trained and stored;

configure the DNN for application to the one or more analog RPU arrays, the configuring including:

applying a set of modifiers to respective output nodes of the set of output nodes, wherein each modifier of the set of modifiers is applied to a respective corresponding output node of the set of output nodes, and wherein each modifier of the set of modifiers a corresponds and applies to a respective corresponding analog RPU array of the one or more analog RPU arrays; and

digitally recreating an array configuration of the one or more analog RPU arrays in one or more layers of the DNN;

train the DNN using a training process, the training yielding an updated set of weight parameters and an updated set of modifiers;

convert the updated set of weight parameters to a set of conductance values corresponding to the one or more analog RPU arrays; and

transfer the updated set of weight parameters and the updated set of modifiers to the one or more analog RPU arrays.

17. The computer system of claim 16 , wherein the converting is completed prior to the transferring the set of conductance values corresponding to the one or more analog RPU arrays.

18. The computer system of claim 16 , the program instructions, when executed by the one or more computer processors, further causing the one or more computer processors to:

instruct the one or more analog RPU arrays to perform an analog inference task utilizing the updated set of weight parameters and the updated set of modifiers.

19. The computer system of claim 16 , wherein one or more modifiers of the set of modifiers are affine transformations.

20. The computer system of claim 19 , wherein the set of modifiers include, for each output node of the set of output nodes, a respective linear scaling factor and a respective bias term.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2021
From: RASCH, MALTE JOHANNES
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056438/0601 →
Continuity (1)
Related Publication 20220391688A1 · Dec 8, 2022
References Cited (28)
US 9646243B1 · Gokmen · 2017 [cited by applicant]
US 20180005110A1 · Gokmen · 2018 [cited by examiner]
US 20190147327A1 · Martin · 2019 [cited by applicant]
US 20200082255A1 · Kataeva · 2020 [cited by applicant]
US 20210279556A1 · Gokmen · 2021 [cited by examiner]
US 20210326685A1 · Lee · 2021 [cited by examiner]
US 20220101142A1 · Tsai · 2022 [cited by examiner]
Rasch, M. J., Gokmen, T., & Haensch, W. (2019). Training large-scale artificial neural networks on simulated resistive crossbar arrays. IEEE Design & Test, 37(2), 19-29. (Year: 2019). [cited by examiner]
Gokmen, T., Onen, M., & Haensch, W. (2017). Training deep convolutional neural networks with resistive cross-point devices. Frontiers in neuroscience, 11, 538. (Year: 2017). [cited by examiner]
He, Z., Lin, J., Ewetz, R., Yuan, J. S., & Fan, D. (Jun. 2019). Noise injection adaption: End-to-end ReRAM crossbar non-ideal effect adaption for neural network mapping. In Proceedings of the 56th Annual Design Automati… [cited by examiner]
Song, Z., Sun, Y., Chen, L., Li, T., Jing, N., Liang, X., & Jiang, L. (Apr. 21, 2020). ITT-RNA: Imperfection tolerable training for RRAM-crossbar-based deep neural-network accelerator. IEEE Transactions on Computer-Aide… [cited by examiner]
Chakraborty, I., Ali, M. F., Kim, D. E., Ankit, A., & Roy, K. (Jul. 2020). Geniex: A generalized approach to emulating non-ideality in memristive xbars using neural networks. In 2020 57th ACM/IEEE Design Automation Conf… [cited by examiner]
Tsai, H., Ambrogio, S., Narayanan, P., Shelby, R. M., & Burr, G. W. (2018). Recent progress in analog memory-based accelerators for deep learning. Journal of Physics D: Applied Physics, 51(28), 283001. (as cited by the … [cited by examiner]
NPL Gokmen Acceleration of DNN Training Resistive Cross Point 2016. [cited by examiner]
NPL Gokmen Training Deep CNNs with Resistive Cross Point Devices 2017. [cited by examiner]
NPL Gokmen Training LSTM Networks with Resistive 2018. [cited by examiner]
NPL Gokmen Training NNs on Resistive Device Arrays 2020. [cited by examiner]
NPL Haensch The Next Generation of Deep Learning Hardware 2019. [cited by examiner]
NPL Jain CxDNN Hardware software Compensation Methods for DNNs 2019. [cited by examiner]
NPL Lopez Coarse Grained Reconfigurable Computing Mar. 2021. [cited by examiner]
NPL Rash Inventor Training Large Scale ANNs on Simulated RCAs Apr. 2020. [cited by examiner]
NPL Tsai Recent progress in analog memory based Accelerators 2018. [cited by examiner]
Gokmen et al., “Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices: Design Considerations”, Frontiers in Neuroscience, vol. 10, Article 333, published: Jul. 21, 2016, 13 pages, <https://doi.… [cited by applicant]
Joshi et al., “Accurate deep neural network inference using computational phase-change memory”, Nature Communications, Published: May 18, 2020, 13 pages, <https://doi.org/10.1038/s41467-020-16108-9>. [cited by applicant]
Kazemi et al., “A Device Non-Ideality Resilient Approach for Mapping Neural Networks to Crossbar Arrays”, In 2020 57th ACM/IEEE Design Automation Conference (DAC), 6 pages, © 2020 IEEE. [cited by applicant]
Kendall et al., “Training End-to-End Analog Neural Networks with Equilibrium Propagation”, arXiv:2006.01981v2 [cs.NE] Jun. 9, 2020, 31 pages. [cited by applicant]
Kim et al., “Zero-Shifting Technique for Deep Neural Network Training on Resistive Cross-point Arrays”, arXiv:1907.10228v2 [cs.ET], last revised Aug. 2, 2019 (this version, v2), 27 pages, <https://arxiv.org/abs/1907.102… [cited by applicant]
Tsai et al., “Recent progress in analog memory-based accelerators for deep learning”, Journal of Physics D: Applied Physics, 51 (2018) 283001, Published Jun. 21, 2018, 28 pages, <https://doi.org/10.1088/1361-6463/aac8a5… [cited by applicant]