IP Library › Granted Patent US 12,499,369
Granted Patent B2
US 12,499,369 · App. 17/411,672 · Granted Dec 16, 2025

Efficient identification of critical faults in neuromorphic hardware of a neural network

Inventors: Ching-Yuan Chen (Chapel Hill, NC); Krishnendu Chakrabarty (Chapel Hill, NC)
Assignee: NVIDIA Corporation
G06N3/084G06F11/1476
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,369
App. No.
17/411,672
Granted
Dec 16, 2025
Kind
B2
Abstract

The disclosure provides misclassification-driven training (MDT) that efficiently identifies critical faults in neuromorphic hardware, such as a memristor crossbar. MDT advantageously identifies whether a hardware fault is a critical fault and can be used to limit fault recovery when a hardware fault is not a critical fault. By applying fault-tolerant techniques directed to critical faults, such as only for critical faults, processing overhead of a neural network can be reduced. In one aspect, the disclosure provides a method of identifying critical faults in neuromorphic hardware of a neural network. In one example the method of identifying includes: (1) determining a significant parameter of a trained neural network that impacts classification of a sample of a dataset, (2) obtaining a location of the significant parameter in the neuromorphic hardware, and (3) identifying the location as a critical fault of the neuromorphic hardware.

Claims (43)

1 . A method of identifying critical faults in neuromorphic hardware of a neural network, comprising:

determining a significant parameter of a trained neural network that impacts classification of a sample of a dataset by introducing deviated weights to parameters of the trained neural network;

obtaining a location of the significant parameter in the neuromorphic hardware, wherein the neuromorphic hardware is a crossbar of memristors; and

identifying the location as a critical fault of the neuromorphic hardware, wherein the location is a memristor location of the crossbar.

2 . The method as recited in claim 1 , wherein the determining, the obtaining, and the identifying the location are performed iteratively on other samples of the dataset.

3 . The method as recited in claim 2 , wherein the method is performed iteratively until all data points of the dataset are impacted.

4 . The method as recited in claim 1 , wherein the determining further includes deriving a gradient for each of the parameters using backward propagation, determining a significance value of each of the parameters based on the gradient, and identifying a parameter of the parameters having a greatest significance value as the significant parameter.

5 . The method as recited in claim 1 , further comprising creating a deviated parameter value at the location by combining a gradient for the significant parameter with the value of the significant parameter and inferring a type of fault at the location.

6 . The method as recited in claim 5 , wherein the combining is a gradient-ascent step wherein a change in direction of deviation is same as the gradient.

7 . The method as recited in claim 5 , wherein the combining is a gradient-descent when the classification is directed to a targeted class.

8 . The method as recited in claim 5 , wherein the type of fault is a stuck-on fault when the deviated parameter value is within a nominal value range of elements of the neuromorphic hardware.

9 . The method as recited in claim 5 , wherein the type of fault is a stuck-open fault when the deviated parameter value is outside of a nominal value range of elements of the neuromorphic hardware.

10 . A computer program product comprising a non-transitory computer-readable medium having a series of operating instructions stored thereon that directs a processor when executed thereby to perform operations to execute at least some of the steps of the method of claim 1 .

11 . A computing system including a processor that performs at least some of the steps of the method of claim 1 .

12 . A method of training a machine learning (ML) model for classifying critical faults, comprising:

identifying critical faults from a dataset using misclassification-driven training (MDT);

identifying benign faults from the dataset using random fault injections and forward inferencing;

creating a training dataset that includes the critical faults and the benign faults; and

training the ML model using the training dataset.

13 . The method of training as recited in claim 12 , wherein input features for the critical faults include a fault location, a fault type, a parameter significance value, and a deviated parameter value.

14 . The method of training as recited in claim 12 , wherein significance of the benign faults are determined by backpropagating a randomly selected data point from an input dataset.

15 . The method of training as recited in claim 12 , wherein the training of the ML model includes multiple training runs using the training dataset.

16 . The method of training as recited in claim 12 , wherein the training of the ML model includes verifying the benign faults are not identified as a critical fault by injecting the critical faults into the ML model and evaluating whether they are critical.

17 . The method of training as recited in claim 12 , wherein the training of the ML model includes verifying the critical faults cause a misclassification.

18 . A method of identifying faults in a neural network (NN), comprising:

classifying a dataset using a NN;

determining if detection of faults is needed, wherein the faults are hardware faults of the NN and the NN is a crossbar of memristors fabricated on a chip;

performing detection of the faults when needed; and

performing fault recovery of the hardware faults when a detected fault is a critical fault.

19 . The method of identifying faults as recited in claim 18 , wherein the determining is periodically based on a number of rounds of the classifying.

20 . The method of identifying faults as recited in claim 18 , wherein the determining is based on a trigger.

21 . The method of identifying faults as recited in claim 18 , wherein the fault recovery includes remapping.

22 . The method of identifying faults as recited in claim 18 , wherein the fault recovery includes retraining.

23 . A method of manufacturing a chip including a neural network (NN) having a memristor crossbar, comprising:

identifying one or more memristor cells of a crossbar as critical;

allocating fault-tolerant hardware for the one or more memristor cells that are critical; and

fabricating the chip with the fault-tolerant hardware.

24 . The method of manufacturing as recited in claim 23 , wherein the fault-tolerant hardware is at least one spare memristor column corresponding to the one or more memristor cells.

25 . The method of manufacturing as recited in claim 23 , wherein the identifying is based on a classification of critical faults for the neural network and a target application.

26 . The method of manufacturing as recited in claim 23 , wherein the identifying is based on a classification of critical faults for at least one of another neural network and another target application.

27 . A neural network including a chip manufactured according to the method of claim 23 .

28 . An autonomous driving system including a chip manufactured according to the method of claim 23 .

29 . A vision system including a chip manufactured according to the method of claim 23 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2021
From: CHEN, CHING-YUAN; CHAKRABARTY, KRISHNENDU
To: NVIDIA CORPORATION
Reel/Frame 057286/0803 →
Continuity (2)
Provisional Application 63070419 · Aug 26, 2020
Related Publication 20220067531A1 · Mar 3, 2022
References Cited (28)
US 20210012974A1 · Zhou · 2021 [cited by examiner]
US 20210318922A1 · Roberts · 2021 [cited by examiner]
Madry, et al.; “Towards Deep Learning Models Resistant to Adversarial Attacks”; Int. Conf. Learning Representations; 2018; 23 pgs. [cited by applicant]
Gao, et al.; “STRIP: A Defence Against Trojan Attacks on Deep Neural Networks”; Proc. Computer Security Applications Conf.; 2019; 13 pgs. [cited by applicant]
Hu, et al.; “Dot-Product Engine for Neuromorphic Computing: Programming 1T1M Crossbar to Accelerate Matrix-Vector Multiplication”; DAC '16, Austin, TX; Jun. 2016; 6 pgs. [cited by applicant]
Chaudhuri, et al.; “Analysis of Process Variations, Defects, and Design-Induced Coupling in Memristors”; International Test Conference; 2018; 10 pgs. [cited by applicant]
Sun, et al.; “Impact of Non-Ideal Characteristics of Resistive Synaptic Devices on Implementing Convolutional Neural Networks”; IEEE Journal of Emerging and Selected Topics in Circuits and Systems; Sep. 2019; 10 pgs. [cited by applicant]
Liu, et al.; “Design of Fault-Tolerant Neuromorphic Computing Systems”; 23rd IEEE European Test Symposium; 2018; 9 pgs. [cited by applicant]
Gebregiorgis, et al.; “Testing of Neuromorphic Circuits: Structural vs Functional”; International Test Conference; 2019; 10 pgs. [cited by applicant]
Li, et al.; “Understanding Error Propagation in Deep Learning Neural Network (DNN) Accelerators and Applications”; Proc. Int. Conf. High Performance Computing, Networking, Storage and Analysis; 2017; 12 pgs. [cited by applicant]
Fieback, et al.; “Device-Aware Test: A New Test Approach Towards DPPB Level”; International Test Conference; 2019; 10 pgs. [cited by applicant]
Kannan, et al.; “Sneak-Path Testing of Crossbar-Based Nonvolatile Random Access Memories”; IEEE Transactions on Nanotechnology; May 2013; 14 pgs. [cited by applicant]
Xu, et al.; “Safety Design of a Convolutional Neural Network Accelerator with Error Localization and Correction”; International Test Conference; 2019; 10 pgs. [cited by applicant]
Li, et al.; “ICE: Inline Calibration for Memristor Crossbar-based Computing Engine”; Date; 2014; 4 pgs. [cited by applicant]
Liu, et al.; “Rescuing Memristor-based Neuromorphic Design with High Defects”; DAC '17, Austin, TX; 2017; 6 pgs. [cited by applicant]
Xia, et al.; “Stuck-at Fault Tolerance in RRAM Computing Systems”; IEEE Journal on Emerging and Selected Topics in Circuits and Systems; Mar. 2018; 14 pgs. [cited by applicant]
Krizhevsky, et al.; “ImageNet Classification with Deep Convolutional Neural Networks”; Proc. Int. Conf. Neural Information Processing Systems; 2012; 9 pgs. [cited by applicant]
Simonyan, et al.; “Very Deep Convolutional Networks for Large-Scale Image Recognition”; Conference paper at ICLR; 2015; 14 pgs. [cited by applicant]
Szegedy, et al.; “Going deeper with convolutions”; IEEE Conf. Computer Vision and Pattern Recognition; 2015; 12 pgs. [cited by applicant]
He, et al.; “Deep Residual Learning for Image Recognition”; CVPR; 2016; 12 pgs. [cited by applicant]
Xia, et al.; “Fault-Tolerant Training with On-Line Fault Detection for RRAM-Based Neural Computing Systems”; DAC 17, Austin, TX; Jun. 2017; 6 pgs. [cited by applicant]
Krizhevsky, et al.; “The CIFAR-10 dataset”; http://www.cs.toronto.edu/˜kriz/cifar.html; 2009; 4 pgs. [cited by applicant]
Deng, et al.; “ImageNet: A Large-Scale Hierarchical Image Database”; CVPR; 2009; 8 pgs. [cited by applicant]
Hamdioui, et al.; “Testing Open Defects in Memristor-Based Memories”; IEEE Transactions on Computers; Jan. 2015; 13 pgs. [cited by applicant]
Kannan, et al.; “Sneak path Testing and Fault Modeling for Multilevel Memristor-based Memories”; IEEE Int. Conf. Computer Design; 2013; 6 pgs. [cited by applicant]
Chen, et al.; “RRAM Defect Modeling and Failure Analysis Based on March Test and a Novel Squeeze-Search Scheme”; IEEE Transactions on Computers; Jan. 2015; 11 pgs. [cited by applicant]
Chen, et al.; “Fault Modeling and Testing of 1T1R Memristor Memories”; IEEE 33rd VLSI Test Symposium; 2015; 6 pgs. [cited by applicant]
Goodfellow, et al.; “Deep Learning”; Chapters 6, 7, and 12; MIT Press; http://www.deeplearningbook.org; 2016; 150 pgs. [cited by applicant]