IP Library › Granted Patent US 12,748,962
Granted Patent B2
US 12,748,962 · App. 19/030,176 · Granted Sep 29, 2026

Crossbar array apparatuses based on compressed-truncated singular value decomposition (C-TSVD) and analog multiply-accumulate (MAC) operation methods using the same

Inventors: Youngnam Hwang (Hwaseong-si, KR); Jaewon Yang (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06N3/065G06F7/5443G06F17/16G06N3/049
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,962
App. No.
19/030,176
Granted
Sep 29, 2026
Kind
B2
Abstract

A compressed-truncated singular value decomposition (C-TSVD) based crossbar array apparatus is provided. The C-TSVD based crossbar array apparatus may include an original crossbar array in an m×n matrix having row input lines and column output lines and including cells of a resistance memory device, or two partial crossbar arrays obtained by decomposing the original crossbar array based on C-TSVD, an analog to digital converter (ADC) that converts output values of column output lines of sub-arrays obtained through array partitioning, an adder that sums up results of the ADC to correspond to the column output lines, and a controller that controls application of the original crossbar array or the two partial crossbar arrays. Input values are input to the row input lines, a weight is multiplied by the input values and accumulated results are output as output values of the column output lines.

Claims (43)

1 . An analog multiply-accumulate (MAC) operation method comprising:

calculating an original crossbar array in an m×n matrix having n row input lines and m column output lines, the matrix including cells of a resistance memory device, where n and m are natural numbers;

selectively performing compressed-truncated singular value decomposition (C-TSVD) decomposing the original crossbar array into two partial crossbar arrays;

partitioning the original crossbar array into sub-arrays, or partitioning the two partial crossbar arrays into sub-arrays, in accordance with a result of selectively performing the C-TSVD;

inputting input values to row input lines of the sub-arrays;

multiplying a weight by the input values and accumulating multiplication results in the sub-arrays and outputting output values to the column output lines of the sub-arrays;

analog to digital (AD) converting the output values of the column output lines by using an analog to digital converter (ADC); and

summing up the ADC results to correspond to the column output lines by using an adder,

wherein a crossbar array apparatus corresponding to one layer of a neural network is used in neuromorphic computing,

wherein the crossbar array apparatus comprises the original crossbar array or the two partial crossbar arrays, the ADC, the adder, and a controller, and

wherein, in the selectively performing of the C-TSVD, application of the original crossbar array or the two partial crossbar arrays is controlled by the controller.

2 . The analog MAC operation method of claim 1 , wherein a matrix W of the original crossbar array is decomposed into multiplication of three matrices U, S, and V T based on SVD so that the matrix U has a size of m×m, the matrix S has a size of m×n, and the matrix V T has a size of n×n,

wherein, when a rank of the matrix W is k (k≤n & k≤m), the sizes of the matrices U, S, and V T are respectively reduced to m×k, k×k, and k×n and a number of components of the matrices U, S, and V T are further reduced to k′ (k′≤k) via selection of at least partial components among the k components in large order based on TSVD,

wherein the two partial crossbar arrays comprise a first partial crossbar array and a second partial crossbar array, the first partial crossbar array is represented as the matrix U and has a size of m×k, and the second partial crossbar array is represented as a matrix S*V T and has a size of k*n,

wherein, in a first case the C-TSVD is not performed in the selectively performing of the C-TSVD and the original crossbar array is partitioned into the sub-arrays,

wherein, in a second case, the C-TSVD is performed in the selectively performing of the C-TSVD and the two resultant partial crossbar arrays are partitioned into the sub-arrays, and

wherein a number of sub-arrays, a number of operations of the ADS, a number of operations of the adder, and an operating current are reduced in the second case as compared to the first case.

3 . The analog MAC operation method of claim 2 , wherein, for the number of sub-arrays, the number of operations of the ADC, the number of operations of the adder, and the operating current, as a size of a matrix of the original crossbar array increases, a reduction in the second case relative to the first case increases.

4 . The analog MAC operation method of claim 3 , wherein, when a time for performing an operation of the crossbar array apparatus based on the original crossbar array is a 1-stage latency and a time for performing an operation of the crossbar array apparatus based on the two partial crossbar arrays is 2-stage latency,

as a size of a matrix of the original crossbar array increases, an amount of increase in the 2-stage latency relative to the 1-stage latency decreases, and

a degree of the increase is less than a degree of the reduction.

5 . The analog MAC operation method of claim 2 , wherein, when a matrix of the sub-arrays has a size of s×s, where s is a natural number,

in a third case in which the first partial crossbar array is partitioned into the sub-arrays, the number of sub-arrays is Nu=ceil(m/s)*ceil(k/s)*b,

in a fourth case in which the second partial crossbar array is partitioned into the sub-arrays, the number of sub-arrays is Nv=ceil(k/s)*ceil(n/s)*b,

wherein ceil(x) represents a minimum integer of no less than x and b represents the number of cells required per weight element,

wherein, in the first case, the number of sub-arrays is Ns(W)=ceil(m/s)*ceil(n/s)*b, and

wherein, in the second case, the number of all sub-arrays is Ns(Wt)=Nu+Nv=ceil(m/s)*ceil(k/s)*b+ceil(k/s)*ceil(n/s)*b.

6 . The analog MAC operation method of claim 5 , wherein, in the first case, the number of operations of the ADC is Ns(W)*s,

the number of operations of the adder is {Ns(W)−ceil(n/s)*b}*s,

a number of adder stages is ceil (log 2 ceil(m/s)), and

the operating current is Ns(W)*s*s times a cell current, and

wherein, in the second case, the number of operations of the ADC is Ns(Wt)*s,

the number of operations of the adder is {Ns(Wt)−(ceil(k/s)+ceil(n/s))*b}*s,

the number of adder stages is ceil (log 2 ceil(m/s)), and

the operating current is Ns(Wt)*s*s times a cell current.

7 . The analog MAC operation method of claim 2 , wherein, when k′/(m or n)*100 is a taken ratio,

wherein, in a first learning method in which, after calculating the original crossbar array through learning, inference is performed by decomposing the original crossbar array into partial crossbar arrays through the SVD and selecting the k′, inference accuracy in accordance with the taken ratio is affected by a regularization parameter of a regularization term used for preventing over-fitting in the learning, and

wherein, in a second learning method in which, after selecting an arbitrary first original crossbar array before learning and selecting the k′ by decomposing the first original crossbar array into partial crossbar arrays through the SVD, inference is performed by learning an integrated second original crossbar array, inference accuracy in accordance with the taken ratio is less affected by the regularization parameter than in the first learning method.

8 . The analog MAC operation method of claim 7 , wherein the inference accuracy is maintained at no less than 90% although the taken ratio is taken at no more than 10% by performing the second learning method or by selecting a predetermined regularization parameter in the first learning method.

9 . The analog MAC operation method of claim 2 , wherein an adder tree including the adder is arranged in an output portion of the original crossbar array or an output portion of each of the first partial crossbar array and the second partial crossbar array,

wherein an output of the adder tree of the first partial crossbar array is input to the second partial crossbar array through a first circuit of an identity function, and

wherein the output of the adder tree of the original crossbar array or the output of the adder tree of the second partial crossbar array is activated by a second circuit of an activation function and the active output is input to a next layer.

10 . The analog MAC operation method of claim 1 , wherein the weight corresponds to conductance, wherein voltages are applied as input values of the row input lines, and wherein currents are output as output values of the column output lines.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE ADDRESS PREVIOUSLY RECORDED AT REEL: 69920 FRAME: 724. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 21, 2025
From: HWANG, YOUNGNAM; YANG, JAEWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 069964/0601 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: HWANG, YOUNGNAM; YANG, JAEWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 069920/0724 →
Priority Claims (1)
KR 10-2020-0129509 · Oct 7, 2020 · national
Continuity (2)
Continuation 17319679 · May 13, 2021
Related Publication 20250165767A1 · May 22, 2025
References Cited (23)
US 9262724B2 · Kingsbury et al. · 2016 [cited by applicant]
US 9728184B2 · Xue et al. · 2017 [cited by applicant]
US 10127494B1 · Cantin et al. · 2018 [cited by applicant]
US 10268232B2 · Harris et al. · 2019 [cited by applicant]
US 10554272B2 · Tong et al. · 2020 [cited by applicant]
US 20040190636A1 · Oprea · 2004 [cited by applicant]
US 20140019388A1 · Kingsbury et al. · 2014 [cited by applicant]
US 20160313796A1 · Kim et al. · 2016 [cited by applicant]
US 20190065943A1 · Cantin et al. · 2019 [cited by applicant]
US 20190122108A1 · Kliegl et al. · 2019 [cited by applicant]
US 20190294199A1 · Carolan et al. · 2019 [cited by applicant]
US 20190332941A1 · Towal et al. · 2019 [cited by applicant]
US 20190392316A1 · Kim et al. · 2019 [cited by applicant]
US 20200279169A1 · Hoskins et al. · 2020 [cited by applicant]
US 20200401373A1 · Lesso · 2020 [cited by examiner]
US 20210390382A1 · Kwon · 2021 [cited by examiner]
CN 108536422A · 2018 [cited by applicant]
KR 20190059033A · 2019 [cited by applicant]
KR 20200000686A · 2020 [cited by applicant]
KR 1020200057475A · 2020 [cited by applicant]
KR 1020200088955A · 2020 [cited by applicant]
“Chong Li, et al., “Constrained Optimization Based Low-Rank Approximation of Deep Neural Networks,” European Conference on Computer Vision 2018, pp. 746-761”. [cited by applicant]
“Page et al., “FPGA-Based Reduction Techniques for Efficient Deep Neural Network Deployment,” 2016 IEEE 24th Annual International Symposium on Field-Programmable Custom Computing Machines, May 1-3, 2016”. [cited by applicant]