IP Library Granted Patent US 12,299,576
Granted Patent B2
US 12,299,576 · App. 17/343,001 · Granted May 13, 2025

Neural network-based inference method and apparatus

Inventors: Chang Gao (Zurich, CH); Shih-Chii Liu (Zurich, CH); Tobi Delbruck (Zurich, CH); Xi Chen (Zurich, CH)
Assignees: Samsung Electronics Co., Ltd.; University of Zurich
G06N3/082G06F9/5083G06N3/04G06N5/04G06F7/76G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,576
App. No.
17/343,001
Granted
May 13, 2025
Kind
B2
Abstract

Disclosed is a neural network-based inference method and apparatus. The neural network-based inference method includes compressing a matrix comprising processing elements corresponding to an operation of a neural network, balancing workloads related to the operation by reordering the compressed matrix based on the workloads, and performing inference based on the reordered matrix.

Claims (44)

1. A neural network-based inference method, comprising:

compressing a matrix comprising processing elements corresponding to an operation of a neural network;

balancing workloads related to the operation by reordering the compressed matrix based on the workloads; and

performing inference based on the reordered matrix,

wherein the compressing comprises:

generating subcolumns by splitting the matrix; and

compressing the matrix by pruning elements of the subcolumns based on a target sparsity related to compression of the neural network,

wherein the compressing of the matrix by pruning elements of the subcolumns comprises replacing the elements of the subcolumns with a predetermined value such that a ratio of elements other than the predetermined value in the elements of the subcolumns is the target sparsity, and

wherein the pruning of the elements of the subcolumns comprises pruning the elements of the subcolumns by sequentially replacing the elements other than the predetermined value with the predetermined value, starting from an element having a smallest absolute value.

2. The neural network-based inference method of claim 1 , wherein the generating comprises:

generating a first subcolumn corresponding to a first processing element of the matrix; and

generating a second subcolumn corresponding to a second processing element of the matrix.

3. The neural network-based inference method of claim 1 , wherein the balancing of the workloads related to the operation comprises:

calculating the workloads corresponding to neurons of the neural network;

assigning the workloads to the compressed matrix; and

balancing the workloads by reordering the compressed matrix based on the workloads assigned to the compressed matrix.

4. The neural network-based inference method of claim 3 , wherein the calculating of the workloads comprises calculating the workloads by accumulating a numbers of elements other than a predetermined value in the neurons of the neural network.

5. The neural network-based inference method of claim 3 , wherein the assigning comprises assigning the workloads to columns of the compressed matrix.

6. The neural network-based inference method of claim 3 , wherein the balancing comprises:

comparing workloads corresponding to columns of the compressed matrix; and

performing balancing by swapping the columns of the compressed matrix based on a result of the comparing.

7. The neural network-based inference method of claim 1 , further comprising:

encoding the reordered matrix based on any one or any combination of components of the processing elements, positions of the processing elements, and a size of the reordered matrix.

8. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the neural network-based inference method of claim 1 .

9. A neural network-based inference apparatus, comprising:

a processor configured to compress a matrix comprising processing elements corresponding to an operation of a neural network, to balance workloads related to the operation by reordering the compressed matrix based on the workloads, and to perform inference based on the reordered matrix,

wherein the processor is further configured:

to generate subcolumns by splitting the matrix, and

to compress the matrix by pruning elements of the subcolumns based on a target sparsity related to compression of the neural network,

wherein the processor is further configured to prune the elements of the subcolumns by replacing the elements of the subcolumns with a predetermined value such that a ratio of elements other than the predetermined value in the elements of the subcolumns is the target sparsity,

wherein the processor is further configured to prune the elements of the subcolumns by sequentially replacing the elements other than the predetermined value with the predetermined value, starting from an element having a smallest absolute value.

10. The neural network-based inference apparatus of claim 9 , wherein the processor is further configured:

to generate a first subcolumn corresponding to a first processing element of the matrix, and

to generate a second subcolumn corresponding to a second processing element of the matrix.

11. The neural network-based inference apparatus of claim 9 , wherein the processor is further configured:

to calculate the workloads corresponding to neurons of the neural network,

to assign the workloads to the compressed matrix, and

to balance the workloads by reordering the compressed matrix based on the workloads assigned to the compressed matrix.

12. The neural network-based inference apparatus of claim 11 , wherein the processor is further configured to calculate the workloads by accumulating numbers of elements other than a predetermined value in the neurons of the neural network.

13. The neural network-based inference apparatus of claim 11 , wherein the processor is further configured to assign the workloads to columns of the compressed matrix.

14. The neural network-based inference apparatus of claim 11 , wherein the processor is further configured:

to compare workloads corresponding to columns of the compressed matrix, and

to perform balancing by swapping the columns of the compressed matrix based on a result of the comparing.

15. The neural network-based inference apparatus of claim 9 , wherein the processor is further configured to encode the reordered matrix based on any one or any combination of components of the processing elements, positions of the processing elements, and a size of the reordered matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2021
From: GAO, CHANG; LIU, SHIH-CHII; DELBRUCK, TOBI; CHEN, XI
To: SAMSUNG ELECTRONICS CO., LTD.; UNIVERSITY OF ZURICH
Reel/Frame 056486/0795 →
Priority Claims (1)
KR 10-2021-0019755 · Feb 15, 2021 · national
Continuity (1)
Related Publication 20220261649A1 · Aug 18, 2022
References Cited (16)
US 10366322B2 · David · 2019 [cited by examiner]
US 10891538B2 · Dally · 2021 [cited by examiner]
US 10936947B1 · Flunkert · 2021 [cited by examiner]
US 12056614B2 · Gorokhov · 2024 [cited by examiner]
US 20180046914A1 · Li et al. · 2018 [cited by applicant]
US 20180082181A1 · Brothers · 2018 [cited by examiner]
US 20190130271A1 · Narang et al. · 2019 [cited by applicant]
US 20190205746A1 · Nurvitadhi et al. · 2019 [cited by applicant]
US 20210125070A1 · Wang · 2021 [cited by examiner]
CN 110110851A · 2019 [cited by applicant]
Dong, Peiyan, et al. “Rtmobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech Recognition.” [cited by applicant]
Shi, Runbin, et al. “Csb-rnn: A faster-than-realtime rnn acceleration framework with compressed structured blocks.” [cited by applicant]
Extended European Search Report issued on Feb. 22, 2022 in counterpart European Patent Application No. 21191444.5 (9 pages in English). [cited by applicant]
Han, Song, et al. “ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA.” [cited by applicant]
Cao, Shijie, et al. “Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity.” [cited by applicant]
Park, Junki, et al. “Balancing Computation Loads and Optimizing Input Vector Loading in LSTM Accelerators.” [cited by applicant]