IP Library Granted Patent US 12,499,357
Granted Patent B2
US 12,499,357 · App. 17/138,421 · Granted Dec 16, 2025

Data compression method, data compression system and operation method of deep learning acceleration chip

Inventors: Kai-Jiun Yang (Zhubei, TW); Gwo-Giun Lee (Tainan, TW); Chi-Tien Sun (Hsinchu, TW)
Assignee: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
G06N3/065G06F9/30036G06F9/3836G06F9/3877G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,357
App. No.
17/138,421
Granted
Dec 16, 2025
Kind
B2
Abstract

A data compression method, a data compression system and an operation method of a deep learning acceleration chip are provided. The data compression method includes the following steps. A filter coefficient tensor matrix of a deep learning model is obtained. A matrix decomposition procedure is performed according to the filter coefficient tensor matrix to obtain a sparse tensor matrix and a transformation matrix, which is an orthonormal matrix. The product of the transformation matrix and the filter coefficient tensor matrix is the sparse tensor matrix. The sparse tensor matrix is compressed. The sparse tensor matrix and the transformation matrix, or the sparse tensor matrix and a restoration matrix, are stored in a memory. A convolution operation result is obtained by the deep learning acceleration chip using the sparse tensor matrix. The convolution operation result is restored by the deep learning acceleration chip using the restoration matrix.

Claims (23)

1 . A data compression method of a deep learning acceleration chip, comprising:

obtaining a filter coefficient tensor matrix of a deep learning model;

using a relationship that product of a transformation matrix and the filter coefficient tensor matrix is a sparse tensor matrix to convert the filter coefficient tensor matrix to the sparse tensor matrix and the transformation matrix is an orthonormal matrix;

compressing the sparse tensor matrix; and

storing the sparse tensor matrix and the transformation matrix, or storing the sparse tensor matrix and a restoration matrix, in a memory, wherein after the deep learning acceleration chip obtains a recognition data, the deep learning acceleration chip uses a relationship that product of the sparse tensor matrix and the recognition data is the convolution operation result and a relationship that product of the restoration matrix and the convolution operation result is a final convolution operation result, to obtain the final convolution operation result.

2 . The data compression method of the deep learning acceleration chip according to claim 1 , wherein the restoration matrix is a transpose matrix of the transformation matrix.

3 . The data compression method of the deep learning acceleration chip according to claim 1 , wherein the convolution operation result is restored without loss of information.

4 . The data compression method of the deep learning acceleration chip according to claim 1 , wherein in the step of decomposing the filter coefficient tensor matrix, the filter coefficient tensor matrix is partitioned into M parts, the at least one sparse tensor matrix has a quantity of M, the at least one transformation matrix has a quantity of M, and M is a natural number.

5 . The data compression method of the deep learning acceleration chip according to claim 4 , wherein M is 2N, and N is a natural number.

6 . The data compression method of the deep learning acceleration chip according to claim 4 , wherein the filter coefficient tensor matrix is equally partitioned.

7 . The data compression method of the deep learning acceleration chip according to claim 4 , wherein the sparse tensor matrixes have the same size.

8 . The data compression method of the deep learning acceleration chip according to claim 4 , wherein the transformation matrixes have the same size.

9 . A data compression system of a deep learning acceleration chip, wherein the data compression system is used to reduce data movement for a deep learning model and comprises:

a decomposition unit configured to use a relationship that product of a transformation matrix and the filter coefficient tensor matrix is a sparse tensor matrix to convert the filter coefficient tensor matrix to the sparse tensor matrix and the transformation matrix is an orthonormal matrix;

a compression unit configured to compress the sparse tensor matrix; and

a transfer unit configured to store the sparse tensor matrix and the transformation matrix, or store the sparse tensor matrix and a restoration matrix, in a memory, wherein after the deep learning acceleration chip obtains a recognition data, the deep learning acceleration chip uses a relationship that product of the sparse tensor matrix and the recognition data is the convolution operation result and a relationship that product of the restoration matrix and the convolution operation result is a final convolution operation result, to obtain the final convolution operation result.

10 . The data compression system of the deep learning acceleration chip according to claim 9 , wherein the restoration matrix is a transpose matrix of the transformation matrix.

11 . The data compression system of the deep learning acceleration chip according to claim 9 , wherein the convolution operation result is restored without loss of information.

12 . The data compression system of the deep learning acceleration chip according to claim 9 , wherein the decomposition unit further partitions the filter coefficient tensor matrix into M parts, the at least one sparse tensor matrix has a quantity of M, the at least one transformation matrix has a quantity of M, and M is a natural number.

13 . The data compression system of the deep learning acceleration chip according to claim 12 , wherein M is 2N, and N is a natural number.

14 . The data compression system of the deep learning acceleration chip according to claim 12 , wherein the filter coefficient tensor matrix is equally partitioned.

15 . The data compression system of the deep learning acceleration chip according to claim 12 , wherein the sparse tensor matrixes have the same size.

16 . The data compression system of the deep learning acceleration chip according to claim 12 , wherein the transformation matrixes have the same size.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2020
From: YANG, KAI-JIUN; LEE, GWO-GIUN; SUN, CHI-TIEN
To: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
Reel/Frame 054786/0440 →
Continuity (1)
Related Publication 20220207342A1 · Jun 30, 2022
References Cited (39)
US 5754456A · Eitan · 1998 [cited by examiner]
US 10411727B1 · Lan et al. · 2019 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 11086968B1 · Baskaran · 2021 [cited by examiner]
US 11392829B1 · Pool · 2022 [cited by examiner]
US 20170235848A1 · Van Dusen et al. · 2017 [cited by applicant]
US 20180032857A1 · Lele · 2018 [cited by examiner]
US 20180121796A1 · Deisher · 2018 [cited by examiner]
US 20190190538A1 · Park et al. · 2019 [cited by applicant]
US 20190244080A1 · Li et al. · 2019 [cited by applicant]
US 20190266485A1 · Singh · 2019 [cited by examiner]
US 20190303757A1 · Wang et al. · 2019 [cited by applicant]
US 20190311253A1 · Chung et al. · 2019 [cited by applicant]
US 20190311254A1 · Turek et al. · 2019 [cited by applicant]
US 20190340510A1 · Li et al. · 2019 [cited by applicant]
US 20200082268A1 · Chiu et al. · 2020 [cited by applicant]
US 20200410327A1 · Chinya · 2020 [cited by examiner]
US 20220108156A1 · Hunter · 2022 [cited by examiner]
US 20220108157A1 · Hunter · 2022 [cited by examiner]
CN 106877876A · 2017 [cited by applicant]
CN 110390622A · 2019 [cited by applicant]
CN 110443354A · 2019 [cited by applicant]
CN 110874631B · 2020 [cited by applicant]
CN 107609641B · 2020 [cited by applicant]
TW I669921B · 2019 [cited by applicant]
TW 201944745A · 2019 [cited by applicant]
TW 202004658A · 2020 [cited by applicant]
TW I687063B · 2020 [cited by applicant]
WO WO2020233130A1 · 2020 [cited by applicant]
Chen et al., “Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices,” IEEE Journal On Emerging and Selected Topics in Circuits and Systems, vol. 9, No. 2, Jun. 2019, pp. 292-308. [cited by applicant]
Chen et al., “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” IEEE Journal of Solid-State Circuits, 2016, pp. 1-12. [cited by applicant]
Han et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” Stanford University, May 3, 2016, 12 pages. [cited by applicant]
Jiao et al., “A 12nm Programmable Convolution-Efficient Neural-Processing-Unit Chip Achieving 825TOPS,” IEEE International Solid-State Circuits Conference, SESSION 7 / High-Performance Machine Learning / 7.2, 2020, pp. … [cited by applicant]
Kang et al., “GANPU: A 135TFLOPS/W Multi-DNN Training Processor for GANs with Speculative Dual-Sparsity Exploitation,” IEEE International Solid-State Circuits Conference, SESSION 7 / High-Performance Machine Learning / … [cited by applicant]
Lin et al., “A 3.4-to-13.3TOPS/W 3.6TOPS Dual-Core Deep-Learning Accelerator for Versatile AI Applications in 7nm 56 Smartphone Soc,” IEEE International Solid-State Circuits Conference, SESSION 7 / High-Performance Mach… [cited by applicant]
Valavi et al., “A Mixed-Signal Binarized Convolutional-Neural-Network Accelerator Integrating Dense Weight Storage and Multiplication for Reduced Data Movement,” Symposium on VLSI Circuits Digest of Technical Papers, 20… [cited by applicant]
Yue et al., “A 65nm 0.39-to-140.3TOPS/W 1-to-12b Unified Neural-Network Processor Using Block-Circulant-Enabled Transpose-Domain Acceleration with 8.1 x Higher TOPS/mm2 and 6T HBST-TRAM-Based 2D Data-Reuse Architecture,… [cited by applicant]
Taiwanese Office Action and Search Report for Taiwanese Application No. 110100978, dated Mar. 28, 2022. [cited by applicant]
Chinese Office Action and Search Report for Chinese Application No. 202110419932.X, dated Aug. 1, 2024. [cited by applicant]