IP Library Granted Patent US 12,224,774
Granted Patent B2
US 12,224,774 · App. 18/096,551 · Granted Feb 11, 2025

Runtime reconfigurable compression format conversion

Inventors: Jong Hoon Shin (San Jose, CA); Ardavan Pedram (Santa Clara, CA); Joseph Hassoun (Los Gatos, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H03M7/3059G06F8/4432G06F8/4434G06N3/0495H03M7/3079H03M7/6011H03M7/6064H03M7/6094G06N3/063H03M7/46H03M7/6017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,224,774
App. No.
18/096,551
Filed
Jan 12, 2023
Granted
Feb 11, 2025
Kind
B2
Art Unit
2845
USPC
341/50
Abstract

A runtime data-format optimizer for a processing element includes a sparsity-detector and a compression-converter. The sparsity-detector selects a first compression-conversion format during a runtime of the processing element based on a performance model that is based on a first sparsity pattern of first data stored in a first memory that is exterior to the processing element and a second sparsity pattern of second data that is to be stored in a second memory within the processing element. The second sparsity pattern is based on a runtime configuration of the processing element. The first data is stored in the first memory using a first compression format and the second data is to be stored in the second memory using a second compression format. The compression-conversion circuit converts the first compression format of the first data to be the second compression format of the second data based on the first compression-conversion format.

Claims (23)

1. A data-format optimizer for a processing element, comprising:

a sparsity-detection circuit configured to select a first compression-conversion format during a runtime of the processing element based, at least in part, on a first performance model that is based on a first sparsity pattern of first data stored in a first memory that is exterior to the processing element and a second sparsity pattern of second data that is to be stored in a second memory within the processing element, the second sparsity pattern being based on a runtime configuration of the processing element, the first data being stored in the first memory using a first compression format and the second data to be stored in the second memory using a second compression format; and

a compression-conversion circuit configured to convert the first compression format of the first data to be the second compression format of the second data based on the first compression-conversion format.

2. The data-format optimizer of claim 1 , wherein the first performance model further includes a latency associated with converting the first data to the second data.

3. The data-format optimizer of claim 2 , wherein the first performance model further includes an energy consumption associated with converting the first data to the second data.

4. The data-format optimizer of claim 1 , wherein the first performance model further includes an energy consumption associated with converting the first data to the second data.

5. The data-format optimizer of claim 1 , wherein the sparsity-detection circuit is further configured to select a second compression-conversion format during runtime of the processing element based, at least in part, on a second performance model that is based on a third sparsity pattern of output data stored in a third memory within the processing element; and

wherein the compression-conversion circuit configured to convert the output data to a third compression format that is based on the second compression-conversion format for storage in the first memory.

6. The data-format optimizer of claim 1 , wherein the sparsity-detection circuit is further configured to generate two or more performance models from which the first performance model is selected.

7. The data-format optimizer of claim 6 , wherein the sparsity-detection circuit is configured to sample the first sparsity pattern of first data to generate the two or more performance models.

8. The data-format optimizer of claim 7 , wherein the sparsity-detection circuit is configured to generate the two or more performance models further based on sparsity operational modes of the processing element.

9. The data-format optimizer of claim 1 , wherein at least one of the first compression format and the second compression format comprises a Run Length Encoding (RLE), a Zero Value Compression (ZVC), a Compressed Sparse Column (CSC) encoding, or a Compressed Sparse Row (CSR) encoding.

10. A data-format optimizer for a processing element, comprising:

a sparsity-detection circuit configured to generate, during a runtime of the processing element, at least one or more performance models that are based on a first sparsity pattern of first data stored in a first memory and a second sparsity pattern of second data that is to be stored in a second memory, the first memory being exterior to the processing element and the second memory being within the processing element, and to determine a first compression-conversion format based, at least in part, a selection of the one or more performance models; and

a compression-conversion circuit configured to apply the first compression-conversion format to the first data to become the second data.

11. The data-format optimizer of claim 10 , wherein the sparsity-detection circuit is configured to sample the first sparsity pattern of the first data to generate the one or more performance models.

12. The data-format optimizer of claim 11 , wherein the sparsity-detection circuit is configured to generate the one or more performance models further based on sparsity operational modes of the processing element.

13. The data-format optimizer of claim 10 , wherein each performance model includes a latency associated with converting the first data to the second data.

14. The data-format optimizer of claim 13 , wherein at least one performance model further includes an energy consumption associated with converting the first data to the second data.

15. The data-format optimizer of claim 10 , wherein at least one performance model further includes an energy consumption associated with converting the first data to the second data.

16. The data-format optimizer of claim 10 , wherein the sparsity-detection circuit is further configured to select a second compression-conversion format during the runtime of the processing element based, at least in part, on a second performance model that is based on a third sparsity pattern of output data stored in a third memory within the processing element; and

wherein the compression-conversion circuit configured to convert the output data to a third compression format that is based on the second compression-conversion format for storage in the first memory.

17. The data-format optimizer of claim 10 , wherein the first compression format comprises a Run Length Encoding (RLE), a Zero Value Compression (ZVC), a Compressed Sparse Column (CSC) encoding, or a Compressed Sparse Row (CSR) encoding.

Continuity (2)
Provisional Application 63426027 · Nov 16, 2022
Related Publication 20240162916A1 · May 16, 2024
References Cited (20)
US 10186011B2 · Nurvitadhi et al. · 2019 [cited by applicant]
US 10346944B2 · Nurvitadhi et al. · 2019 [cited by applicant]
US 10546393B2 · Ray et al. · 2020 [cited by applicant]
US 10951897B2 · Taubman et al. · 2021 [cited by applicant]
US 11489541B2 · Latorre · 2022 [cited by examiner]
US 11757469B2 · Kulkarni · 2023 [cited by examiner]
US 20180189675A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20200364508A1 · Gurel et al. · 2020 [cited by applicant]
US 20210103550A1 · Appu et al. · 2021 [cited by applicant]
US 20210146531A1 · Tremblay et al. · 2021 [cited by applicant]
US 20220129521A1 · Surti et al. · 2022 [cited by applicant]
US 20220318557A1 · Mohseni et al. · 2022 [cited by applicant]
Cavigelli, Lukas et al., “EBPC: Extended Bit-Plane Compression for Deep Neural Network Inference and Training Accelerators” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 9, No. 4, 2019, 13 p… [cited by applicant]
Dong, Shi et al., “Spartan: a Sparsity-Adaptive Framework to Accelerate Deep Neural Network Training on GPUs” IEEE Transactions on Parallel and Distributed Systems, vol. 32, No. 10, 2021, 17 pages. [cited by applicant]
Ko, Yousun et al., “Lane Compression: a Lightweight Lossless Compression Method for Machine Learning on Embedded Systems” ACM Transactions on Embedded Computing Systems, vol. 20, No. 2, 2021, 26 pages. [cited by applicant]
Kwon, Jisu et al., “Sparse Convolutional Neural Network Acceleration with Lossless Input Feature Map Compression for Resource-Constrained Systems”, IET Computers & Digital Techniques, vol. 16, No. 1, 2021, pp. 29-43. [cited by applicant]
Parashar, Angshuman et al., “Timeloop: a Systematic Approach to DNN Accelerator Evaluation” 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2019, 12 pages. [cited by applicant]
Qin, Eric et al., “Extending Sparse Tensor Accelerators to Support Multiple Compression Formats” 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2021, 11 pages. [cited by applicant]
Qin, Eric et al., “MINT: Microarchitecture for Efficient and Interchangeable CompressioN Formats on Tensor Algebra” MICRO 2020, Sandia National Lab.(SNL-NM), No. SAND2020-4760C, 2020, 13 pages. [cited by applicant]
Yazdanbakhsh, Amir et al., “Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation”, 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2022, 19 pages. [cited by applicant]