IP Library › Granted Patent US 12,231,152
Granted Patent B2
US 12,231,152 · App. 18/096,557 · Granted Feb 18, 2025

Runtime reconfigurable compression format conversion with bit-plane granularity

Inventors: Jong Hoon Shin (San Jose, CA); Ardavan Pedram (Santa Clara, CA); Joseph Hassoun (Los Gatos, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H03M7/6088G06F15/781G06N3/0495G06N3/063H03M7/3079H03M7/42H03M7/6082H03M7/46H03M7/6017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,231,152
App. No.
18/096,557
Granted
Feb 18, 2025
Kind
B2
Abstract

A runtime bit-plane data-format optimizer for a processing element includes a sparsity-detector and a compression-converter. The sparsity-detector selects a bit-plane compression-conversion format during a runtime of the processing element using a performance model that is based on a first sparsity pattern of first bit-plane data stored in a memory exterior to the processing element and a second sparsity pattern of second bit-plane data that is to be stored in a memory within the processing element. The second sparsity pattern is based on a runtime configuration of the processing element. The first bit-plane data is stored using a first bit-plane compression format and the bit-plane second data is to be stored using a second bit-plane compression format. The compression-conversion circuit converts the first bit-plane compression format of the first data to be the second bit-plane compression format of the second data.

Claims (23)

1. A data-format optimizer for a processing element, comprising:

a sparsity-detection circuit configured to select a first bit-plane compression-conversion format during a runtime of the processing element based, at least in part, on a first performance model that is based on a first sparsity pattern of a first bit-plane data stored in a first memory that is exterior to the processing element and a second sparsity pattern of a second bit-plane data that is to be stored in a second memory within the processing element, the second sparsity pattern being based on a runtime configuration of the processing element, the first bit-plane data being stored in the first memory using a first bit-plane compression format and is to be converted to a second bit-plane compression format that is to be stored in the second memory using a second compression format; and

a compression-conversion circuit configured to convert the first bit-plane data in the first bit-plane compression format to be the second bit-plane data in the second bit-plane compression format based on the first bit-plane compression-conversion format.

2. The data-format optimizer of claim 1 , wherein the first performance model further includes a latency associated with converting the first bit-plane data to the second bit-plane data.

3. The data-format optimizer of claim 2 , wherein the first performance model further includes an energy consumption associated with converting the first bit-plane data to the second bit-plane data.

4. The data-format optimizer of claim 1 , wherein the first performance model further includes an energy consumption associated with converting the first bit-plane data to the second bit-plane data.

5. The data-format optimizer of claim 1 , wherein the sparsity-detection circuit is further configured to select a second compression-conversion format during runtime of the processing element based, at least in part, on a second performance model that is based on a third sparsity pattern of an output bit-plane data stored in a third memory within the processing element; and

wherein the compression-conversion circuit configured to convert the output bit-plane data to a third compression format that is based on the second compression-conversion format for storage in the first memory.

6. The data-format optimizer of claim 1 , wherein the sparsity-detection circuit is further configured to generate two or more performance models from which the first performance model is selected.

7. The data-format optimizer of claim 6 , wherein the sparsity-detection circuit is configured to sample the first sparsity pattern of the first bit-plane data to generate the two or more performance models.

8. The data-format optimizer of claim 7 , wherein the sparsity-detection circuit is configured to generate the two or more performance models further based on sparsity operational modes of the processing element.

9. The data-format optimizer of claim 1 , wherein at least one of the first bit-plane compression format and the second bit-plane compression format comprises a Run Length Encoding (RLE), a Zero Value Compression (ZVC), a Compressed Sparse Column (CSC) encoding, or a Compressed Sparse Row (CSR) encoding.

10. A data-format optimizer for a processing element, comprising:

a sparsity-detection circuit configured to generate, during a runtime of the processing element, at least one or more performance models that are based on a first sparsity pattern of first bit-plane data stored in a first memory in a first bit-plane compression format and a second sparsity pattern of second bit-plane data that is to be stored in a second memory in a second bit-plane compression format, the first memory being exterior to the processing element and the second memory being within the processing element, and to determine a first bit-plane compression-conversion format based, at least in part, on a selection of one of the one or more performance models; and

a compression-conversion circuit configured to convert the first bit-plane data in the first bit-plane compression format to the second bit-plane data in the second bit-plane compression format based on the first bit-plane compression-conversion format.

11. The data-format optimizer of claim 10 , wherein the sparsity-detection circuit is configured to sample the first sparsity pattern of the first bit-plane data to generate the one or more performance models.

12. The data-format optimizer of claim 11 , wherein the sparsity-detection circuit is configured to generate the one or more performance models further based on sparsity operational modes of the processing element.

13. The data-format optimizer of claim 10 , wherein each performance model includes a latency associated with converting the first bit-plane data to the second bit-plane data.

14. The data-format optimizer of claim 13 , wherein at least one performance model further includes an energy consumption associated with converting the first bit-plane data to the second bit-plane data.

15. The data-format optimizer of claim 10 , wherein at least one performance model further includes an energy consumption associated with converting the first bit-plane data to the second bit-plane data.

16. The data-format optimizer of claim 10 , wherein the sparsity-detection circuit is further configured to select a second compression-conversion format during the runtime of the processing element based, at least in part, on a second performance model that is based on a third sparsity pattern of an output bit-plane data stored in a third memory within the processing element; and

wherein the compression-conversion circuit configured to convert the output bit-plane data to a third compression format that is based on the second compression-conversion format for storage in the first memory.

17. The data-format optimizer of claim 10 , wherein the bit-plane first compression format comprises a Run Length Encoding (RLE), a Zero Value Compression (ZVC), a Compressed Sparse Column (CSC) encoding, or a Compressed Sparse Row (CSR) encoding.

Continuity (2)
Provisional Application 63426029 · Nov 16, 2022
Related Publication 20240162917A1 · May 16, 2024
References Cited (21)
US 6195465B1 · Zandi et al. · 2001 [cited by applicant]
US 10346944B2 · Nurvitadhi et al. · 2019 [cited by applicant]
US 10951897B2 · Taubman et al. · 2021 [cited by applicant]
US 11392829B1 · Pool · 2022 [cited by examiner]
US 20180189675A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20200364508A1 · Gurel et al. · 2020 [cited by applicant]
US 20210004668A1 · Moshovos et al. · 2021 [cited by applicant]
US 20210065052A1 · Muralidharan · 2021 [cited by examiner]
US 20220261649A1 · Gao et al. · 2022 [cited by applicant]
CN 114118394A · 2022 [cited by applicant]
CN 114723033A · 2022 [cited by applicant]
EP 3907601A1 · 2021 [cited by applicant]
GB 2325584A · 1998 [cited by applicant]
WO 2021226720A1 · 2021 [cited by applicant]
Cavigelli, Lukas et al., “EBPC: Extended Bit-Plane Compression for Deep Neural Network Inference and Training Accelerators” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 9, No. 4, 2019, 13 p… [cited by applicant]
Dong, Shi et al., “Spartan: A Sparsity-Adaptive Framework to Accelerate Deep Neural Network Training on GPUs” IEEE Transactions on Parallel and Distributed Systems, vol. 32, No. 10, 2021, 17 pages. [cited by applicant]
Ko, Yousun et al., “Lane Compression: A Lightweight Lossless Compression Method for Machine Learning on Embedded Systems” ACM Transactions on Embedded Computing Systems, vol. 20, No. 2, 2021, 26 pages. [cited by applicant]
Parashar, Angshuman et al., “Timeloop: A Systematic Approach to DNN Accelerator Evaluation” 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2019, 12 pages. [cited by applicant]
Qin, Eric et al., “Extending Sparse Tensor Accelerators to Support Multiple Compression Formats” 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2021, 11 pages. [cited by applicant]
Qin, Eric et al., “MINT: Microarchitecture for Efficient and Interchangeable CompressioN Formats on Tensor Algebra” MICRO 2020, Sandia National Lab.(SNL-NM), No. SAND2020-4760C, 2020, 13 pages. [cited by applicant]
Sun, Gongjin, “Scalable Scientific Computation Acceleration Using Hardware-Accelerated Compression”, Dissertation UC Irvine, 2022, 140 pages. [cited by applicant]
Cited By (1)
US 12,700,877