IP Library › Granted Patent US 12,361,266
Granted Patent B2
US 12,361,266 · App. 16/990,970 · Granted Jul 15, 2025

Hierarchical weight preprocessing for neural network accelerator

Inventors: Jong Hoon Shin (San Jose, CA); Ali Shafiee Ardestani (San Jose, CA); Hamzah Ahmed Ali Abdelaziz (San Jose, CA); Joseph H. Hassoun (Los Gatos, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/06G06F17/16G06N3/0464G06N3/063G06N3/08G06N7/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,266
App. No.
16/990,970
Granted
Jul 15, 2025
Kind
B2
Abstract

A system and method for weight preprocessing. In some embodiments, the method includes performing intra-tile preprocessing of a first weight tensor to form a first pre-processed weight tensor, and performing inter-tile preprocessing of the first pre-processed weight tensor, to form a second pre-processed weight tensor. The intra-tile preprocessing may include moving a first element of a first weight tile of the first weight tensor by one position, within the first weight tile, in a lookahead direction or in a lookaside direction. The inter-tile preprocessing may include moving a first row of a weight tile of the first pre-processed weight tensor by one position in a lookahead direction or by one position in a lookaside direction.

Claims (67)

1. A method, comprising:

performing intra-tile preprocessing of a first weight tensor to form a first pre-processed weight tensor,

performing inter-tile preprocessing of the first pre-processed weight tensor, to form a second pre-processed weight tensor,

performing, by a processing circuit, a first computation cycle on a first row of the second pre-processed weight tensor,

skipping, by the processing circuit, a second row of the second pre-processed weight tensor based on the intra-tile preprocessing and the inter-tile preprocessing, and

performing, by the processing circuit, a second computation cycle on a third row of the second pre-processed weight tensor in place of the second row of the second pre-processed weight tensor,

such that the first row of the second pre-processed weight tensor and the third row of the second pre-processed weight tensor are processed in two computation cycles, instead of three computational cycles,

the intra-tile preprocessing comprising causing a first row of a first weight tile of the first weight tensor to be empty by moving a first element of the first weight tile of the first weight tensor by one position, within the first weight tile, in a lookahead direction or in a lookaside direction, and

the inter-tile preprocessing comprising causing the second row of the second pre-processed weight tensor to be empty by moving a second row of a first weight tile of the first pre-processed weight tensor by one position, in a lookahead direction or by one position in a lookaside direction, based on a second row of a second weight tile of the first pre-processed weight tensor being empty.

2. The method of claim 1 , wherein the intra-tile preprocessing comprises moving the first element of the first weight tile of the first weight tensor by one position, within the first weight tile, in the lookahead direction.

3. The method of claim 1 , wherein the inter-tile preprocessing comprises moving the second row by one position in a lookahead direction.

4. The method of claim 1 , wherein the intra-tile preprocessing further comprises moving a first element of a second weight tile of the first weight tensor by one position, within the second weight tile, in a lookaside direction.

5. The method of claim 1 , wherein

the inter-tile preprocessing comprises moving a second row of the second weight tile to the first weight tile, in a lookaside direction.

6. The method of claim 5 , wherein the inter-tile preprocessing further comprises creating a tile sparsity map corresponding to the first pre-processed weight tensor, the tile sparsity map having:

a column for each weight tile of the first pre-processed weight tensor, and

a row for each row of the weight tiles,

the tile sparsity map indicating positions of empty rows of the weight tiles of the first pre-processed weight tensor.

7. The method of claim 6 , wherein the tile sparsity map has one fewer dimension than the first pre-processed weight tensor.

8. The method of claim 6 , further comprising identifying the first row of the first weight tile based on the tile sparsity map.

9. The method of claim 6 , further comprising:

multiplying a third row of the first weight tile by a first vector of activations, to form a first dot product,

wherein the multiplying comprises fetching the first vector of activations from a column of an activations buffer, the column of the activations buffer being second in the activations buffer.

10. The method of claim 6 , further comprising:

multiplying, in a first processing element circuit, a third row of the first weight tile by a first vector of activations, to form a first dot product,

multiplying, in a second processing element circuit, a third row of weights, of the first pre-processed weight tensor, by a second vector of activations, to form a second dot product, and

adding the first dot product and the second dot product.

11. The method of claim 6 , wherein the inter-tile preprocessing further comprises moving a third row of the first pre-processed weight tensor by one position in a lookahead direction.

12. The method of claim 11 , further comprising identifying the third row based on the tile sparsity map.

13. A system, comprising:

a first processing circuit, the first processing circuit being configured to:

perform intra-tile preprocessing of a first weight tensor to form a first pre-processed weight tensor,

perform inter-tile preprocessing of the first pre-processed weight tensor, to form a second pre-processed weight tensor,

perform a first computation cycle on a first row of the second pre-processed weight tensor,

skip a second row of the second pre-processed weight tensor based on the intra-tile preprocessing and the inter-tile preprocessing, and

perform a second computation cycle on a third row of the second pre-processed weight tensor in place of the second row of the second pre-processed weight tensor,

such that the first row of the second pre-processed weight tensor and the third row of the second pre-processed weight tensor are processed in two computation cycles, instead of three computational cycles,

the intra-tile preprocessing comprising causing a first row of a first weight tile of the first weight tensor to be empty by moving a first element of the first weight tile of the first weight tensor by one position, within the first weight tile, in a lookahead direction or in a lookaside direction, and

the inter-tile preprocessing comprising causing the second row of the second pre-processed weight tensor to be empty by moving a second row of a first weight tile of the first pre-processed weight tensor by one position, in a lookahead direction or by one position in a lookaside direction, based on a second row of a second weight tile of the first pre-processed weight tensor being empty.

14. The system of claim 13 , wherein the intra-tile preprocessing comprises moving the first element of the first weight tile of the first weight tensor by one position, within the first weight tile, in the lookahead direction.

15. The system of claim 13 , wherein the inter-tile preprocessing comprises moving the second row by one position in a lookahead direction.

16. The system of claim 13 , wherein the intra-tile preprocessing further comprises moving a first element of a second weight tile of the first weight tensor by one position, within the second weight tile, in a lookaside direction.

17. The system of claim 13 , wherein

the inter-tile preprocessing comprises moving a second row of the second weight tile to the first weight tile, in a lookaside direction.

18. The system of claim 17 , wherein the inter-tile preprocessing further comprises creating a tile sparsity map corresponding to the first pre-processed weight tensor, the tile sparsity map having:

a column for each weight tile of the first pre-processed weight tensor, and

a row for each row of the weight tiles,

the tile sparsity map indicating positions of empty rows of the weight tiles of the first pre-processed weight tensor.

19. The system of claim 18 , further comprising a second processing circuit comprising:

a first processing element circuit, and

a second processing element circuit,

wherein:

the first processing element circuit is configured to multiply a third row of the first weight tile by a first vector of activations, to form a first dot product; and

the second processing element circuit is configured to:

multiply a fourth row of weights, of the first pre-processed weight tensor, by a second vector of activations, to form a second dot product, and

add the first dot product and the second dot product.

20. A system, comprising:

a processing circuit; and

a memory communicatively coupled to the processing circuit, the memory storing instructions, which, based on being executed by the processing circuit, cause the processing circuit to:

perform intra-tile preprocessing of a first weight tensor to form a first pre-processed weight tensor,

perform inter-tile preprocessing of the first pre-processed weight tensor, to form a second pre-processed weight tensor,

perform a first computation cycle on a first row of the second pre-processed weight tensor,

skip a second row of the second pre-processed weight tensor based on the intra-tile preprocessing and the inter-tile preprocessing, and

perform a second computation cycle on a third row of the second pre-processed weight tensor in place of the second row of the second pre-processed weight tensor,

such that the first row of the second pre-processed weight tensor and the third row of the second pre-processed weight tensor are processed in two computation cycles, instead of three computational cycles,

the intra-tile preprocessing comprising causing a first row of a first weight tile of the first weight tensor to be empty by moving a first element of a first weight tile of the first weight tensor by one position, within the first weight tile, in a lookahead direction or in a lookaside direction, and

the inter-tile preprocessing comprising causing the second row of the second pre-processed weight tensor to be empty by moving a second row of a first weight tile of the first pre-processed weight tensor by one position, in a lookahead direction or by one position in a lookaside direction, based on a second row of a second weight tile of the first pre-processed weight tensor being empty.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2020
From: SHIN, JONG HOON; SHAFIEE ARDESTANI, ALI; ABDELAZIZ, HAMZAH AHMED ALI; HASSOUN, JOSEPH H.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 054241/0616 →
Continuity (2)
Provisional Application 63024676 · May 14, 2020
Related Publication 20210357748A1 · Nov 18, 2021
References Cited (73)
US 10963787B2 · Zlateski · 2021 [cited by examiner]
US 11422781B1 · Neuendorffer · 2022 [cited by examiner]
US 11429852B2 · Lu · 2022 [cited by examiner]
US 20100312735A1 · Knoblauch · 2010 [cited by applicant]
US 20170124464A1 · Crabtree et al. · 2017 [cited by applicant]
US 20170228645A1 · Wang et al. · 2017 [cited by applicant]
US 20180046916A1 · Dally · 2018 [cited by examiner]
US 20180150721A1 · Mostafa · 2018 [cited by examiner]
US 20180253816A1 · John · 2018 [cited by applicant]
US 20180285254A1 · Baum · 2018 [cited by examiner]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190205696A1 · Owechko · 2019 [cited by examiner]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190332941A1 · Towal · 2019 [cited by examiner]
US 20190370631A1 · Fais · 2019 [cited by examiner]
US 20190378011A1 · Iwakura · 2019 [cited by examiner]
US 20190379396A1 · Bajic · 2019 [cited by examiner]
US 20200104692A1 · Hill · 2020 [cited by examiner]
US 20200159809A1 · Catthoor · 2020 [cited by examiner]
US 20200226473A1 · Sharma · 2020 [cited by examiner]
US 20200228137A1 · Chinya · 2020 [cited by examiner]
US 20200278888A1 · Connor · 2020 [cited by examiner]
US 20200285949A1 · Baum · 2020 [cited by examiner]
US 20200285950A1 · Baum · 2020 [cited by examiner]
US 20200293770A1 · Chen · 2020 [cited by examiner]
US 20210004668A1 · Moshovos · 2021 [cited by examiner]
US 20210103550A1 · Appu · 2021 [cited by examiner]
US 20210125070A1 · Wang · 2021 [cited by examiner]
US 20210141571A1 · Lew · 2021 [cited by examiner]
US 20210191724A1 · Pal · 2021 [cited by examiner]
US 20210191733A1 · Gunnam · 2021 [cited by examiner]
US 20210209450A1 · Cassidy · 2021 [cited by examiner]
US 20210303909A1 · Gunnam · 2021 [cited by examiner]
US 20210303976A1 · Gunnam · 2021 [cited by examiner]
US 20210303993A1 · Saeedi · 2021 [cited by examiner]
US 20210326683A1 · Narayanaswami · 2021 [cited by examiner]
US 20220012593A1 · Huang · 2022 [cited by examiner]
US 20220245453A1 · Majnemer · 2022 [cited by examiner]
US 20230052942A1 · Chauhan · 2023 [cited by examiner]
WO WO2019157599A1 · 2019 [cited by examiner]
WO WO2020014590A1 · 2020 [cited by examiner]
WO WO2020190772A1 · 2020 [cited by examiner]
WO WO2020190808A1 · 2020 [cited by examiner]
WO WO2020190809A1 · 2020 [cited by examiner]
Stuart, Dylan “An Efficient Hardware Architecture for Exploiting Sparsity in Neural Networks” Nov. 27, 2019, pp. 1-66. (Year: 2019). [cited by examiner]
Srivastava et al., “Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations” Apr. 16, 2020, pp. 689-702. (Year: 2020). [cited by examiner]
Gondimalla et al., “SparTen: A Sparse Tensor Accelerator for Convolutional Neural Networks” Oct. 2019, pp. 151-165. (Year: 2019). [cited by examiner]
Moreau et al., “A Hardware-Software Blueprint for Flexible Deep Learning Specialization” Apr. 23, 2019, arXiv: 1807.04188v3, pp. 1-7. (Year: 2019). [cited by examiner]
Zheng et al., “FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System” Mar. 2020, pp. 859-873. (Year: 2020). [cited by examiner]
Hedge et al., “ExTensor: An Accelerator for Sparse Tensor Algebra” Oct. 2019, pp. 319-333. (Year: 2019). [cited by examiner]
Xi et al., “SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads” Dec. 11, 2019, arXiv: 1912.04481v2, pp. 1-14. (Year: 2019). [cited by examiner]
Zhang et al., “Lookahead Optimizer: k steps forward, 1 step back” Dec. 3, 2019, arXiv: 1907.08610v2, pp. 1-19. (Year: 2019). [cited by examiner]
Aimar et al., “NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps” Mar. 2019, pp. 646-656. (Year: 2019). [cited by examiner]
Zhu et al., “An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGA” Jan. 7, 2020, arXiv: 2001.01955v1, pp. 1-12. (Year: 2020). [cited by examiner]
Kim et al., “torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models” Apr. 21, 2020, arXiv: 2004.09910v1, pp. 1-10. (Year : 2020). [cited by examiner]
Lu et al., “Tetris: Re-architecting Convolutional Neural Network Computation for Machine Learning Accelerators” Jan. 3, 2019, pp. 1-8. (Year: 2019). [cited by examiner]
Kim et al., “FPGA Prototyping of Low-Precision Zero-Skipping Accelerator for Neural Networks” 2018, IEEE, pp. 1-7. (Year: 2018). [cited by examiner]
Qin et al., “SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training” Apr. 16, 2020, pp. 58-70. (Year: 2020). [cited by examiner]
Zhu et al., “Sparse Tensor Core: Algorithm and Hardware Co-Design for Vector-wise Sparse Neural Networks on Modern GPUs” Oct. 2019, pp. 359-371. (Year: 2019). [cited by examiner]
Wu et al., “Compute-Efficient Neural-Network Acceleration” Feb. 2019, pp. 191-200. (Year: 2019). [cited by examiner]
Struhark et al., “CoNNa-Hardware accelerator for compressed convolutional neural networks” Jan. 11, 2020, pp. 1-28. (Year: 2020). [cited by examiner]
Hojabr et al., “SkippNN: An Embedded Sotchastic-Computing Accelerator for Convolutional Neural Networks” Jun. 2019, pp. 1-6. (Year: 2019). [cited by examiner]
Song et al., “AccPar: Tensor Partitioning for Heterogeneous Deep Learning Accelerators” Apr. 16, 2020, pp. 342-355. (Year: 2020). [cited by examiner]
Boutros et al., “You Cannot Improve What You Do Not Measure: FPGA vs. ASIC Efficiency Gaps for Convolutional Neural Network Inference” Dec. 2018, pp. 1-22. (Year: 2018). [cited by examiner]
Chou et al., “Automatic Generation of Efficient Sparse Tensor Format Conversion Routines” Apr. 8, 2020, arXiv: 2001.02609v2, pp. 1-16. (Year: 2020). [cited by examiner]
Zhang et al., “SpArch: Efficient Architecture for Sparse Matrix Multiplication” Feb. 20, 2020, arXiv: 2002.08947v1, pp. 1-15. (Year: 2020). [cited by examiner]
Soltaniyeh et al., “Synergistic CPU-FPGA Acceleration of Sparse Linear Algebra” Apr. 29, 2020, arXiv: 2004.13907v1, pp. 1-12. (Year: 2020). [cited by examiner]
Lascorz, A.D. et al., “Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks”, ASPLOS '19, Providence, RI, USA, Apr. 13-17, 2019, 15 pages, Association for Computing Machiner… [cited by applicant]
Lascorz, A.D. et al., “Bit-Tactical: Exploiting Ineffectual Computations in Convolutional Neural Networks: Which, Why, and How”, arXiv:1803.03688v1, Mar. 9, 2018, pp. 1-14. [cited by applicant]
Mahmoud, M. et al., “Accelerating Image-Sensor-Based Deep Learning Applications”, IEEE Micro, Sep./Oct. 2019, pp. 26-35, IEEE Computer Society. [cited by applicant]
Stuart, Dylan J.M., “An Efficient Hardware Architecture for Exploiting Sparsity in Neural Networks,” University of Toronto, ProQuest Dissertations & Theses, 2019, 63 pages. [cited by applicant]
Srivastava, Nitish, et al., “Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations,” 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2020, pp. 689-702. [cited by applicant]
Korean Office Action dated Apr. 23, 2025, issued in Korean Patent Application No. 10-2021-0032408, 5 pages. [cited by applicant]