IP Library › Granted Patent US 11,663,746
Granted Patent B2
US 11,663,746 · App. 17/095,544 · Granted May 30, 2023

Systolic arithmetic on sparse data

Inventors: Abhishek R. Appu (El Dorado Hills, CA); Prasoonkumar Surti (Folsom, CA); Jill Boyce (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Michael Apodaca (Folsom, CA); Adam T. Lake (Portland, OR); James Holland (Folsom, CA); Vasanth Ranganathan (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Lidong Xu (Beijing, CN); Nikos Kaburlasos (Folsom, CA)
Assignee: Intel Corporation
G06T9/002G06N3/045G06T9/007G06T9/008G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,663,746
App. No.
17/095,544
Granted
May 30, 2023
Kind
B2
Abstract

Embodiments described herein provided for an instruction and associated logic to enable a processing resource including a tensor accelerator to perform optimized computation of sparse submatrix operations. One embodiment provides hardware logic to apply a numerical transform to matrix data to increase the sparsity of the data. Increasing the sparsity may result in a higher compression ratio when the matrix data is compressed.

Claims (40)

1. A computing device comprising:

one or more processors including a graphics processor, the graphics processor including:

a processing resource including a tensor accelerator, wherein the tensor accelerator is to perform numerical operations to train a neural network model and generate a first matrix of weights associated with the neural network model; and

a weight transformer to apply a numerical transform to the first matrix of weights to generate a set of transformed weights and a transform type, wherein the transform type identifies the numerical transform applied to the first matrix of weights, wherein the first matrix of weights is a sparse matrix and the transformed weights compress to a higher compression ratio than the first matrix of weights.

2. The computing device as in claim 1 , wherein the weight transformer is a fixed-function hardware unit of the graphics processor.

3. The computing device as in claim 1 , wherein the weight transformer is implemented via programmable hardware of the graphics processor.

4. The computing device as in claim 1 , wherein the weight transformer includes a transform selector and a transform tester, the transform tester to:

apply a first numerical transform to at least a portion of the first matrix of weights to generate first test transform data;

apply a second numerical transform to at least a portion of the first matrix of weights to generate second test transform data;

determine compressibility metrics based on analysis of the first test transform data and the second test transform data; and

send a recommended transform to the transform selector.

5. The computing device as in claim 4 , the transform selector further to select a numerical transform to apply to the first matrix of weights based on the recommended transform.

6. The computing device as in claim 5 , wherein the numerical transform is selected from a set of numerical transforms including a discrete cosine transform, a discrete sine transform, a bit-flip transform, and a bit-rotate transform.

7. The computing device as in claim 1 , the graphics processor further comprising:

an inverse weight transformer to apply a numerical inverse transform to the set of transformed weights to generate a second matrix of weights, wherein the numerical inverse transform to perform is identified via the transform type associated with the set of transformed weights.

8. The computing device as in claim 1 , wherein the tensor accelerator includes a systolic array of processing elements.

9. A method comprising:

on a graphics processor including a tensor accelerator:

loading matrix data and metadata associated with the matrix data into the tensor accelerator, wherein the metadata indicates a numerical transform applied to the matrix data;

performing, by the tensor accelerator, a numerical inverse transform on the matrix data, the numerical inverse transform indicated by the metadata associated with the matrix data;

performing one or more compute operations within the tensor accelerator after performing the numerical inverse transform; and

writing output of the one or more compute operations to a memory of the graphics processor.

10. The method as in claim 9 , additionally comprising decompressing or decoding the matrix data within the tensor accelerator before performing the numerical inverse transform on the matrix data.

11. The method as in claim 10 , additionally comprising compressing or encoding the output within the tensor accelerator after performing the numerical transform on the matrix data.

12. The method as in claim 9 , wherein writing output of the one or more compute operations to the memory of the graphics processor includes writing the output to a cache memory of the graphics processor.

13. The method as in claim 9 , wherein writing output of the one or more compute operations to the memory of the graphics processor includes writing the output to a main memory of the graphics processor.

14. The method as in claim 9 , further comprising applying a numerical transform to the output of the one or more compute operations before writing the output to the memory of the graphics processor.

15. The method as in claim 14 , wherein applying the numerical transform to the output of the one or more compute operations comprises applying a specified numerical transform to the output.

16. The method as in claim 14 , wherein applying the numerical transform to the output of the one or more compute operations comprises generating, by the tensor accelerator, transform test metrics for multiple numerical transforms, selecting a numerical transform based on the transform test metrics, and applying the selected numerical transform to the output.

17. The method as in claim 16 , wherein the selected numerical transform is selected from a set of numerical transforms including a discrete cosine transform, a discrete sine transform, a bit-flip transform, and a bit-rotate transform.

18. A graphics processor comprising:

a processing resource including a tensor accelerator, wherein the tensor accelerator is to perform numerical operations to train a neural network model and generate a first matrix of weights associated with the neural network model;

a weight transformer to apply a numerical transform to the first matrix of weights to generate a set of transformed weights and a transform type, wherein the transform type identifies the numerical transform applied to the first matrix of weights, wherein the first matrix of weights is a sparse matrix and the transformed weights compress to a higher compression ratio than the first matrix of weights; and

an inverse weight transformer to apply a numerical inverse transform to the transformed weights to generate a second matrix of weights, wherein the numerical inverse transform to perform is identified via the transform type associated with the set of transformed weights.

19. The graphics processor as in claim 18 , wherein the weight transformer includes a transform selector and a transform tester, the transform tester to:

apply a first numerical transform to at least a portion of the first matrix of weights to generate first test transform data;

apply a second numerical transform to at least a portion of the first matrix of weights to generate second test transform data;

determine compressibility metrics based on analysis of the first test transform data and the second test transform data; and

send a recommended transform to the transform selector, the transform selector to select a numerical transform to apply to the first matrix of weights based on the recommended transform.

20. The graphics processor as in claim 19 , wherein the numerical transform is selected from a set of numerical transforms including a discrete cosine transform, a discrete sine transform, a bit-flip transform, and a bit-rotate transform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2021
From: APPU, ABHISHEK R.; SURTI, PRASOONKUMAR; BOYCE, JILL; MAIYURAN, SUBRAMANIAM; APODACA, MICHAEL; LAKE, ADAM T.; HOLLAND, JAMES; RANGANATHAN, VASANTH; KOKER, ALTUG; XU, LIDONG; KABURLASOS, NIKOS
To: INTEL CORPORATION
Reel/Frame 055492/0105 →
Continuity (2)
Provisional Application 62935670 · Nov 15, 2019
Related Publication 20210150770A1 · May 20, 2021
Cited By (13)
US 12,217,053 US 12,242,414 US 12,271,767 US 12,493,922 US 12,554,674 US 12,561,276 US 12,561,277 US 12,572,997 US 12,670,121 US 12,688,146 US 12,730,759 US 12,737,317 US 12,737,318