IP Library Granted Patent US 11,615,297
Granted Patent B2
US 11,615,297 · App. 16/879,780 · Granted Mar 28, 2023

Structured weight based sparsity in an artificial neural network compiler

Inventors: Avi Baum (Givat Shmuel, IL); Or Danon (Kiryat Ono, IL); Daniel Chibotero (Ramat Gan, IL)
G06N3/063G06F8/41G06N3/04G06F17/16G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,615,297
App. No.
16/879,780
Granted
Mar 28, 2023
Kind
B2
Abstract

A novel and useful system and method of improved power performance and lowered memory requirements for an artificial neural network based on packing memory utilizing several structured sparsity mechanisms. The invention applies to neural network (NN) processing engines adapted to implement mechanisms to search for structured sparsity in weights and activations, resulting in a considerably reduced memory usage. The sparsity guided training mechanism synthesizes and generates structured sparsity weights. A compiler mechanism within a software development kit (SDK), manipulates structured weight domain sparsity to generate a sparse set of static weights for the NN. The structured sparsity static weights are loaded into the NN after compilation and utilized by both the structured weight domain sparsity mechanism and the structured activation domain sparsity mechanism. The application of structured sparsity lowers the span of search options and creates a relatively loose coupling between the data and control planes.

Claims (28)

1. A method of weight domain sparsity for use in an artificial neural network (ANN) compiler, the method comprising:

searching a plurality of weight tensors within a weight space for one or more sparsity patterns within a predefined codebook of valid sparsity patterns;

packing a weight memory with packed weight tensors to reduce memory usage in accordance with one or more found sparsity patterns; and

generating one or more weight sparsity instructions corresponding to a skip sequence based on said found sparsity patterns for use in subsequent retrieval of weights and input data from said weight memory and input memory, respectively, during an inference mode of operation.

2. The method according to claim 1 , wherein said packed weight tensors each are configured to represent one or more predetermined weight sparsity patterns that effectively reduce memory usage and power consumption.

3. The method according to claim 1 , wherein said weight sparsity patterns are known a priori and are selected from a group consisting of a vertical column, a horizontal row, a diagonal, an ‘X’ shape, a ‘+’ shape, a triangular block, a single weight, and any combination or superposition thereof.

4. The method according to claim 1 , wherein said one or more weight sparsity instructions are adapted to be subsequently stored in a neural network processor as one or more microcode instructions.

5. The method according to claim 4 , wherein each microcode instruction comprises a plurality of opcodes operative to generate a synchronized sequence of operations on said neural network processor including weights and input data.

6. The method according to claim 1 , wherein said one or more found sparsity patterns are known a priori and stored in one or more configuration registers.

7. A method of weight domain sparsity for use in an artificial neural network (ANN) compiler, the method comprising:

searching a plurality of weights stored in a weight memory for one or more sparsity patterns known a priori;

generating scoring for said plurality of weights for a minimum possible memory size;

reordering said plurality of weights in said weight memory as one or more weight tensors in accordance with corresponding scores;

repeating said steps of searching, scoring, and reordering to reduce weight memory usage in accordance with one or more found patterns;

maximally packing said weight tensors in said weight memory in accordance with one or more found sparsity patterns; and

generating one or more weight sparsity instructions based on said one or more found sparsity patterns for use in subsequent retrieval of weights and input data from said weight memory and input memory, respectively.

8. The method according to claim 7 , wherein said packed weight tensors each are configured to represent one or more predetermined weight sparsity patterns that effectively reduce memory usage and power consumption.

9. The method according to claim 7 , wherein said weight sparsity patterns are known a priori and are selected from a group consisting a vertical column, a horizontal row, a diagonal, an ‘X’ shape, a ‘+’ shape, a triangular block, a single weight, and any combination or superposition thereof.

10. The method of claim 9 wherein said weight sparsity patterns comprise an argument operative to shift weight data vertically, horizontally, and/or to shorten or lengthen one of said weight sparsity patterns.

11. The method according to claim 7 , wherein reordering said plurality of weights comprises rearranging said tensor dimensions.

12. The method according to claim 7 , wherein reordering of said plurality of weights comprises applying a transpose operation.

13. The method according to claim 7 , wherein reordering of said plurality of weights comprises swapping a plurality of axes.

14. The method according to claim 7 , wherein reordering of said plurality of weight comprises unrolling one of said weights into a vector using a row-major order.

15. The method according to claim 7 , wherein reordering of said plurality of weight comprises flipping one or more input data memory locations.

16. The method according to claim 7 , wherein said packed weight memory comprises one or more weight memory tensors representing a predetermined plurality of weight sparsity patterns that effectively reduce memory usage and power consumption.

17. The method according to claim 7 , wherein said one or more weight sparsity instructions are implemented in hardwired circuitry.

18. The method according to claim 7 , wherein said one or more weight sparsity instructions are adapted to be subsequently stored in a neural network processor as one or more microcode instructions.

19. The method according to claim 18 , wherein said microcode instructions comprise a plurality of opcodes operative to generate and effect, on said neural network processor, said subsequent retrieval of weights and input data synchronized appropriately taking into account said one or more memory address skipping operations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: BAUM, AVI; DANON, OR; CHIBOTERO, DANIEL
To: HAILO TECHNOLOGIES LTD.
Reel/Frame 052719/0591 →
Continuity (4)
Continuation In Part 15943992 · Apr 3, 2018
Provisional Application 62531372 · Jul 12, 2017
Provisional Application 62481492 · Apr 4, 2017
Related Publication 20200285950A1 · Sep 10, 2020
Cited By (4)
US 12,346,803 US 12,353,987 US 12,443,571 US 12,602,576