IP Library › Granted Patent US 12,566,939
Granted Patent B2
US 12,566,939 · App. 17/130,690 · Granted Mar 3, 2026

Quantization for neural network computation

Inventors: Andreas Moshovos (Toronto, CA); Ali Hadi Zadeh (Toronto, CA); Isak Edo Vivancos (Cambridge, GB); Omar Mohamed Awad (Toronto, CA)
Assignee: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,939
App. No.
17/130,690
Granted
Mar 3, 2026
Kind
B2
Abstract

A method for memory storage including storing a neural network by storing values of the neural network each as a reference to a representative value; and, in some embodiments, storing additional values of the neural network. Each of the representative values can be generated by assigning each of the values of the neural network to a cluster; and for each cluster, selecting a centroid from the cluster. The method can include performing one or more multiply-accumulate operations A 1 B 1 + . . . +A n B n on input vectors A and input vectors B, by accumulating input vectors A to an accumulated sum of input vectors A per input vector B having the same representative value and subsequently multiplying each of the accumulated sums of input vectors A by the representative value of the input vector B. A system is also described, as well as a method for configuring memory according to a data structure.

Claims (113)

1 . A system for computation of layers to accelerate computation in a neural network stored in a computer memory, comprising:

a Tensor Processing Unit (TPU) or a Graphics Processing Unit (GPU);

a hardware accelerator comprising one or more tiles, each tile comprising a global buffer, a quantized weight buffer, and one or more processing elements coupled to a Shared Processing Unit (SPU) configured to optimize on-chip memory;

one or more processing elements each configured to accumulate one or more values of the neural network for each of one or more references to an identical representative value to generate an output for each accumulation; and

the SPU configured to, for each of the outputs from each processing element, multiply the output with the identical representative value respective to the output to generate a final output;

wherein the SPU is further configured to accumulate one or more of the final outputs with one or more additional values if present to generate an accumulated output, the one or more additional values being determined using a probability density function defining a probability that each additional value belongs to a distribution;

wherein the one or more values are stored as indexes and provided to the one or more processing elements via the quantized weight buffer;

wherein the SPU is activated whenever the one or more additional values are encountered, and the SPU reads the one or more additional values from the global buffer in preparation for a decompression computation;

wherein each of the additional values has a probability density function (pdf) less than a threshold value, the pdf defined by

pdf

⁢

(

x

|

μ

,

σ

2

)

=

1

2

⁢

π

⁢

σ

2

⁢

e

-

(

x

-

μ

)

2

2

⁢

σ

2

,

wherein x is the additional value, u is a mean of one or more parameters in the component of the neural network having x, and o is a standard deviation of one or more parameters in the component of the neural network having x; and

upon a computation being required of the neural network, the SPU is configured to:

retrieve from a data structure in the memory at least one final output, at least one accumulated output, or both from the memory using a reference to a memory location where the at least one representative value, at least one additional value of the one or more additional values, or both are stored; and

perform a decompression computation of multiplying the additional value with a corresponding representative value of the at least one representative value and adding multiple additional values and corresponding representative value products by summing additional values to be multiplied with corresponding representative values having the same reference stored in memory for that corresponding representative value and subsequently retrieving the reference from memory and multiplying the reference with the corresponding accumulation of additional values.

2 . The system of claim 1 , further comprising one or more tiles, each tile comprising the one or more accumulators and the shared multiplier-accumulator.

3 . The system of claim 2 , wherein the one or more accumulators are 16 accumulators.

4 . The system of claim 2 , wherein each of the one or more accumulators comprise a 32-bit floating point (FP32) adder and an 8-entry register file.

5 . The system of claim 3 , wherein the shared multiplier-accumulator has a 16-entry output activation register file.

6 . Use of the system of claim 1 for natural language processing.

7 . The system of claim 2 , further comprising a global buffer which supplies the one or more tiles with data to be processed.

8 . A computer system for implementing memory compression on a neural network stored in a computer memory, comprising:

the memory;

at least one Tensor Processing Unit (TPU) or Graphics Processing Unit (GPU) and a hardware accelerator comprising one or more tiles, each tile comprising a global buffer, a quantized weight buffer, and one or more processing elements coupled to a Shared Processing Unit (SPU) containing a decompression engine, which is configured to optimize on-chip memory in communication with the computer memory, the memory comprising instructions which, when executed by the at least one TPU or the GPU, carries out the steps of:

storing a neural network in the memory by:

determining if one or more additional values of the neural network exist;

storing, in the memory, one or more values of the neural network each as a reference to a representative value; and

if the one or more additional values exist, storing, in the memory, the one or more additional values of the neural network;

wherein the one or more values are stored as indexes and provided to the one or more processing elements via the quantized weight buffer;

wherein the SPU is activated whenever one or more additional values are encountered, and the SPU reads the one or more additional values from the global buffer in preparation for the activation of the decompression engine;

wherein the one or more additional values are determined using a probability density function defining a probability that each additional value belongs to a distribution;

wherein each of the additional values is in a component of the neural network and has a probability density function pdf less than a threshold value, the pdf defined by

pdf

⁡

(

x

|

μ

,

σ

2

)

=

1

2

⁢

π

⁢

σ

2

⁢

e

-

(

x

-

μ

)

2

2

⁢

σ

2

,

wherein x is the additional value, u is a mean of one or more parameters in the component of the neural network having x, and o is a standard deviation of one or more parameters in the component of the neural network having x;

upon a computation being required of the neural network, retrieving from a data structure at least one representative value, at least one additional value, or both from the memory using a reference to a memory using a reference to a memory location where the one or more representative values, one or more additional values or both are stored; and

the decompression engine performing a decompression computation of multiplying the additional value with a corresponding representative value of the at least one representative value and adding multiple additional values and corresponding representative value products by summing additional values to be multiplied with corresponding representative values having the same reference stored in memory for that corresponding representative value and subsequently retrieving the reference from memory and multiplying the reference with the corresponding accumulation of additional values.

9 . The computer system of claim 8 , each of the one or more representative values being generated by:

assigning each of the one or more values of the neural network to a cluster; and

for each cluster, selecting a selected value from the cluster as the representative value for each of the one or more values of the neural network of the cluster.

10 . The computer system of claim 9 , the steps further comprising: minimizing a sum of each distance between a) each of the one or more values of the neural network and b) the selected value of the cluster that the value is assigned to by iteratively:

performing the assigning, the assigning further comprising reassigning a first value of the one or more values of the neural network from an original cluster of the clusters to a new cluster of the clusters where an original distance of the first value to the selected of the original cluster is greater than a new distance of the first value to the selected value of the new cluster; and

subsequently performing the selecting on at least the original cluster and the new cluster;

wherein the first value is a different value of the one or more values of the neural network upon each iteration.

11 . The computer system of claim 8 , wherein performing the decompression computation includes:

generating an output from performing one or more multiply-accumulate operations A1 B1+ . . . +AnBn on input vectors A and input vectors B, wherein n is the n-th input vector and wherein one or more of input vectors B are each one of the representative values, by accumulating input vectors A to an accumulated sum of input vectors A per input vector B having the same representative value and subsequently multiplying each of the accumulated sums of input vectors A by the representative value of the input vector B.

12 . The computer system of claim 8 , wherein storing the neural network in the memory includes configuring the memory according to a data structure comprising:

a header containing metadata;

a quantized weights section, the quantized weights section storing a reference to a representative value for each parameter of a neural network, each of the references stored in a same order as a corresponding parameter in the neural network; and

an outliers section storing one or more outlier parameters of the parameters of the neural network.

13 . The computer system of claim 12 , wherein each reference for each of the one or more outlier parameters of the parameters of the neural network is a dummy index.

14 . The computer system of claim 13 , wherein the dummy index is used to designate use of the actual value of a value without reference to a reference or representative value.

15 . Use of the system of claim 8 for natural language processing.

Assignments (2)
NUNC PRO TUNC ASSIGNMENT Recorded Nov 10, 2022
From: MOSHOVOS, ANDREAS; ZADEH, ALI HADI; VIVANCOS, ISAK EDO; AWAD, OMAR MOHAMED
To: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
Reel/Frame 061722/0300 →
NUNC PRO TUNC ASSIGNMENT Recorded May 5, 2021
From: MOSHOVOS, ANDREAS; HADI ZADEH, ALI; MOHAMED AWAD, OMAR; EDO VIVANCOS, ISAK
To: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
Reel/Frame 056139/0819 →
Continuity (2)
Provisional Application 63082009 · Sep 23, 2020
Related Publication 20220092382A1 · Mar 24, 2022
References Cited (17)
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180260690A1 · Young et al. · 2018 [cited by applicant]
US 20180293758A1 · Bar-On et al. · 2018 [cited by applicant]
US 20200057934A1 · Yoo · 2020 [cited by examiner]
WO 2018193361A1 · 2018 [cited by applicant]
Garland et al., “Low Complexity Multiply Accumulate Unit for Weight-Sharing Convolutional Neural Networks”, Jan. 22, 2017, IEEE Computer Architecture Letters, vol. 16, No. 2, pp. 132-135 (Year: 2017). [cited by examiner]
Han et al., “Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding”, Feb. 15, 2016, ICLR 2016, pp. 1-14 (Year: 2016). [cited by examiner]
Park et al., “Energy-efficient Neural Network Accelerator Based on Outlier-aware Low-precision Computation”, Jul. 23, 2018, 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA), pp. 688-698 … [cited by examiner]
Teknomo, “K-Means Clustering Tutorial”, Jul. 2007, pp. 1-12 (Year: 2007). [cited by examiner]
Banner et al., “ACIQ: Analytical Clipping for Integer Quantization of neural networks”, Sep. 27, 2018, ICLR 2019 Conference Blind Submission, pp. 1-11 (Year: 2019). [cited by examiner]
Shen et al., “Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT”, Apr. 3, 2020, The Thirty-Fourth AAAI Conference on Artificial Intelligence, pp. 8815-8821 (Year: 2020). [cited by examiner]
Xu et al., “Deep Neural Network Compression with Single and Multiple Level Quantization”, 2018, The Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18), pp. 4335-4342 (Year: 2018). [cited by examiner]
Miljkovic et al., “Review of Novelty Detection Methods”, May 28, 2010, MIPRO 2010, pp. 593-598. (Year: 2010). [cited by examiner]
Hodge et al., “A Survey of Outlier Detection Methodologies”, 2004, Artificial Intelligence Review 22, pp. 85-126. (Year: 2004). [cited by examiner]
S. Han, H. Mao, and W. J. Dally, Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding in 4th International Conference on Learning Representations, ICLR 2016, San Juan, … [cited by applicant]
E. Park, S. Yoo, and P. Vajda, Value-aware quantization for training and inference of neural networks in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 58 / 595. [cited by applicant]
E. Park, D. Kim, and S. Yoo, Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation. IEEE Press, 2018, p. 68x 698. [Online]. Available: https://doi.org/10.1109/ISCA.2018.00063. [cited by applicant]