IP Library Patent Application 17622954
Patent Application
App. No. 17/622,954

CLUSTERING-BASED QUANTIZATION FOR NEURAL NETWORK COMPRESSION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/622,954
Abstract

Systems, methods, and instrumentalities are disclosed for clustering-based quantization for neural network (NN) compression. A distribution of weights in weight tensors in NN layers may be analyzed to identify cluster outliers. Cluster inliers may be coded from cluster outliers, for example, using scalar and/or vector quantization. Weight-rearrangement may rearrange weights for higher dimensional weight tensors into lower dimensional matrices. For example, weight rearrangement may flatten a convolutional kernel into a vector. Correlation between kernels may be preserved, for example, by treating a filter or kernels across a channel as a point. A tensor may be split into multiple subspaces, for example, along an input and/or an output channel. Predictive coding may be performed for a current block of weights or weight matrix based on a reshaped or previously coded block or matrix. Arrangement, inlier, outlier, and/or prediction information may be signaled to a decoder for reconstruction of a compressed NN.

Claims (42)

1 - 14 . (canceled)

15 . A method of encoding comprising:

obtaining a neural network (NN) model, wherein the NN model comprises an NN layer, and wherein the NN layer is associated with a weight matrix;

identifying a dimensionality of the weight matrix;

based on the identified dimensionality of the weight matrix, reshaping the weight matrix to reduce the dimensionality of the weight matrix; and

coding the NN layer based on the reshaped weight matrix.

16 . The method of claim 15 , wherein reshaping the weight matrix comprises flattening or rearranging the dimensionality of the weight matrix.

17 . The method of claim 15 , wherein the dimensionality of the weight matrix comprises a two-dimension, a three-dimension, or a higher dimension, and the weight matrix is reshaped to a one-dimension weight vector.

18 . The method of claim 15 , wherein the method comprises at least one of:

transmitting the identified dimensionality and the reduced dimensionality of the weight matrix in a bitstream; or

performing prediction based on the reshaped weight matrix.

19 . The method of claim 15 , wherein coding the NN layer comprises performing a quantization on the NN layer, and wherein the quantization comprises vector quantization.

20 . An apparatus for encoding comprising:

a processor configured to:

obtain a neural network (NN) model, wherein the NN model comprises an NN layer, and wherein the NN layer is associated with a weight matrix;

identify a dimensionality of the weight matrix;

based on the identified dimensionality of the weight matrix, reshape the weight matrix to reduce the dimensionality of the weight matrix; and

coding the NN layer based on the reshaped weight matrix.

21 . The apparatus of claim 20 , wherein to reshape the weight matrix comprises being configured to flatten or rearrange the dimensionality of the weight matrix.

22 . The apparatus of claim 20 , wherein the dimensionality of the weight matrix comprises a two-dimension, a three-dimension, or a higher dimension, and the weight matrix is reshaped to a one-dimension weight vector.

23 . The apparatus of claim 20 , wherein the processor is configured to:

transmit the identified dimensionality and the reduced dimensionality of the weight matrix in a bitstream.

24 . The apparatus of claim 20 , wherein coding the NN layer comprises performing a quantization on the NN layer, and wherein, the quantization comprises a vector quantization.

25 . The apparatus of claim 20 , the processor is configured to:

perform prediction based on the reshaped weight matrix.

26 . A method of decoding comprising:

obtaining a compressed neural network (NN) model, wherein the compressed NN model comprises a quantized NN layer, and wherein the quantized NN layer is associated with a weight matrix having a first dimensionality;

obtaining a weight matrix shape indication, wherein the weight matrix shape indication indicates a weight matrix shape having a second dimensionality;

based on the weight matrix shape indication, reshaping the weight matrix to the second dimensionality; and

decoding the NN layer based on the reshaped weight matrix.

27 . The method of claim 26 , wherein reshaping the weight matrix comprises restoring the weight matrix having the first dimensionality to the weight matrix having the second dimensionality.

28 . The method of claim 26 , wherein the weight matrix shape having the second dimensionality comprises the weight matrix having an original dimensionality prior to the quantization, and wherein the weight matrix shape indication comprises a number of columns and a number of rows associated with the original dimensionality.

29 . The method of claim 26 , wherein the second dimensionality of the weight matrix comprises a two-dimension, a three-dimension, or a higher dimension, and the weight matrix is reshaped by increasing the first dimensionality of the weight matrix to the second dimensionality of the weight matrix.

30 . An apparatus for decoding comprising:

a processor configured to:

obtain a compressed neural network (NN) model, wherein the compressed NN model comprises a quantized NN layer, and wherein the quantized NN layer is associated with a weight matrix having a first dimensionality;

obtain a weight matrix shape indication, wherein the weight matrix shape indication indicates a weight matrix shape having a second dimensionality;

based on the weight matrix shape indication, reshape the weight matrix to the second dimensionality; and

decode the NN layer based on the reshaped weight matrix.

31 . The apparatus of claim 30 , wherein to reshape the weight matrix comprises being configured to restore the weight matrix having the first dimensionality to the weight matrix having the second dimensionality.

32 . The apparatus of claim 30 , wherein the weight matrix shape having the second dimensionality comprises the weight matrix having an original dimensionality prior to the quantization, and wherein the weight matrix shape indication comprises a number of columns and a number of rows associated with the original dimensionality.

33 . The apparatus of claim 30 , wherein the second dimensionality of the weight matrix comprises a two-dimension, a three-dimension, or a higher dimension, and the weight matrix is reshaped by increasing the first dimensionality of the weight matrix to the second dimensionality of the weight matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →