IP Library Patent Application 17811186
Patent Application
App. No. 17/811,186

CLUSTER COMPRESSION FOR COMPRESSING WEIGHTS IN NEURAL NETWORKS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/811,186
Abstract

A method for instantiating a convolutional neural network on a computing system. The convolutional neural network includes a plurality of layers, and instantiating the convolutional neural network includes training the convolutional neural network using a first loss function until a first classification accuracy is reached, clustering a set of F×K kernels of the first layer into a set of C clusters, training the convolutional neural network using a second loss function until a second classification accuracy is reached, creating a dictionary which maps each of a number of centroids to a corresponding centroid identifier, quantizing and compressing F filters of the first layer, storing F quantized and compressed filters of the first layer in a memory of the computing system, storing F biases of the first layer in the memory, and classifying data received by the convolutional neural network.

Claims (20)

1 . A method, comprising:

instantiating a neural network on a computing system, the neural network including a plurality of layers, wherein instantiating the neural network comprises:

training the neural network using a loss function until a classification accuracy is reached, wherein the loss function calculates a classification error of the neural network, wherein training the neural network with the loss function comprises optimizing, for a first one of the layers, a set of F filters and a set of F biases so as to minimize the loss function, wherein each of the F filters is formed from K kernels, wherein K and F are each greater than one, wherein each of the kernels comprises nine parameters and wherein each of the biases is scalar;

clustering the set of F×K kernels of the first layer into a set of C clusters, wherein each of the clusters is characterized by a centroid, thereby the C clusters being characterized by C centroids, wherein each of the centroids comprises nine parameters, and wherein C is less than F×K;

creating a dictionary which maps each of the centroids to a corresponding scalar centroid identifier;

quantizing and compressing the F filters of the first layer by, for each of the F×K kernels, replacing the nine parameters of the kernel with one of the scalar centroid identifiers from the dictionary;

storing the F quantized and compressed filters of the first layer in a memory of the computing system, the F quantized and compressed filters comprising F×K scalar centroid identifiers; and

storing the F biases of the first layer in the memory; and

classifying data received by the neural network, wherein the classification comprises:

retrieving the F quantized and compressed filters of the first layer from the memory, the F quantized and compressed filters comprising the F×K scalar centroid identifiers;

decompressing, using the dictionary, the F quantized and compressed filters of the first layer into F quantized filters by mapping the F×K scalar centroid identifiers into F×K corresponding quantized kernels, the F×K corresponding quantized kernels each comprising nine parameters and forming the F quantized filters;

retrieving the F biases of the first layer from the memory; and

for the first layer, processing the received data or data output from a layer previous to the first layer with the F quantized filters and the F biases, wherein a number of channels of the received data or data output from the layer previous to the first layer is equal to K.

2 . The method of claim 1 , further comprising:

during the instantiation of the neural network, further storing the dictionary in the memory; and

during the classification of the received data, further retrieving the dictionary from the memory.

3 . The method of claim 1 , wherein C is equal to 2 n , where n is a natural number.

4 . The method of claim 3 , wherein each of the scalar centroid identifiers is expressed using n-bits.

5 . The method of claim 1 , wherein the F×K kernels are clustered into the C clusters using the k-means algorithm.

6 . The method of claim 1 , wherein each of the F×K kernels is represented by a 3×3 matrix of parameters.

Assignments (2)
CHANGE OF NAME Recorded Sep 11, 2025
From: RECOGNI INC.
To: TENSORDYNE, INC.
Reel/Frame 072859/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2022
From: BACKHUS, GILLES J. C. A.; FEINBERG, EUGENE M.
To: RECOGNI INC.
Reel/Frame 060624/0335 →