IP Library › Granted Patent US 11,593,609
Granted Patent B2
US 11,593,609 · App. 16/794,062 · Granted Feb 28, 2023

Vector quantization decoding hardware unit for real-time dynamic decompression for parameters of neural networks

Inventors: Giuseppe Desoli (San Fermo Della Battaglia, IT); Carmine Cappetta (Battipaglia, IT); Thomas Boesch (Rovio, CH); Surinder Pal Singh (Noida, IN); Saumya Suneja (New Delhi, IN)
Assignees: STMicroelectronics S.r.l.; STMicroelectronics International N.V.
G06N3/04G06F16/2282G06K9/6262G06N3/063G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,593,609
App. No.
16/794,062
Granted
Feb 28, 2023
Kind
B2
Abstract

Embodiments of an electronic device include an integrated circuit, a reconfigurable stream switch formed in the integrated circuit along with a plurality of convolution accelerators and a decompression unit coupled to the reconfigurable stream switch. The decompression unit decompresses encoded kernel data in real time during operation of convolutional neural network.

Claims (36)

1. A convolutional neural network processing system, comprising:

an input layer configured to receive input data;

a decompressor unit configured to receive encoded kernel data encoded with a vector quantization process and to generate decompressed kernel data based on the encoded kernel data, wherein the decompressor unit includes:

a lookup table configured to store codebook data associated with the encoded kernel data, wherein the encoded kernel data includes the codebook data;

an index stream buffer, wherein the encoded kernel data includes index data for retrieving code vectors from the lookup table, wherein the decompressor unit generates the decompressed kernel data by retrieving code vectors from the lookup table based on the index data, wherein the decompressor unit is configured to receive the codebook data, store the codebook data in the lookup table, receive the index data, and retrieve the code vectors from the codebook data with the index data;

a convolutional accelerator configured to receive the decompressed kernel data, to receive feature data based on the input data, and to perform a convolution operation on the feature data and the decompressed kernel data; and

a fully connected layer configured to receive convolved data from the convolutional accelerator and to generate prediction data based on the convolved data.

2. The system of claim 1 , wherein a first vector quantization codebook and a first index data are generated with the vector quantization process.

3. The system of claim 1 , wherein the decompressor unit is configured to store first and second vector quantization codebooks simultaneously.

4. The system of claim 1 , wherein the input data is image data from an image sensor.

5. The system of claim 4 , wherein the prediction data identifies features in the image data.

6. The system of claim 1 , wherein a first convolutional accelerator defines a first convolutional layer of the convolutional neural network, wherein a second convolutional accelerator defines a second convolutional layer of the convolutional neural network.

7. A method, comprising:

receiving, with a decompression unit of a convolutional neural network implemented in a system on chip, encoded kernel data from a source external to the system on chip, wherein the encoded kernel data includes a first vector quantization codebook for a first convolutional accelerator of the convolutional neural network, first index data for the first vector quantization codebook, a second vector quantization codebook for a second convolutional accelerator of the convolutional neural network, and second index data for the second vector quantization codebook;

storing the first vector quantization codebook in a lookup table of the decompression unit;

generating first decompressed kernel data with the decompression unit by retrieving code vectors from the lookup table with the first index data;

receiving feature data at the first convolutional accelerator;

receiving the first decompressed kernel data with the first convolutional accelerator from the decompression unit;

performing convolution operations on the first decompressed kernel data and the feature data with the first convolutional accelerator;

storing the second vector quantization codebook in the lookup table of the decompression unit;

generating second decompressed kernel data by retrieving code vectors from the second vector quantization codebook with the second index data; and

providing the second decompressed kernel data to the second convolutional accelerator.

8. The method of claim 7 , further comprising generating prediction data with the convolutional neural network based on the feature data and the first decompressed kernel data.

9. A method, comprising:

providing, during operation of a convolutional neural network implemented in a system on chip, encoded kernel data from a source external to the system on chip to a decompression unit of the convolutional neural network, wherein the encoded kernel data includes a first vector quantization codebook for a first convolutional accelerator of the convolutional neural network, first index data for the first vector quantization codebook, a second vector quantization codebook for a second convolutional accelerator of the convolutional neural network, and second index data for the second vector quantization codebook;

storing the first vector quantization codebook in a lookup table of the decompression unit;

generating first decompressed kernel data with the decompression unit by retrieving code vectors from the lookup table with the first index data; and

providing the first decompressed kernel data to the first convolutional accelerator;

storing the second vector quantization codebook in the lookup table of the decompression unit;

generating second decompressed kernel data by retrieving code vectors from the second vector quantization codebook with the second index data; and

providing the second decompressed kernel data to the second convolutional accelerator.

10. The method of claim 9 , further comprising:

receiving feature data at the first convolutional accelerator of the convolutional neural network; and

performing convolution operations on the first decompressed kernel data and the feature data with the first convolutional accelerator.

11. The method of claim 10 , further comprising generating prediction data with the convolutional neural network based on the feature data and the first decompressed kernel data.

12. The method of claim 11 , wherein the feature data is generated from image data from an image sensor.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2020
From: DESOLI, GIUSEPPE; CAPPETTA, CARMINE
To: STMICROELECTRONICS S.R.L.
Reel/Frame 052040/0551 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2020
From: BOESCH, THOMAS; SINGH, SURINDER PAL; SUNEJA, SAUMYA
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 052119/0385 →
Continuity (1)
Related Publication 20210256346A1 · Aug 19, 2021
Cited By (1)
US 12,353,971