IP Library Granted Patent US 11,496,775
Granted Patent B2
US 11,496,775 · App. 17/088,075 · Granted Nov 8, 2022

Neural network model compression with selective structured weight unification

Inventors: Wei Jiang (Palo Alto, CA); Wei Wang (Palo Alto, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
H04N19/96G06N3/04G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,496,775
App. No.
17/088,075
Granted
Nov 8, 2022
Kind
B2
Abstract

A method, computer program, and computer system is provided for compressing a neural network model. One or more coding tree units are identified corresponding to a multi-dimensional tensor associated with a neural network. A set of weight coefficients associated with the coding tree units is unified. A model of the neural network is compressed based on the unified set of weight coefficients.

Claims (37)

1. A method for compressing a neural network model, executable by a processor, comprising:

identifying one or more coding tree units corresponding to a multi-dimensional weight tensor associated with a neural network;

unifying a set of weight coefficients, associated with the coding tree units, by reducing a number of dimensions of the weight tensor from five dimensions to three dimensions; and

compressing a model of the neural network based on the unified set of weight coefficients.

2. The method of claim 1 , wherein unifying the set of weight coefficients comprises:

quantizing the weight coefficients; and

selecting the subset of weight coefficients based on minimizing a unification loss value associated with the weight coefficients.

3. The method of claim 2 , further comprising training the deep neural network based on back-propagating the minimized unification loss value.

4. The method of claim 2 , wherein one or more weight coefficients from among the subset of weight coefficients are fixed to one or more values based on back-propagating the minimized unification loss value.

5. The method of claim 4 , further comprising updating one or more non-fixed weight coefficients from among the subset of weight coefficients based on determining a gradient and a unifying mask associated with the set of weight coefficients.

6. The method of claim 1 , further comprising compressing the set of weight coefficients by quantizing and entropy-coding the subset of weight coefficients.

7. The method of claim 1 , wherein the unified set of weight coefficients comprises one or more weight coefficients having a same absolute value.

8. A computer system for compressing a neural network model, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

identifying code configured to cause the one or more computer processors to identify one or more coding tree units corresponding to a multi-dimensional weight tensor associated with a neural network;

unifying code configured to cause the one or more computer processors to unify a set of weight coefficients, associated with the coding tree units, by reducing a number of dimensions of the multi-dimensional weight tensor from five dimensions to three dimensions; and

compressing code configured to cause the one or more computer processors to compress a model of the neural network based on the unified set of weight coefficients.

9. The computer system of claim 8 , wherein the unifying code comprises:

quantizing code configured to cause the one or more computer processors to quantize the weight coefficients; and

selecting code configured to cause the one or more computer processors to select the subset of weight coefficients based on minimizing a unification loss value associated with the weight coefficients.

10. The computer system of claim 9 , further comprising training code configured to cause the one or more computer processors to train the deep neural network based on back-propagating the minimized unification loss value.

11. The computer system of claim 9 , wherein one or more weight coefficients from among the subset of weight coefficients are fixed to one or more values based on back-propagating the minimized unification loss value.

12. The computer system of claim 11 , further comprising updating code configured to cause the one or more computer processors to update one or more non-fixed weight coefficients from among the subset of weight coefficients based on determining a gradient and a unifying mask associated with the set of weight coefficients.

13. The computer system of claim 8 , further comprising compressing code configured to cause the one or more computer processors to compress the set of weight coefficients by quantizing and entropy-coding the subset of weight coefficients.

14. The computer system of claim 8 , wherein the unified set of weight coefficients comprises one or more weight coefficients having a same absolute value.

15. A non-transitory computer readable medium having stored thereon a computer program for compressing a neural network model, the computer program configured to cause one or more computer processors to:

identify one or more coding tree units corresponding to a multi-dimensional weight tensor associated with a neural network;

unify a set of weight coefficients, associated with the coding tree units, by reducing a number of dimensions of the multi-dimensional weight tensor from five dimensions to three dimensions; and

compress a model of the neural network based on the unified set of weight coefficients.

16. The computer readable medium of claim 15 , wherein the unifying code comprises:

quantizing code configured to cause the one or more computer processors to quantize the weight coefficients; and

selecting code configured to cause the one or more computer processors to select the subset of weight coefficients based on minimizing a unification loss value associated with the weight coefficients.

17. The computer readable medium of claim 16 , further comprising training code configured to cause the one or more computer processors to train the deep neural network based on back-propagating the minimized unification loss value.

18. The computer readable medium of claim 16 , wherein one or more weight coefficients from among the subset of weight coefficients are fixed to one or more values based on back-propagating the minimized unification loss value.

19. The computer readable medium of claim 18 , further comprising updating code configured to cause the one or more computer processors to update one or more non-fixed weight coefficients from among the subset of weight coefficients based on determining a gradient and a unifying mask associated with the set of weight coefficients.

20. The computer readable medium of claim 15 , further comprising compressing code configured to cause the one or more computer processors to compress the set of weight coefficients by quantizing and entropy-coding the subset of weight coefficients.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2020
From: JIANG, WEI; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 054258/0336 →
Continuity (3)
Provisional Application 62984107 · Mar 2, 2020
Provisional Application 62979038 · Feb 20, 2020
Related Publication 20210266607A1 · Aug 26, 2021
Cited By (1)
US 12,197,532