IP Library Granted Patent US 11,811,429
Granted Patent B2
US 11,811,429 · App. 17/034,739 · Granted Nov 7, 2023

Variational dropout with smoothness regularization for neural network model compression

Inventors: Wei Jiang (Palo Alto, CA); Wei Wang (Palo Alto, CA); Shan Liu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
H03M7/70G06F18/211G06F18/2155G06N3/084G06N5/046G06V10/764G06V10/82H03M7/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,811,429
App. No.
17/034,739
Granted
Nov 7, 2023
Kind
B2
Abstract

A method, computer program, and computer system is provided for compressing a deep neural network model. Weight coefficients associated with a deep neural network are quantize and entropy-coded. The quantized and entropy-coded weight coefficients are locally smoothed. The smoothed weight coefficients are compressed based on applying a variational dropout to the weight coefficients.

Claims (31)

1. A method of compressing a deep neural network model, executable by a processor, the method comprising:

quantizing and entropy-coding weight coefficients associated with the deep neural network;

locally smoothing the quantized and entropy-coded weight coefficients; and

compressing the smoothed weight coefficients based on applying a variational dropout to the weight coefficients.

2. The method of claim 1 , wherein the weight coefficients correspond to dimensions associated with a multi-dimensional tensor.

3. The method of claim 1 , wherein the weight coefficients are locally smoothed based on determining a gradient associated with the weight coefficients.

4. The method of claim 1 , wherein the weight coefficients are locally smoothed based on minimizing a loss value associated with a smoothness metric of the weight coefficients.

5. The method of claim 4 , wherein the loss value is minimized based on selecting a subset of weight coefficients from among the weight coefficients.

6. The method of claim 4 , wherein the smoothness metric corresponds to an absolute value of gradients of the weight coefficients at one or more axes measured at a point based on balancing contributions of the gradients at the one or more axes.

7. The method of claim 4 , further comprising training the deep neural network based on the minimized loss value.

8. A computer system for compressing a deep neural network model, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

quantizing and entropy-coding code configured to cause the one or more computer processors to quantize and entropy-code weight coefficients associated with the deep neural network;

smoothing code configured to cause the one or more computer processors to locally smooth the quantized and entropy-coded weight coefficients; and

compressing code configured to cause the one or more computer processors to compress the smoothed weight coefficients based on applying a variational dropout to the weight coefficients.

9. The computer system of claim 8 , wherein the weight coefficients correspond to dimensions associated with a multi-dimensional tensor.

10. The computer system of claim 8 , wherein the weight coefficients are locally smoothed based on determining a gradient associated with the weight coefficients.

11. The computer system of claim 8 , wherein the weight coefficients are locally smoothed based on minimizing a loss value associated with a smoothness metric of the weight coefficients.

12. The computer system of claim 11 , wherein the loss value is minimized based on selecting a subset of weight coefficients from among the weight coefficients.

13. The computer system of claim 11 , wherein the smoothness metric corresponds to an absolute value of gradients of the weight coefficients at one or more axes measured at a point based on balancing contributions of the gradients at the one or more axes.

14. The computer system of claim 11 , further comprising training code configured to cause the one or more computer processors to train the deep neural network based on the minimized loss value.

15. A non-transitory computer readable medium having stored thereon a computer program for compressing a deep neural network model, the computer program configured to cause one or more computer processors to:

quantize and entropy-code weight coefficients associated with the deep neural network;

locally smooth the quantized and entropy-coded weight coefficients; and

compress the smoothed weight coefficients based on applying a variational dropout to the weight coefficients.

16. The computer readable medium of claim 15 , wherein the weight coefficients correspond to dimensions associated with a multi-dimensional tensor.

17. The computer readable medium of claim 15 , wherein the weight coefficients are locally smoothed based on determining a gradient associated with the weight coefficients.

18. The computer readable medium of claim 15 , wherein the weight coefficients are locally smoothed based on minimizing a loss value associated with a smoothness metric of the weight coefficients.

19. The computer readable medium of claim 18 , wherein the loss value is minimized based on selecting a subset of weight coefficients from among the weight coefficients.

20. The computer readable medium of claim 18 , wherein the smoothness metric corresponds to an absolute value of gradients of the weight coefficients at one or more axes measured at a point based on balancing contributions of the gradients at the one or more axes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2020
From: JIANG, WEI; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 053905/0291 →
Continuity (3)
Provisional Application 62939060 · Nov 22, 2019
Provisional Application 62915337 · Oct 15, 2019
Related Publication 20210111736A1 · Apr 15, 2021