IP Library › Patent Application 19379619
Patent Application
App. No. 19/379,619

DYNAMIC QUANTIZATION FOR ENERGY EFFICIENT DEEP LEARNING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/379,619
Abstract

A method performed by a deep neural network (DNN) includes receiving, at a layer of the DNN during an inference stage, a layer input comprising content associated with a DNN input received at the DNN. The method also includes quantizing one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input. The method further includes performing a task corresponding to the DNN input, the task performed with the one or more one quantized parameters.

Claims (57)

1 . A method performed by a deep neural network (DNN), comprising:

receiving, at a layer of the DNN during an inference stage, a layer input comprising content associated with a DNN input received at the DNN;

quantizing one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input; and

performing a task corresponding to the DNN input, the task performed with the one or more one quantized parameters.

2 . The method of claim 1 , in which the plurality of parameters comprise a set of weights and a set of activations.

3 . The method of claim 2 , in which quantizing the one or more parameters comprises quantizing one or both of the respective set of weights or the respective set of activations of one or more output channels associated with the layer.

4 . The method of claim 1 , in which a first amount of quantization of a first parameter of the plurality of parameters is different than a second amount of quantization of a second parameter of the plurality of parameters.

5 . The method of claim 1 , in which quantizing the one or more parameters comprises generating an adjusted bit-width by adjusting a size of an original bit-width associated with the one or more parameters.

6 . The method of claim 5 , in which generating the adjusted bit-width comprising discarding bits of the original bit-width from least significant bits to most significant bits until the size of the original bit-width equals a size of the adjusted bit-width determined based on the content of the layer input.

7 . The method of claim 6 , further comprising training the DNN to determine the size for adjusting the original bit-width based on a total loss, the total loss being a function of a performance loss and a regularization loss.

8 . The method of claim 7 , in which the performance loss determines a cross-entropy loss or a mean-squared error.

9 . The method of claim 8 , in which the regularization loss is a bitwise L0 regularization loss that penalizes one of:

the adjusted bit-width and a complexity metric associated with a bit-level operation; or

a number of bits allocated to the adjusted bit-width.

10 . The method of claim 9 , in which the complexity metric comprises one or more of a number of binary operations of the DNN, a memory footprint of the DNN, or a computer power of the DNN.

11 . The method of claim 9 , further comprising:

reformulating the bitwise L0 regularization loss as a Bernoulli distribution;

relaxing the reformulated bitwise L0 regularization loss based on a sigmoid function; and

minimizing the performance loss and the regularization loss based on the number of bits selected for the adjusted bit-width.

12 . The method of claim 1 , in which the layer is one layer of a plurality of layers of the DNN, and

the method further comprises quantizing one or more parameters of a respective plurality of parameters of each layer of the plurality of layers based on content of a respective layer input.

13 . The method of claim 12 , in which a quantization amount is different for each layer of the plurality of layers.

14 . An apparatus for implementing a deep neural network (DNN), comprising:

a processor;

a memory coupled with the processor; and

instructions stored in the memory and operable, when executed by the processor, to cause the apparatus to:

receive, at a layer of the DNN during an inference stage, a layer input comprising content associated with a DNN input received at the DNN;

quantize one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input; and

perform a task corresponding to the DNN input, the task performed with the one or more one quantized parameters.

15 . The apparatus of claim 14 , in which the plurality of parameters comprise a set of weights and a set of activations.

16 . The apparatus of claim 15 , in which the instructions further cause the apparatus to quantize the one or more parameters by quantizing one or both of the respective set of weights or the respective set of activations of one or more output channels associated with the layer.

17 . The apparatus of claim 14 , in which a first amount of quantization of a first parameter of the plurality of parameters is different than a second amount of quantization of a second parameter of the plurality of parameters.

18 . The apparatus of claim 14 , in which the instructions further cause the apparatus to quantize the one or more parameters by generating an adjusted bit-width by adjusting a size of an original bit-width associated with the one or more parameters.

19 . The apparatus of claim 18 , in which the instructions further cause the apparatus to generate the adjusted bit-width by discarding bits of the original bit-width from least significant bits to most significant bits until the size of the original bit-width equals a size of the adjusted bit-width determined based on the content of the layer input.

20 . The apparatus of claim 19 , in which the instructions further cause the apparatus to determine, during a training stage, the size for adjusting the original bit-width based on a total loss, the total loss being a function of a performance loss and a regularization loss.

21 . The apparatus of claim 20 , in which the regularization loss is a bitwise L0 regularization loss that penalizes one of:

the adjusted bit-width and a complexity metric associated with a bit-level operation; or

a number of bits allocated to the adjusted bit-width.

22 . The apparatus of claim 21 , in which the complexity metric comprises one or more of a number of binary operations of the DNN, a memory footprint of the DNN, or a computer power of the DNN.

23 . The DNN of claim 21 , in which the instructions further cause the DNN to:

reformulate the bitwise L0 regularization loss as a Bernoulli distribution;

relax the reformulated bitwise L0 regularization loss based on a sigmoid function; and

minimize the performance loss and the regularization loss based on the number of bits selected for the adjusted bit-width.

24 . The DNN of claim 14 , in which the layer is one layer of a plurality of layers of the DNN, and

the instructions further cause the DNN to quantize one or more parameters of a respective plurality of parameters of each layer of the plurality of layers based on content of a respective layer input.

25 . The DNN of claim 24 , in which a quantization amount is different for each layer of the plurality of layers.

26 . A non-transitory computer-readable medium having program code recorded thereon for a deep neural network (DNN), the program code executed by a processor and comprising:

program code to receive, at a layer of the DNN during an inference stage, a layer input comprising content associated with a DNN input received at the DNN;

program code to quantize one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input; and

program code to perform a task corresponding to the DNN input, the task performed with the one or more one quantized parameters.

27 . The non-transitory computer-readable medium of claim 26 , in which the plurality of parameters comprise a set of weights and a set of activations.

28 . The non-transitory computer-readable medium of claim 27 , in which the program code to quantize the one or more parameters comprises program code to quantize one or both of the respective set of weights or the respective set of activations of one or more output channels associated with the layer.

29 . An apparatus implementing a deep neural network (DNN), comprising:

means for receiving, at a layer of the DNN during an inference stage, a layer input comprising content associated with a DNN input received at the DNN;

means for quantizing one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input; and

means for performing a task corresponding to the DNN input, the task performed with the one or more one quantized parameters.

30 . The apparatus of claim 29 , in which the means for quantizing the one or more parameters comprises means for quantizing one or both of the respective set of weights or the respective set of activations of one or more output channels associated with the layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2025
From: ARDYWIBOWO, RANDY; DAYANA, VENKATA RAVI KIRAN; HWANG, HAU
To: QUALCOMM INCORPORATED
Reel/Frame 073102/0963 →