IP Library › Granted Patent US 12,400,103
Granted Patent B2
US 12,400,103 · App. 17/194,158 · Granted Aug 26, 2025

Variable quantization for neural networks

Inventors: Chirag Sureshbhai Patel (San Diego, CA); Tijmen Pieter Frederik Blankevoort (Amsterdam, NL); Jonathan Dewitt Wolfe (Austin, TX); Erich Plondke (Austin, TX)
Assignee: QUALCOMM Incorporated
G06N3/04G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,103
App. No.
17/194,158
Granted
Aug 26, 2025
Kind
B2
Abstract

A method for an artificial neural network includes receiving an input. A quantization threshold is determined based on the input, or a characteristic or type of the input. Neural network values, such as weights or activations, of one or more layers of the artificial neural network are quantized according to the quantization threshold. The artificial neural network generates an output based on the quantized neural network values.

Claims (56)

1. A processor-implemented method performed by one or more processors, the processor-implemented method comprising:

receiving, by a first layer of an artificial neural network (ANN), a set of input values;

determining a first quantization threshold for at least a second layer of the ANN based at least in part on a first activation range for the first layer of the ANN;

quantizing neural network values of one or more layers of the ANN according to the first quantization threshold; and

generating an output based on the quantized neural network values.

2. The processor-implemented method of claim 1 , in which each of the neural network values comprises a weight or an activation.

3. The processor-implemented method of claim 1 , further comprising:

determining a selection variable based on the set of input values; and

determining the first quantization threshold based on the selection variable.

4. The processor-implemented method of claim 3 , in which the selection variable comprises an input type or an input characteristic associated with the set of input values.

5. The processor-implemented method of claim 4 , in which the input type comprises an end-of-sentence (EOS) token.

6. The processor-implemented method of claim 4 , in which the input characteristic includes one or more of an image scene type, an image brightness, or an image contrast.

7. The processor-implemented method of claim 1 , further comprising determining a second quantization threshold for a third layer of the ANN based on a second activation range for the second layer of the ANN.

8. The processor-implemented method of claim 1 , in which the first quantization threshold is dynamically determined during runtime.

9. The processor-implemented method of claim 1 , in which the output generated by the ANN relates to at least one of image processing, audio processing, or sensor-data processing.

10. An apparatus for an artificial neural network (ANN), comprising:

at least one memory; and

at least one processor coupled to the at least one memory, the at least one processor being configured to:

receive, via a first layer of the ANN, a set of input values;

determine a first quantization threshold for at least a second layer of the ANN based at least in part on a first activation range for the first layer of the ANN;

quantize neural network values of one or more layers of the ANN according to the first quantization threshold; and

generate an output based on the quantized neural network values.

11. The apparatus of claim 10 , in which each of the neural network values comprises a weight or an activation.

12. The apparatus of claim 10 , in which the at least one processor is further configured to:

determine a selection variable based on the set of input values; and

determine the first quantization threshold based on the selection variable.

13. The apparatus of claim 12 , in which the selection variable comprises an input type or an input characteristic associated with the set of input values.

14. The apparatus of claim 13 , in which the input type comprises an end-of-sentence (EOS) token.

15. The apparatus of claim 13 , in which the input characteristic includes one or more of an image scene type, an image brightness, or an image contrast.

16. The apparatus of claim 10 , in which the at least one processor is further configured to determine a second quantization threshold for a third layer of the ANN based on a second activation range for the second layer of the ANN.

17. The apparatus of claim 10 , in which the at least one processor is further configured to dynamically determine the first quantization threshold during runtime.

18. The apparatus of claim 10 , in which the output relates to at least one of image processing, audio processing, or sensor-data processing.

19. An apparatus for an artificial neural network (ANN), comprising:

means for receiving, via a first layer of the ANN, a set of input values;

means for determining a first quantization threshold for at least a second layer of the ANN based at least in part on a first activation range for the first layer of the ANN;

means for quantizing neural network values of one or more layers of the ANN according to the first quantization threshold; and

means for generating an output based on the quantized neural network values.

20. The apparatus of claim 19 , in which each of the neural network values comprises a weight or an activation.

21. The apparatus of claim 19 , further comprising:

means for determining a selection variable based on the set of input values; and

means for determining the first quantization threshold based on the selection variable.

22. The apparatus of claim 21 , in which the selection variable comprises an input type or an input characteristic associated with the set of input values.

23. The apparatus of claim 19 , further comprising means for determining a second quantization threshold for a third layer of the ANN based on a second activation range for the second layer of the ANN.

24. The apparatus of claim 19 , further comprising means for dynamically determining the first quantization threshold during runtime.

25. A non-transitory computer readable medium having encoded thereon program code for an artificial neural network (ANN), the program code being executed by a processor and comprising:

program code to receive, via a first layer of the ANN, a set of input values;

program code to determine a first quantization threshold for at least a second layer of the ANN based at least in part on a first activation range for the first layer of the ANN;

program code to quantize neural network values of one or more layers of the ANN according to the first quantization threshold; and

program code to generate an output based on the quantized neural network values.

26. The non-transitory computer readable medium of claim 25 , in which each of the neural network values comprises a weight or an activation.

27. The non-transitory computer readable medium of claim 25 , in which the at least one processor is further configured:

to determine a selection variable based on the set of input values; and

to determine the first quantization threshold based on one or more of the selection variable.

28. The non-transitory computer readable medium of claim 27 , in which the selection variable comprises an input type or an input characteristic associated with the set of input values.

29. The non-transitory computer readable medium of claim 25 , further comprising program code to determine a second quantization threshold for a third layer of the ANN based on a second activation range for the second layer of the ANN.

30. The non-transitory computer readable medium of claim 25 , further comprising program code to dynamically determine the first quantization threshold during runtime.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE 4TH INVENTOR'S EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 057125 FRAME: 0866. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT . Recorded Sep 1, 2021
From: PATEL, CHIRAG SURESHBHAI; BLANKEVOORT, TIJMEN PIETER FREDERIK; WOLFE, JONATHAN DEWITT; PLONDKE, ERICH
To: QUALCOMM INCORPORATED
Reel/Frame 057771/0970 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2021
From: PATEL, CHIRAG SURESHBHAI; BLANKEVOORT, TIJMEN PIETER FREDERIK; WOLFE, JONATHAN DEWITT; PLONDKE, ERICH
To: QUALCOMM INCORPORATED
Reel/Frame 057125/0866 →
Continuity (1)
Related Publication 20220284260A1 · Sep 8, 2022
References Cited (20)
US 11861492B1 · Hsu · 2024 [cited by examiner]
US 20140380466A1 · Schultz · 2014 [cited by examiner]
US 20160073271A1 · Schultz · 2016 [cited by examiner]
US 20190012559A1 · Desappan · 2019 [cited by examiner]
US 20190050733A1 · Bopardikar · 2019 [cited by examiner]
US 20190340499A1 · Burger · 2019 [cited by examiner]
US 20200125947A1 · Park · 2020 [cited by examiner]
US 20200210830A1 · Shen · 2020 [cited by examiner]
US 20210073635A1 · Sasagawa · 2021 [cited by examiner]
US 20210150334A1 · Wolfe · 2021 [cited by examiner]
US 20220027126A1 · Cao · 2022 [cited by examiner]
US 20220092384A1 · Kim · 2022 [cited by examiner]
Hubara et al., “Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activation,” arXiv: 1609.07061v1, Sep. 22, 2016, 29 pgs. (Year: 2016). [cited by examiner]
Zhu et al., “Trained Ternary Quantization,” arXiv:1612.01064v3, Feb. 23, 2017, 10 pgs. (Year: 2017). [cited by examiner]
Jacob et al., “Quantization and Training of Neural Networks for Efficient Inter-Arithmetic-Only Inference,” 2018 IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 2704-2713 (Year: 2018). [cited by examiner]
Xu et al., “DNQ: Dynamic Netowrk Quantization,” arXiv:1812.02375v1, Dec. 6, 2018, 10 pgs. (Year: 2018). [cited by examiner]
O'Connor et al., “Sigma-Delta Quantized Networks,” published Nov. 10, 2016, 13 pgs. (Year: 2016). [cited by examiner]
Zhou et al., “Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks,” published Jun. 22, 2017, 34 pgs. (Year: 2017). [cited by examiner]
Athar, Ali, “An Overview of Datatype Quantization Techniques for Convolutional Neural Networks,” published Aug. 22, 2018, 4 pgs. (Year: 2018). [cited by examiner]
Ando et al., “Dither NN: An Accurate Neural Network with Dithering for Low Bit-Precision Hardware,” 2018 International Conference on Field-Programmable Technology, pp. 9-16. (Year: 2018). [cited by examiner]