IP Library › Granted Patent US 12,406,169
Granted Patent B2
US 12,406,169 · App. 17/814,957 · Granted Sep 2, 2025

Optimally clipped tensors and vectors

Inventors: Charbel Sakr (Mountain View, CA); Steve Haihang Dai (Union City, CA); Brucek Kurdo Khailany (Austin, TX); William James Dally (Incline Village, NV); Rangharajan Venkatesan (San Jose, CA); Brian Matthew Zimmer (Berkeley, CA)
Assignee: NVIDIA Corporation
G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,169
App. No.
17/814,957
Granted
Sep 2, 2025
Kind
B2
Abstract

Quantizing tensors and vectors processed within a neural network reduces power consumption and may accelerate processing. Quantization reduces the number of bits used to represent a value, where decreasing the number of bits used can decrease the accuracy of computations that use the value. Ideally, quantization is performed without reducing accuracy. Quantization-aware training (QAT) is performed by dynamically quantizing tensors (weights and activations) using optimal clipping scalars. “Optimal” in that the mean squared error (MSE) of the quantized operation is minimized and the clipping scalars define the degree or amount of quantization for various tensors of the operation. Conventional techniques that quantize tensors during training suffer from high amounts of noise (error). Other techniques compute the clipping scalars offline through a brute force search to provide high accuracy. In contrast, the optimal clipping scalars can be computed online and provide the same accuracy as the clipping scalars computed offline.

Claims (46)

1. A computer-implemented method for quantizing tensors of a neural network model comprising multiple processing layers, comprising:

computing first clipping scalars for quantizing first tensors of a first processing layer that is coupled between two processing layers of the multiple processing layers;

processing an input by the neural network model, according to quantized tensors that include the quantized first tensors, by each processing layer of the multiple processing layers in sequence to produce intermediate tensors and an output of the neural network model;

adjusting the first tensors based on a loss gradient; and

updating the first clipping scalars based on a mean squared error to reduce differences between the adjusted first tensors and quantized adjusted first tensors.

2. The computer-implemented method of claim 1 , wherein the first tensors are at least one of weights or activations.

3. The computer-implemented method of claim 1 , wherein the first clipping scalars are recursively computed within the first processing layer according to a Newton-Raphson algorithm.

4. The computer-implemented method of claim 1 , wherein the first clipping scalars are computed by minimizing a quantization mean squared error.

5. The computer-implemented method of claim 1 , repeating the processing, adjusting, and updating for additional inputs.

6. The computer-implemented method of claim 1 , wherein the first clipping scalars and additional clipping scalars for the intermediate tensors are dynamically or statically updated during training.

7. The computer-implemented method of claim 1 , wherein the first clipping scalars and additional clipping scalars for the intermediate tensors are dynamically or statically updated during inference.

8. The computer-implemented method of claim 1 , wherein the first tensors are quantized from a floating-point format to an integer format.

9. The computer-implemented method of claim 1 , wherein the first tensors are quantized from a floating-point format to a lower precision floating-point format.

10. The computer-implemented method of claim 1 , wherein updating the first clipping scalars minimizes a mean squared error of the differences.

11. The computer-implemented method of claim 1 , wherein adjusting the first tensors comprises applying the loss gradient and the first clipping scalars to update the first tensors.

12. The computer-implemented method of claim 1 , further comprising estimating the loss gradient using a magnitude attenuation operation.

13. The computer-implemented method of claim 1 , wherein the first clipping scalars include a separate scalar for each channel of the first tensors.

14. The computer-implemented method of claim 1 , wherein the first tensors are decomposed into sub-tensors and the first clipping scalars include a separate scalar for each sub-tensor of the first tensors.

15. The computer-implemented method of claim 14 , wherein the first tensors are decomposed into vectors and the first clipping scalars include a separate scalar for each vector of the first tensors.

16. The computer-implemented method of claim 1 , further comprising:

adjusting the intermediate tensors based on the loss gradient; and

updating second clipping scalars of a second processing layer of the multiple processing layers based on the mean squared error to reduce differences between the adjusted intermediate tensors and quantized adjusted second tensors.

17. The computer-implemented method of claim 1 , wherein at least one of the steps of computing, processing, adjusting, and updating are performed on a server or in a data center and the output is streamed to a user device.

18. The computer-implemented method of claim 1 , wherein at least one of the steps of computing, processing, adjusting, and updating are performed within a cloud computing environment.

19. The computer-implemented method of claim 1 , wherein at least one of the steps of computing, processing, adjusting, and updating are performed for training, testing, or certifying the neural network employed in a machine, robot, or autonomous vehicle.

20. The computer-implemented method of claim 1 , wherein at least one of the steps of computing, processing, adjusting, and updating is performed on a virtual machine comprising a portion of a graphics processing unit.

21. A system, comprising:

a processor configured to implement a neural network model comprising multiple processing layers by:

computing first clipping scalars for quantizing first tensors of a first processing layer that is coupled between two processing layers of the multiple processing layers;

processing an input by the neural network model, according to quantized tensors that include the quantized first tensors, by each layer of the multiple layers in sequence to produce intermediate tensors and an output of the neural network model;

adjusting the first tensors based on a loss gradient; and

updating the first clipping scalars based on a mean squared error to reduce differences between the adjusted first tensors and quantized adjusted first tensors.

22. The system of claim 21 , wherein the first clipping scalars are recursively computed within the first processing layer according to a Newton-Raphson algorithm.

23. A non-transitory computer-readable media storing computer instructions for quantizing tensors of a neural network model comprising multiple processing layers that, when executed by one or more processors, cause the one or more processors to perform the steps of:

computing first clipping scalars for quantizing first tensors of a first processing layer that is coupled between two processing layers of the multiple processing layers;

processing an input by the neural network model, according to quantized tensors that include the quantized first tensors, by each layer of the multiple layers in sequence to produce intermediate tensors and an output of the neural network model;

adjusting the first tensors based on a loss gradient; and

updating the first clipping scalars based on a mean squared error to reduce differences between the adjusted first tensors and quantized adjusted first tensors.

24. The non-transitory computer-readable media of claim 23 , wherein the first clipping scalars are computed by minimizing a quantization mean squared error.

25. A computer-implemented method for reducing power consumption of a neural network model comprising multiple processing layers, comprising:

computing first clipping scalars for quantizing first tensors of a first processing layer of the multiple processing layers;

processing an input by the neural network model, according to quantized tensors that include the quantized first tensors, by each processing layer of the multiple processing layers in sequence, to produce an output of the neural network model by consuming a first amount of power;

adjusting the first tensors based on a loss gradient;

updating the first clipping scalars based on a mean squared error to reduce differences between the adjusted first tensors and quantized adjusted first tensors; and

quantizing the first tensors using the updated first clipping scalars to produce second quantized first tensors, wherein processing the input by the neural network model according to the second quantized first tensors to produce a second output of the neural network model consumes a second amount of power that is less than the first amount of power.

26. The computer-implemented method of claim 25 , wherein the first tensors are at least one of weights or activations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2022
From: SAKR, CHARBEL; DAI, STEVE HAIHANG; KHAILANY, BRUCEK KURDO; DALLY, WILLIAM JAMES; VENKATESAN, RANGHARAJAN; ZIMMER, BRIAN MATTHEW
To: NVIDIA CORPORATION
Reel/Frame 060624/0171 →
Continuity (2)
Provisional Application 63303899 · Jan 27, 2022
Related Publication 20230237308A1 · Jul 27, 2023
References Cited (42)
US 20220092426A1 · Pang · 2022 [cited by examiner]
Abdolrashidi, A., et al., “Pareto-optimal quantized resnet is mostly 4-bit,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3091-3099, 2021. [cited by applicant]
Bianco, S., et al., “Benchmark analysis of representative deep neural network architectures,” IEEE Access, 6:64270-64277, 2018. [cited by applicant]
Choi, J., et al., “PACT: Parameterized clipping activation for quantized neural networks,” arXiv preprint arXiv:1805.06085, 2018. [cited by applicant]
Choi, Y., et al., “Data-free network quantization with adversarial knowledge distillation,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 710-711, 2020. [cited by applicant]
Courbariaux, M., et al., “BinaryConnect: Training deep neural networks with binary weights during propagations,” In Advances in Neural Information Processing Systems, pp. 3123-3131, 2015. [cited by applicant]
Dai, S., et al., “Vs-Quant: Per-vector scaled quantization for accurate low-precision neural network inference,” Proceedings of Machine Learning and Systems, 3, 2021. [cited by applicant]
Dbouk, H., et al., “DBQ: A differentiable branch quantizer for lightweight deep neural networks,” In European Conference on Computer Vision, pp. 90-106, Springer, 2020. [cited by applicant]
Deng, J., et al., “ImageNet: A large-scale hierarchical image database,” In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248-255, IEEE, 2009. [cited by applicant]
Devlin, J., et al., “BERT: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018. [cited by applicant]
Goel, M., et al., “Finite-precision analysis of the pipelined strength-reduced adaptive filter,” Signal Processing, IEEE Transactions on, 46(6):1763-1769, 1998. [cited by applicant]
Gonugondla, S., et al., “Fundamental limits on the precision of in-memory architectures,” In Proceedings of the 39th International Conference on Computer-Aided Design, pp. 1-9, 2020. [cited by applicant]
Gupta, S., et al., “Deep learning with limited numerical precision,” In International Conference on Machine Learning, pp. 1737-1746, 2015. [cited by applicant]
Han, S., et al., “Eie: Efficient inference engine on compressed deep neural network,” ACM SIGARCH Computer Architecture News, 44(3):243-254, 2016. [cited by applicant]
He, K., et al., “Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification,” In Proceedings of the IEEE International Conference on Computer Vision, pp. 1026-1034, 2015. [cited by applicant]
Howard, A., et al., “Searching for mobilenetv3,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1314-1324, 2019. [cited by applicant]
Hubara, I., et al., “Binarized neural networks,” In Advances in Neural Information Processing Systems, pp. 4107-4115, 2016. [cited by applicant]
Jain, S., et al., “BiScaled-DNN: Quantizing long-tailed datastructures with two scale factors for deep neural networks,” In 2019 56th ACM/IEEE Design Automation Conference (DAC), pp. 1-6, IEEE, 2019. [cited by applicant]
Koster, U., et al., “Flexpoint: An adaptive numerical format for efficient training of deep neural networks,” In Advances in Neural Information Processing Systems, pp. 1740-1750, 2017. [cited by applicant]
Krizhevsky, A., et al., “ImageNet classification with deep convolutional neural networks,” In Advances in Neural Information Processing Systems, pp. 1097-1105, 2012. [cited by applicant]
Lecun, Y., et al., “Deep learning,” Nature, 521(7553):436-444, 2015. [cited by applicant]
Lee, E.H., et al., “Lognet: Energy-efficient neural networks using logarithmic computation,” In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5900-5904, IEEE, 2017. [cited by applicant]
Lin, Y., et al., “PredictiveNet: an energy-efficient convolutional neural network via zero prediction,” In Circuits and Systems (ISCAS), 2017 IEEE International Symposium on. IEEE, 2017. [cited by applicant]
Lloyd, S., “Least squares quantization in PCM,” IEEE Transations on Information Theory, 28(2):129-137, 1982. [cited by applicant]
Nagel, M., et al., “Data-free quantization through weight equalization and bias correction,” In Proceedings of the IEEE International Conference on Computer Vision, pp. 1325-1334, 2019. [cited by applicant]
Park, E., et al., “Profit: A novel training method for sub-4-bit mobilenet models,” In European Conference on Computer Vision, pp. 430-446, Springer 2020. [cited by applicant]
Paszke, A., et al., “Automatic differentiation in PyTorch,” In NeurIPS Workshop on Automatic Differentiation, 2017. [cited by applicant]
Rajpurkar, P., et al., “SQuAD: 100,000+ questions for machine comprehension of text,” In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383-2392, 2016. [cited by applicant]
Sakr, C., et al., “Per-tensor fixed-point quantization of the back-propagation algorithm,” In 7th International Conference on Learning Representations, ICLR 2019. [cited by applicant]
Sakr, C., et al., “Accumulation bit-width scaling for ultra-low precision training of deep networks,” In 7th International Conference on Learning Representations, ICLR 2019. [cited by applicant]
Sandler, M., et al., “Mobilenetv2: Inverted residuals and linear bottlenecks,” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510-4520, 2018. [cited by applicant]
Srivastava, N., et al., “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, 15(1):1929-1958, 2014. [cited by applicant]
Sun, X., et al., “Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks,” In NeurIPS, 2019. [cited by applicant]
Taigman, Y., et al., “Deepface: Closing the gap to human-level performance in face verification,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1701-1708, 2014. [cited by applicant]
Tambe, T., et al., “Adaptivfloat: A floating-point based data type for resilient deep learning inference,” arXiv preprint arXiv:1909.13271, 2019. [cited by applicant]
Wang, N., et al., “Training deep neural networks with 8-bit floating point numbers,” In Advances in Neural Information Processing Systems, 2018. [cited by applicant]
Widrow, B., et al., “Quantization noise,” Cambridge University Press, 2:5, 2008. [cited by applicant]
Wu, H., et al., “Integer quantization for deep learning inference: Principles and empirical evaluation,” arXiv preprint arXiv:2004.09602, 2020. [cited by applicant]
Zhang, D., et al., “LQ-Nets: Learned quantization for highly accurate and compact deep neural networks,” In Proceedings of the European Conference on Computer Vision (ECCV), pp. 365-382, 2018. [cited by applicant]
Zhao, J., et al., “Low-precision training in logarithmic number system using multiplicative weight update,” arXiv breprint arXiv:2106. 13914, 2021. [cited by applicant]
Zhou, S., et al., “DoReFa-Net: Training low bandwidth convolutional neural networks with low bandwidth gradients,” arXiv preprint arXiv:1606.06160, 2016. [cited by applicant]
Zhu, Y., et al., “Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,” In Proceedings of the IEEE International conference on computer vision, pp. 19-27, 2015. [cited by applicant]
Cited By (1)
US 12,634,116