IP Library › Granted Patent US 12,412,096
Granted Patent B2
US 12,412,096 · App. 17/032,248 · Granted Sep 9, 2025

Efficient weight clipping for neural networks

Inventors: Arun Coimbatore Ramachandran (Bangalore, IN); Chandra Kumar Ramasamy (Bangalore, IN); Keerthan S. Shagrithaya (Bangalore, IN); Prakash Sathyanath Raghavendra (Bangalore, IN); Vasanthakumar Rajagopal (Bangalore, IN)
Assignee: Advanced Micro Devices, Inc.
G06N3/082G06F7/50G06F7/523G06F7/535G06F7/5443G06N3/063G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,096
App. No.
17/032,248
Granted
Sep 9, 2025
Kind
B2
Abstract

Systems, apparatuses, and methods for implementing one-sided per-kernel clipping and weight transformation for neural networks are disclosed. Various parameters of a neural network are quantized from higher-bit representations to lower-bit representations to reduce memory utilization and power consumption. To exploit the effective range of quantized representations, positively biased weights are clipped and negated before convolution. Then, the results are rescaled back after convolution. A one-sided clipping technique is used for transforming weights to exploit the quantization range effectively, with the side chosen to be clipped being the biased side. This technique uses a global strategy for clipping without requiring skilled expertise. This approach allows the system to retain as much information as possible without losing unnecessary accuracy when quantizing parameters from higher-bit representations to lower-bit representations.

Claims (68)

1. A system comprising:

a processing circuit, wherein the processing circuit is configured to:

receive a plurality of weights of a given kernel corresponding to a layer of a neural network comprising a plurality of layers, each of the plurality of weights comprising a signed value; and

generate a classification of an input dataset with the neural network, wherein to generate the classification, the processing circuit is configured to:

negate each weight of the plurality of weights, responsive to a maximum weight of the plurality of weights of the given kernel being greater than an absolute value of a minimum weight of the plurality of weights of the given kernel;

quantize the negated weights to generate quantized versions of the plurality of weights; and

execute, using the generated quantized versions of the plurality of weights, the layer of the neural network to generate an output based on the input dataset.

2. The system as recited in claim 1 , wherein the processing circuit is further configured to perform one-sided clipping of the plurality of weights prior to negating the plurality of weights, and wherein quantizing negated weights comprises rounding to a nearest quantization level of a reduced precision data format.

3. The system as recited in claim 2 , wherein performing one-sided clipping of the plurality of weights comprises:

calculating a scale threshold based on the maximum weight and the absolute value of a minimum weight of the given kernel; and

scaling each weight of the plurality of weights using the scale threshold.

4. The system as recited in claim 3 , wherein calculating the scale threshold comprises:

multiplying a tunable parameter by the maximum weight to generate a first product;

subtracting the tunable parameter from one to generate a difference value;

multiplying the difference value by the absolute value of the minimum weight to generate a second product; and

adding the first product to the second product.

5. The system as recited in claim 3 , wherein each weight is multiplied by a maximum positive value of a range of the reduced precision data format, and wherein each weight is divided by the scale threshold.

6. The system as recited in claim 3 , wherein the processing circuit is further configured to perform convolution between activation data and negated scaled versions of the plurality of weights, wherein negated scaled versions of the plurality of weights are quantized to an integer 4 (INT4) representation with a range of −8 to 7.

7. The system as recited in claim 1 , further comprising:

a memory storing the plurality of weights of a plurality of kernels of the plurality of layers of the neural network; and

wherein for each kernel of the plurality of kernels of the plurality of layers of the neural network, the processing circuit is configured to:

compare a maximum weight of the kernel to an absolute value of a minimum weight of the kernel;

responsive to determining that the maximum weight of the kernel is greater than the absolute value of the minimum weight of the kernel:

negate each weight of the plurality of weights of the kernel;

quantize, to a reduced precision data format, the negated weights of the kernel; and

perform convolution using quantized negated versions of the plurality of weights of the kernel.

8. A method comprising:

receiving, by a processing circuit, a plurality of weights of a given kernel corresponding to a layer of a neural network comprising a plurality of layers, each of the plurality of weights comprising a signed value; and

generating, by the processing circuit, a classification of an input dataset with the neural network, wherein generating the classification comprises:

negating, by the processing circuit, each weight of the plurality of weights, responsive to a maximum weight of the plurality of weights of the given kernel being greater than an absolute value of a minimum weight of the plurality of weights of the given kernel;

quantizing, by the processing circuit, the negated weights to generate quantized versions of the plurality of weights; and

executing, by the processing circuit using the generated quantized versions of the plurality of weights, the layer of the neural network to generate an output based on the input dataset.

9. The method as recited in claim 8 , further comprising performing, by the processing circuit, one-sided clipping of the plurality of weights prior to negating the plurality of weights, and wherein quantizing negated weights comprises rounding to a nearest quantization level of a reduced precision data format.

10. The method as recited in claim 9 , wherein performing one-sided clipping of the plurality of weights comprises:

calculating, by the processing circuit, a scale threshold based on the maximum weight and the absolute value of a minimum weight of the given kernel; and

scaling, by the processing circuit, each weight of the plurality of weights using the scale threshold.

11. The method as recited in claim 10 , wherein calculating the scale threshold comprises performing, by the processing circuit, one or more operations, at least including:

multiplying a tunable parameter by the maximum weight to generate a first product;

subtracting the tunable parameter from one to generate a difference value;

multiplying the difference value by the absolute value of the minimum weight to generate a second product; and

adding the first product to the second product.

12. The method as recited in claim 10 , wherein each weight is multiplied by a maximum positive value of a range of a reduced precision data format, and wherein each weight is divided by the scale threshold.

13. The method as recited in claim 10 , further comprising performing, by the processing circuit, convolution between activation data and negated scaled versions of the plurality of weights, and wherein the negated scaled versions of the plurality of weights are quantized to an integer 4 (INT4) representation with a range of −8 to 7.

14. The method as recited in claim 8 , wherein for each kernel of a plurality of kernels of a plurality of layers of a neural network, the method further comprising:

comparing, by the processing circuit, a maximum weight of the kernel to an absolute value of a minimum weight of the kernel;

responsive to determining that the maximum weight of the kernel is greater than the absolute value of the minimum weight of the kernel:

negating, by the processing circuit, each weight of the plurality of weights of the kernel;

quantizing, by the processing circuit, negated versions of the plurality of weights of the kernel; and

performing, by the processing circuit, a convolution operation using quantized negated versions of the plurality of weights of the kernel.

15. An apparatus comprising:

a memory comprising circuitry configured to store a plurality of weights of a given kernel corresponding to a layer of a neural network comprising a plurality of layers, each of the plurality of weights comprising a signed value; and

a processing circuit coupled to the memory, wherein the processing circuit is configured to:

retrieve, from the memory, the plurality of weights of the given kernel; and

generate a classification of an input dataset with the neural network, wherein to generate the classification, the processing circuit is configured to:

negate each weight of the plurality of weights, responsive to a maximum weight of the plurality of weights of the given kernel being greater than an absolute value of a minimum weight of the plurality of weights of the given kernel;

quantize the negated weights to generate quantized versions of the plurality of weights; and

execute, using the generated quantized versions of the plurality of weights, the layer of the neural network to generate an output based on the input dataset.

16. The apparatus as recited in claim 15 , wherein the processing circuit is further configured to perform one-sided clipping of the plurality of weights prior to negating the plurality of weights, and wherein quantizing negated weights comprises rounding to a nearest quantization level of a reduced precision data format.

17. The apparatus as recited in claim 16 , wherein performing one-sided clipping of the plurality of weights comprises:

calculating a scale threshold based on the maximum weight and the absolute value of a minimum weight of the given kernel; and

scaling each weight of the plurality of weights using the scale threshold.

18. The apparatus as recited in claim 17 , wherein calculating the scale threshold comprises:

multiplying a tunable parameter by the maximum weight to generate a first product;

subtracting the tunable parameter from one to generate a difference value;

multiplying the difference value by the absolute value of the minimum weight to generate a second product; and

adding the first product to the second product.

19. The apparatus as recited in claim 17 , wherein each weight is multiplied by a maximum positive value of a range of the reduced precision data format, and wherein each weight is divided by the scale threshold.

20. The apparatus as recited in claim 17 , wherein responsive to the maximum weight being less than the absolute value of the minimum weight, the quantized versions of the plurality of weights are based on non-negated versions of the plurality of weights.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE EXECUTION DATE OF THE 5TH INVENTOR PREVIOUSLY RECORDED AT REEL: 054125 FRAME: 0102. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 21, 2021
From: RAMACHANDRAN, ARUN COIMBATORE; RAMASAMY, CHANDRA KUMAR; SHAGRITHAYA, KEERTHAN S.; RAGHAVENDRA, PRAKASH SATHYANATH; RAJAGOPAL, VASANTHAKUMAR
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 055347/0663 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2020
From: RAMACHANDRAN, ARUN COIMBATORE; RAMASAMY, CHANDRA KUMAR; SHAGRITHAYA, KEERTHAN S.; RAGHAVENDRA, PRAKASH SATHYANATH; RAJAGOPAL, VASANTHAKUMAR
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 054125/0102 →
Continuity (2)
Provisional Application 63044610 · Jun 26, 2020
Related Publication 20210406690A1 · Dec 30, 2021
References Cited (22)
US 5058034A · Murphy · 1991 [cited by examiner]
US 6192360B1 · Dumais · 2001 [cited by examiner]
US 20180225116A1 · Henry · 2018 [cited by examiner]
US 20190347550A1 · Jung · 2019 [cited by examiner]
US 20200134376A1 · Sallee · 2020 [cited by examiner]
Xu, Jiawei, et al. “A low-power arithmetic element for multi-base logarithmic computation on deep neural networks.” 2018 31st IEEE International System-on-Chip Conference (SOCC). IEEE, 2018. (Year: 2018). [cited by examiner]
Choukroun, Yoni, et al. “Low-bit quantization of neural networks for efficient inference.” 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). IEEE, 2019. (Year: 2019). [cited by examiner]
Helwegen, Koen, et al. “Latent weights do not exist: Rethinking binarized neural network optimization.” Advances in neural information processing systems 32 (2019). (Year: 2019). [cited by examiner]
Thierens, Dirk. “Non-redundant genetic coding of neural networks.” Proceedings of IEEE International Conference on Evolutionary Computation. IEEE, 1996. (Year: 1996). [cited by examiner]
Courbariaux, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights During Propagations”, 9 pages. [cited by applicant]
Jacob, et al., “Quantization and Training of Neural Networks for Effcient Integer-Arithmetic-Only Inference”, Google Inc., Dec. 15, 2017, 14 pages. [cited by applicant]
Choi, et al., “PACT: Parameterized Clipping Activation for Quantized Neural Networks”, IBM Research AI, Jul. 17, 2018, 15 pages. [cited by applicant]
Krishnamoorthi, “Quantizing Deep Convultional Networks for Effcient Inference: A Whitepaper”, Jun. 21, 2018, 36 pages. [cited by applicant]
Zhao, et al., “Improving Neural Network Quantization without Retraining Using Outlier Channel Splitting”, May 22, 2019, 10 pages. [cited by applicant]
Choukroun, et al., “Low-Bit Quantization of Neural Networks for Efficient Inference”, Mar. 25, 2019, 10 pages. [cited by applicant]
Anonymous Authors, ACIQ: Analyiical Clipping for Integer Quantization of Neural Networks, 11 pages. [cited by applicant]
Migacz, “8-Bit Inference with TensorRT”, NVIDIA, May 8, 2017, 41 pages. [cited by applicant]
Wu, “Low Precision Inference on GPU”, NVIDIA, 60 pages. [cited by applicant]
Colbert et al., U.S. Appl. No. 18/065,393, entitled “Quantization—Aware Training With Numerical Overflow Avoidance for Neural Networks”, filed Dec. 13, 2022, 45 pages. [cited by applicant]
International Search Report and Written Opinion of International Application No. PCT/S2024/033580, mailed Oct. 2, 2024, 12 pp. [cited by applicant]
Nilesh Prasad Pandey et al: “A Practical Mixed Precision Algorithm for Post-Training Quantization”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Feb. 10, 2023 (Feb. 10, 20… [cited by applicant]
Sangeetha Siddegowda et al: “Neural Network Quantization with AI Model Efficiency Toolkit (AIMET)”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 20, 2022 (Jan. 20, 20… [cited by applicant]