IP Library › Granted Patent US 12,450,490
Granted Patent B2
US 12,450,490 · App. 17/830,827 · Granted Oct 21, 2025

Neural network construction method and apparatus having average quantization mechanism

Inventor: Yu-Che Kao (Hsinchu, TW)
Assignee: REALTEK SEMICONDUCTOR CORPORATION
G06N3/082G06N3/0495
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,490
App. No.
17/830,827
Granted
Oct 21, 2025
Kind
B2
Abstract

The present invention discloses a neural network construction method having average quantization mechanism that includes steps outlined below. A weight combination included in each of network layers of a neural network is retrieved. A loss function is generated according to the weight combination of all the network layers and target values. Corresponding to each network layers, a Gini coefficient of the weight combination is calculated and the Gini coefficients corresponding to all the network layers are accumulated as a regularized correction term. The loss function and the regularized correction term are merged as a regularized loss function to perform training on the neural network accordingly to generate a trained weight combination of each of the network layers. A quantization is performed on the trained weight combination of each of the network layers to generate a quantized neural network, in which each of the network layers thereof includes the trained weight combination.

Claims (22)

1. A neural network construction method having average quantization mechanism, comprising:

retrieving a weight combination comprised in each of network layers of a neural network, wherein the weight combination comprises a plurality of floating-point weights;

generating a loss function according to the weight combination of all the network layers and a plurality of target values, wherein the loss function is a first function of the weight combination and the target values;

corresponding to each of the network layers, calculating a Gini coefficient of the weight combination and accumulating the Gini coefficient of each of the network layers as a regularized correction term, wherein the regularized correction term is a second function of the weight combination;

merging the loss function and the regularized correction term to generate a regularized loss function to perform training on the neural network according to the regularized loss function to keep modifying the floating-point weights in the weight combination, so as to generate a trained weight combination of each of the network layers;

performing quantization on the trained weight combination of each of the network layers to generate a quantized neural network, in which each of the network layers of the quantized neural network comprises the quantization weight combination, wherein the quantization weight combination comprises a plurality of integer weights, and a first distribution of the floating-point weights in the trained weight combination is even such that a second distribution of the integer weights in the quantization weight combination is even; and

implementing the quantized neural network by an embedded system chip.

2. The neural network construction method of claim 1 , further comprising:

multiplying the regularized correction term by an order correction parameter such that the regularized correction term and the loss function have the same order.

3. The neural network construction method of claim 1 , wherein the Gini coefficient has a larger value when a distribution of the weight combination is more uneven and has a smaller value when the distribution of the weight combination is more even.

4. A neural network construction apparatus having average quantization mechanism comprising:

a storage circuit configured to store a computer executable command; and

a processing circuit configured to retrieve and execute the computer executable command to execute a neural network construction method, comprising:

retrieving a weight combination comprised in each of network layers of a neural network, wherein the weight combination comprises a plurality of floating-point weights;

generating a loss function according to the weight combination of all the network layers and a plurality of target values, wherein the loss function is a first function of the weight combination and the target values;

corresponding to each of the network layers, calculating a Gini coefficient of the weight combination and accumulating the Gini coefficient of each of the network layers as a regularized correction term, wherein the regularized correction term is a second function of the weight combination;

merging the loss function and the regularized correction term to generate a regularized loss function to perform training on the neural network according to the regularized loss function to keep modifying the floating-point weights in the weight combination, so as to generate a trained weight combination of each of the network layers;

performing quantization on the trained weight combination of each of the network layers to generate a quantized neural network, in which each of the network layers of the quantized neural network comprises the quantization weight combination, wherein the quantization weight combination comprises a plurality of integer weights, and a first distribution of the floating-point weights in the trained weight combination is even such that a second distribution of the integer weights in the quantization weight combination is even; and

implementing the quantized neural network by an embedded system chip.

5. The neural network construction apparatus of claim 4 , wherein the neural network construction method further comprising:

multiplying the regularized correction term by an order correction parameter such that the regularized correction term and the loss function have the same order.

6. The neural network construction apparatus of claim 4 , wherein the Gini coefficient has a larger value when a distribution of the weight combination is more uneven and has a smaller value when the distribution of the weight combination is more even.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2022
From: KAO, YU-CHE
To: REALTEK SEMICONDUCTOR CORPORATION
Reel/Frame 060087/0183 →
Priority Claims (1)
TW 110136791 · Oct 1, 2021 · national
Continuity (1)
Related Publication 20230114610A1 · Apr 13, 2023
References Cited (9)
Guest, Olivia. “Using the Gini coefficient to evaluate deep neural network layer representations.” Blog post. <Neuroplausible.com/gini> (2017). (Year: 2017). [cited by examiner]
Long, Xin, et al. “Learning sparse convolutional neural network via quantization with low rank regularization.” IEEE Access 7 (2019): 51866-51876. (Year: 2017). [cited by examiner]
B. L.Deng, G.Li, S.Han, L.Shi, andY.Xie, “Model Compression and Hardware Acceleration for Neural Networks: a Comprehensive Survey,” Proc. IEEE, vol. 108, No. 4, pp. 485-532, 2020. [cited by applicant]
R.Krishnamoorthi, “Quantizing deep convolutional networks for efficient inference: a whitepaper,” arXiv Prepr. arXiv1806.08342, 2018. [cited by applicant]
B.Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 2704-2713. [cited by applicant]
T.Sheng, C.Feng, S.Zhuo, X.Zhang, L.Shen, andM.Aleksic, “A Quantization-Friendly Separable Convolution for MobileNets,” in 2018 1st Workshop on Energy Efficient Machine Learning and Cognitive Computing for Embedded Appl… [cited by applicant]
M.Alizadeh, A.Behboodi, M.vanBaalen, C.Louizos, T.Blankevoort, andM.Welling, “Gradient L1 Regularization for Quantization Robustness,” in International Conference on Learning Representations, 2020. [cited by applicant]
Y.Choi, M.El-Khamy, andJ.Lee, “Learning Sparse Low-Precision Neural Networks With Learnable Regularization,” IEEE Access, vol. 8, pp. 96963-96974, 2020. [cited by applicant]
O.Guest, “Using the Gini coefficient to evaluate deep neural network layer representations,” Blog post, 2017. [cited by applicant]
Cited By (1)
US 12,634,116