IP Library › Granted Patent US 12,632,711
Granted Patent B2
US 12,632,711 · App. 17/928,713 · Granted May 19, 2026

Method, apparatus, computing device and medium for quantizing neutral network model

Inventor: Yongsen Jiang (Beijing, CN)
Assignee: DOUYIN VISION CO., LTD.
G06N3/0495
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,711
App. No.
17/928,713
Granted
May 19, 2026
Kind
B2
Abstract

A method comprises updating a neural network model based on a training dataset; adjusting a first set of parameters of a first part of the updated neural network model to being within a first range, and adjusting a second set of parameters of a second part of the updated neural network model to being within a second range, a size of the second range exceeding that of the first range. The method further comprises quantizing the adjusted first set of parameters with a first number of bits and the adjusted second set of parameters with a second number of bits, the second number being greater than the first number. In this way, the parameters are quantified in differentiated manners in combination with a training process, and the compression efficiency and execution efficiency of the neural network model may be improved while maintaining the parameter precision and model performance.

Claims (55)

1 . A method comprising:

updating a neural network model based on a training dataset, wherein the neural network model is applicable for voice recognition and includes a Transformer-based model;

adjusting a first set of parameters of a first part of the updated neural network model to be within a first range;

adjusting a second set of parameters of a second part of the updated neural network model to be within a second range, wherein a size of the second range is greater than a size of the first range;

quantizing the adjusted first set of parameters with a first number of bits; and

quantizing the adjusted second set of parameters with a second number of bits, the second number being greater than the first number,

wherein the first part comprises a linear transformation-based network, and the first set of parameters comprises weights for a linear transformation, and

the second part comprises a convolutional neural network, and the second set of parameters comprises parameters of a convolution kernel of the convolutional neural network.

2 . The method according to claim 1 , further comprising: before quantizing the adjusted second set of parameters with the second number of bits,

iteratively updating the adjusted second set of parameters based on the training dataset and the quantized first set of parameters.

3 . The method according to claim 2 , wherein iteratively updating the adjusted second set of parameters based on the training dataset and the quantized first set of parameters comprises:

iteratively updating the neural network model based on the training dataset while keeping the quantized first set of parameters unchanged, wherein the iteratively updated neural network model comprises an iteratively updated second set of parameters.

4 . The method according to claim 1 , wherein adjusting the first set of parameters of the first part of the updated neural network model to be within the first range comprises:

adjusting a parameter in the first set of parameters that is greater than an upper limit value of the first range to the upper limit value; and

adjusting a parameter in the first set of parameters that is smaller than a lower limit value of the first range to the lower limit value.

5 . The method according to claim 1 , wherein quantizing the adjusted first set of parameters with the first number of bits comprises:

dividing the first range into a plurality of sub-ranges based on the first number; and

determining a quantized value of each parameter in the first set of parameters based on the plurality of sub-ranges and a value of each parameter.

6 . The method according to claim 1 , further comprising:

generating a quantized neural network model based on the quantized first set of parameters and the quantized second set of parameters; and

transmitting the quantized neural network model to a terminal device.

7 . The method according to claim 1 , wherein the first number is 4 and the second number is 8.

8 . The method according to claim 1 , wherein updating the neural network model comprises:

updating the neural network model based on a floating point format.

9 . A device, comprising:

at least one processing unit; and

at least one memory coupled to the at least one processing unit and having instructions stored thereon, which when executed by the at least one processing unit, cause the device to

update a neural network model based on a training dataset, wherein the neural network model is applicable for voice recognition and includes a Transformer-based model;

adjust a first set of parameters of a first part of the updated neural network model to be within a first range;

adjust a second set of parameters of a second part of the updated neural network model to be within a second range, wherein a size of the second range is greater than a size of the first range;

quantize the adjusted first set of parameters with a first number of bits; and

quantize the adjusted second set of parameters with a second number of bits, the second number being greater than the first number,

wherein the first part comprises a linear transformation-based network, and the first set of parameters comprises weights for a linear transformation; and

the second part comprises a convolutional neural network, and the second set of parameters comprises parameters of a convolution kernel of the convolutional neural network.

10 . The device of claim 9 , wherein the instructions, when executed by the at least one processing unit, further cause the device to

iteratively update the second set of parameters based on the training dataset and the quantized first set of parameters.

11 . The device according to claim 10 , wherein the instructions, when executed by the at least one processing unit, further cause the device to

iteratively update the neural network model based on the training dataset while keeping the quantized first set of parameters unchanged, wherein the iteratively updated neural network model comprises an iteratively updated second set of parameters.

12 . The device according to claim 9 , wherein the instructions, when executed by the at least one processing unit, further cause the device to

adjust a parameter in the first set of parameters that is greater than an upper limit value of the first range to the upper limit value; and

adjust a parameter in the first set of parameters that is smaller than a lower limit value of the first range to the lower limit value.

13 . The device according to claim 9 , wherein the instructions, when executed by the at least one processing unit, further cause the device to

divide the first range into a plurality of sub-ranges based on the first number; and

determine a quantized value of each parameter in the first set of parameters based on the plurality of sub-ranges and a value of each parameter.

14 . The device according to claim 9 , wherein the first number is 4 and the second number is 8.

15 . The device according to claim 9 , wherein the instructions, when executed by the at least one processing unit, further cause the device to

update the neural network model based on a floating point format.

16 . A non-transitory computer-readable storage medium comprising machine-executable instructions that, when executed by a device, cause the device to

update a neural network model based on a training dataset, wherein the neural network model is applicable for voice recognition and includes a Transformer-based model;

adjust a first set of parameters of a first part of the updated neural network model to be within a first range;

adjust a second set of parameters of a second part of the updated neural network model to be within a second range, wherein a size of the second range is greater than a size of the first range;

quantize the adjusted first set of parameters with a first number of bits; and

quantize the adjusted second set of parameters with a second number of bits, the second number being greater than the first number,

wherein the first part comprises a linear transformation-based network, and the first set of parameters comprises weights for a linear transformation; and

the second part comprises a convolutional neural network, and the second set of parameters comprises parameters of a convolution kernel of the convolutional neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2023
From: JIANG, YONGSEN
To: JIANGSU BYTEDANCE TCHNOLOGY CO., LTD.
Reel/Frame 064103/0104 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2023
From: JIANGSU BYTEDANCE TCHNOLOGY CO., LTD.
To: DOUYIN VISION CO., LTD.
Reel/Frame 064103/0133 →
Priority Claims (1)
CN 202210910943.2 · Jul 29, 2022 · national
Continuity (1)
Related Publication 20240220782A1 · Jul 4, 2024
References Cited (22)
US 11615304B1 · Palkar · 2023 [cited by examiner]
US 11645493B2 · Burger · 2023 [cited by examiner]
US 12099915B2 · Boo · 2024 [cited by examiner]
US 20190122100A1 · Kang · 2019 [cited by examiner]
US 20190340492A1 · Burger et al. · 2019 [cited by applicant]
US 20200257960A1 · Gabriel · 2020 [cited by examiner]
US 20200311552A1 · A · 2020 [cited by examiner]
US 20210286688A1 · Liu · 2021 [cited by examiner]
US 20210326710A1 · Wang · 2021 [cited by examiner]
US 20230229895A1 · Nunes Coelho, Jr. · 2023 [cited by examiner]
US 20230306255A1 · Charlaix · 2023 [cited by examiner]
US 20240220782A1 · Jiang · 2024 [cited by examiner]
CN 110717585A · 2020 [cited by applicant]
CN 110799994A · 2020 [cited by applicant]
CN 111582432A · 2020 [cited by applicant]
CN 113610709A · 2021 [cited by applicant]
CN 114139683A · 2022 [cited by applicant]
Gulati, A. et al., “Conformer: Convolution-augmented Transformer for Speech Recognition,” arXiv:2005.08100v1, May 16, 2020, 5 pages. [cited by applicant]
European Patent Office, Extended European Search Report Issued in Application No. 22808581.7, Apr. 4, 2024, Germany, 9 pages. [cited by applicant]
Gholami, A. et al., “A Survey of Quantization Methods for Efficient Neural Network Inference,” arXiv:2103.13630v3, Jun. 21, 2021, 33 pages. [cited by applicant]
Ding, S. et al., “4-bit Conformer with Native Quantization Aware Training for Speech Recognition,” arXiv:2203.15952v1, Mar. 29, 2022, 5 pages. [cited by applicant]
ISA China National Intellectual Property Administration, International Search Report Issued in Application No. PCT/CN2022/130432, Apr. 19, 2023, 3 pages. [cited by applicant]