IP Library › Granted Patent US 12,620,388
Granted Patent B2
US 12,620,388 · App. 18/632,237 · Granted May 5, 2026

Robustness aware norm decay for quantization aware training and generalization

Inventors: David Qiu (Fremont, CA); David Rim (Mountain View, CA); Shaojin Ding (Mountain View, CA); Yanzhang He (Mountain View, CA)
Assignee: Google LLC
G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,388
App. No.
18/632,237
Granted
May 5, 2026
Kind
B2
Abstract

A method includes obtaining a plurality of training samples, determining a minimum integer fixed-bit width representing a maximum quantization of an automatic speech recognition (ASR) model, and training the ASR model on the plurality of training samples using a quantity of random noise. The ASR model includes a plurality of weights that each include a respective float value. The quantity of random noise is based on the minimum integer fixed-bit value. After training the ASR model, the method also includes selecting a target integer fixed-bit width greater than or equal to the minimum integer fixed-bit width, and for each respective weight of the plurality of weights, quantizing the respective weight from the respective float value to a respective integer associated with a value of the selected target integer fixed-bit width. The operations also include providing the quantized trained ASR model to a user device.

Claims (56)

1 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a plurality of training samples, each respective training sample of the plurality of training samples comprising:

a respective speech utterance; and

a respective textual utterance representing a transcription of the respective speech utterance;

determining a minimum integer fixed-bit width representing a maximum quantization of an automatic speech recognition (ASR) model, the ASR model comprising a plurality of weights, each respective weight of the plurality of weights comprising a respective float value;

training the ASR model on the plurality of training samples using a quantity of random noise, the quantity of random noise based on the minimum integer fixed-bit width;

after training the ASR model, selecting a target integer fixed-bit width greater than or equal to the minimum integer fixed-bit width;

for each respective weight of the plurality of weights, quantizing the respective weight from the respective float value to a respective integer associated with a value of the selected target integer fixed-bit width; and

providing the quantized trained ASR model to a user device.

2 . The method of claim 1 , wherein the maximum quantization level comprises 4-bit quantization.

3 . The method of claim 1 , wherein the maximum quantization level comprises 2-bit quantization.

4 . The method of claim 1 , wherein the random noise is drawn from a uniform distribution of noise.

5 . The method of claim 1 , wherein training the ASR model using the quantity of random noise comprises, for each respective channel of each respective tensor of the ASR model:

determining a respective maximum value for the respective channel of the respective tensor; and

adding, to the respective channel of the respective tensor, a uniform distribution of noise based on the respective maximum value and the minimum integer fixed-bit width.

6 . The method of claim 5 , wherein the uniform distribution of noise represents the entire range of noise the ASR model receives due to quantization up to the minimum integer fixed-bit width.

7 . The method of claim 5 , wherein adding the uniform distribution of noise comprises scaling the uniform distribution of noise based on the respective maximum value.

8 . The method of claim 7 , wherein scaling the uniform distribution of noise is further based on a sensitivity of the respective channel to scaling.

9 . The method of claim 1 , wherein training the ASR model using the quantity of random noise comprises adding, during forward propagation of the ASR model, the quantity of random noise.

10 . The method of claim 1 , wherein:

the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising a respective float value; and

for each respective activation of the plurality of activations, the operations further comprise quantizing the respective activation from the respective float value to the respective integer with the value of the selected target fixed-bit width.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining a plurality of training samples, each respective training sample of the plurality of training samples comprising:

a respective speech utterance; and

a respective textual utterance representing a transcription of the respective speech utterance;

determining a minimum integer fixed-bit width representing a maximum quantization of an automatic speech recognition (ASR) model, the ASR model comprising a plurality of weights, each respective weight of the plurality of weights comprising a respective float value;

training the ASR model on the plurality of training samples using a quantity of random noise, the quantity of random noise based on the minimum integer fixed-bit width;

after training the ASR model, selecting a target integer fixed-bit width greater than or equal to the minimum integer fixed-bit width;

for each respective weight of the plurality of weights, quantizing the respective weight from the respective float value to a respective integer associated with a value of the selected target integer fixed-bit width; and

providing the quantized trained ASR model to a user device.

12 . The system of claim 11 , wherein the maximum quantization level comprises 4-bit quantization.

13 . The system of claim 11 , wherein the maximum quantization level comprises 2-bit quantization.

14 . The system of claim 11 , wherein the random noise is drawn from a uniform distribution of noise.

15 . The system of claim 11 , wherein training the ASR model using the quantity of random noise comprises, for each respective channel of each respective tensor of the ASR model:

determining a respective maximum value for the respective channel of the respective tensor; and

adding, to the respective channel of the respective tensor, a uniform distribution of noise based on the respective maximum value and the minimum integer fixed-bit width.

16 . The system of claim 15 , wherein the uniform distribution of noise represents the entire range of noise the ASR model receives due to quantization up to the minimum integer fixed-bit width.

17 . The system of claim 15 , wherein adding the uniform distribution of noise comprises scaling the uniform distribution of noise based on the respective maximum value.

18 . The system of claim 17 , wherein scaling the uniform distribution of noise is further based on a sensitivity of the respective channel to scaling.

19 . The system of claim 11 , wherein training the ASR model using the quantity of random noise comprises adding, during forward propagation of the ASR model, the quantity of random noise.

20 . The system of claim 11 , wherein:

the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising a respective float value; and

for each respective activation of the plurality of activations, the operations further comprise quantizing the respective activation from the respective float value to the respective integer associated with the value of the selected target fixed-bit width.

21 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a plurality of training samples;

determining a minimum integer fixed-bit width representing a maximum quantization of a model, the model comprising a plurality of weights, each respective weight of the plurality of weights comprising a respective float value;

training the model on the plurality of training samples using a quantity of random noise, the quantity of random noise based on the minimum integer fixed-bit width;

after training the model, selecting a target integer fixed-bit width greater than or equal to the minimum integer fixed-bit width;

for each respective weight of the plurality of weights, quantizing the respective weight from the respective float value to a respective integer associated with a value of the selected target integer fixed-bit width; and

providing the quantized trained model to a user device.

22 . The method of claim 21 , wherein the model comprises an automated speech recognition model.

23 . The method of claim 21 , wherein the model comprises a large language model.

24 . The method of claim 21 , wherein the model comprises an image processing model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2024
From: QIU, DAVID; RIM, DAVID; DING, SHAOJIN; HE, YANZHANG
To: GOOGLE LLC
Reel/Frame 067147/0052 →
Continuity (2)
Provisional Application 63495310 · Apr 11, 2023
Related Publication 20240347043A1 · Oct 17, 2024
References Cited (91)
US 9275638B2 · Meloney · 2016 [cited by examiner]
US 9865254B1 · Filimonov · 2018 [cited by examiner]
US 10127905B2 · Lee · 2018 [cited by examiner]
US 12007867B1 · Park · 2024 [cited by examiner]
US 12475388B2 · Yu · 2025 [cited by examiner]
US 20100318354A1 · Seltzer · 2010 [cited by examiner]
US 20140257813A1 · Mortensen · 2014 [cited by examiner]
US 20180268244A1 · Moazzami · 2018 [cited by examiner]
US 20190050710A1 · Wang · 2019 [cited by examiner]
US 20190340492A1 · Burger · 2019 [cited by examiner]
US 20190340499A1 · Burger · 2019 [cited by examiner]
US 20190347550A1 · Jung · 2019 [cited by examiner]
US 20200160079A1 · Reda · 2020 [cited by examiner]
US 20200193270A1 · Wu · 2020 [cited by examiner]
US 20200202213A1 · Darvish Rouhani · 2020 [cited by examiner]
US 20200202218A1 · Csefalvay · 2020 [cited by examiner]
US 20200211154A1 · Ng · 2020 [cited by examiner]
US 20200272882A1 · Lo · 2020 [cited by examiner]
US 20200302289A1 · Ren · 2020 [cited by examiner]
US 20200302298A1 · Van Baalen · 2020 [cited by examiner]
US 20200302299A1 · Nagel · 2020 [cited by examiner]
US 20200320385A1 · Yang · 2020 [cited by examiner]
US 20200364552A1 · Guo · 2020 [cited by examiner]
US 20200380357A1 · Yao · 2020 [cited by examiner]
US 20210014039A1 · Zhang · 2021 [cited by examiner]
US 20210019076A1 · Kim · 2021 [cited by examiner]
US 20210019616A1 · Chen · 2021 [cited by examiner]
US 20210019630A1 · Yao · 2021 [cited by examiner]
US 20210065011A1 · Liu · 2021 [cited by examiner]
US 20210089906A1 · Lazovich · 2021 [cited by examiner]
US 20210110273A1 · Kwon · 2021 [cited by examiner]
US 20210117623A1 · Aly · 2021 [cited by examiner]
US 20210141799A1 · Steedman Henderson · 2021 [cited by examiner]
US 20210209822A1 · Yang · 2021 [cited by examiner]
US 20210216902A1 · Sutcher-Shepard · 2021 [cited by examiner]
US 20210274496A1 · Chen · 2021 [cited by examiner]
US 20210279635A1 · Gadelrab · 2021 [cited by examiner]
US 20210370993A1 · Qian · 2021 [cited by examiner]
US 20220044114A1 · Sriram · 2022 [cited by examiner]
US 20220058530A1 · Lin · 2022 [cited by examiner]
US 20220083855A1 · Choi · 2022 [cited by examiner]
US 20220092382A1 · Moshovos · 2022 [cited by examiner]
US 20220101034A1 · Ekin · 2022 [cited by examiner]
US 20220114457A1 · Schradin, III · 2022 [cited by examiner]
US 20220124433A1 · Kupryjanow · 2022 [cited by examiner]
US 20220245444A1 · Golub · 2022 [cited by examiner]
US 20220335293A1 · Choi · 2022 [cited by examiner]
US 20220398430A1 · Boo · 2022 [cited by examiner]
US 20230004816A1 · Lee · 2023 [cited by examiner]
US 20230042397A1 · Yu · 2023 [cited by examiner]
US 20230068381A1 · Gollanapalli · 2023 [cited by examiner]
US 20230143985A1 · Han · 2023 [cited by examiner]
US 20230222334A1 · Clement · 2023 [cited by examiner]
US 20230281423A1 · Bijalwan · 2023 [cited by examiner]
US 20230298569A1 · Ding · 2023 [cited by examiner]
US 20230351207A1 · Yan · 2023 [cited by examiner]
US 20230360065A1 · Huang · 2023 [cited by examiner]
US 20230360358A1 · Zhao · 2023 [cited by examiner]
US 20230368227A1 · Chien · 2023 [cited by examiner]
US 20230401834A1 · Liang · 2023 [cited by examiner]
US 20240028887A1 · Tenace · 2024 [cited by examiner]
US 20240046426A1 · Jung · 2024 [cited by examiner]
US 20240070266A1 · Chai · 2024 [cited by examiner]
US 20240080788A1 · Wang · 2024 [cited by examiner]
US 20240129577A1 · Umakanth · 2024 [cited by examiner]
US 20240135153A1 · Csefalvay · 2024 [cited by examiner]
US 20240135174A1 · Zhou · 2024 [cited by examiner]
US 20240160406A1 · Venkatesan · 2024 [cited by examiner]
US 20240160898A1 · Lee · 2024 [cited by examiner]
US 20240185046A1 · Li · 2024 [cited by examiner]
US 20240185086A1 · Hou · 2024 [cited by examiner]
US 20240185498A1 · Francis · 2024 [cited by examiner]
US 20240202501A1 · Zan · 2024 [cited by examiner]
US 20240205563A1 · Beerel · 2024 [cited by examiner]
US 20240212673A1 · Zheng · 2024 [cited by examiner]
US 20240220782A1 · Jiang · 2024 [cited by examiner]
US 20240233511A1 · Golbon Haghighi · 2024 [cited by examiner]
US 20240256842A1 · Ashfaq · 2024 [cited by examiner]
US 20240257800A1 · Michieli · 2024 [cited by examiner]
US 20240273399A1 · Yang · 2024 [cited by examiner]
US 20240303505A1 · Ko · 2024 [cited by examiner]
US 20240386704A1 · Zhang · 2024 [cited by examiner]
US 20250045573A1 · Yao · 2025 [cited by examiner]
US 20250086459A1 · Hu · 2025 [cited by examiner]
US 20250200348A1 · Lin · 2025 [cited by examiner]
US 20250209274A1 · Yang · 2025 [cited by examiner]
Alexandre D'Efossez et al: “Differentiable Model Compression via Pseudo Quantization Noise”, arvix.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 17, 2022 (Oct. 17, 2022), XP… [cited by examiner]
International Search Report and Written Opinion issued in related PCT Application No. PCT/US2024/023948, dated Jun. 6, 2024. [cited by applicant]
Alexandre D'Efossez et al: “Differentiable Model Compression via Pseudo Quantization Noise”, arvix.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 17, 2022 (Oct. 17, 2022), XP… [cited by applicant]
Andrea Fasoli et al: “4-bit Quantization of LSTM-based Speech Recognition Models”, arvix.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 27, 2021 (Aug. 27, 2021), XP091040290. [cited by applicant]
Kai Zhen et al: “Sub-8-Bit Quantization Aware Training for 8-Bit Neural Network Accelerator with On-Device Speech Recognition”, arvix.org, Cornell University Library, 201 Olin Library Cornell Univeristy Ithaca, NY 14853… [cited by applicant]