IP Library › Granted Patent US 12,374,323
Granted Patent B2
US 12,374,323 · App. 18/186,774 · Granted Jul 29, 2025

4-bit conformer with accurate quantization training for speech recognition

Inventors: Shaojin Ding (Mountain View, CA); Oleg Rybakov (Mountain View, CA); Phoenix Meadowlark (Mountain View, CA); Shivani Agrawal (Mountain View, CA); Yanzhang He (Palo Alto, CA); Lukasz Lew (Mountain View, CA)
Assignee: Google LLC
G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,323
App. No.
18/186,774
Granted
Jul 29, 2025
Kind
B2
Abstract

A method for training a model includes obtaining a plurality of training samples. Each respective training sample of the plurality of training samples includes a respective speech utterance and a respective textual utterance representing a transcription of the respective speech utterance. The method includes training, using quantization aware training with native integer operations, an automatic speech recognition (ASR) model on the plurality of training samples. The method also includes quantizing the trained ASR model to an integer target fixed-bit width. The quantized trained ASR model includes a plurality of weights. Each weight of the plurality of weights includes an integer with the target fixed-bit width. The method includes providing the quantized trained ASR model to a user device.

Claims (42)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a plurality of training samples, each respective training sample of the plurality of training samples comprising:

a respective speech utterance; and

a respective textual utterance representing a transcription of the respective speech utterance;

training, using quantization aware training with native integer operations, an automatic speech recognition (ASR) model on the plurality of training samples;

quantizing the trained ASR model to an integer target fixed-bit width, the quantized trained ASR model comprising a plurality of weights, each weight of the plurality of weights comprising an integer with the target fixed-bit width; and

providing the quantized trained ASR model to a user device.

2. The method of claim 1 , wherein the target fixed-bit width is four.

3. The method of claim 1 , wherein the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising an integer with the target fixed-bit width.

4. The method of claim 1 , wherein the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising an integer with a fixed bit width greater than the target fixed-bit width.

5. The method of claim 1 , wherein the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising a float value.

6. The method of claim 1 , wherein quantizing the trained ASR model comprises determining a scale factor based on an estimated max value of an axis to be quantized and the target fixed-bit width.

7. The method of claim 1 , wherein the ASR model comprises one or more multi-head attention layers.

8. The method of claim 7 , wherein the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers.

9. The method of claim 1 , wherein:

the ASR model comprises a plurality of encoders and a plurality of decoders; and

quantizing the ASR model comprises quantizing the plurality of encoders and not quantizing the plurality of decoders.

10. The method of claim 1 , wherein:

the ASR model comprises an audio encoder; and

the audio encoder comprises a cascaded encoder comprising a first causal encoder and a second non-causal encoder.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining a plurality of training samples, each respective training sample of the plurality of training samples comprising:

a respective speech utterance; and

a respective textual utterance representing a transcription of the respective speech utterance;

training, using quantization aware training with native integer operations, an automatic speech recognition (ASR) model on the plurality of training samples;

quantizing the trained ASR model to an integer target fixed-bit width, the quantized trained ASR model comprising a plurality of weights, each weight of the plurality of weights comprising an integer with the target fixed-bit width; and

providing the quantized trained ASR model to a user device.

12. The system of claim 11 , wherein the target fixed-bit width is four.

13. The system of claim 11 , wherein the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising an integer with the target fixed-bit width.

14. The system of claim 11 , wherein the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising an integer with a fixed bit width greater than the target fixed-bit width.

15. The system of claim 11 , wherein the ASR model further comprises a plurality of activations, each activation of the plurality of activations comprising a float value.

16. The system of claim 11 , wherein quantizing the trained ASR model comprises determining a scale factor based on an estimated max value of an axis to be quantized and the target fixed-bit width.

17. The system of claim 11 , wherein the ASR model comprises one or more multi-head attention layers.

18. The system of claim 17 , wherein the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers.

19. The system of claim 11 , wherein:

the ASR model comprises a plurality of encoders and a plurality of decoders; and

quantizing the ASR model comprises quantizing the plurality of encoders and not quantizing the plurality of decoders.

20. The system of claim 11 , wherein:

the ASR model comprises an audio encoder; and

the audio encoder comprises a cascaded encoder comprising a first causal encoder and a second non-causal encoder.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2023
From: DING, SHAOJIN; RYBAKOV, OLEG; MEADOWLARK, PHOENIX; AGRAWAL, SHIVANI; HE, YANZHANG; LEW, LUKASZ
To: GOOGLE LLC
Reel/Frame 064135/0146 →
Continuity (2)
Provisional Application 63269705 · Mar 21, 2022
Related Publication 20230298569A1 · Sep 21, 2023
References Cited (8)
US 11132988B1 · Steedman Henderson · 2021 [cited by examiner]
US 20200380215A1 · Kannan · 2020 [cited by examiner]
WO 2021258752A1 · 2021 [cited by applicant]
International Search Report and Written Opinion (EPO) for Application No. PCT/US2023/015695 dated May 23, 2023. [cited by applicant]
Andrea Fasoli et al: “4-bit Quantization of LSTM-based Speech Recognition Models”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 27, 2021 (Aug. 27, 2021), XP091040290. [cited by applicant]
Shaojin Ding et al: “4-bit Conformer with Native Quantization Aware Training for Speech Recognition”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 29, 2022 (Mar. 29, … [cited by applicant]
Anmol Gulati et al: “Conformer: Convolution-augmented Transformer for Speech Recognition”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, May 16, 2020 (May 16, 2020), XP0816… [cited by applicant]
Niccolvo Nicodemo et al: “Memory Requirement Reduction of Deep Neural Networks Using Low-bit Quantization of Parameters”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov.… [cited by applicant]
Cited By (1)
US 12,586,579