IP Library › Granted Patent US 12,387,718
Granted Patent B1
US 12,387,718 · App. 18/311,849 · Granted Aug 12, 2025

Removing bias from automatic speech recognition models using internal language model estimates

Inventors: Nilaksh Das (Seattle, WA); Monica Lakshmi Sunkara (San Jose, CA); Sravan Babu Bodapati (Fremont, CA); Jinglun Cai (Seattle, WA); Devang Kulshreshtha (Montreal, CA); Jeffrey John Farris (Crystal Lake, IL); Nicholas G Aldridge (Seattle, WA); Srikanth Ronanki (San Jose, CA); Katrin Kirchhoff (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/16G10L2015/081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,718
App. No.
18/311,849
Granted
Aug 12, 2025
Kind
B1
Abstract

Bias may be removed from automatic speech recognition model predictions using internal language model estimates. Audio data may be received for speech recognition. The audio data may be processed both through an automatic speech recognition model to produce original word token predictions and masked in different portions of the audio data to produce other word token predictions for the masked audio. A comparison of the original word token predictions and the other word token predictions may provide an estimate of an internal language model for the automatic speech recognition model. This estimate can be used to modify the original word token predictions to remove the lexical bias and produce a speech prediction.

Claims (41)

1. A system, comprising:

at least one processor; and

a memory, storing program instructions that when executed by the at least one processor, cause the at least one processor to implement a speech recognition system, configured to:

receive audio data to recognize speech in the audio data;

cause the audio data to be processed through a machine learning model trained for speech recognition that outputs original token predictions corresponding to respective words recognized in the audio data by the machine learning model;

generate a plurality of masked versions of the audio data, wherein individual ones of plurality of masked version of the audio data comprise different respective masked portions of the audio data;

cause the plurality of masked versions of the audio data to be processed through the machine learning model to output internal language model estimation token predictions corresponding to respective words recognized in the plurality of masked versions of the audio data by the machine learning model;

compare the internal language model estimation token predictions with the original token predictions to determine modifications to one or more of the original token predictions according to differences between the internal language model estimation token predictions with the original token predictions that are above difference threshold;

apply the modifications to the one or more original token predictions; and

generate a speech prediction in the audio data according to the original token predictions including the modified one or more original token predictions.

2. The system of claim 1 , wherein to generate the speech prediction in the audio data according to the original token predictions including the modified one or more original token predictions, the speech recognition system is configured to apply a beam search technique using the original token predictions fused with an external machine learning language model.

3. The system of claim 1 , wherein the different respective masked portions correspond to different ones of equal partitions of the audio data.

4. The system of claim 1 , wherein the speech recognition system is implemented as part of a natural language processing service offered by a provider network and wherein the speech prediction is provided to one or more other resources of the natural language processing service to perform a natural language processing task on the speech prediction.

5. A method, comprising:

receiving, at a speech recognition system, audio data to recognize speech in the audio data;

processing, by the speech recognition system, the audio data through a machine learning model trained for speech recognition that outputs original token predictions corresponding to respective words recognized in the audio data by the machine learning model;

generating, by the speech recognition system, a plurality of masked versions of the audio data, wherein individual ones of plurality of masked version of the audio data comprise different respective masked portions of the audio data;

processing, by the speech recognition system, the plurality of masked versions of the audio data through the machine learning model to output internal language model estimation token predictions corresponding to respective words recognized in the plurality of masked versions of the audio data by the machine learning model;

comparing, by the speech recognition system, the internal language model estimation token predictions with the original token predictions to apply modifications to one or more of the original token predictions according to differences between the internal language model estimation token predictions with the original token predictions that are above difference threshold; and

generating, by the speech recognition system, a speech prediction in the audio data according to the original token predictions including the modified one or more original token predictions.

6. The method of claim 5 , wherein the machine learning model is a non-autoregressive model.

7. The method of claim 5 , wherein generating the speech prediction in the audio data according to the original token predictions including the modified one or more original token predictions comprises applying a beam search technique using the original token predictions fused with an external machine learning language model.

8. The method of claim 5 , wherein generating the plurality of masked versions of the audio data is performed according to a masking scheme to select the different respective masked portions, and wherein the masking scheme is specified according to a request received via an interface of the speech recognition system.

9. The method of claim 5 , wherein the different respective masked portions correspond to different ones of equal partitions of the audio data.

10. The method of claim 5 , wherein the machine learning model is a deep neural network trained according to Connectionist Temporal Classification (CTC) technique, and wherein a blank token of the original token predictions output by the machine learning model is not modified according to the comparing.

11. The method of claim 5 , wherein the difference threshold is specified according to a request received via an interface of the speech recognition system.

12. The method of claim 5 , wherein the speech prediction is one of a plurality of possible speech predictions generated using the original token predictions including the modified one or more original token predictions.

13. The method of claim 5 , wherein the speech recognition system is implemented as part of a natural language processing service offered by a provider network and wherein the speech prediction is provided to one or more other resources of the natural language processing service to perform a natural language processing task on the speech prediction.

14. One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices cause the one or more computing devices to implement:

receiving audio data to recognize speech in the audio data;

causing the audio data to be processed through a machine learning model trained for speech recognition that outputs original token predictions corresponding to respective words recognized in the audio data by the machine learning model;

generating a plurality of masked versions of the audio data, wherein individual ones of plurality of masked version of the audio data comprise different respective masked portions of the audio data;

causing the plurality of masked versions of the audio data to be processed through the machine learning model to output internal language model estimation token predictions corresponding to respective words recognized in the plurality of masked versions of the audio data by the machine learning model;

comparing the internal language model estimation token predictions with the original token predictions to apply modifications to one or more of the original token predictions according to differences between the internal language model estimation token predictions with the original token predictions that are above difference threshold; and

generating a speech prediction in the audio data according to the original token predictions including the modified one or more original token predictions.

15. The one or more non-transitory, computer-readable storage media of claim 14 , wherein the machine learning model is an autoregressive model.

16. The one or more non-transitory, computer-readable storage media of claim 14 , wherein, in generating the speech prediction in the audio data according to the original token predictions including the modified one or more original token predictions, the program instructions cause the one or more computing devices to implement applying a beam search technique using the original token predictions.

17. The one or more non-transitory, computer-readable storage media of claim 14 , wherein generating the plurality of masked versions of the audio data is performed according to a masking scheme to select the different respective masked portions, and wherein the masking scheme is specified according to a request received via an interface of a speech recognition system.

18. The one or more non-transitory, computer-readable storage media of claim 14 , wherein the different respective masked portions correspond to different words in the audio data.

19. The one or more non-transitory, computer-readable storage media of claim 14 , wherein the difference threshold is specified according to a request received via an interface of the speech recognition system.

20. The one or more non-transitory, computer-readable storage media of claim 14 , wherein the one or more computing devices are implemented as part of a natural language processing service offered by a provider network and wherein the speech prediction is provided to one or more other resources of the natural language processing service to perform a natural language processing task on the speech prediction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2024
From: DAS, NILAKSH; SUNKARA, MONICA LAKSHMI; BODAPATI, SRAVAN BABU; CAI, JINGLUN; KULSHRESHTHA, DEVANG; FARRIS, JEFFREY JOHN; ALDRIDGE, NICHOLAS G; RONANKI, SRIKANTH; KIRCHHOFF, KATRIN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 066953/0195 →
References Cited (36)
US 12045568B1 · Shrivastava · 2024 [cited by examiner]
US 12254864B1 · Lajszczak · 2025 [cited by examiner]
US 20190228763A1 · Czarnowski · 2019 [cited by examiner]
US 20220139380A1 · Meng · 2022 [cited by examiner]
US 20220310062A1 · Sainath et al. · 2022 [cited by applicant]
US 20230103722A1 · Rosenberg · 2023 [cited by examiner]
US 20230186898A1 · Weisz · 2023 [cited by examiner]
US 20240203399A1 · Stooke · 2024 [cited by examiner]
US 20240203406A1 · Khorram · 2024 [cited by examiner]
US 20240290321A1 · Wang · 2024 [cited by examiner]
Higuchi, Yosuke, et al. “Mask CTC: Non-autoregressive end-to-end ASR with CTC and mask predict.” arXiv preprint arXiv: 2005.08700 (2020). (Year: 2020). [cited by examiner]
Higuchi, Yosuke, et al. “Improved Mask-CTC for non-autoregressive end-to-end ASR.” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021. (Year: 2021). [cited by examiner]
S. Toshniwal, A. Kannan, C.-C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A comparison of techniques for language model integration in encoder-decoder speech recognition,” in IEEE Spoken Language Technology Workshop (… [cited by applicant]
Liu, Y. Gu, A. Gourav, A. Gandhe, S. Kalmane, D. Filimonov, A. Rastrow, and I. Bulyko, “Domain-aware neural language models for speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Proce… [cited by applicant]
E. McDermott, H. Sak, and E. Variani, “A density ratio approach to language model fusion in end-to-end automatic speech recognition,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, retriev… [cited by applicant]
C. Choudhury, A. Gandhe, X. Ding, and I. Bulyko, “A likelihood ratio based domain adaptation method for e2e models,” In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 1-5. [cited by applicant]
W. Zhou, Z. Zheng, R. Schl{umlaut over ( )}uter, and H. Ney, “On language model integration for RNN transducer based speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICAS… [cited by applicant]
Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal language model estimation for domain-adaptive end-to-end speech recognition,” in IEEE Spoken Language Technolo… [cited by applicant]
M. Zeineldeen, A. Glushko, W. Michel, A. Zeyer, R. Schluter, and H. Ney, “Investigating methods to improve language model integration for attention-based encoder-decoder ASR models,” arXiv preprint arXiv:2104.05544, 202… [cited by applicant]
Y. Liu, R. Ma, H. Xu, Y. He, Z. Ma, and W. Zhang, “Internal language model estimation through explicit context vector earning for attention-based encoder-decoder ASR,” arXiv preprint arXiv:2201.11627, 2022, pp. 1-5. [cited by applicant]
J. Lee and S. Watanabe, “Intermediate loss regularization for CTC-based speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, retrieved from arXiv:2102.03216, p… [cited by applicant]
R. Fan, W. Chu, P. Chang, and J. Xiao, “CASS-NAT: CTC alignment-based single step non-autoregressive transformer for speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICAS… [cited by applicant]
X. Zhang, F. Zhang, C. Liu, K. Schubert, J. Chan, P. Prakash, J. Liu, C.-F. Yeh, F. Peng, Y. Saraf et al., “Benchmarking LFMMI, CTC and RNN-T criteria for streaming ASR,” in IEEE Spoken Language Technology Workshop (SLT… [cited by applicant]
S. Dingliwal, M. Sunkara, S. Ronanki, J. Farris, K. Kirchhoff, and S. Bodapati, “Towards Personalization of CTC Speech Recognition Models With Contextual Adapters and Adaptive Boosting,” arXiv preprint arXiv:2210.09510,… [cited by applicant]
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates et al., “Deep speech: Scaling up end-to-end speech recognition,” arXiv preprint arXiv:1412.5567, 2014, pp… [cited by applicant]
D. Le, G. Keren, J. Chan, J. Mahadeokar, C. Fuegen, and M. L. Seltzer, “Deep shallow fusion for RNN-T personalization,” in IEEE Spoken Language Technology Workshop (SLT), 2021, retrieved from arXiv:2011.07754, pp. 1-7. [cited by applicant]
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid Autoregressive Transducer (HAT),” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, retrieved from arXiv:2003.07705, pp… [cited by applicant]
A. Zeyer, R. Schl{umlaut over ( )}uter, and H. Ney, “Why does CTC result in peaky behavior?” arXiv preprint arXiv:2105.14849, 2021, pp. 1-10. [cited by applicant]
A. Gulati, J. Qin, C. C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proceedings of the Annual Conference of… [cited by applicant]
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing, 2017, pp. 2623-2627. [cited by applicant]
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Proceedings of th… [cited by applicant]
K. Heafield, “Kenlm: Faster and smaller language model queries,” in Proceedings of the Sixth Workshop on Statistical Machine Translation, 2011, pp. 187-197. [cited by applicant]
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in IEEE International Conference on Acoustics, Speech and Signal Processing, 2015, pp. 1-5. [cited by applicant]
C. Wang, M. Rivi{grave over ( )}ere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A largescale multilingual speech corpus for representation learning, semi-supervised learnin… [cited by applicant]
N. Das, M. Sunkara, D. Bekal, D. H. Chau, S. Bodapati, and K. Kirchhoff, “Listen, know and spell: Knowledge-infused subword modeling for improving ASR performance of OOV named entities,” in IEEE International Conference… [cited by applicant]
J. S. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete LDC93S6A,” Linguistic Data Consortium, 1993. [Online]. Available: https://catalog.ldc.upenn.edu/LDC93S6A, pp. 1-2. [cited by applicant]