IP Library › Granted Patent US 12,505,317
Granted Patent B2
US 12,505,317 · App. 18/737,621 · Granted Dec 23, 2025

Multi-task automatic speech recognition system

Inventors: Alec Radford (San Francisco, CA); Jong Wook Kim (San Francisco, CA); Tao Xu (San Francisco, CA); Greg Brockman (San Francisco, CA); Christine McLeavey-Payne (San Francisco, CA); Ilya Sutskever (San Francisco, CA)
Assignee: OpenAI Opco, LLC
G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,317
App. No.
18/737,621
Granted
Dec 23, 2025
Kind
B2
Abstract

Disclosed herein are methods, systems, and computer-readable media for generating an output transcript from an input audio segment using a multi-task transformer model. In some embodiments, the transformer model can be trained to transcribe or translate audio data in multiple languages using labeled audio data. The labeled audio data can include first audio segments associated with first same-language transcripts of the first audio segments and second audio segments associated with second different-language transcripts of the second audio segments. In some embodiments, a vocabulary of the model can include special purpose and time stamp tokens. The special purpose tokens can specify tasks for the model to perform.

Claims (45)

1 . A system comprising:

at least one memory storing instructions; and

at least one processor configured to execute the instructions to perform operations for multi-language, multi-task speech recognition, the operations comprising:

obtaining a transformer model including an encoder and a decoder, the transformer model trained to transcribe or translate audio data in multiple languages using labeled audio data, the labeled audio data including first audio segments associated with first same-language transcripts of the first audio segments and second audio segments associated with second different-language transcripts of the second audio segments; and

generating an output transcript from an input audio segment using the transformer model, generation including:

configuring a decoder input with an autoregressively generated token that specifies a task, an autoregressively generated token that specifies an input language, or an autoregressively generated token that specifies an absence of speech.

2 . The system of claim 1 , wherein configuring the decoder input comprises additionally configuring the decoder input with an assigned transcribe token or an assigned translate token.

3 . The system of claim 1 , wherein configuring the decoder input comprises configuring the decoder input with an autoregressively generated input language token corresponding to the first language and an assigned translate token; and

generating the output transcript further includes:

autoregressively configuring the decoder input with a textual token predicted using the assigned translate token and the autoregressively generated input language token, the textual token associated with the second language differing from the first language.

4 . The system of claim 1 , wherein the decoder is configured with a vocabulary including timestamp tokens.

5 . The system of claim 1 , wherein generating the output transcript includes:

autoregressively configuring the decoder input with a token that specifies an absence of speech, followed by an end of transcript token.

6 . The system of claim 1 , wherein the transformer model is configured to perform inverse text normalization.

7 . The system of claim 1 , wherein generating the output transcript further includes:

applying a first subsegment of the input audio segment to generate one or more predicted timestamp tokens for the first subsegment; and

generating a second subsegment of the input audio segment using the one or more predicted timestamp tokens for the first subsegment.

8 . The system of claim 1 , wherein generating the output transcript further includes performing a beam search using an output softmax temperature dependent on at least one of:

log probabilities of previously generated tokens of the output transcript; or

a compression rate of the previously generated tokens of the output transcript.

9 . The system of claim 1 , wherein generating the output transcript further includes prepending a transcript generated for a preceding input audio segment to the decoder input.

10 . A method for multi-language, multi-task speech recognition, comprising:

obtaining a transformer model including an encoder and a decoder, the transformer model trained to transcribe or translate audio data in multiple languages using labeled audio data, the labeled audio data including first audio segments associated with first same-language transcripts of the first audio segments and second audio segments associated with second different-language transcripts of the second audio segments; and

generating an output transcript from an input audio segment using the transformer model, generation including:

configuring a decoder input with an autoregressively generated token that specifies a task, an autoregressively generated token that specifies an input language, or an autoregressively generated token that specifies an absence of speech.

11 . The method of claim 10 , wherein configuring the decoder input further comprises configuring the decoder input with an assigned transcribe token or an assigned translate token.

12 . The method of claim 10 , wherein configuring the decoder input comprises configuring the decoder input with an autoregressively generated input language token corresponding to the first language and an assigned translate token; and

generating the output transcript further includes:

autoregressively configuring the decoder input with a textual token predicted using the assigned translate token and the autoregressively generated input language token, the textual token associated with the second language differing from the first language.

13 . The method of claim 10 , wherein the decoder is configured with a vocabulary including timestamp tokens.

14 . The method of claim 10 , wherein generating the output transcript includes:

autoregressively configuring the decoder input with a token that specifies an absence of speech, followed by an end of transcript token.

15 . The method of claim 10 , wherein the transformer model is configured to perform inverse text normalization.

16 . The method of claim 10 , wherein generating the output transcript further includes:

applying a first subsegment of the input audio segment to generate one or more predicted timestamp tokens for the first subsegment; and

generating a second subsegment of the input audio segment using the one or more predicted timestamp tokens for the first subsegment.

17 . The method of claim 10 , wherein generating the output transcript further includes:

performing a beam search using an output softmax temperature dependent on at least one of:

log probabilities of previously generated tokens of the output transcript; or

a compression rate of the previously generated tokens of the output transcript; or

prepending a transcript generated for a preceding input audio segment to the decoder input.

18 . A non-transitory computer-readable medium including instructions that are executable by one or more processors to perform operations comprising:

obtaining a transformer model including an encoder and a decoder, the transformer model trained to transcribe or translate audio data in multiple languages using labeled audio data, the labeled audio data including first audio segments associated with first same-language transcripts of the first audio segments and second audio segments associated with second different-language transcripts of the second audio segments; and

generating an output transcript from an input audio segment using the transformer model, generation including:

configuring a decoder input an autoregressively generated token that specifies a task, an autoregressively generated token that specifies an input language, or an autoregressively generated token that specifies an absence of speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2024
From: RADFORD, ALEC; KIM, JONG WOOK; XU, TAO; BROCKMAN, GREG; MCLEAVEY-PAYNE, CHRISTINE; SUTSKEVER, ILYA
To: OPENAI OPCO LLC
Reel/Frame 067661/0090 →
Continuity (2)
Continuation 18302289 · Apr 18, 2023
Related Publication 20240354521A1 · Oct 24, 2024
References Cited (148)
US 7630892B2 · Wu · 2009 [cited by examiner]
US 10146789B1 · Lakshmanan et al. · 2018 [cited by applicant]
US 10747962B1 · Fuerstenau et al. · 2020 [cited by applicant]
US 10949230B2 · Hawkins et al. · 2021 [cited by applicant]
US 11107462B1 · Fuegen et al. · 2021 [cited by applicant]
US 11315569B1 · Talieh et al. · 2022 [cited by applicant]
US 11367433B2 · Sypniewski et al. · 2022 [cited by applicant]
US 11482214B1 · Arieli et al. · 2022 [cited by applicant]
US 11545134B1 · Federico et al. · 2023 [cited by applicant]
US 11741298B1 · Polavaram · 2023 [cited by examiner]
US 20090157385A1 · Tian · 2009 [cited by applicant]
US 20090193013A1 · Hazlewood et al. · 2009 [cited by applicant]
US 20100318356A1 · Hamaker et al. · 2010 [cited by applicant]
US 20140067374A1 · Wilkins et al. · 2014 [cited by applicant]
US 20150067459A1 · Lester · 2015 [cited by applicant]
US 20150154185A1 · Waibel · 2015 [cited by examiner]
US 20180174576A1 · Soltau et al. · 2018 [cited by applicant]
US 20200211530A1 · Zass et al. · 2020 [cited by applicant]
US 20200213680A1 · Ingel · 2020 [cited by examiner]
US 20200226327A1 · Matusov et al. · 2020 [cited by applicant]
US 20200243094A1 · Thomson et al. · 2020 [cited by applicant]
US 20200410045A1 · Vozila et al. · 2020 [cited by applicant]
US 20210056162A1 · Dehghani et al. · 2021 [cited by applicant]
US 20210056956A1 · Dimitriadis et al. · 2021 [cited by applicant]
US 20210165973A1 · Kofman et al. · 2021 [cited by applicant]
US 20210192140A1 · Galley et al. · 2021 [cited by applicant]
US 20210312944A1 · Yamada et al. · 2021 [cited by applicant]
US 20210334299A1 · Sonntag et al. · 2021 [cited by applicant]
US 20210375291A1 · Zeng et al. · 2021 [cited by applicant]
US 20220084273A1 · Pan et al. · 2022 [cited by applicant]
US 20220108688A1 · Wang et al. · 2022 [cited by applicant]
US 20220115006A1 · Hori et al. · 2022 [cited by applicant]
US 20220122586A1 · Yu · 2022 [cited by examiner]
US 20220254331A1 · Jafari · 2022 [cited by applicant]
US 20220301543A1 · Elias et al. · 2022 [cited by applicant]
US 20220343894A1 · Doutre et al. · 2022 [cited by applicant]
US 20220343898A1 · Fu · 2022 [cited by applicant]
US 20220366157A1 · Lee et al. · 2022 [cited by applicant]
US 20220366898A1 · Qian et al. · 2022 [cited by applicant]
US 20220382999A1 · Gunasekara et al. · 2022 [cited by applicant]
US 20230089902A1 · Arkhangorodsky · 2023 [cited by examiner]
US 20230094511A1 · Guha et al. · 2023 [cited by applicant]
US 20230104228A1 · Li et al. · 2023 [cited by applicant]
US 20230114834A1 · Chadwick et al. · 2023 [cited by applicant]
US 20230154172A1 · Wasnik et al. · 2023 [cited by applicant]
US 20230169281A1 · Zheng et al. · 2023 [cited by applicant]
US 20230177282A1 · Emelyanenko · 2023 [cited by examiner]
US 20230237990A1 · Wu et al. · 2023 [cited by applicant]
US 20230245649A1 · Singh et al. · 2023 [cited by applicant]
US 20230260531A1 · Srivastava et al. · 2023 [cited by applicant]
US 20230274096A1 · Bohra et al. · 2023 [cited by applicant]
US 20230281248A1 · Schalkwyk et al. · 2023 [cited by applicant]
US 20230289536A1 · Guar et al. · 2023 [cited by applicant]
US 20230289538A1 · Goel et al. · 2023 [cited by applicant]
US 20240256793A1 · Maschmeyer · 2024 [cited by examiner]
US 20240256795A1 · Ross · 2024 [cited by examiner]
US 20240320444A1 · Maschmeyer · 2024 [cited by examiner]
US 20240330579A1 · Saxena · 2024 [cited by examiner]
CA 3089001A1 · 2021 [cited by examiner]
CA 3239059A1 · 2024 [cited by examiner]
CN 111081259B · 2022 [cited by applicant]
CN 115101050A · 2022 [cited by applicant]
Le, Hang, et al. “Dual-decoder transformer for joint automatic speech recognition and multilingual speech translation.” arXiv preprint arXiv:2011.00747 (2020). (Year: 2020). [cited by examiner]
Woodward, Alejandro, et al. “Confidence Measures in Encoder-Decoder Models for Speech Recognition.” Interspeech. 2020. ( Year: 2020). [cited by examiner]
Alcorn, M. A. et al., “Strike With a (Pose): Neural Networks are Easily Fooled by Strange Poses of Familiar Objects”, 2019 IEEE/CVE Conference on Computer Vision and Patter Recognition (CVPR), IEEE Computer Society, (20… [cited by applicant]
Amodei, D. et al., “Deep Speech 2: End-to-End Speech Recognition in English and Mandarin”, arxiv. arXivpreprint arXiv: 1512.02595, (2015), 28 pages. [cited by applicant]
Ardila, R. et al., “Common Voice: A Massively-Multilingual Speech Corpus”, arXiv preprint arXiv: 1912.06670v2, (2020), 5 pages. [cited by applicant]
Babu, A. et al., “XLS-R: Self-Supervised Cross-Lingual Speed Representation Learning to Scale”, arXiv preprint arXiv:2111.09296v3, (2021), 23 pages. [cited by applicant]
Baevski, A. et al., “Unsupervised Speech Recognition”, Advances in Neural Information Processing Systems, 35 [cited by applicant]
Baevski, A. et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”, arXiv preprint arXiv:2006.11477v3, (2020), 19 pages. [cited by applicant]
Bapna, A. et al., “mSLAM: Massively Multilingual joint pre-training for speech and text”, arXiv preprint arXiv:2206.01374v1, (2022), 15 pages. [cited by applicant]
Barbu, A. et al., “ObjectNet: A Large-Scale Bias-Controlled Dataset for Pushing the Limits of Object Recognition Models”, Advances in Neural Information Processing Systems, vol. 32, (2019), 11 pages. [cited by applicant]
Caruana, R., “Multitask Learning”, Machine Learning, Kluwer Academic Publishers, vol. 28, No. 1, (1997), pp. 41-75. [cited by applicant]
Chan, W. et al., “SpeechStew: Simply Mix All Available Speech Recognition data to Train One Large Neural Network”, arXiv preprint arXiv:2104.02133v3, (2021), 5 pages. [cited by applicant]
Chen, S. et al., “UNISPEECH-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training”, In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Procesing (ICASSP), IEEE, (20… [cited by applicant]
Chen, T. et al., “Training Deep Nets with Sublinear Memory Cost”, arXiv preprint arXiv:1604.06174v2, (2016), 12 pages. [cited by applicant]
Chen, Z. et al., “MAESTRO: Matched Speech Text Representations Through Modality Matching”, arXiv preprint arXiv:2204.03409v2, (2022), 5 pages. [cited by applicant]
Child, R. et al., “Generating Long Sequences with Sparse Transformers”, arXiv preprint arXiv:1904.10509v1, (2019), 10 pages. [cited by applicant]
Collobert, R. et al., “Natural Language Processing (Almost) From Scratch”, Journal of Machine Learning Research, vol. 12, (2011), pp. 2493-2537. [cited by applicant]
Conneau, A. et al., “FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech”, arXiv preprint arXiv:2205.12446v1, (2022), 10 pages. [cited by applicant]
Del Rio, M. et al., “Earnings-21: A Practical Benchmark for ASR in the Wild”, arXiv preprint arXiv:2104.11348v3, (2021), 5 pages. [cited by applicant]
Galvez, D. et al., “The People's Speech: A Large-Scale Diverse English Speech Recognition Datasheet for Commercial Usage”, arXiv preprint arXiv:2111.09344v1, (2021), 12 pages. [cited by applicant]
Geirhos, R. et al., “Shortcut Learning in Deep Neural Networks”, Nature Machine Intelligence, vol. 2, (2020), pp. 665-673. [cited by applicant]
Ghorbani, B. et al., “Scaling Laws for Neural Machine Translation”, arXiv preprint arXiv:2109.07740v1, (2021), 31 pages. [cited by applicant]
Griewank, A. et al., “Algorithm 799: Revolve: An Implementation of Checkpointing for the Reverse or Adjoint Mode of Computational Differentiation”, ACM Transactions of Mathematical Software (TOMS), vol. 26, No. 1, (Mar.… [cited by applicant]
Gunter, K. et al., “Contextualizing /s/ retraction: Sibilant variation and change in Washington, D.C. African American Language”, Language Variation and Change, vol. 33, (2021), pp. 331-357. [cited by applicant]
Harris, C. R. et al., “Array Programming with NumPy”, Nature, vol. 585, (2020), pp. 357-362. [cited by applicant]
Hendrycks, D. et al., “Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units”, arXiv preprint arXiv:1606.08415v1, (2016), 6 pages. [cited by applicant]
Hendrycks, D. et al., “Pretrained Transformers Improve Out-of-Distribution Robustness”, arXiv preprint arXiv:2004.06100v2, (2020), 11 pages. [cited by applicant]
Hernandez, F. et al., “TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation”, In SPECOM, Springer, Nature Switzerland AG, (2018), 198-208. [cited by applicant]
Hsu, W.-N., et al., HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 29, (2021), pp. 3451-3460. [cited by applicant]
Hsu, W.-N., et al., “Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre- Training”, arXiv preprint arXiv:2104.01027v2, (2021), 9 pages. [cited by applicant]
Huang, G. et al., “Deep Networks with Stochastic Depth”, In European Conference on Computer Vision, Springer International Publishing AG, (2016), pp. 646-661. [cited by applicant]
Jia, R. et al., “Adversarial Examples for Evaluating Reading Comprehension Systems”, arXiv preprint arXiv:1707.07328v1, (2017), 11 pages. [cited by applicant]
Johnson, M. et al., Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation, Transactions of The Association for Computational Linguistics, vol. 5, (2017), pp. 339-351. [cited by applicant]
Kendall, T. et al., “The Corpus of Regional African American Language”, Version Jul. 2021, The Online Resources for African American Language Project, Eugene, Oregon, [Retrieved from the Internet—Feb. 23, 2024], agon.ed… [cited by applicant]
Koenecke, A. et al., Racial Disparities in Automated Speech Recognition, Proceedings of the National Academy of Sciences, vol. 117, No. 14, (2020), pp. 7684-7689. [cited by applicant]
Kolesnikov, A. et al., “Big Transfer (BiT): General Visual Representation Learning”, In European Conference on Computer Vision, Springer Nature Switzerland AG, (2020), pp. 491-507. [cited by applicant]
Kuchaiev, O. et al., “NeMo: a toolkit for building AI applications using Neural Modules”, arXiv preprint arXiv:1909.09577v1, (2019), 8 pages. [cited by applicant]
Lake, B. M. et al., “Building Machines that Learn and Think Like People”, Behavioral and Brain Sciences, Cambridge University Press, vol. 40, (2017), 72 pages. [cited by applicant]
Liao, H. et al., “Large Scale Deep Neural Network Acoustic Modeling with Semi-Supervised Training Data for YouTube Video Transcription”, In 2013 IEEE Workshop on Automatic Speech Recognition and Understanding, IEEE, (20… [cited by applicant]
Likhomanenko, T. et al., “Rethinking Evaluation in ASR: Are Our Models Robust Enough?”, arXiv preprint arXiv:2010.11745v3, (2020), 11 pages. [cited by applicant]
Loshchilov, I. et al., “Decoupled Weight Decay Regularization”, arXiv preprint arXiv:1711.05101v3, (2019), 19 pages. [cited by applicant]
Luong, M.-T. et al., “Multi-Task Sequence to Sequence Learning”, arXiv preprint arXi:1511.06114v4, Published as a conference paper at ICLR 2016, (2016), 10 pages. [cited by applicant]
Mahajan, D. et al., “Exploring the Limits of Weakly Supervised Pretraining”, In Proceedings of the European Conference on Computer Vision (EVVC), Springer Nature Switzerland AG, (2018), pp. 185-201. [cited by applicant]
Mauch, M. et al., “The Audio Degradation Toolbox and its Application to Robustness Evaluation”, In Proceedings of the 14th International Society for Music Information Retrieval Conference (ISMIR 2013), Curitiba, Brazil,… [cited by applicant]
McCann, B et al., “The Natural Language Decathlon: Multitask Learning as Question Answering”, arXiv preprint arXiv:1806.08730v1, (2018), 23 pages. [cited by applicant]
Meyer, J. et al., “Artie Bias Corpus: An Open Dataset For Detecting Demographic Bias in Speech Applications”, In Proceedings of the 12th Language Resources and Evaluation Conference, European Language Resources Associat… [cited by applicant]
Miller, J. et al., “The Effect of Natural Distribution Shift on Question Answering Models”, In ICML, arXiv:2004.14444v1, (2020), 73 pages. [cited by applicant]
Mohamed, A.-r. et al., “Deep Belief Networks for Phone Recognition”, In Nips Workshop on Deep Learning for Speech Recognition and Related Applications, vol. 1, (2009), 9 pages. [cited by applicant]
Narayanan, A. et al., “Toward Domain-Invariant Speech Recognition Via Large Scale Training”, In 2018 IEEE Spoken Language Technology Workshop (SLT), IEEE, (2018), pp. 441-447. [cited by applicant]
Panayotov, V. et al., “LIBRISPEECH: An ASR Corpus Based on Public Domain Audio Books”, In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, (2015), pp. 5206-5210. [cited by applicant]
Pandas Development Team, T. pandas-dev/pandas: Pandas, [Retrieved from the Internet Feb. 22, 2024], Published Apr. 3, 2023 URL https://doi.org/10.5281/zenodo.350914, 5 pages. [cited by applicant]
Park, D. S. et al., “SpecAugment: A simple data augmentation method for automatic speech recognition”, arXiv preprint arXiv:1904.08779v3, (2019), 6 pages. [cited by applicant]
Pascanu, R. et al., “On the Difficulty of Training Recurrent Neural Networks”, In International Conference on Machine Learning, PMLR, (2013), pp. 1310-1318. [cited by applicant]
Paszke, A. et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library”, In Advances in Neural Information Processing Systems, vol. 32, (2019), pp. 8024-8035. [cited by applicant]
Pedregosa, F. et al., Scikit-learn: Machine Learning in Python, Journal of Machine Learning Research, vol. 12, (2011), pp. 2825-2830. [cited by applicant]
Polyak, B. T. et al., “Acceleration of Stochastic Approximation by Averaging”, SIAM Journal on Control and Optimization, vol. 30, No. 4, (Jul. 1992), pp. 838-855. [cited by applicant]
Pratap, V. et al., “Massively multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters”, ArXiv, abs/2007.03001v2, (2020), 5 pages. [cited by applicant]
Pratap, V. et al., “Mls: A Large-Scale Multilingual Dataset for Speech Research”, arXiv preprint arXiv:2012.03411v2, (2020), 10 pages. [cited by applicant]
Press, O. et al., “Using the Output Embedding to Improve Language Models”, In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: vol. 2, Short Papers, Valencia, … [cited by applicant]
Provilko, I. et al., “BPE-Dropout: Simple and Effective Subword regularization”, arXiv preprint arXiv:1910.13267v2, (2019), 11 pages. [cited by applicant]
Radford, A. et al., “Language Models are Unsupervised Multitask Learners”, (2019), 24 pages. [cited by applicant]
Radford, A. et al., “Learning Transferable Visual Models from Natural Language Supervision”, arXiv preprint arXiv:2103.00020v1, (2021), 48 pages. [cited by applicant]
Raffel, C. et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” J. Mach. Learn. Res., vol. 21, No. 140, (2020), pp. 1-67. [cited by applicant]
Ravanelli, M. et al., “SpeechBrain: A general-purpose speech toolkit”, arXiv:2106.04624, (2021). [cited by applicant]
Recht, B. et al., “Do ImageNet Classifiers Generalize to ImageNet?”, Proceedings of the 36th International Conference on Machine Learning, vol. 97 of Proceedings of Machine Learning Research, PMLR, (Jun. 9-15, 2019), pp… [cited by applicant]
Russakovsky, O. et al., “Imagenet Large Scale Visual Recognition Challenge”, International Journal of Computer Vision, vol. 115, (2015), pp. 211-252. [cited by applicant]
Seide, F. et al., Feature Engineering in Context-Dependent Deep Neural Networks of Conversational Speech Transcription, In 2011 IEEE Workshop on Automatic Speech Recognition and Understanding, IEEE, (2011), pp. 24-29. [cited by applicant]
Sennrich, R. et al., “Neural Machine Translation or Rare Words with Subword Units”, arXiv preprint arXiv:1508.07909v1, (2015), 11 pages. [cited by applicant]
Speer, R., “ftfy: fixes text for you”. Zenodo, [Retrieved from the Internet Feb. 23, 2024] URL https://doi.org/10.5281/zenodo.2591652, Version 5.5., (2019), 3 pages. [cited by applicant]
Sutskever, I. et al., “Sequence to Sequence Learning with Neural Networks”, Advances in Neural Information Processing Systems, vol. 27, (2014), 9 pages. [cited by applicant]
Taori, R. et al., “Measuring Robustness to Natural Distribution Shifts in Image Classification”, Advances in Neural Information Processing Systems, vol. 33, pp. 18583-18599. [cited by applicant]
Torralba, A. et al., “Unbiased Look at Dataset Bias”, CVPR 2011, (2011), pp. 1521-1528. [cited by applicant]
Toshniwal, S., “Multilingual Speech Recognition with a Single End-to-End Model”, 2018 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), (2018), pp. 4904-4908. [cited by applicant]
Valk, J. et al., “Voxlingual107: A Dataset for Spoken Language Recognition”, In 2021 Spoken Language Technology Workshop (SLT), IEEE, (2021), pp. 652-658. [cited by applicant]
Vaswani, A. et al., “Attention is All You Need”, In Advances in Neural Information Processing Systems, 31st Conference on Neural Information Processing Systems, Long Beach, CA, (2017), 11 pages. [cited by applicant]
Virtanen, P. et al., “SciPy 1.0: fundamental algorithms for scientific computing in python”, Nature Methods, vol. 17, (2020), pp. 261-272. [cited by applicant]
Wang, C. et al., “CoVoST 2 and Massively Multilingual Speech-to-Text Translation”, arXiv preprint arXiv:2007.10310v3, (2020), 6 pages. [cited by applicant]
Wang, C. et al., “VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation”, arXiv preprint arXiv:2101.00390v2, (2021). 11 pages. [cited by applicant]
Wang, P. et al., “Multitask Training with Text Data for End-to-End Speech Recognition”, arXiv preprint arXiv:2010.14318v2, (2020), 5 pages. [cited by applicant]
Watanabe, S. et al., “CMiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings”, arXiv preprint arXiv:2004.09249v2, (2020), 7 pages. [cited by applicant]
Xu, Q. et al., “Self-Training and Pre-Training are Complementary for Speech Recognition”, In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, (2021), pp. 3030-303… [cited by applicant]
Zhang, Y. et al., “Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition”, arXiv preprint arXiv:2010.10504v21, (2020), 11 pages. [cited by applicant]
Zhang, Y. et al., “BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition”, arXiv preprint arXiv:2109.13226v3, (2021), 14 pages. [cited by applicant]
Chen, Guoguo, et al. “Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio.” arXiv preprint arXiv:2106.06909 (2021). (Year: 2021). [cited by applicant]
Chorowski, Jan, and Navdeep Jaitly. “Towards better decoding and language model integration in sequence to sequence models.” arXiv preprint arXiv:1612.02695 (2016). (Year: 2016). [cited by applicant]
Wicks, Rachel, and Kevin Duh. “The Effects of Language Token Prefixing for Multilingual Machine Translation.” Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistic… [cited by applicant]