IP Library Granted Patent US 12,223,948
Granted Patent B2
US 12,223,948 · App. 17/649,810 · Granted Feb 11, 2025

Token confidence scores for automatic speech recognition

Inventors: Pranav Singh (Sunnyvale, CA); Saraswati Mishra (San Jose, CA); Eunjee Na (Busan, KR)
Assignee: SoundHound, Inc.
G10L15/1815G10L15/02G10L15/26G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,948
App. No.
17/649,810
Granted
Feb 11, 2025
Kind
B2
Abstract

Methods and systems for correction of a likely erroneous word in a speech transcription are disclosed. By evaluating token confidence scores of individual words or phrases, the automatic speech recognition system can replace a low-confidence score word with a substitute word or phrase. Among various approaches, neural network models can be used to generate individual confidence scores. Such word substitution can enable the speech recognition system to automatically detect and correct likely errors in transcription. Furthermore, the system can indicate the token confidence scores on a graphic user interface for labeling and dictionary enhancement.

Claims (50)

1. A computer-implemented method for speech recognition, comprising:

receiving, at an automatic speech recognition system, an utterance;

generating, based on an acoustic model, a phoneme sequence of the utterance;

generating, based on the acoustic model, a plurality of phoneme sequences of the utterance that comprise the phoneme sequence;

assigning sentence-level acoustic scores to individual phoneme sequences to indicate the likelihood of correctness to represent the utterance;

segmenting the phoneme sequence into a token sequence that represents the phoneme sequence based on a pronunciation dictionary;

assigning respective token confidence scores to individual tokens in the token sequence, wherein a token confidence score represents a level of confidence of the correct representation of a tokenized word;

determining that a first confidence score associated with a token is lower than a predetermined threshold;

determining a substitute token associated with a second confidence score that is higher than the first confidence score; and

updating the token sequence by replacing the token with the substitute token.

2. The computer-implemented method of claim 1 , following updating the token sequence, further comprising:

reassigning the respective sentence-level acoustic scores to individual phoneme sequences; and

determining a phoneme hypothesis from the plurality of phoneme sequences to represent the utterance.

3. The computer-implemented method of claim 1 , wherein a confidence score model assigns the respective token confidence scores to individual tokens in the token sequence.

4. The computer-implemented method of claim 1 , wherein a translation model determines the substitute token and update the token sequence by replacing the token with the substitute token.

5. The computer-implemented method of claim 1 , wherein the token confidence scores are based on one or more of a token sequence probability analysis, acoustic probability analysis, and semantic analysis.

6. The computer-implemented method of claim 1 , further comprising:

generating, based on the phoneme sequence, a text transcription of the utterance.

7. A computer-implemented method for speech recognition, comprising:

generating, based on an acoustic model, a plurality of phoneme sequences of an utterance;

assigning sentence-level acoustic scores to individual phoneme sequences to indicate the likelihood of correctness to represent the utterance;

receiving, at an acoustic model of an automatic speech recognition system, a plurality of tokens and phrases representing the utterance;

assigning, by a confidence score model, respective token confidence scores and phrase confidence scores to the plurality of tokens and phrases, wherein a token confidence score and phrase confidence score represent a level of confidence of the correct representation of a word or phrase;

determining, by a translation model, that a first confidence score associated with a phrase is lower than a predetermined threshold;

determining a substitute phrase associated with a second confidence score that is higher than the first confidence score; and

replacing the phrase with the substitute phrase to generate an updated phoneme sequence.

8. The computer-implemented method of claim 7 , wherein the translation model determines a substitute token and update the token sequence by replacing the token with the substitute token.

9. The computer-implemented method of claim 7 , further comprising:

receiving the utterance of a query sentence;

generating a phoneme sequence of the utterance; and

segmenting the phoneme sequence into a plurality of tokens and phrases based on a pronunciation dictionary, wherein a phrase comprises one or more tokens.

10. The computer-implemented method of claim 7 , following replacing the phrase with the substitute phrase, further comprising:

reassigning the respective sentence-level acoustic scores to individual phoneme sequences; and

determining a phoneme hypothesis from the plurality of phoneme sequences to represent the utterance.

11. A computer-implemented method for speech recognition, comprising:

receiving, at an automatic speech recognition system, an utterance;

generating, based on an acoustic model, a phoneme sequence of the utterance;

generating, based on the acoustic model, a plurality of phoneme sequences of the utterance that comprise the phoneme sequence;

assigning sentence-level acoustic scores to individual phoneme sequences to indicate the likelihood of correctness to represent the utterance;

segmenting the phoneme sequence into a token sequence that represents the phoneme sequence based on a pronunciation dictionary;

assigning respective token confidence scores to individual tokens in the token sequence, wherein a token confidence score represents a level of confidence of the correct representation of a tokenized word;

determining that a first confidence score associated with a token is lower than a predetermined threshold;

determining a substitute token associated with a second confidence score that is higher than the first confidence score;

updating the token sequence by replacing the token with the substitute token; and

determining a phoneme hypothesis from the plurality of phoneme sequences to represent the utterance.

12. The computer-implemented method of claim 11 , wherein a confidence score model assigns the respective token confidence scores to individual tokens in the token sequence.

13. The computer-implemented method of claim 11 , wherein a translation model determines the substitute token and update the token sequence by replacing the token with the substitute token.

14. The computer-implemented method of claim 11 , wherein the token confidence scores are based on one or more of a token sequence probability analysis, acoustic probability analysis, and semantic analysis.

15. The computer-implemented method of claim 11 , further comprising:

generating, based on the phoneme sequence, a text transcription of the utterance.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2022
From: SINGH, PRANAV; MISHRA, SARASWATI; NA, EUNJEE
To: SOUNDHOUND, INC.
Reel/Frame 059041/0156 →
Continuity (1)
Related Publication 20230245649A1 · Aug 3, 2023
References Cited (91)
US 8812316B1 · Chen · 2014 [cited by examiner]
US 10452782B1 · Kumar et al. · 2019 [cited by applicant]
US 10515155B2 · Bachrach et al. · 2019 [cited by applicant]
US 11580959B2 · Freed · 2023 [cited by examiner]
US 20020082833A1 · Marasek · 2002 [cited by examiner]
US 20020133340A1 · Basson · 2002 [cited by examiner]
US 20060129381A1 · Wakita · 2006 [cited by applicant]
US 20060195318A1 · Stanglmayr · 2006 [cited by examiner]
US 20070033026A1 · Bartosik · 2007 [cited by examiner]
US 20070208567A1 · Amento · 2007 [cited by examiner]
US 20100106505A1 · Shu · 2010 [cited by examiner]
US 20110218802A1 · Bouganim · 2011 [cited by examiner]
US 20130018649A1 · Deshmukh et al. · 2013 [cited by applicant]
US 20150179169A1 · John · 2015 [cited by examiner]
US 20160155436A1 · Choi · 2016 [cited by examiner]
US 20160180242A1 · Byron et al. · 2016 [cited by applicant]
US 20160196820A1 · Williams et al. · 2016 [cited by applicant]
US 20170200458A1 · Kang · 2017 [cited by examiner]
US 20170286869A1 · Zarosim et al. · 2017 [cited by applicant]
US 20180061408A1 · Andreas et al. · 2018 [cited by applicant]
US 20180068653A1 · Trawick · 2018 [cited by examiner]
US 20180121419A1 · Lee et al. · 2018 [cited by applicant]
US 20180329883A1 · Leidner et al. · 2018 [cited by applicant]
US 20190108257A1 · Lefebure · 2019 [cited by examiner]
US 20190147104A1 · Wu et al. · 2019 [cited by applicant]
US 20190155905A1 · Bachrach et al. · 2019 [cited by applicant]
US 20190251165A1 · Bachrach et al. · 2019 [cited by applicant]
US 20190278841A1 · Pusateri · 2019 [cited by examiner]
US 20190370323A1 · Davidson · 2019 [cited by examiner]
US 20200004787A1 · Gupta et al. · 2020 [cited by applicant]
US 20200065334A1 · Rodriguez et al. · 2020 [cited by applicant]
US 20200142888A1 · Alakuijala et al. · 2020 [cited by applicant]
US 20200142959A1 · Mallinar et al. · 2020 [cited by applicant]
US 20200143247A1 · Jonnalagadda et al. · 2020 [cited by applicant]
US 20200167379A1 · Faruqui et al. · 2020 [cited by applicant]
US 20200334334A1 · Keskar et al. · 2020 [cited by applicant]
US 20210042614A1 · Walters et al. · 2021 [cited by applicant]
US 20210056266A1 · Ma et al. · 2021 [cited by applicant]
US 20210110816A1 · Choi et al. · 2021 [cited by applicant]
US 20210118436A1 · Kim · 2021 [cited by examiner]
US 20210141798A1 · Henderson · 2021 [cited by applicant]
US 20210141799A1 · Steedman Henderson · 2021 [cited by applicant]
US 20210142164A1 · Liu et al. · 2021 [cited by applicant]
US 20210142789A1 · Gurbani · 2021 [cited by examiner]
US 20210149993A1 · Torres · 2021 [cited by applicant]
US 20210151039A1 · Wu · 2021 [cited by examiner]
US 20210209304A1 · Yang et al. · 2021 [cited by applicant]
US 20210264115A1 · Wang et al. · 2021 [cited by applicant]
US 20210397610A1 · Singh et al. · 2021 [cited by applicant]
US 20220004717A1 · Park et al. · 2022 [cited by applicant]
US 20220059095A1 · Faria · 2022 [cited by examiner]
US 20220093088A1 · Rangarajan Sridhar et al. · 2022 [cited by applicant]
US 20220165257A1 · Singh et al. · 2022 [cited by applicant]
US 20220215159A1 · Qian et al. · 2022 [cited by applicant]
US 20230044079A1 · Bala · 2023 [cited by applicant]
US 20230143110A1 · Han et al. · 2023 [cited by applicant]
US 20230186898A1 · Weisz · 2023 [cited by examiner]
EP 3486842A1 · 2019 [cited by applicant]
Extended European Search Report of EP21180858.9 by EPO dated Apr. 5, 2022. [cited by applicant]
Agnihotri, Souparni. “Hyperparameter Optimization on Neural Machine Translation.” (2019). [cited by applicant]
Cer, Daniel, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant et al. “Universal Sentence Encoder for English.” EMNLP 2018 (2018): 169. [cited by applicant]
Hashemi, Homa B., Amir Asiaee, and Reiner Kraft. “Query intent detection using convolutional neural networks.” In International Conference on Web Search and Data Mining, Workshop on Query Understanding. 2016. [cited by applicant]
Karagkiozis, Nikolaos. “Clustering Semantically Related Questions.” (2019). [cited by applicant]
Klein, Guillaume, François Hernandez, Vincent Nguyen, and Jean Senellart. “The OpenNMT neural machine translation toolkit: 2020 edition.” In Proceedings of the 14th Conference of the Association for Machine Translation … [cited by applicant]
Kriz, Reno, Joao Sedoc, Marianna Apidianaki, Carolina Zheng, Gaurav Kumar, Eleni Miltsakaki, and Chris Callison-Burch. “Complexity-weighted loss and diverse reranking for sentence simplification.” arXiv preprint arXiv:1… [cited by applicant]
Lin, Ting-En, Hua Xu, and Hanlei Zhang. “Discovering new intents via constrained deep adaptive clustering with cluster refinement.” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 05, pp. … [cited by applicant]
OpenNMT, Models—OpenNMT, https://opennmt.net/OpenNMT/training/models/. [cited by applicant]
Reimers, Nils, and Iryna Gurevych. “Sentence-bert: Sentence embeddings using siamese bert-networks.” arXiv preprint arXiv:1908.10084 (2019). [cited by applicant]
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. “Attention is all you need.” arXiv preprint arXiv:1706.03762 (2017). [cited by applicant]
Wang, Peng, Bo Xu, Jiaming Xu, Guanhua Tian, Cheng-Lin Liu, and Hongwei Hao. “Semantic expansion using word embedding clustering and convolutional neural network for improving short text classification.” Neurocomputing … [cited by applicant]
Xu, Jiaming, Bo Xu, Peng Wang, Suncong Zheng, Guanhua Tian, and Jun Zhao. “Self-taught convolutional neural hetworks for short text clustering.” Neural Networks 88 (2017): 22-31. [cited by applicant]
U.S. Appl. No. 17/455,727, filed Nov. 19, 2021, Pranav Singh. [cited by applicant]
Dyer, Chris, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. “Recurrent neural network grammars.” arXiv preprint arXiv:1602.07776 (2016). [cited by applicant]
Geitgey, A. “Faking the News with Natural Language Processing and GPT-2.” Medium. Sep. 27, 2019. [cited by applicant]
Guu, Kelvin, Tatsunori B. Hashimoto, Yonatan Oren, and Percy Liang. “Generating sentences by editing prototypes.” Transactions of the Association for Computational Linguistics 6 (2018): 437-450. [cited by applicant]
Juraska, Juraj, and Marilyn Walker. “Characterizing variation in crowd-sourced data for training neural language generators to produce stylistically varied outputs.” arXiv preprint arXiv:1809.05288 (2018). [cited by applicant]
Lebret, Rémi, Pedro O. Pinheiro, and Ronan Collobert. “Simple image description generator via a linear phrase-based approach.” arXiv preprint arXiv:1412.8419 (2014). [cited by applicant]
Lewis, Mike, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. “Bart: Denoising sequence-to-sequence pre-training for natural language generation, transla… [cited by applicant]
Li, Zichao, Xin Jiang, Lifeng Shang, and Qun Liu. “Decomposable neural paraphrase generation.” arXiv preprint arXiv:1906.09741 (2019). [cited by applicant]
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. “Language models are unsupervised multitask learners.” OpenAI blog 1, No. 8 (2019): 9. [cited by applicant]
Rajapakse, T. “To Distil or Not to Distil: BERT, ROBERTa, and XLNet” towardsdatascience.com. Feb. 7, 2020. [cited by applicant]
Shaw, Andrew. “A multitask music model with bert, transformer-xl and seq2seq.” [cited by applicant]
Tran, Van-Khanh, and Le-Minh Nguyen. “Semantic refinement gru-based neural language generation for spoken dialogue systems.” In International Conference of the Pacific Association for Computational Linguistics, pp. 63-7… [cited by applicant]
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention is all you need.” In Advances in neural information processing systems, pp. 5998-… [cited by applicant]
Vinyals, Oriol, Alexander Toshev, Samy Bengio, and Dumitru Erhan. “Show and tell: A neural image caption generator.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3156-3164. 2015. [cited by applicant]
Wen, Tsung-Hsien, Milica Gasic, Nikola Mrksic, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, and Steve Young. “Multi-domain neural network language generation for spoken dialogue systems.” arXiv preprint arXiv:1603… [cited by applicant]
Wen, Tsung-Hsien, Milica Gašic, Nikola Mrkšic, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, and Steve Young. “Toward multi-domain language generation using recurrent neural networks.” In NIPS Workshop on Machine L… [cited by applicant]
Zhang, Yaoyuan, Zhenxu Ye, Yansong Feng, Dongyan Zhao, and Rui Yan. “A constrained sequence-to-sequence neural model for sentence simplification.” arXiv preprint arXiv:1704.02312 (2017). [cited by applicant]
Zheng, Renjie, Mingbo Ma, and Liang Huang. “Multi-reference training with pseudo-references for neural translation and text generation.” arXiv preprint arXiv:1808.09564 (2018). [cited by applicant]
Examination Report by EPO of European counterpart patent application No. 21180858.9, dated Sep. 10, 2024. [cited by applicant]
Li, Zichao, et al., “Decomposable neural paraphrase generation.” arXiv preprint arXiv:1906.09741 (Year: 2019). [cited by applicant]