IP Library › Granted Patent US 12,198,681
Granted Patent B1
US 12,198,681 · App. 17/937,297 · Granted Jan 14, 2025

Personalized batch and streaming speech-to-text transcription of audio

Inventors: Monica Lakshmi Sunkara (San Jose, CA); Srikanth Ronanki (San Jose, CA); Sravan Babu Bodapati (Redmond, WA); Jeffrey John Farris (Crystal Lake, IL); Katrin Kirchhoff (Seattle, WA); Vivek Govindan (Redmond, WA); Yide Zou (Aachen, DE); Mohit Narendra Gupta (Seattle, WA); Silviu Mihai Burz (Sinking Spring, PA)
Assignee: Amazon Technologies, Inc.
G10L15/16G10L15/30G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,681
App. No.
17/937,297
Granted
Jan 14, 2025
Kind
B1
Abstract

Techniques for personalized batch and streaming speech-to-text transcription of audio reduce the error rate of automatic speech recognition (ASR) systems in transcribing rare and out-of-vocabulary words. The techniques achieve personalization of connectionist temporal classification (CT) models by using adaptive boosting to perform biasing at the level of sub-words. In addition to boosting, the techniques encompass a phone alignment network to bias sub-word predictions towards rare long-tail words and out-of-vocabulary words. A technical benefit of the techniques is that the accuracy of speech-to-text transcription of rare and out-of-vocabulary words in a custom vocabulary by automatic speech recognition (ASR) system can be improved without having to train the ASR system on the custom vocabulary. Instead, the techniques allow the same ASR system trained on a base vocabulary to realize the accuracy improvements for different custom vocabularies spanning different domains.

Claims (74)

1. A method comprising:

receiving, at a transcription service in a provider network, media data comprising audio data that encodes a spoken utterance, the transcription service implemented by one or more electronic devices in the provider network;

obtaining, at the transcription service, an output of a connectionist temporal classification (CTC) decoder, the output comprising a set of probability values for a sequence of time steps and a set of sub words, the set of sub words comprising a blank sub word, each probability value in the set of probability values for one corresponding time step of the sequence of time steps and one corresponding sub word of the set of sub words;

beam search decoding, at the transcription service, the output to yield a set of N-best candidate sub word sequences;

wherein beam search decoding the output comprises boosting a probability of a candidate sub word sequence based on determining that the candidate sub word sequence contains a prefix of a sub word sequence representing a rare or out-of-vocabulary word;

selecting, at the transcription service, the candidate sub word sequence for inclusion in the set of N-best candidate sub word sequences based on determining the boosted probability; and

generating, at the transcription service, a text transcription of the spoken utterance based on selecting a 1-best sub word sequence of the set of N-best candidate sub word sequences.

2. The method of claim 1 , further comprising:

receiving, at the transcription service, the media data from a storage service in the provider network, the storage service implemented by one or more electronic devices in the provider network.

3. The method of claim 1 , further comprising:

receiving, at a streaming endpoint of the transcription service, the media data from a user device; and

sending, from the streaming endpoint, the text transcription to the user device.

4. A method comprising:

receiving media data comprising audio data that encodes a spoken utterance;

obtaining an output of a connectionist temporal classification (CTC) decoder, the output comprising a set of probability values for a sequence of time steps and a set of sub words, the set of sub words comprising a blank sub word, each probability value in the set of probability values for one corresponding time step of the sequence of time steps and one corresponding sub word of the set of sub words;

beam search decoding the output to yield a set of best candidate sub word sequences;

wherein beam search decoding the output comprises boosting a probability of a candidate sub word sequence based on determining that the candidate sub word sequence contains a prefix of a sub word sequence representing a rare or out-of-vocabulary word;

selecting the candidate sub word sequence for inclusion in the set of best candidate sub word sequences based on determining the boosted probability; and

generating a text transcription of the spoken utterance based on selecting a 1-best sub word sequence of the set of best candidate sub word sequences.

5. The method of claim 4 , wherein beam search decoding the output comprises:

determining a boost amount based on determining a difference between (a) the probability of the particular candidate sub word sequence and (b) a probability of a best candidate sub word sequence.

6. The method of claim 5 , wherein beam search decoding the output comprises:

determining a boosting scale based on determining the difference between (a) the probability of the particular candidate sub word sequence and (b) the probability of the best candidate sub word sequence for the sub sequence of time steps; and

determining the boost amount as a product of:

the boosting scale, and

the difference between (a) the probability of the particular candidate sub word sequence and (b) the probability of the best candidate sub word sequence.

7. The method of claim 4 , further comprising:

converting the set of best candidate sub word sequences into a corresponding set of phone sequences;

using a dynamic time warping algorithm to compute a set of alignment distances, each alignment distance of the set of alignment distances between a respective phone sequence of the set of phone sequences and a phone sequence representing the spoken utterance, the phone sequence generated by a phone alignment network;

reordering the set of best candidate sub word sequences based on determining the set of alignment distances; and

selecting the 1-best sub word sequence from the reordered set of best candidate sub word sequences.

8. The method of claim 4 , wherein:

generating the text transcription of the spoken utterance based on selecting the 1-best sub word sequence comprises: tokenizing the 1-best sub word sequence into a set of words; and replacing a particular word in the set of words with a custom word from a custom vocabulary based on determining that (a) a pronunciation of the particular word according to corresponding phone predictions derived from the audio data matches (b) a lexicon-derived pronunciation of the custom word; and

the text transcription contains the custom word in place of the particular word.

9. The method of claim 4 , wherein:

the method further comprises determining a set of one or more words that are phonetically similar to the rare or out-of-vocabulary word; and

generating the text transcription of the spoken utterance based on selecting the 1-best sub word sequence comprises: recovering a word sequence from the 1-best sub word sequence; and replacing a particular word in the word sequence with the rare or out-of-vocabulary word based on determining that the particular word is in the set of one or more words that are phonetically similar to the rare or out-of-vocabulary word.

10. The method of claim 4 , further comprising:

receiving, at a transcription service in a provider network, the media data from a storage service in the provider network, the transcription service implemented by a first one or more electronic devices in the provider network, the storage service implemented by a second one or more electronic devices in the provider network.

11. The method of claim 4 , further comprising:

receiving, at a streaming endpoint of a transcription service in a provider network, the media data from a user device, the transcription service implemented by a first one or more electronic devices in the provider network; and

sending, from the streaming endpoint, the text transcription to the user device.

12. The method of claim 4 , further comprising:

a shared conformer encoder generating an acoustic embedding representation of the spoken utterance; and

the connectionist temporal classification (CTC) decoder generating the output based on determining the acoustic embedding representation.

13. The method of claim 4 , further comprising:

receiving, from a user device, a custom vocabulary comprising the rare or out-of-vocabulary word.

14. The method of claim 4 , wherein the method is performed by an edge electronic device.

15. A system comprising:

one or more electronic devices to implement a transcription service in a provider network, the transcription service comprising instructions which when executed cause the transcription service to:

receive media data comprising audio data that encodes a spoken utterance;

obtain an output of a connectionist temporal classification (CTC) decoder, the output comprising a set of probability values for a sequence of time steps and a set of sub words, the set of sub words comprising a blank sub word, each probability value in the set of probability values for one corresponding time step of the sequence of time steps and one corresponding sub word of the set of sub words;

beam search decode the output to yield a set of best candidate sub word sequences;

wherein the instructions to beam search decode the output comprise instructions which when executed cause the transcription service to: boost a probability of a candidate sub word sequence based on determining that the candidate sub word sequence contains a prefix of a sub word sequence representing a rare or out-of-vocabulary word;

select the candidate sub word sequence for inclusion in the set of best candidate sub word sequences based on determining the boosted probability; and

generate a text transcription of the spoken utterance based on selecting a 1-best sub word sequence of the set of best candidate sub word sequences.

16. The system of claim 15 , wherein the instructions to beam search decode the output comprise instructions which when executed cause the transcription service to:

determine a boost amount based on determining a difference between (a) the probability of the particular candidate sub word sequence and (b) a probability of a best candidate sub word sequence.

17. The system of claim 16 , wherein the instructions to beam search decode the output comprise instructions which when executed cause the transcription service to:

determine a boosting scale based on determining the difference between (a) the probability of the particular candidate sub word sequence and (b) the probability of the best candidate sub word sequence; and

determine the boost amount as a product of:

the boosting scale, and

the difference between (a) the probability of the particular candidate sub word sequence and (b) the probability of the best candidate sub word sequence.

18. The system of claim 15 , wherein the instructions when executed further cause the transcription service to:

convert the set of best candidate sub word sequences into a corresponding set of phone sequences;

use a dynamic time warping algorithm to compute a set of alignment distances, each alignment distance of the set of alignment distances between a respective phone sequence of the set of phone sequences and a phone sequence representing the spoken utterance, the phone sequence generated by a phone alignment network;

reorder the set of best candidate sub word sequences based on determining the set of alignment distances; and

select the 1-best sub word sequence from the reordered set of best candidate sub word sequences.

19. The system of claim 15 , wherein:

the instructions for generating the text transcription of the spoken utterance based on selecting the 1-best sub word sequence comprise instructions which when executed cause the transcription service to: tokenize the 1-best sub word sequence into a set of words; and replace a particular word in the set of words with a custom word from a custom vocabulary based on determining that (a) a pronunciation of the particular word according to corresponding phone predictions derived from the audio data matches (b) a lexicon-derived pronunciation of the custom word; and

the text transcription contains the custom word in place of the particular word.

20. The system of claim 15 , wherein:

the instructions when executed further cause the transcription service to determine a set of one or more words that are phonetically similar to the rare or out-of-vocabulary word; and

the instructions for generating the text transcription of the spoken utterance based on selecting the 1-best sub word sequence comprise instructions which when executed cause the transcription service to: recover a word sequence from the 1-best sub word sequence; and replace a particular word in the word sequence with the rare or out-of-vocabulary word based on the determining that the particular word is in the set of one or more words that are phonetically similar to the rare or out-of-vocabulary word.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2024
From: SUNKARA, MONICA LAKSHMI; RONANKI, SRIKANTH; BODAPATI, SRAVAN BABU; FARRIS, JEFFREY JOHN; KIRCHHOFF, KATRIN; GOVINDAN, VIVEK; ZOU, YIDE; GUPTA, MOHIT NARENDRA; BURZ, SILVIU MIHAI
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 066791/0617 →
References Cited (11)
US 20190189111A1 · Watanabe · 2019 [cited by examiner]
US 20200242197A1 · Srinivasan · 2020 [cited by examiner]
Gulati et al., “Conformer: Convolution-augmented Transformer for Speech Recognition”, Electrical Engineering and Systems Science, May 16, 2020, 5 pages. [cited by applicant]
Heafield et al., “Scalable Modified Kneser-Ney Language Model Estimation”, Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, Aug. 4-9, 2013, pp. 690-696. [cited by applicant]
Kim et al., “Joint CTC-attention based end-to-end speech recognition using multi-task learning”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Jan. 31, 2017, 5 pages. [cited by applicant]
Le et al., “Deep Shallow Fusion for RNN-T Personalization”, IEEE Spoken Language Technology Workshop (SLT), Nov. 16, 2021, 7 pages. [cited by applicant]
Le et al., “G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Feb. 13, 2020, 5 pages. [cited by applicant]
Liptchinsky et al., “Letter-Based Speech Recognition with Gated ConvNets”, Computer Science, Feb. 16, 2019, 10 pages. [cited by applicant]
Saxon et al., “End-to-End Spoken Language Understanding for Generalized Voice Assistants”, Computer Science, Jul. 19, 2021, 5 pages. [cited by applicant]
Wu et al., “Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation”, Computer Science, Oct. 8, 2016, pp. 1-23. [cited by applicant]
Zhao et al., “On Addressing Practical Challenges for RNN-Transducer”, IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Jul. 18, 2021, 8 pages. [cited by applicant]
Cited By (2)
US 12,400,659 US 12,620,396