IP Library Granted Patent US 12,154,549
Granted Patent B2
US 12,154,549 · App. 17/644,416 · Granted Nov 26, 2024

Lattice speech corrections

Inventors: Ágoston Weisz (Mountain View, CA); Leonid Velikovich (New York, NY)
Assignee: Google LLC
G10L15/063G10L15/22G10L15/26G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,154,549
App. No.
17/644,416
Granted
Nov 26, 2024
Kind
B2
Abstract

A method includes receiving audio data corresponding to a query spoken and processing the audio data to generate multiple candidate hypotheses each represented by a respective sequence of hypothesized terms. For each candidate hypothesis, the method includes determining whether the sequence of hypothesized terms includes a source phrase from a list of phrase correction pairs. Each phrase correction pair includes a corresponding source phrase that was misrecognized and a corresponding target phrase replacing the source phrase. When the respective sequence of hypothesized terms includes the source phrase, the method includes generating a corresponding additional candidate hypothesis that replaces the source phrase. The method also includes ranking the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated and generating a transcription of the query spoken by the user by selecting the highest ranking one of the multiple candidate hypotheses and each additional candidate hypothesis.

Claims (72)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to a query spoken by a user;

processing, using a speech recognizer, the audio data to generate multiple candidate hypotheses, each candidate hypothesis corresponding to a candidate transcription for the query and represented by a respective sequence of hypothesized terms;

for each candidate hypothesis:

determining whether the respective sequence of hypothesized terms includes a source phrase from a list of phrase correction pairs, each phrase correction pair in the list of phrase correction pairs comprising:

a corresponding source phrase that was misrecognized in a corresponding previous transcription transcribed by the speech recognizer for a previous utterance spoken by the user; and

a corresponding target phrase that corresponds to a user correction replacing the source phrase misrecognized in the corresponding previous transcription transcribed by the speech recognizer; and

when the respective sequence of hypothesized terms includes the source phrase, generating a corresponding additional candidate hypothesis that replaces the source phrase in the respective sequence of hypothesized terms with the corresponding target phrase;

ranking the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated;

generating a transcription of the query spoken by the user by selecting the highest ranking one of the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated; and

for each phrase correction pair in the list of phrase correction pairs:

obtaining an original sequence of n-grams representing the corresponding previous transcription that misrecognized the corresponding source phrase, the original sequence of n-grams including the corresponding source phrase and one or more other terms that precede and/or are subsequent to the corresponding source phrase in the corresponding previous transcription;

obtaining a corrected sequence of n-grams that replaces the source phrase in the original sequence of n-grams with the corresponding target phrase, the corrected sequence of n-grams including the target phrase and the same one or more other terms that precede and/or are subsequent to the corresponding source phrase in the corresponding previous transcription;

modifying a language model by:

adding the original sequence of n-grams and the corrected sequence of n-grams to the language model; and

conditioning the language model to determine a higher prior likelihood score for a number of n-grams from the corrected sequence of n-grams that includes the target phrase than for a same number of n-grams from the original sequence of n-grams that includes the source phrase; and

determining, using the modified language model configured to receive each additional candidate hypothesis as input, a corresponding prior likelihood score for each additional candidate hypothesis,

wherein ranking the multiple candidate hypotheses and each additional candidate hypothesis is based on the corresponding prior likelihood score determined for each additional candidate hypothesis.

2. The computer-implemented method of claim 1 , wherein a margin between the prior likelihood scores determined by the language model for the number of n-grams from the corrected sequence of n-grams and the same number of n-grams from the original sequence of n-grams increases as the number of n-grams from the corrected and original sequences of n-grams increases.

3. The computer-implemented method of claim 1 , wherein conditioning the language model further comprises conditioning the language model to determine a lower prior likelihood score for a first number of n-grams from the sequence of n-grams that includes the target phrase than for a greater second number of n-grams from the sequence of n-grams that includes the target phrase.

4. The computer-implemented method of claim 1 , wherein the original and corrected sequences of n-grams each further includes n-grams representing sentence boundaries of the corresponding previous transcription transcribed by the speech recognizer for the previous utterance spoken by the user.

5. The computer-implemented method of claim 1 , wherein the operations further comprise:

for each of the multiple candidate hypotheses generated by the speech recognizer, obtaining a corresponding likelihood score that the speech recognizer assigned to the corresponding candidate hypothesis; and

after generating each additional candidate hypothesis, determining, using an additional hypothesis scorer, a corresponding likelihood score for each additional candidate hypothesis generated,

wherein ranking the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated is based on the corresponding likelihood scores assigned to the multiple candidate hypothesis by the speech recognizer and the corresponding likelihood score determined for each additional candidate hypothesis using the additional hypothesis scorer.

6. The computer-implemented method of claim 5 , wherein:

the additional hypothesis scorer comprises at least one of:

an acoustic model configured to process audio data to determine an acoustic modeling score for a portion of the audio data that includes either the source phrase or the target phrase; or

the language model configured to receive each additional candidate hypothesis as input and determine the corresponding prior likelihood score for each additional candidate hypothesis; and

the corresponding likelihood score determined for each additional candidate hypothesis is based on at least one of the acoustic modeling score or the corresponding prior likelihood score determined for the additional candidate hypothesis.

7. The computer-implemented method of claim 6 , wherein the language model comprises an auxiliary language model external to the speech recognizer or an internal language model integrated with the speech recognizer.

8. The computer-implemented method of claim 6 , wherein the speech recognizer comprises an end-to-end speech recognition model configured to generate the corresponding likelihood score for each of the multiple candidate hypotheses.

9. The computer-implemented method of claim 5 , wherein:

the speech recognizer comprises an acoustic model and the language model; and

the corresponding likelihood score that the speech recognizer assigned to each of the multiple candidate hypotheses is based on at least one of an acoustic modeling score output by the acoustic model or the corresponding prior likelihood score output by the language model.

10. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data corresponding to a query spoken by a user;

processing, using a speech recognizer, the audio data to generate multiple candidate hypotheses, each candidate hypothesis corresponding to a candidate transcription for the query and represented by a respective sequence of hypothesized terms;

for each candidate hypothesis:

determining whether the respective sequence of hypothesized terms includes a source phrase from a list of phrase correction pairs, each phrase correction pair in the list of phrase correction pairs comprising:

a corresponding source phrase that was misrecognized in a corresponding previous transcription transcribed by the speech recognizer for a previous utterance spoken by the user; and

a corresponding target phrase that corresponds to a user correction replacing the source phrase misrecognized in the corresponding previous transcription transcribed by the speech recognizer; and

when the respective sequence of hypothesized terms includes the source phrase, generating a corresponding additional candidate hypothesis that replaces the source phrase in the respective sequence of hypothesized terms with the corresponding target phrase;

ranking the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated;

generating a transcription of the query spoken by the user by selecting the highest ranking one of the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated; and

for each phrase correction pair in the list of phrase correction pairs:

obtaining an original sequence of n-grams representing the corresponding previous transcription that misrecognized the corresponding source phrase, the original sequence of n-grams including the corresponding source phrase and one or more other terms that precede and/or are subsequent to the corresponding source phrase in the corresponding previous transcription;

obtaining a corrected sequence of n-grams that replaces the source phrase in the original sequence of n-grams with the corresponding target phrase, the corrected sequence of n-grams including the target phrase and the same one or more other terms that precede and/or are subsequent to the corresponding source phrase in the corresponding previous transcription;

modifying a language model by:

adding the original sequence of n-grams and the corrected sequence of n-grams to the language model; and

conditioning the language model to determine a higher prior likelihood score for a number of n-grams from the corrected sequence of n-grams that includes the target phrase than for a same number of n-grams from the original sequence of n-grams that includes the source phrase; and

determining, using the modified language model configured to receive each additional candidate hypothesis as input, a corresponding prior likelihood score for each additional candidate hypothesis,

wherein ranking the multiple candidate hypotheses and each additional candidate hypothesis is based on the corresponding prior likelihood score determined for each additional candidate hypothesis.

11. The system of claim 10 , wherein a margin between the prior likelihood scores determined by the language model for the number of n-grams from the corrected sequence of n-grams and the same number of n-grams from the original sequence of n-grams increases as the number of n-grams from the corrected and original sequences of n-grams increases.

12. The system of claim 10 , wherein conditioning the language model further comprises conditioning the language model to determine a lower prior likelihood score for a first number of n-grams from the sequence of n-grams that includes the target phrase than for a greater second number of n-grams from the sequence of n-grams that includes the target phrase.

13. The system of claim 10 , wherein the original and corrected sequences of n-grams each further includes n-grams representing sentence boundaries of the corresponding previous transcription transcribed by the speech recognizer for the previous utterance spoken by the user.

14. The system of claim 10 , wherein the operations further comprise:

for each of the multiple candidate hypotheses generated by the speech recognizer, obtaining a corresponding likelihood score that the speech recognizer assigned to the corresponding candidate hypothesis; and

after generating each additional candidate hypothesis, determining, using an additional hypothesis scorer, a corresponding likelihood score for each additional candidate hypothesis generated,

wherein ranking the multiple candidate hypotheses and each corresponding additional candidate hypothesis generated is based on the corresponding likelihood scores assigned to the multiple candidate hypothesis by the speech recognizer and the corresponding likelihood score determined for each additional candidate hypothesis using the additional hypothesis scorer.

15. The system of claim 14 , wherein:

the additional hypothesis scorer comprises at least one of:

an acoustic model configured to process audio data to determine an acoustic modeling score for a portion of the audio data that includes either the source phrase or the target phrase; or

the language model configured to receive each additional candidate hypothesis as input and determine the corresponding prior likelihood score for each additional candidate hypothesis; and

the corresponding likelihood score determined for each additional candidate hypothesis is based on at least one of the acoustic modeling score or the corresponding prior likelihood score determined for the additional candidate hypothesis.

16. The system of claim 15 , wherein the language model comprises an auxiliary language model external to the speech recognizer or an internal language model integrated with the speech recognizer.

17. The system of claim 15 , wherein the speech recognizer comprises an end-to-end speech recognition model configured to generate the corresponding likelihood score for each of the multiple candidate hypotheses.

18. The system of claim 14 , wherein:

the speech recognizer comprises an acoustic model and the language model; and

the corresponding likelihood score that the speech recognizer assigned to each of the multiple candidate hypotheses is based on at least one of an acoustic modeling score output by the acoustic model or the corresponding prior likelihood score output by the language model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2022
From: WEISZ, AGOSTON; VEILKOVICH, LEONID
To: GOOGLE LLC
Reel/Frame 059060/0917 →
Continuity (2)
Provisional Application 63265366 · Dec 14, 2021
Related Publication 20230186898A1 · Jun 15, 2023