IP Library Granted Patent US 11,189,264
Granted Patent B2
US 11,189,264 · App. 16/614,241 · Granted Nov 30, 2021

Speech recognition hypothesis generation according to previous occurrences of hypotheses terms and/or contextual data

Inventors: Ágoston Weisz (Zurich, CH); Alexandru Dovlecel (Zurich, CH); Gleb Skobeltsyn (Kilchberg, CH); Evgeny Cherepanov (Adliswil, CH); Justas Klimavicius (Zurich, CH); Yihui Ma (Zurich, CH); Lukas Lopatovsky (Kilchberg, CH)
Assignee: GOOGLE LLC
G10L15/07G10L15/083G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,189,264
App. No.
16/614,241
Granted
Nov 30, 2021
Kind
B2
Abstract

Implementations set forth herein relate to speech recognition techniques for handling variations in speech among users (e.g. due to different accents) and processing features of user context in order to expand a number of speech recognition hypotheses when interpreting a spoken utterance from a user. In order to adapt to an accent of the user, terms common to multiple speech recognition hypotheses can be filtered out in order to identify inconsistent terms apparent in a group of hypotheses. Mappings between inconsistent terms can be stored for subsequent users as term correspondence data. In this way, supplemental speech recognition hypotheses can be generated and subject to probability-based scoring for identifying a speech recognition hypothesis that most correlates to a spoken utterance provided by a user. In some implementations, prior to scoring, hypotheses can be supplemented based on contextual data, such as on-screen content and/or application capabilities.

Claims (39)

1. A method implemented by one or more processors, the method comprising:

processing, at a computing device, audio data corresponding to a spoken utterance provided by a user;

generating, based on processing the audio data, a plurality of speech recognition hypotheses,

wherein each speech recognition hypothesis of the plurality of speech recognition hypotheses includes corresponding natural language content predicted to characterize original natural language content of the spoken utterance from the user;

determining, based on processing the audio data, whether a first term, of a first speech recognition hypothesis of the plurality of speech recognition hypotheses, is different from a second term, of a second speech recognition hypothesis of the plurality of speech recognition hypotheses; and

when the first term of the first speech recognition hypothesis is different from the second term of the second speech recognition hypothesis:

generating, based on determining that the first term is different from the second term, term correspondence data that characterizes a relationship between the first term and the second term;

subsequent to generating the term correspondence data:

processing the term correspondence data in furtherance of supplementing subsequent speech recognition hypotheses that identify the first term, but not the second term,

generating a supplemental speech recognition hypothesis for the subsequent speech recognition hypotheses, wherein the supplemental speech recognition hypothesis includes the second term,

determining a prioritized speech recognition hypothesis for the spoken utterance from the plurality of speech recognition hypotheses and the supplemental speech recognition hypothesis, and

causing one or more applications and/or devices to initialize performance of one or more actions according to the prioritized speech recognition hypothesis;

processing, at the computing device, additional audio data corresponding to an additional spoken utterance provided by the user;

generating, based on processing the additional audio data, an additional plurality of speech recognition hypotheses,

determining, based on processing the additional audio data, whether an additional first term, of an additional first speech recognition hypothesis of the additional plurality of speech recognition hypotheses, is different from an additional second term, of an additional second speech recognition hypothesis of the additional plurality of additional speech recognition hypotheses; and

when the additional first term of the additional first speech recognition hypothesis is not different from the additional second term of the additional second speech recognition hypothesis:

determining, based on existing term correspondence data, whether the additional first term and/or the additional second term are with a related term in the existing term correspondence data.

2. The method of claim 1 , further comprising:

determining whether the first term and the second term are each predicted based at least in part on a same segment of the audio data,

wherein generating the term correspondence data is performed when the first term and the second term are each predicted based at least in part on the same segment of audio data.

3. The method of claim 1 , further comprising:

determining whether the first term of the first speech recognition hypothesis shares a common position with the second term of the second speech recognition hypothesis,

wherein generating the term correspondence data is performed when the first term of the first speech recognition hypothesis shares the common position with the second term of the second speech recognition hypothesis.

4. The method of claim 3 , wherein determining whether the first term of the first speech recognition hypothesis shares the common position with the second term of the second speech recognition hypothesis includes:

determining that the first term is directly adjacent to a particular natural language term within the first speech recognition hypothesis of the plurality of speech recognition hypotheses, and

determining that the second term is also directly adjacent to the particular natural language term within the second speech recognition hypothesis of the plurality of speech recognition hypotheses.

5. The method of claim 3 , wherein determining whether the first term of the first speech recognition hypothesis shares the common position with the second term of the second speech recognition hypothesis includes:

determining that the first term is directly between two natural language terms within the first speech recognition hypothesis of the plurality of speech recognition hypotheses, and

determining that the second term is also directly between the two natural language terms within the second speech recognition hypothesis of the plurality of speech recognition hypotheses.

6. The method of claim 1 , further comprising:

determining, subsequent to generating the term correspondence data, a prioritized speech recognition hypothesis from the plurality of speech recognition hypotheses based on contextual data that characterizes a context in which the user provided the spoken utterance; and

causing the computing device to render an output based on the prioritized speech recognition hypothesis.

7. The method of claim 6 , wherein the contextual data characterizes graphical content being rendered at a graphical user interface of the computing device when the user provided the spoken utterance.

8. The method of claim 6 , wherein the contextual data further characterizes one or more applications that are accessible via the computing device, and determining the prioritized speech recognition hypothesis includes:

prioritizing each speech recognition hypothesis of the plurality of speech recognition hypotheses according to whether each speech recognition hypothesis corresponds to an action that is capable of being initialized via the one or more applications that are accessible via the computing device.

9. The method of claim 1 , further comprising:

when the first term of the first speech recognition hypothesis is not different from the second term of the second speech recognition hypothesis, and when the first term and/or the second term are correlated with the related term in the existing term correspondence data:

generating, based on the existing term correspondence data, another supplemental speech recognition hypothesis that includes the related term.

10. The method of claim 9 , wherein the other supplemental speech recognition hypothesis is void of the first term and the second term.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2020
From: WEISZ, ÁGOSTON; DOVLECEL, ALEXANDRU; SKOBELTSYN, GLEB; CHEREPANOV, EVGENY; KLIMAVICIUS, JUSTAS; MA, YIHUI; LOPATOVSKY, LUKAS
To: GOOGLE LLC
Reel/Frame 051961/0148 →
Continuity (2)
Provisional Application 62871571 · Jul 8, 2019
Related Publication 20210012765A1 · Jan 14, 2021
Cited By (1)
US 12,462,793