IP Library Granted Patent US 12,462,793
Granted Patent B2
US 12,462,793 · App. 17/536,938 · Granted Nov 4, 2025

Speech recognition hypothesis generation according to previous occurrences of hypotheses terms and/or contextual data

Inventors: Ágoston Weisz (Zurich, CH); Alexandru Dovlecel (Zurich, CH); Gleb Skobeltsyn (Kilchberg, CH); Evgeny Cherepanov (Adliswil, CH); Justas Klimavicius (Zurich, CH); Yihui Ma (Zurich, CH); Lukas Lopatovsky (Kilchberg, CH)
Assignee: GOOGLE LLC
G10L15/07G10L15/083G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,793
App. No.
17/536,938
Granted
Nov 4, 2025
Kind
B2
Abstract

Implementations set forth herein relate to speech recognition techniques for handling variations in speech among users (e.g. due to different accents) and processing features of user context in order to expand a number of speech recognition hypotheses when interpreting a spoken utterance from a user. In order to adapt to an accent of the user, terms common to multiple speech recognition hypotheses can be filtered out in order to identify inconsistent terms apparent in a group of hypotheses. Mappings between inconsistent terms can be stored for subsequent users as term correspondence data. In this way, supplemental speech recognition hypotheses can be generated and subject to probability-based scoring for identifying a speech recognition hypothesis that most correlates to a spoken utterance provided by a user. In some implementations, prior to scoring, hypotheses can be supplemented based on contextual data, such as on-screen content and/or application capabilities.

Claims (42)

1 . A method for performing speech recognition on a spoken utterance from a user, the method implemented by one or more processors and comprising:

processing, at a speech recognition engine of a computing device, audio data corresponding to the spoken utterance;

generating, based on processing the audio data and by the speech recognition engine of the computing device, a plurality of current speech recognition hypotheses,

wherein each current speech recognition hypothesis of the plurality of current speech recognition hypotheses includes corresponding terms that are predicted to correspond to original natural language content of the spoken utterance from the user;

identifying term correspondence data that characterizes relationships between previous terms provided in previous speech recognition hypotheses generated based on previous spoken utterances from the user;

determining, based on the term correspondence data, that a given term, of at least a given current hypothesis of the plurality of current speech recognition hypotheses:

is included in the term correspondence data, and

corresponds to a related term of at least a given previous hypothesis of the previous speech recognition hypotheses, in the term correspondence data, that is not included in any of the plurality of current speech recognition hypotheses based at least in part on the given term sharing a common position in the given current hypothesis with the related term in the given previous speech recognition hypothesis;

based on determining that the given term is included in the term correspondence data and corresponds to the related term that is not included in any of the plurality of current speech recognition hypotheses:

generating, via the speech recognition engine of the computing device, a supplemental current speech recognition hypothesis that conforms to the given current hypothesis, but replaces the given term with the related term; and

selecting the supplemental current speech recognition hypothesis as an actual speech recognition result; and

in response to the selecting, causing the computing device to render an output based on the supplemental current speech recognition hypothesis.

2 . The method of claim 1 , wherein selecting the supplemental current speech recognition hypothesis as the actual speech recognition result is based on contextual data that characterizes a context in which the user provided the spoken utterance.

3 . The method of claim 2 , wherein the contextual data characterizes graphical content being rendered at a graphical user interface of the computing device when the user provided the spoken utterance.

4 . The method of claim 2 , wherein the contextual data characterizes one or more applications that are accessible via the computing device.

5 . The method of claim 4 , wherein selecting the supplemental current speech recognition hypothesis as the actual speech recognition result based on contextual data comprises:

selecting the supplemental current speech recognition hypothesis based on determining that the supplemental current speech recognition hypothesis corresponds to an action that is capable of being initialized via the one or more applications that are accessible via the computing device.

6 . The method of claim 1 , wherein the given term corresponds to the related term, in the term correspondence data, based on the given term occurring in a first previous speech recognition hypothesis generated based on a first previous spoken utterance from the user and the related term occurring in a second previous speech recognition hypothesis generated based on the same first previous spoken utterance.

7 . The method of claim 1 , further comprising, prior to processing the audio data:

generating, in the term correspondence data, the correspondence between the given term and the related term, wherein generating the correspondence is in response to the given term and the related term both being included, in one or more corresponding previous speech recognition hypotheses, for one or more same of the previous spoken utterances.

8 . A computing device comprising:

one or more microphones;

memory storing term correspondence data that characterizes relationships between previous terms provided in previous speech recognition hypotheses generated based on previous spoken utterances from a user; and

one or more processors configured to:

process, via a speech recognition engine, audio data that is captured via the one or more microphones and that captures a spoken utterance of the user;

generate, based on processing the audio data and via the speech recognition engine, a plurality of current speech recognition hypotheses,

wherein each current speech recognition hypothesis of the plurality of current speech recognition hypotheses includes corresponding terms that are predicted to correspond to original natural language content of the spoken utterance from the user;

determine, based on the term correspondence data, that a given term, of at least a given current hypothesis of the plurality of current speech recognition hypotheses:

is included in the term correspondence data, and

corresponds to a related term of at least a given previous hypothesis of the previous speech recognition hypotheses, in the term correspondence data, that is not included in any of the plurality of current speech recognition hypotheses based at least in part on the given term sharing a common position in the given current hypothesis with the related term in the given previous speech recognition hypothesis;

based on determining that the given term is included in the term correspondence data and corresponds to the related term that is not included in any of the plurality of current speech recognition hypotheses:

generate, via the speech recognition engine, a supplemental current speech recognition hypothesis that conforms to the given current hypothesis, but replaces the given term with the related term; and

select the supplemental current speech recognition hypothesis as an actual speech recognition result; and

in response to the selecting, cause an output to be rendered based on the supplemental current speech recognition hypothesis.

9 . The computing device of claim 8 , wherein in selecting the supplemental current speech recognition hypothesis as the actual speech recognition result one or more of the processors are to select the supplemental current speech recognition hypothesis as the actual speech recognition result based on contextual data that characterizes a context in which the user provided the spoken utterance.

10 . The computing device of claim 9 , wherein the contextual data characterizes graphical content being rendered at a graphical user interface of the computing device when the user provided the spoken utterance.

11 . The computing device of claim 9 , wherein the contextual data characterizes one or more applications that are accessible via the computing device.

12 . The computing device of claim 11 , wherein in selecting the supplemental current speech recognition hypothesis as the actual speech recognition result based on contextual data one or more of the processors are to:

select the supplemental current speech recognition hypothesis based on determining that the supplemental current speech recognition hypothesis corresponds to an action that is capable of being initialized via the one or more applications that are accessible via the computing device.

13 . The computing device of claim 8 , wherein the given term corresponds to the related term, in the term correspondence data, based on the given term occurring in a first previous speech recognition hypothesis generated based on a first previous spoken utterance from the user and the related term occurring in a second previous speech recognition hypothesis generated based on the same first previous spoken utterance.

14 . The computing device of claim 8 , wherein one or more of the processors are further to, prior to processing the audio data:

generate, in the term correspondence data, the correspondence between the given term and the related term, wherein one or more of the processors are to generate the correspondence in response to the given term and the related term both being included, in one or more corresponding previous speech recognition hypotheses, for one or more same of the previous spoken utterances.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2022
From: WEISZ, ÁGOSTON; DOVLECEL, ALEXANDRU; SKOBELTSYN, GLEB; CHEREPANOV, EVGENY; KLIMAVICIUS, JUSTAS; MA, YIHUI; LOPATOVSKY, LUKAS
To: GOOGLE LLC
Reel/Frame 059297/0554 →
Continuity (3)
Continuation 16614241
Provisional Application 62871571 · Jul 8, 2019
Related Publication 20220084503A1 · Mar 17, 2022
References Cited (29)
US 9224386B1 · Weber · 2015 [cited by applicant]
US 9460708B2 · Zweig · 2016 [cited by examiner]
US 9558740B1 · Mairesse · 2017 [cited by examiner]
US 9576578B1 · Skobeltsyn et al. · 2017 [cited by applicant]
US 10152298B1 · Salvador · 2018 [cited by examiner]
US 10176802B1 · Ladhak et al. · 2019 [cited by applicant]
US 10210267B1 · Lloyd · 2019 [cited by examiner]
US 10504512B1 · Sarikaya · 2019 [cited by examiner]
US 10522133B2 · Labsky · 2019 [cited by examiner]
US 10896681B2 · Aleksic · 2021 [cited by examiner]
US 11189264B2 · Weisz et al. · 2021 [cited by applicant]
US 20020184019A1 · Hartley et al. · 2002 [cited by applicant]
US 20040186714A1 · Baker · 2004 [cited by applicant]
US 20050033574A1 · Kim et al. · 2005 [cited by applicant]
US 20100305947A1 · Schwarz · 2010 [cited by examiner]
US 20150106082A1 · Ge et al. · 2015 [cited by applicant]
US 20170229124A1 · Strohman · 2017 [cited by examiner]
US 20180096678A1 · Zhou et al. · 2018 [cited by applicant]
US 20210064822A1 · Velikovich et al. · 2021 [cited by applicant]
CN 103730115 · 2014 [cited by applicant]
CN 107977356 · 2018 [cited by applicant]
CN 108091328 · 2018 [cited by applicant]
CN 108399914 · 2018 [cited by applicant]
JP 5906615 · 2016 [cited by applicant]
European Patent Office; Communication pursuant to Article 94(3) issued in Application No. 19746407.6, 3 pages, dated Aug. 17, 2022. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of Ser. No. PCT/US2019/042204; 16 pages; dated Jan. 29, 2020. [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 19746407.6, 44 pages, dated Jun. 4, 2024. [cited by applicant]
European Patent Office, Communication issued in Application No. 24209931.5; 8 pages; dated Jan. 9, 2025. [cited by applicant]
Chinese Patent Office; Notice of First Office Action issued in Application No. 201980097908.X, 36 pages, dated May 23, 2025. [cited by applicant]