IP Library Granted Patent US 12,499,871
Granted Patent B2
US 12,499,871 · App. 18/312,587 · Granted Dec 16, 2025

Acoustic model training using corrected terms

Inventors: Olga Kapralova (Bern, CH); Evgeny A. Cherepanov (Adliswil, CH); Dmitry Osmakov (Zurich, CH); Martin Baeuml (Hedingen, CH); Gleb Skobeltsyn (Kilchberg, CH)
Assignee: Google LLC
G10L15/063G10L15/01G10L15/06G10L15/10G10L15/22G10L15/32G10L2015/0635G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,871
App. No.
18/312,587
Granted
Dec 16, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for speech recognition. One of the methods includes receiving first audio data corresponding to an utterance; obtaining a first transcription of the first audio data; receiving data indicating (i) a selection of one or more terms of the first transcription and (ii) one or more of replacement terms; determining that one or more of the replacement terms are classified as a correction of one or more of the selected terms; in response to determining that the one or more of the replacement terms are classified as a correction of the one or more of the selected terms, obtaining a first portion of the first audio data that corresponds to one or more terms of the first transcription; and using the first portion of the first audio data that is associated with the one or more terms of the first transcription to train an acoustic model for recognizing the one or more of the replacement terms.

Claims (40)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving a voice input from a user of a user device;

processing, using a trained voice recognition engine, audio data representing the voice input to generate a first complete transcription of the voice input and one or more complete alternative transcriptions of the voice input, the voice recognition engine trained by a training process that comprises:

obtaining a plurality of training samples, each training sample comprising corresponding audio data representing a corresponding utterance and a corresponding transcription of the corresponding utterance; and

for each training sample, training the voice recognition engine to predict, based on the corresponding audio data representing the corresponding utterance, the corresponding transcription;

displaying, in a graphical user interface of the user device, the first complete transcription of the voice input;

displaying, in the graphical user interface, the one or more complete alternative transcriptions of the voice input;

receiving, via the graphical user interface, a selection indication indicating a user selection of a selected transcription of the first complete transcription or the one or more complete alternative transcriptions displayed in the graphical user interface to provide a corrected transcription of the voice input;

in response to receiving the selection indication indicating user selection of the selected transcription, displaying, in the graphical user interface, one or more links to resources determined to be responsive to the selected transcription; and

receiving, via the graphical user interface, another selection indication indicating user selection of one of the one or more links displayed in the graphical user interface to cause the graphical user interface to present a new screen based on the user selection of the one of the one or more links.

2 . The method of claim 1 , wherein the operations further comprise, after receiving the voice input, displaying, in the graphical user interface, one or more other links to resources determined to be responsive to the first complete transcription.

3 . The method of claim 1 , wherein the operations further comprise, prior to receiving the selection indication indicating the user selection of the selected transcription, receiving, at the user device, a selection indication indicating a user selection in the graphical user interface to correct the first complete transcription.

4 . The method of claim 3 , wherein the operations further comprise, in response to receiving the selection indication indicating the user selection in the graphical user interface to correct the first complete transcription, displaying a keyboard in the graphical user interface.

5 . The method of claim 1 , wherein the operations further comprise providing the selected transcription to a search engine, the search engine configured to generate search results that are determined to be responsive to the selected transcription.

6 . The method of claim 5 , wherein the generated search results comprise a particular resource determined by the search engine to be responsive to a type of request in the voice input.

7 . The method of claim 1 , wherein the one or more links to resources determined to be responsive to the selected transcription displayed in the graphical user interface comprise one or more suggested terms.

8 . The method of claim 1 , wherein the first complete transcription comprises a mis-transcribed term that is replaced in at least one of the one or more complete alternative transcriptions.

9 . The method of claim 8 , wherein the at least one of the one or more complete alternative transcriptions spells out a corresponding replacement term for the mis-transcribed term.

10 . The method of claim 9 , wherein the operations further comprise isolating at least a portion of the voice input as training data for training a model to recognize the corresponding replacement term.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving a voice input from a user of a user device;

processing, using a trained voice recognition engine, audio data representing the voice input to generate a complete first transcription of the voice input and one or more complete alternative transcriptions of the voice input, the voice recognition engine trained by a training process that comprises:

obtaining a plurality of training samples, each training sample comprising corresponding audio data representing a corresponding utterance and a corresponding transcription of the corresponding utterance; and

for each training sample, training the voice recognition engine to predict, based on the corresponding audio data representing the corresponding utterance, the corresponding transcription;

displaying, in a graphical user interface of the user device, the first complete transcription of the voice input;

displaying, in the graphical user interface, the one or more complete alternative transcriptions of the voice input;

receiving, via the graphical user interface, a selection indication indicating a user selection of a selected transcription of the first complete transcription or the one or more complete alternative transcriptions displayed in the graphical user interface to provide a corrected transcription of the voice input;

in response to receiving the selection indication indicating user selection of the selected transcription, displaying, in the graphical user interface, one or more links to resources determined to be responsive to the selected transcription; and

receiving, via the graphical user interface, another selection indication indicating user selection of one of the one or more links displayed in the graphical user interface to cause the graphical user interface to present a new screen based on the user selection of the one of the one or more links.

12 . The system of claim 11 , wherein the operations further comprise, after receiving the voice input, displaying, in the graphical user interface, one or more other links to resources determined to be responsive to the first complete transcription.

13 . The system of claim 11 , wherein the operations further comprise, prior to receiving the selection indication indicating the user selection of the selected transcription, receiving, at the user device, a selection indication indicating a user selection in the graphical user interface to correct the first complete transcription.

14 . The system of claim 13 , wherein the operations further comprise, in response to receiving the selection indication indicating the user selection in the graphical user interface to correct the first complete transcription, displaying a keyboard in the graphical user interface.

15 . The system of claim 11 , wherein the operations further comprise providing the selected transcription to a search engine, the search engine configured to generate search results that are determined to be responsive to the selected transcription.

16 . The system of claim 15 , wherein the generated search results comprise a particular resource determined by the search engine to be responsive to a type of request in the voice input.

17 . The system of claim 11 , wherein the one or more links to resources determined to be responsive to the selected transcription displayed in the graphical user interface comprise one or more suggested terms.

18 . The system of claim 11 , wherein the first complete transcription comprises a mis-transcribed term that is replaced in at least one of the one or more complete alternative transcriptions.

19 . The system of claim 18 , wherein the at least one of the one or more complete alternative transcriptions spells out a corresponding replacement term for the mis-transcribed term.

20 . The system of claim 19 , wherein the operations further comprise isolating at least a portion of the voice input as training data for training a model to recognize the corresponding replacement term.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2024
From: KAPRALOVA, OLGA; CHEREPANOV, EVGENY A.; OSMAKOV, DMITRY; BAEUMI, MARTIN; SKOBELTSYN, GLEB
To: GOOGLE, INC.
Reel/Frame 067014/0682 →
CHANGE OF NAME Recorded Apr 5, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 067025/0697 →
Continuity (5)
Continuation 17457421 · Dec 2, 2021
Continuation 16837393 · Apr 1, 2020
Continuation 16023658 · Jun 29, 2018
Continuation 15224104 · Jul 29, 2016
Related Publication 20230274729A1 · Aug 31, 2023
References Cited (27)
US 6735565B2 · Gschwendtner · 2004 [cited by applicant]
US 6912498B2 · Stevens et al. · 2005 [cited by applicant]
US 7356467B2 · Kemp · 2008 [cited by applicant]
US 8185392B1 · Strope et al. · 2012 [cited by applicant]
US 8494853B1 · Mengíbar et al. · 2013 [cited by applicant]
US 8719014B2 · Wagner · 2014 [cited by applicant]
US 9263033B2 · Siohan et al. · 2016 [cited by applicant]
US 9378731B2 · Kapralova et al. · 2016 [cited by applicant]
US 20020123893A1 · Woodward · 2002 [cited by applicant]
US 20020138265A1 · Stevens · 2002 [cited by examiner]
US 20050203751A1 · Stevens et al. · 2005 [cited by applicant]
US 20060015338A1 · Poussin · 2006 [cited by applicant]
US 20060074656A1 · Mathias et al. · 2006 [cited by applicant]
US 20060178886A1 · Braho · 2006 [cited by examiner]
US 20080104055A1 · Segel · 2008 [cited by applicant]
US 20090138266A1 · Nagae · 2009 [cited by applicant]
US 20100023331A1 · Duta et al. · 2010 [cited by applicant]
US 20140019127A1 · Park et al. · 2014 [cited by applicant]
US 20140056475A1 · Jang et al. · 2014 [cited by applicant]
US 20150243278A1 · Kibre et al. · 2015 [cited by applicant]
US 20150286634A1 · Shin · 2015 [cited by examiner]
US 20160155436A1 · Choi et al. · 2016 [cited by applicant]
US 20170004120A1 · Eck et al. · 2017 [cited by applicant]
PCT International Preliminary Report on Patentability in International AppIn. PCT/US2017 /038336, dated Feb. 7, 2019, 10 pages. [cited by applicant]
Invitation to Pay Additional Fees and Where Applicable Protest Fee, with Partial Search Report, issued in International Application No. PCT/US2017/038336, mailed on Oct. 2, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, issued in International Application No. PCT/US2017/038336, mailed on Nov. 23, 2017, 17 pages. [cited by applicant]
Korean Intellectual Property Office—Notice of Office Action for Application No. 10-2019-7005232, Dated Feb. 21, 2019. [cited by applicant]