IP Library Granted Patent US 11,200,887
Granted Patent B2
US 11,200,887 · App. 16/837,393 · Granted Dec 14, 2021

Acoustic model training using corrected terms

Inventors: Olga Kapralova (Bern, CH); Evgeny A. Cherepanov (Adliswil, CH); Dmitry Osmakov (Zurich, CH); Martin Baeuml (Hedingen, CH); Gleb Skobeltsyn (Kilchberg, CH)
Assignee: Google LLC
G10L15/063G10L15/01G10L15/06G10L15/10G10L15/22G10L15/32G10L2015/0635G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,200,887
App. No.
16/837,393
Granted
Dec 14, 2021
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for speech recognition. One of the methods includes receiving first audio data corresponding to an utterance; obtaining a first transcription of the first audio data; receiving data indicating (i) a selection of one or more terms of the first transcription and (ii) one or more of replacement terms; determining that one or more of the replacement terms are classified as a correction of one or more of the selected terms; in response to determining that the one or more of the replacement terms are classified as a correction of the one or more of the selected terms, obtaining a first portion of the first audio data that corresponds to one or more terms of the first transcription; and using the first portion of the first audio data that is associated with the one or more terms of the first transcription to train an acoustic model for recognizing the one or more of the replacement terms.

Claims (29)

1. A method comprising:

receiving, at a user device, a voice input from a user of the user device;

after receiving the voice input, displaying, by the user device, in a graphical user interface, a first transcription of the voice input, the first transcription comprising a plurality of recognized terms;

receiving, at the user device, a typed input indicating one or more text characters inputted by the user through the graphical user interface, the one or more text characters replacing one of the plurality of recognized terms of the first transcription of the voice input to provide a corrected transcription of the voice input; and

displaying, by the user device, in the graphical user interface, one or more links to resources determined to be responsive to the corrected transcription of the voice input,

wherein the one or more links to the resources determined to be responsive to the corrected transcription of the voice input displayed in the graphical user interface comprise the one or more text characters inputted by the user through the graphical user interface.

2. The method of claim 1 , further comprising, after receiving the voice input, displaying, by the user device, in the graphical user interface, one or more other links to other resources determined to be responsive to the first transcription of the voice input.

3. The method of claim 1 , further comprising, prior to receiving the typed input, receiving, at the user device, a selection indication indicating a user selection in the graphical user interface to correct the first transcription.

4. The method of claim 3 , further comprising, in response to receiving the selection indication, displaying, by the user device, a keyboard in the graphical user interface.

5. The method of claim 4 , wherein the user uses the keyboard displayed in the graphical user interface to provide the typed input.

6. The method of claim 1 , wherein the one or more text characters comprise one or more letters inputted by the user through the graphical user interface.

7. The method of claim 1 , wherein the one of the plurality of recognized terms of the first transcription replaced by the one or more text characters inputted by the user through the graphical user interface comprises a mis-transcribed term.

8. The method of claim 1 , wherein the one or more text characters inputted by the user through the graphical user interface spell out a replacement term.

9. The method of claim 8 , further comprising isolating, by the user device, at least a portion of the voice input as training data for training a model to recognize the replacement term.

10. A user device comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a voice input from a user of the user device;

after receiving the voice input, displaying, in a graphical user interface, a first transcription of the voice input, the first transcription comprising a plurality of recognized terms;

receiving a typed input indicating one or more text characters inputted by the user through the graphical user interface, the one or more text characters replacing one of the plurality of recognized terms of the first transcription of the voice input to provide a corrected transcription of the voice input; and

displaying, in the graphical user interface, one or more links to resources determined to be responsive to the corrected transcription of the voice input, wherein the one or more links to the resources determined to be responsive to the corrected transcription of the voice input displayed in the graphical user interface comprise the one or more text characters inputted by the user through the graphical user interface.

11. The user device of claim 10 , wherein the operations further comprise, after receiving the voice input, displaying, in the graphical user interface, one or more other links to other resources determined to be responsive to the first transcription of the voice input.

12. The user device of claim 10 , wherein the operations further comprise, prior to receiving the typed input, receiving a selection indication indicating a user selection in the graphical user interface to correct the first transcription.

13. The user device of claim 12 , wherein the operations further comprise, in response to receiving the selection indication, displaying a keyboard in the graphical user interface.

14. The user device of claim 13 , wherein the user uses the keyboard displayed in the graphical user interface to provide the typed input.

15. The user device of claim 10 , wherein the one or more text characters comprise one or more letters inputted by the user through the graphical user interface.

16. The user device of claim 10 , wherein the one of the plurality of recognized terms of the first transcription replaced by the one or more text characters inputted by the user through the graphical user interface comprises a mis-transcribed term.

17. The user device of claim 10 , wherein the one or more text characters inputted by the user through the graphical user interface spell out a replacement term.

18. The user device of claim 17 , wherein the operations further comprise isolating at least a portion of the voice input as training data for training a model to recognize the replacement term.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2024
From: KAPRALOVA, OLGA; CHEREPANOV, EVGENY A.; OSMAKOV, DMITRY; BAEUMI, MARTIN; SKOBELTSYN, GLEB
To: GOOGLE, INC.
Reel/Frame 067014/0682 →
CHANGE OF NAME Recorded Apr 5, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 067025/0697 →
Continuity (3)
Continuation 16023658 · Jun 29, 2018
Continuation 15224104 · Jul 29, 2016
Related Publication 20200243070A1 · Jul 30, 2020