IP Library Granted Patent US 11,682,381
Granted Patent B2
US 11,682,381 · App. 17/457,421 · Granted Jun 20, 2023

Acoustic model training using corrected terms

Inventors: Olga Kapralova (Bern, CH); Evgeny A. Cherepanov (Adliswil, CH); Dmitry Osmakov (Zurich, CH); Martin Baeuml (Hedingen, CH); Gleb Skobeltsyn (Kilchberg, CH)
Assignee: Google LLC
G10L15/063G10L15/01G10L15/06G10L15/10G10L15/22G10L15/32G10L2015/0635G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,682,381
App. No.
17/457,421
Granted
Jun 20, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for speech recognition. One of the methods includes receiving first audio data corresponding to an utterance; obtaining a first transcription of the first audio data; receiving data indicating (i) a selection of one or more terms of the first transcription and (ii) one or more of replacement terms; determining that one or more of the replacement terms are classified as a correction of one or more of the selected terms; in response to determining that the one or more of the replacement terms are classified as a correction of the one or more of the selected terms, obtaining a first portion of the first audio data that corresponds to one or more terms of the first transcription; and using the first portion of the first audio data that is associated with the one or more terms of the first transcription to train an acoustic model for recognizing the one or more of the replacement terms.

Claims (32)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving training data comprising a corrected transcription including a specific term that replaced another term in an initial transcription of an utterance;

training, using the training data, a model to learn to recognize pronunciations of the specific term in voice inputs;

receiving a voice query comprising a pronunciation of the specific term;

transcribing, using a speech recognition engine, the voice query to a first recognized query comprising a plurality of recognized terms; and

generating, using the trained model, based on a phonetic similarity between the specific term and one of the plurality of recognized terms, a corrected recognized query that replaces the one of the plurality of recognized terms with the specific term.

2. The computer-implemented method of claim 1 , wherein the operations further comprise displaying, in a graphical user interface, the corrected recognized query.

3. The computer-implemented method of claim 1 , wherein the operations further comprise providing the corrected recognized query to a search engine, the search engine configured to generate search results that are determined to be responsive to the corrected recognized query.

4. The computer-implemented method of claim 3 , wherein the operations further comprise displaying, in a graphical user interface, the search results generated by the search engine that are determined to be responsive to the corrected recognized query.

5. The computer-implemented method of claim 4 , wherein displaying the search results comprises displaying, in the graphical user interface, links to resources determined to be responsive to the corrected recognized query.

6. The computer-implemented method of claim 4 , wherein search results displayed in the graphical user interface comprise the specific term.

7. The computer-implemented method of claim 1 , wherein receiving the voice query comprises receiving the voice query from a user device associated with a user.

8. The computer-implemented method of claim 1 , wherein the one of the plurality of recognized terms of the recognized query replaced by the specific term comprises a mis-transcribed term.

9. The computer-implemented method of claim 1 , wherein the specific term and the one of the recognized terms replaced by the specific term in the corrected recognized query comprise nouns.

10. The computer-implemented method of claim 1 , wherein the speech recognition engine comprises a language model.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

receiving training data comprising a corrected transcription including a specific term that replaced another term in an initial transcription of an utterance;

training, using the training data, a model to learn to recognize pronunciations of the specific term in voice inputs;

receiving a voice query comprising a pronunciation of the specific term;

transcribing, using a speech recognition engine, the voice query to a first recognized query comprising a plurality of recognized terms; and

generating, using the trained model, based on a phonetic similarity between the specific term and one of the plurality of recognized terms, a corrected recognized query that replaces the one of the plurality of recognized terms with the specific term.

12. The system of claim 11 , wherein the operations further comprise displaying, in a graphical user interface, the corrected recognized query.

13. The system of claim 11 , wherein the operations further comprise providing the corrected recognized query to a search engine, the search engine configured to generate search results that are determined to be responsive to the corrected recognized query.

14. The system of claim 13 , wherein the operations further comprise displaying, in a graphical user interface, the search results generated by the search engine that are determined to be responsive to the corrected recognized query.

15. The system of claim 14 , wherein displaying the search results comprises displaying, in the graphical user interface, links to resources determined to be responsive to the corrected recognized query.

16. The system of claim 14 , wherein search results displayed in the graphical user interface comprise the specific term.

17. The system of claim 11 , wherein receiving the voice query comprises receiving the voice query from a user device associated with a user.

18. The system of claim 11 , wherein the one of the plurality of recognized terms of the recognized query replaced by the specific term comprises a mis-transcribed term.

19. The system of claim 11 , wherein the specific term and the one of the recognized terms replaced by the specific term in the corrected recognized query comprise nouns.

20. The system of claim 11 , wherein the speech recognition engine comprises a language model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2024
From: KAPRALOVA, OLGA; CHEREPANOV, EVGENY A.; OSMAKOV, DMITRY; BAEUMI, MARTIN; SKOBELTSYN, GLEB
To: GOOGLE, INC.
Reel/Frame 067014/0682 →
CHANGE OF NAME Recorded Apr 5, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 067025/0697 →
Continuity (4)
Continuation 16837393 · Apr 1, 2020
Continuation 16023658 · Jun 29, 2018
Continuation 15224104 · Jul 29, 2016
Related Publication 20220093080A1 · Mar 24, 2022