IP Library Granted Patent US 11,037,566
Granted Patent B2
US 11,037,566 · App. 16/854,670 · Granted Jun 15, 2021

Word-level correction of speech input

Inventors: Michael J. Lebeau (New York, NY); William J. Byrne (Davis, CA); John Nicholas Jitkoff (Palo Alto, CA); Brandon M. Ballinger (San Franciso, CA); Trausti T. Kristjansson (Mountain View, CA)
Assignee: Google LLC
G10L15/22G06F3/0482G06F3/04842G06F3/04886G06F40/137G06F40/166G06F40/232G06F40/284G10L15/01G10L15/26G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,037,566
App. No.
16/854,670
Granted
Jun 15, 2021
Kind
B2
Abstract

The subject matter of this specification can be implemented in, among other things, a computer-implemented method for correcting words in transcribed text including receiving speech audio data from a microphone. The method further includes sending the speech audio data to a transcription system. The method further includes receiving a word lattice transcribed from the speech audio data by the transcription system. The method further includes presenting one or more transcribed words from the word lattice. The method further includes receiving a user selection of at least one of the presented transcribed words. The method further includes presenting one or more alternate words from the word lattice for the selected transcribed word. The method further includes receiving a user selection of at least one of the alternate words. The method further includes replacing the selected transcribed word in the presented transcribed words with the selected alternate word.

Claims (48)

1. A method comprising:

presenting, by data processing hardware of a mobile computing device, in a region of a display of the mobile computing device, a first transcription of an utterance spoken by a user of the mobile computing device;

receiving, at the data processing hardware, a user input indication indicating selection of the first transcription to correct at least one incorrect word in the first transcription; and

in response to receiving the user input indication indicating selection of the first transcription, presenting, by the data processing hardware, in the region of the display of the mobile computing device, a second transcription of the utterance spoken by the user of the mobile computing device,

wherein the second transcription of the utterance comprises one or more alternate words substituted for the at least one incorrect word in the first transcription.

2 . The method of claim 1 , further comprising, prior to presenting the first transcription of the utterance:

receiving, at the data processing hardware, audio data corresponding to the utterance spoken by the user of the mobile computing device;

transmitting, by the data processing hardware, the audio data to a server-based, automated speech recognizer in communication with the mobile computing device via a network; and

obtaining, by the data processing hardware, the first and second transcriptions from the server-based, automated speech recognizer.

3. The method of claim 2 , wherein:

the first transcription of the utterance comprises one or more words from a word lattice; and

the second transcription of the utterance comprises the one or more alternate words from the word lattice that are substituted for the at least one incorrect word in the first transcription.

4. The method of claim 3 , wherein the word lattice comprises nodes corresponding to words of the first transcription of the utterance and words of the second transcription of the utterance, and edges between the nodes that identify possible paths through the word lattice, each path having an associated probability of being correct.

5. The method of claim 2 , wherein the first transcription of the utterance corresponds to a recognition result from the server-based, automated speech recognizer that has a highest speech recognition confidence score.

6 . The method of claim 2 , wherein the second transcription of the utterance represents an alternative recognition result to the first transcription from the server-based, automated speech recognizer.

7. The method of claim 1 , wherein:

the display of the mobile computing device comprises a touch-sensitive display; and

the user input indication indicating selection of the first transcription of the utterance comprises a user input selecting the first transcription presented in the region of the touch-sensitive display.

8. The method of claim 1 , further comprising:

Obtaining, by the data processing hardware, a word lattice based on automated speech recognition performed on audio data corresponding to the utterance spoken by the user; and

selecting, by the data processing hardware, the first transcription based on the word lattice.

9. The method of claim 8 , further comprising selecting, by the data processing hardware, the second transcription based on the word lattice.

10. The method of claim 1 , wherein presenting the second transcription of the utterance in the region of the display of the mobile computing device comprises presenting the second transcription of the utterance without displaying an alternates list.

11. A mobile computing device comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

presenting, in a region of a display of the mobile computing device, a first transcription of an utterance spoken by a user of the mobile computing device;

receiving a user input indication indicating selection of the first transcription to correct at least one incorrect word in the first transcription; and

in response to receiving the user input indication indicating selection of the first transcription, presenting, in the region of the display of the mobile computing device, a second transcription of the utterance spoken by the user of the mobile computing device,

wherein the second transcription of the utterance comprises one or more alternate words substituted for the at least one incorrect word in the first transcription.

12. The mobile computing device of claim 11 , wherein the operations further comprise, prior to presenting the first transcription of the utterance:

receiving audio data corresponding to the utterance spoken by the user of the mobile computing device;

transmitting the audio data to a server-based, automated speech recognizer in communication with the mobile computing device via a network; and

obtaining the first and second transcriptions from the server-based, automated speech recognizer.

13. The mobile computing device of claim 12 , wherein:

the first transcription of the utterance comprises one or more words from a word lattice; and

the second transcription of the utterance comprises the one or more alternate words from the word lattice that are substituted for the at least one incorrect word in the first transcription.

14. The mobile computing device of claim 13 , wherein the word lattice comprises nodes corresponding to words of the first transcription of the utterance and words of the second transcription of the utterance, and edges between the nodes that identify possible paths through the word lattice, each path having an associated probability of being correct.

15. The mobile computing device of claim 12 , wherein the first transcription of the utterance corresponds to a recognition result from the server-based, automated speech recognizer that has a highest speech recognition confidence score.

16. The mobile computing device of claim 12 , wherein the second transcription of the utterance represents an alternative recognition result to the first transcription from the server-based, automated speech recognizer.

17. The mobile computing device of claim 11 , wherein:

the display of the mobile computing device comprises a touch-sensitive display; and

the user input indication indicating selection of the first transcription of the utterance comprises a user input selecting the first transcription presented in the region of the touch-sensitive display.

18. The mobile computing device of claim 11 , wherein the operations further comprise:

Obtaining, by the data processing hardware, a word lattice based on automated speech recognition performed on audio data corresponding to the utterance spoken by the user; and

selecting, by the data processing hardware, the first transcription based on the word lattice.

19. The mobile computing device of claim 18 , wherein the operations further comprise selecting, by the data processing hardware, the second transcription based on the word lattice.

20. The mobile computing device of claim 11 , wherein presenting the second transcription of the utterance in the region of the display of the mobile computing device comprises presenting the second transcription of the utterance without displaying an alternates list.

Continuity (10)
Continuation 15849967 · Dec 21, 2017
Continuation 15608110 · May 30, 2017
Continuation 15350309 · Nov 14, 2016
Continuation 15045571 · Feb 17, 2016
Continuation 14988201 · Jan 5, 2016
Continuation 14747306 · Jun 23, 2015
Continuation 13947284 · Jul 22, 2013
Continuation 12913407 · Oct 27, 2010
Provisional Application 61292440 · Jan 5, 2010
Related Publication 20200251113A1 · Aug 6, 2020
Cited By (1)
US 12,334,060