IP Library Granted Patent US 10,672,394
Granted Patent B2
US 10,672,394 · App. 15/849,967 · Granted Jun 2, 2020

Word-level correction of speech input

Inventors: Michael J. LeBeau (New York, NY); William J. Byrne (Davis, CA); John Nicholas Jitkoff (Palo Alto, CA); Brandon M. Ballinger (San Francisco, CA); Trausti T. Kristjansson (Mountain View, CA)
Assignee: Google LLC
G10L15/22G06F3/0482G06F3/04842G06F3/04886G06F40/137G06F40/166G06F40/232G06F40/284G10L15/01G10L15/26G10L15/265G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,672,394
App. No.
15/849,967
Granted
Jun 2, 2020
Kind
B2
Abstract

The subject matter of this specification can be implemented in, among other things, a computer-implemented method for correcting words in transcribed text including receiving speech audio data from a microphone. The method further includes sending the speech audio data to a transcription system. The method further includes receiving a word lattice transcribed from the speech audio data by the transcription system. The method further includes presenting one or more transcribed words from the word lattice. The method further includes receiving a user selection of at least one of the presented transcribed words. The method further includes presenting one or more alternate words from the word lattice for the selected transcribed word. The method further includes receiving a user selection of at least one of the alternate words. The method further includes replacing the selected transcribed word in the presented transcribed words with the selected alternate word.

Claims (37)

1. A computer-implemented method comprising:

providing a transcription of an utterance in a region of a touch-sensitive display;

receiving data indicating single touch selection of a particular word in the transcription of the utterance in the region of the touch-sensitive display;

determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a particular type of touch selection; and

in response to determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is the particular type of touch selection, providing an updated transcription of the utterance in the region of the touch-sensitive display.

2. The method of claim 1 , wherein determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a particular type of touch selection comprises determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a long press touch selection.

3. The method of claim 1 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing a new transcription of the utterance in which the particular word is removed.

4. The method of claim 1 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing a new transcription of the utterance in which the particular word is replaced with a different word.

5. The method of claim 1 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing the updated transcription of the utterance without displaying an alternates list.

6. The method of claim 1 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises a new transcription in which a multi-word phrase that includes the particular word is replaced with a different multi-word phrase that does not include the particular word.

7. The method of claim 1 , comprising:

obtaining a word lattice based on performing automated speech recognition on audio data; and

selecting the transcription based on the word lattice.

8. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

providing a transcription of an utterance in a region of a touch-sensitive display;

receiving data indicating single touch selection of a particular word in the transcription of the utterance in the region of the touch-sensitive display;

determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a particular type of touch selection; and

in response to determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is the particular type of touch selection, providing an updated transcription of the utterance in the region of the touch-sensitive display.

9. The system of claim 8 , wherein determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a particular type of touch selection comprises determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a long press touch selection.

10. The system of claim 8 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing a new transcription of the utterance in which the particular word is removed.

11. The system of claim 8 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing a new transcription of the utterance in which the particular word is replaced with a different word.

12. The system of claim 8 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing the updated transcription of the utterance without displaying an alternates list.

13. The system of claim 8 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises a new transcription in which a multi-word phrase that includes the particular word is replaced with a different multi-word phrase that does not include the particular word.

14. The system of claim 8 , wherein the operations comprise:

obtaining a word lattice based on performing automated speech recognition on audio data; and

selecting the transcription based on the word lattice.

15. A computer-readable non-transitory medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

providing a transcription of an utterance in a region of a touch-sensitive display;

receiving data indicating single touch selection of a particular word in the transcription of the utterance in the region of the touch-sensitive display;

determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a particular type of touch selection; and

in response to determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is the particular type of touch selection, providing an updated transcription of the utterance in the region of the touch-sensitive display.

16. The medium of claim 15 , wherein determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a particular type of touch selection comprises determining that the single touch selection of the particular word in the transcription of the utterance in the region of the touch-sensitive display is a long press touch selection.

17. The medium of claim 15 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing a new transcription of the utterance in which the particular word is removed.

18. The medium of claim 15 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing a new transcription of the utterance in which the particular word is replaced with a different word.

19. The medium of claim 15 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises providing the updated transcription of the utterance without displaying an alternates list.

20. The medium of claim 15 , wherein providing the updated transcription of the utterance in the region of the touch-sensitive display comprises a new transcription in which a multi-word phrase that includes the particular word is replaced with a different multi-word phrase that does not include the particular word.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2017
From: LEBEAU, MICHAEL J.; BYRNE, WILLIAM J.; JITKOFF, JOHN NICHOLAS; BALLINGER, BRANDON M.; KRISTJANSSON, TRAUSTI T.
To: GOOGLE INC.
Reel/Frame 044461/0249 →
CHANGE OF NAME Recorded Dec 21, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044938/0822 →
Continuity (9)
Continuation 15608110 · May 30, 2017
Continuation 15350309 · Nov 14, 2016
Continuation 15045571 · Feb 17, 2016
Continuation 14988201 · Jan 5, 2016
Continuation 14747306 · Jun 23, 2015
Continuation 13947284 · Jul 22, 2013
Continuation 12913407 · Oct 27, 2010
Provisional Application 61292440 · Jan 5, 2010
Related Publication 20180114530A1 · Apr 26, 2018
Cited By (1)
US 12,300,217