IP Library Granted Patent US 9,583,107
Granted Patent B2
US 9,583,107 · App. 14/517,720 · Granted Feb 28, 2017

Continuous speech transcription performance indication

Inventors: James Richard Terrell, II (Charlotte, NC); Marc White (Charlotte, NC); Igor Roditis Jablokov (Charlotte, NC)
Assignee: Amazon Technologies, Inc.
G10L15/26G10L15/01G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,583,107
App. No.
14/517,720
Granted
Feb 28, 2017
Kind
B2
Abstract

A method of providing speech transcription performance indication includes receiving, at a user device data representing text transcribed from an audio stream by an ASR system, and data representing a metric associated with the audio stream; displaying, via the user device, said text; and via the user device, providing, in user-perceptible form, an indicator of said metric. Another method includes displaying, by a user device, text transcribed from an audio stream by an ASR system; and via the user device, providing, in user-perceptible form, an indicator of a level of background noise of the audio stream. Another method includes receiving data representing an audio stream; converting said data representing an audio stream to text via an ASR system; determining a metric associated with the audio stream; transmitting data representing said text to a user device; and transmitting data representing said metric to the user device.

Claims (47)

1. A computer-implemented method comprising:

receiving first text, wherein the first text was created by performing automatic speech recognition on a first portion of audio data, wherein the audio data includes a plurality of portions;

receiving a first value associated with the first text;

causing presentation of the first text with a first graphical element indicating the first value;

receiving, during presentation of the first text, second text, wherein the second text was created by performing automatic speech recognition on a second portion of the plurality of portions of audio data and wherein the second portion of the audio data is subsequent to the first portion of the audio data; and

causing presentation of the second text.

2. The computer-implemented method of claim 1 , wherein the first value indicates a volume level of a portion of the first portion of the audio data corresponding to the first text, a level of background noise in the portion of the first portion of the audio data or a confidence level associated with the first text as determined by an automatic speech recognition engine.

3. The computer-implemented method of claim 1 , wherein the first graphical element comprises font color, font grayscale, font weight, font size or underlining.

4. The computer-implemented method of claim 1 , wherein the first portion of the audio data comprises at least part of a voicemail message.

5. The computer-implemented method of claim 1 , further comprising:

receiving a second value associated with the second text; and

causing presentation of the second text with a second graphical element indicating the second value, wherein the second text comprises a modified version of the first text.

6. The computer-implemented method of claim 1 , further comprising:

receiving a third value associated with a third text, wherein the third text was created by performing automatic speech recognition on the first portion of the audio data; and

causing presentation of the third text with a third graphical element indicating the third value.

7. A system comprising:

an electronic data store configured to store transcription information; and

one or more computing devices in communication with the electronic data store, the one or more computing devices configured to at least:

receive first text, wherein the first text was created by performing automatic speech recognition on a first portion of audio data, wherein the audio data includes a plurality of portions;

receive a first value associated with the first text;

cause presentation of the first text with a first graphical element indicating the first value;

receive, during presentation of the first text, second text, wherein the second text was created by performing automatic speech recognition on a second portion of the plurality of portions of audio data and wherein the second portion of the audio data is subsequent to the first portion of the audio data; and

cause presentation of the second text.

8. The system of claim 7 , wherein the first graphical element comprises a lighter color if the first value is in a lower range, or a darker color if the first value is in a higher range.

9. The system of claim 8 , wherein the lighter color is gray and the darker color is black.

10. The system of claim 7 , wherein the first portion of the audio data comprises at least part of a voicemail message.

11. The system of claim 7 , wherein the one or more computing devices are further configured to:

receive a second value associated with the second text; and

cause presentation of the second text with a second graphical element indicating the second value, wherein the second text comprises a modified version of the first text.

12. The system of claim 7 , wherein the one or more computing devices are further configured to:

receive a third value associated with third text, wherein the third text was created by performing automatic speech recognition on the first portion of the audio data; and

cause presentation of the third text with a third graphical element indicating the third value.

13. The system of claim 12 , wherein the third value indicates a volume level of a portion of the first portion of the audio data corresponding to the third text, a level of background noise in the portion of the first portion of the audio data or a confidence level associated with the third text as determined by an automatic speech recognition engine.

14. A non-transitory computer-readable medium storing instructions that, when executed by a processor on a computing device, cause the computing device to at least:

receive first text, wherein the first text was created by performing automatic speech recognition on a first portion of audio data, wherein the audio data includes a plurality of portions;

receive a first value associated with the first text;

cause presentation of the first text with a first graphical element indicating the first value;

receive, during presentation of the first text, second text, wherein the second text was created by performing automatic speech recognition on a second portion of the plurality of portions of audio data and wherein the second portion of the audio data is subsequent to the first portion of the audio data; and

cause presentation of the second text.

15. The non-transitory computer-readable medium of claim 14 , wherein the first graphical element indicates a volume level associated with a portion of the first portion of the audio data corresponding to the first text and is presented substantially simultaneously with the presentation of the first text.

16. The non-transitory computer-readable medium of claim 14 , wherein the first graphical element comprises font color, font grayscale, font weight, font size or underlining.

17. The non-transitory computer-readable medium of claim 14 , wherein the first portion of the audio data comprises at least part of a voicemail message.

18. The non-transitory computer-readable medium of claim 14 , further comprising instructions to filter the first text by replacing one or more words in the first text with corresponding numbers or digits.

19. The non-transitory computer-readable medium of claim 14 , wherein the first portion of the audio data is captured at a first device and the first text is presented at a second device.

20. The non-transitory computer-readable medium of claim 14 , further comprising instructions to:

receive a second value associated with the second text; and

cause presentation of the second text with a second graphical element indicating the second value, wherein the second text comprises a modified version of the first text.

Assignments (2)
ENTITY CONVERSION Recorded Nov 4, 2016
From: YAP INC.
To: YAP LLC
Reel/Frame 040567/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2015
From: CANYON IP HOLDINGS LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 037083/0914 →
Continuity (21)
Continuation 14010433 · Aug 26, 2013
Continuation 13621179 · Sep 15, 2012
Continuation 12197213 · Aug 22, 2008
Provisional Application 60957386 · Aug 22, 2007
Provisional Application 60957393 · Aug 22, 2007
Provisional Application 60957701 · Aug 23, 2007
Provisional Application 60957702 · Aug 23, 2007
Provisional Application 60957706 · Aug 23, 2007
Provisional Application 60972851 · Sep 17, 2007
Provisional Application 60972853 · Sep 17, 2007
Provisional Application 60972854 · Sep 17, 2007
Provisional Application 60972936 · Sep 17, 2007
Provisional Application 60972943 · Sep 17, 2007
Provisional Application 60972944 · Sep 17, 2007
Provisional Application 61016586 · Dec 25, 2007
Provisional Application 61021335 · Jan 16, 2008
Provisional Application 61021341 · Jan 16, 2008
Provisional Application 61034815 · Mar 7, 2008
Provisional Application 61038046 · Mar 19, 2008
Provisional Application 61042219 · Mar 31, 2008
Related Publication 20160027443A1 · Jan 28, 2016