IP Library Granted Patent US 8,209,176
Granted Patent B2
US 8,209,176 · App. 13/219,640 · Granted Jun 26, 2012

System and method for latency reduction for automatic speech recognition using partial multi-pass results

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,209,176
App. No.
13/219,640
Granted
Jun 26, 2012
Kind
B2
Abstract

A system and method is provided for reducing latency for automatic speech recognition. In one embodiment, intermediate results produced by multiple search passes are used to update a display of transcribed text.

Claims (40)

1. A method comprising:

transcribing, via a processor, speech data using a first automatic speech recognition pass, which operates at a first transcription rate, to produce a first transcription data and a first word graph;

displaying a displayed part comprising an indication of a second automatic speech recognition pass which is forthcoming and at least part of the first transcription data corresponding to a portion of the speech data;

after displaying the displayed part, transcribing the speech data using the second automatic speech recognition pass, wherein the second automatic speech recognition pass uses the first word graph to produce second transcription data and a second word graph, and wherein the second automatic speech recognition pass is slower than the first automatic speech recognition pass; and

upon completing the second automatic speech recognition pass, updating the displayed part based at least in part on the second transcription data.

2. The method of claim 1 , wherein the first automatic speech recognition pass operates at real time.

3. The method of claim 1 , wherein the first automatic speech recognition pass operates faster than real time.

4. The method of claim 1 , wherein the indication signifies that more additional transcription data is being generated.

5. The method of claim 1 , wherein the displayed part changes color upon updating.

6. The method of claim 1 , wherein a low confidence portion of the displayed part is distinctly displayed as compared to a high confidence portion of the displayed part.

7. The method of claim 6 , wherein the low confidence portion of the displayed part is displayed in a darker shade as compared to the high confidence portion of the displayed data.

8. The method of claim 1 , further comprising:

adapting an additional model for a third automatic speech recognition pass that uses the second word graph, wherein the third automatic speech recognition pass produces a third transcription data and a third word graph and wherein the third automatic speech recognition pass is slower than the second automatic speech recognition pass; and

updating the displayed part with at least the third transcription data.

9. A system comprising:

a processor; and

a non-transitory computer-readable memory storing instructions which, when executed by the processor, cause the processor to perform a method comprising:

transcribing speech data using a first automatic speech recognition pass, which operates at a first transcription rate, to produce a first transcription data and a first word graph;

displaying a displayed part comprising an indication of a second automatic speech recognition pass which is forthcoming and at least part of the first transcription data corresponding to a portion of the speech data

after displaying the displayed part, transcribing the speech data using the second automatic speech recognition pass, wherein the second automatic speech recognition pass uses the first word graph, wherein the second automatic speech recognition pass produces a second transcription data and a second word graph, and wherein the second automatic speech recognition pass is slower than the first automatic speech recognition pass;

and

upon completing the second automatic speech recognition pass, updating the displayed part based at least in part on the second transcription data.

10. The system of claim 9 , wherein the first automatic speech recognition pass operates at real time.

11. The system of claim 9 , wherein the first automatic speech recognition pass operates faster than real time.

12. The system of claim 9 , wherein the indication that signifies that more additional transcription data is being generated.

13. The system of claim 9 , wherein the displayed part changes color upon updating.

14. The system of claim 9 , wherein a low confidence portion of the displayed part is distinctly displayed as compared to a high confidence portion of the displayed part.

15. The system of claim 14 , wherein the low confidence portion of the displayed part is displayed in a darker shade as compared to the high confidence portion of the displayed data.

16. The system of claim 9 , the instructions further comprising:

adapting an additional model for a third automatic speech recognition pass that uses the second word graph, wherein the third automatic speech recognition pass produces a third transcription data and a third word graph and wherein the third automatic speech recognition pass is slower than the second automatic speech recognition pass; and

updating the displayed part with at least the third transcription data.

17. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing device, cause the computing device to perform a method comprising:

transcribing speech data using a first automatic speech recognition pass, which operates at a first transcription rate, to produce a first transcription data and a first word graph;

displaying a displayed part comprising an indication of a second automatic speech recognition pass which is forthcoming and at least part of the first transcription data corresponding to a portion of the speech data

after displaying the displayed part, transcribing the speech data using the second automatic speech recognition pass, wherein the second automatic speech recognition pass uses the first word graph, wherein the second automatic speech recognition pass produces a second transcription data and a second word graph, and wherein the second automatic speech recognition pass is slower than the first automatic speech recognition pass;

and

upon completing the second automatic speech recognition pass, updating the displayed part based at least in part on the second transcription data.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the first automatic speech recognition pass operates at real time.

19. The non-transitory computer-readable storage medium of claim 17 , wherein the first automatic speech recognition pass operates faster than real time.

20. The non-transitory computer-readable storage medium of claim 17 , wherein the indication signifies that more additional transcription data is being generated.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2011
From: BACCHIANI, MICHIEL ADRIAAN UNICO; AMENTO, BRIAN SCOTT
To: AT&T CORP.
Reel/Frame 026818/0125 →