IP Library Granted Patent US 9,117,450
Granted Patent B2
US 9,117,450 · App. 13/712,032 · Granted Aug 25, 2015

Combining re-speaking, partial agent transcription and ASR for improved accuracy / human guided ASR

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,117,450
App. No.
13/712,032
Granted
Aug 25, 2015
Kind
B2
Abstract

A speech transcription system is described for producing a representative transcription text from one or more different audio signals representing one or more different speakers participating in a speech session. A preliminary transcription module develops a preliminary transcription of the speech session using automatic speech recognition having a preliminary recognition accuracy performance. A speech selection module enables user selection of one or more portions of the preliminary transcription to receive higher accuracy transcription processing. A final transcription module is responsive to the user selection for developing a final transcription output for the speech session having a final recognition accuracy performance for the selected one or more portions which is higher than the preliminary recognition accuracy performance.

Claims (13)

1. A speech transcription system for producing a representative transcription text from one or more audio signals representing one or more speakers participating in a speech session, the system comprising:

a preliminary transcription module for developing a preliminary transcription of the speech session using automatic speech recognition having a preliminary recognition accuracy performance;

a speech selection module for user selection of one or more portions of the preliminary transcription to receive higher accuracy transcription processing; and

a final transcription module responsive to the user selection for developing a final transcription output for the speech session having a final recognition accuracy performance for the selected one or more portions which is higher than the preliminary recognition accuracy performance.

2. The system according to claim 1 , wherein the preliminary transcription module develops the preliminary transcription automatically.

3. The system according to claim 1 , wherein the preliminary transcription module develops the preliminary transcription with human assistance.

4. The system according to claim 1 , wherein the speech selection module makes the user selection based on one more specified selection rules.

5. The system according to claim 1 , wherein the speech selection module makes the user selection by manual user selection.

6. The system according to claim 1 , wherein the final transcription module develops the final transcription output using manual transcription.

7. The system according to claim 1 , wherein the final transcription module develops the final transcription output using automatic speech recognition.

8. The system according to claim 7 , wherein the final transcription module uses transcription models adapted from information developed by the preliminary transcription module.

9. The system according to claim 7 , wherein the final transcription module uses transcription modules adapted from one or more different speakers.

10. The system according to claim 1 , wherein there are a plurality of different audio signals representing a plurality of different speakers participating in the speech session.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2012
From: COOK, GARY DAVID; GANONG, WILLIAM F., III; DABORN, ANDREW JOHNATHON
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 029491/0266 →