IP Library Granted Patent US 7,881,930
Granted Patent B2
US 7,881,930 · App. 11/767,537 · Granted Feb 1, 2011

ASR-aided transcription with segmented feedback training

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,881,930
App. No.
11/767,537
Granted
Feb 1, 2011
Kind
B2
Abstract

An ASR-aided transcription system with segmented feedback training is provided, the system including a transcription process manager configured to extract a first segment and a second segment from an audio input of speech uttered by a speaker, and an ASR engine configured to operate in a first speech recognition mode to convert the first speech segment into a first text transcript using a speaker-independent acoustic model and a speaker-independent language model, operate in a first training mode to create a speaker-specific acoustic model and a speaker-specific language model by adapting the speaker-independent acoustic model and the speaker-independent language model using either of the first segment and a corrected version of the first text transcript, and operate in a second speech recognition mode to convert the second speech segment into a second text transcript using the speaker-specific acoustic model and the speaker-specific language model.

Claims (55)

1. An ASR-aided transcription system with segmented feedback training, the system comprising:

a transcription process manager configured to extract a first segment and a second segment from an audio input of speech uttered by a speaker; and

an ASR engine configured to:

operate in a first speech recognition mode to convert said first speech segment into a first text transcript using a speaker-independent acoustic model and a speaker-independent language model;

operate in a first training mode to create a speaker-specific acoustic model and a speaker-specific language model by:

adapting said speaker-independent acoustic model using said first segment and/or a proofread version of said first text transcript; and

adapting said speaker-independent language model using at least the proofread version of said first text transcript; and

operate in a second speech recognition mode to convert said second speech segment into a second text transcript using said speaker-specific acoustic model and said speaker-specific language model.

2. A system according to claim 1 wherein said ASR engine is configured to operate in a second training mode to:

adapt said speaker-specific acoustic model using said second speech audio segment and/or a proofread version of said second text transcript; and

adapt said speaker-specific language model using at least said proofread version of said second text transcript.

3. A system according to claim 2 wherein said ASR engine is configured to alternately operate in said second speech recognition mode and said second training mode for processing a plurality of segments of said audio input subsequent to processing said first segment.

4. A system according to claim 2 wherein said ASR engine is configured to cyclically process multiple speech segments from multiple speakers, where said ASR engine uses a different set of said speaker-specific models for each of said speakers.

5. A system according to claim 1 wherein said ASR engine is configured when in either of said speech recognition modes to determine a plurality of word time offsets for words in either of said segments, and when in said training mode to adapt said speaker-specific models using said word time offsets.

6. A system according to claim 5 and further comprising an Acoustic Model Training Dataset Preparator operative to

receive any of said segments, said proofread transcript, and said word time offsets,

divide said received segment and said proofread transcript into a plurality of audio and text pieces, and

create a training dataset of corresponding pairs of said audio and text pieces,

wherein said ASR engine is configured to use said training dataset when in any of said training modes to adapt any of said speaker-specific acoustic models.

7. An ASR-aided transcription method with segmented feedback training, the method comprising:

a) extracting a first segment and a second segment from an audio input of speech uttered by a speaker;

b) converting said first speech segment into a first text transcript using a speaker-independent acoustic model and a speaker-independent language model;

c) creating a speaker-specific acoustic model by adapting said speaker-independent acoustic model using said first segment and/or a proofread version of said first text transcript;

d) creating a speaker-specific language model by adapting said speaker-independent language model using at least said proofread version of said first text transcript; and

e) converting said second speech segment into a second text transcript using said speaker-specific acoustic model and said speaker-specific language model.

8. A method according to claim 7 and further comprising

f) adapting said speaker-specific acoustic model using said second speech audio segment and/or a proofread version of said second text transcript; and

g) adapting said speaker-specific language model using at least said proofread version of said second text transcript.

9. A method according to claim 8 and further comprising cyclically performing said e) converting, said f) adapting and said g) adapting for processing a plurality of segments of said audio input subsequent to processing said first segment.

10. A method according to claim 8 and further comprising alternatingly processing multiple speech segments from multiple speakers using separate sets of said speaker-specific models for each of said speakers.

11. A method according to claim 7 and further comprising:

determining a plurality of word time offsets for words in either of said segments; and

adapting said speaker-specific models using said word time offsets.

12. A method according to claim 11 and further comprising:

dividing any of said segments and said proofread transcript corresponding to said segment into a plurality of audio and text pieces;

creating a training dataset of corresponding pairs of said audio and text pieces; and

adapting any of said speaker-specific acoustic models using said training dataset.

13. A computer-readable recordable device encoded with a plurality of instructions that, when executed by at least one computer, perform a method comprising:

extracting a first segment and a second segment from an audio input of speech uttered by a speaker;

converting said first speech segment into a first text transcript using a speaker-independent acoustic model and a speaker-independent language model;

creating a speaker-specific acoustic model by adapting said speaker-independent acoustic model using said first segment and/or a proofread version of said first text transcript;

creating a speaker-specific language model by adapting said speaker-independent language model using at least said proofread version of said first text transcript; and

converting said second speech segment into a second text transcript using said speaker-specific acoustic model and said speaker-specific language model.

14. The computer readable recording device of claim 13 , wherein the method further comprises:

adapting said speaker-specific acoustic model using at least said second speech audio segment and/or a proofread version of said second text transcript; and

adapting said speaker-specific language model using at least said proofread version of said second text transcript.

15. The computer readable recording device of claim 13 , wherein the method comprises cyclically processing a plurality of segments of said audio input subsequent to processing said first segment.

16. The computer readable recording device of claim 13 , wherein the method further comprises:

determining a plurality of word time offsets for words in either of said segments; and

adapting said speaker-specific acoustic model using said word time offsets.

17. The computer readable recording device of claim 13 , wherein the method further comprises:

dividing any of said segments and said proofread transcript corresponding to said segment into a plurality of audio and text pieces;

creating a training dataset of corresponding pairs of said audio and text pieces; and

adapting any of said speaker-specific acoustic models using said training dataset.

18. The computer readable recording device of claim 13 , wherein the method comprises cyclically processing multiple speech segments from multiple speakers using separate sets of said speaker-specific models for each of said speakers.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →