IP Library Granted Patent US 8,335,688
Granted Patent B2
US 8,335,688 · App. 10/922,513 · Granted Dec 18, 2012

Document transcription system training

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,335,688
App. No.
10/922,513
Granted
Dec 18, 2012
Kind
B2
Abstract

A system is provided for training an acoustic model for use in speech recognition. In particular, such a system may be used to perform training based on a spoken audio stream and a non-literal transcript of the spoken audio stream. Such a system may identify text in the non-literal transcript which represents concepts having multiple spoken forms. The system may attempt to identify the actual spoken form in the audio stream which produced the corresponding text in the non-literal transcript, and thereby produce a revised transcript which more accurately represents the spoken audio stream. The revised, and more accurate, transcript may be used to train the acoustic model, thereby producing a better acoustic model than that which would be produced using conventional techniques, which perform training based directly on the original non-literal transcript.

Claims (82)

1. In a system including a first document, the document tangibly stored in a computer-readable medium and containing at least some information in common with a spoken audio stream, a method performed by a computer processor executing instructions tangibly stored in a first computer-readable medium, the method comprising steps of:

(A) identifying text tangibly stored in the first document on a second computer-readable medium, wherein the text represents a concept;

(B) identifying, based on the identified text, a plurality of at least three spoken forms of the concept, including at least one spoken form not contained in the first document, wherein all of the plurality of spoken forms have the same content as each other, wherein (B) comprises:

(B) (1) identifying a name of the identified text; and

(B) (2) using the identified name to identify a corresponding context-free grammar in a grammar repository, wherein the corresponding context-free grammar specifies the plurality of spoken forms of the concept;

(C) replacing the identified text with the corresponding context-free grammar to produce a second document tangibly stored in a third computer-readable medium; and

(D) generating a first language model, tangibly stored in a fourth computer-readable medium, based on the second document.

2. The method of claim 1 , wherein the concept comprises a semantic concept.

3. The method of claim 2 , wherein the concept comprises a date.

4. The method of claim 1 , wherein the concept comprises a syntactic concept.

5. The method of claim 4 , wherein the concept comprises a sentence.

6. The method of claim 4 , wherein the concept comprises the entire second document.

7. The method of claim 1 , further comprising a step of:

(E) using the first language model in a speech recognition process to recognize the spoken audio stream and thereby to produce a third document tangibly stored in a fifth computer-readable medium.

8. The method of claim 7 , further comprising a step of:

(F) using the third document and the spoken audio stream to train an acoustic model tangibly stored in a sixth computer-readable medium.

9. The method of claim 8 , wherein the step (F) comprises steps of:

(F)(1) filtering text from the third document by reference to the second document to produce a filtered document tangibly stored in a seventh computer-readable medium; and

(F)(2) using the filtered document and the spoken audio stream to train the acoustic model.

10. The method of claim 9 , wherein the step (F)(1) comprises applying a robust parser to the second and third documents to produce the filtered document.

11. The method of claim 7 , wherein the step (E) comprises steps of:

(E) (1) interpolating the first language model with a second language model to produce a third language model tangibly stored in a sixth computer-readable medium; and

(E) (2) using the third language model in the speech recognition process to recognize the spoken audio stream and thereby to produce the third document.

12. The method of claim 1 , wherein the first document comprises a document generated based on the spoken audio stream.

13. The method of claim 1 , further comprising a step of:

(E) prior to step (D), normalizing the second document to produce a normalized document tangibly stored in a fifth computer-readable medium.

14. The method of claim 1 , further comprising a step of:

(E) prior to step (D), repeating steps (A), (B), and (C) for each of a plurality of texts in the first document.

15. The method of claim 1 , wherein the step (C) comprises steps of:

(C)(1) generating probabilities for the plurality of spoken forms specified by the context-free grammar; and

(C)(2) including the probabilities in the context-free grammar.

16. The method of claim 15 , wherein the step (C) further comprises a step of:

(C)(3) including the plurality of spoken forms in the context-free grammar.

17. The method of claim 1 , wherein the context-free grammar comprises a finite state grammar.

18. The method of claim 1 , wherein the first document comprises one of a first plurality of documents tangibly stored in the second computer-readable medium, and wherein the method further comprises a step of:

(E) repeating steps (A), (B), and (C) for each of the plurality of documents to produce a second plurality of documents, including the second document, tangibly stored in a fifth computer-readable medium; and

wherein the step (D) comprises a step of generating the first language model based on the second plurality of documents.

19. The method of claim 1 , wherein all of the plurality of spoken forms have the same semantic meaning as each other.

20. A system comprising:

a first computer-readable medium tangibly storing a first document containing at least some information in common with a spoken audio stream;

a second computer-readable medium tangibly storing computer program instructions for identifying text in the first document representing a concept;

a third computer-readable medium tangibly storing computer program instructions for identifying, based on the identified text, a plurality of at least three spoken forms of the concept, including at least one spoken form not contained in the first document, wherein all of the plurality of spoken forms have the same content as each other, wherein identifying the plurality of spoken forms of the concept comprises:

identifying a name of the identified text; and

using the identified name to identify a corresponding context-free grammar in a grammar repository, wherein the corresponding context-free grammar specifies the plurality of spoken forms of the concept;

a fourth computer-readable medium tangibly storing computer program instructions for replacing the identified text with the corresponding context-free grammar to produce a second document tangibly stored on a fifth computer-readable medium; and

a sixth computer-readable medium tangibly storing computer program instructions for generating a first language model based on the second document.

21. The system of claim 20 , wherein the concept comprises a syntactic concept.

22. The system of claim 20 , further comprising:

a seventh computer-readable medium tangibly storing computer program instructions for using the first language model in a speech recognition process to recognize the spoken audio stream and thereby to produce a third document tangibly stored in an eighth computer-readable medium.

23. The system of claim 22 , further comprising:

a ninth computer-readable medium tangibly storing computer program instructions for using the third document and the spoken audio stream to train an acoustic model tangibly stored in a tenth computer-readable medium.

24. The system of claim 22 , wherein the means for using the first language model comprises:

means for interpolating the first language model with a second language model to produce a third language model tangibly stored in a ninth computer-readable medium; and

means for using the third language model in the speech recognition process to recognize the spoken audio stream and thereby to produce the third document.

25. The system of claim 20 , wherein the first document comprises a document generated based on the spoken audio stream.

26. The system of claim 20 , further comprising:

means for normalizing the second document to produce a normalized document tangibly stored in a seventh computer-readable medium.

27. The system of claim 20 , wherein the first document comprises one of a first plurality of documents tangibly stored in the first computer-readable medium, and wherein the system further comprises:

a seventh computer-readable medium comprising tangibly storing computer program instructions for repeatedly activating the instructions for identifying text, the instructions for identifying the plurality of spoken forms, and the instructions for replacing the identified text for each of the plurality of documents to produce a second plurality of documents, tangibly stored in the fifth computer-readable medium, including the second document; and

wherein the means for generating the first language model comprises means for generating the first language model based on the second plurality of documents.

28. The system of claim 20 , wherein the concept comprises a semantic concept.

29. The system of claim 20 , wherein all of the plurality of spoken forms have the same semantic meaning as each other.

30. A method performed by a computer processor executing instructions tangibly stored in a first computer-readable medium, the method comprising steps of:

(A) identifying a first document, tangibly stored in a second computer-readable medium, containing at least some information in common with a spoken audio stream;

(B)

(B)(1) identifying text in the first document representing a concept;

(B)(2) identifying a name of the identified text;

(B)(3) using the identified name to identify a corresponding context-free grammar in a grammar repository, wherein the corresponding context-free grammar specifies a plurality of at least three spoken forms of the concept, and wherein the corresponding context-free grammar includes at least one spoken form not contained in the first document, wherein all of the plurality of spoken forms have the same content as each other;

(C) replacing the identified text with the corresponding context-free grammar to produce a second document tangibly stored in a third computer-readable medium; and

(D) replacing text in the second document with normalized text to produce a normalized document, tangibly stored in a fourth computer-readable medium, the normalized document including samples of text belonging to spoken forms in the finite state grammar.

31. The method of claim 30 , wherein the concept comprises a semantic concept.

32. The method of claim 30 , wherein the concept comprises a syntactic concept.

33. The method of claim 30 , wherein the context-free grammar comprises a finite state grammar.

34. The method of claim 30 , wherein all of the plurality of spoken forms have the same semantic meaning as each other.

35. A device comprising:

a first computer-readable medium tangibly storing computer program instructions for identifying a first document, tangibly stored in a second computer-readable medium, the first document containing at least some information in common with a spoken audio stream;

a third computer-readable medium tangibly storing computer program instructions for: (1) identifying text in the first document representing a concept; (2) identifying a name of the identified text; and (3) using the identified name to identify a corresponding context-free grammar in a grammar repository, wherein the corresponding context-free grammar specifies a plurality of at least three spoken forms of the concept, and wherein the corresponding context-free grammar includes at least one spoken form not contained in the first document, wherein all of the plurality of spoken forms have the same content as each other;

a fourth computer-readable medium tangibly storing computer program instructions for replacing the identified text with the corresponding context-free grammar to produce a second document tangibly stored in a fifth computer-readable medium; and

a sixth computer-readable medium tangibly storing computer program instructions for replacing text in the second document with normalized text to produce a normalized document, tangibly stored in a seventh computer-readable medium, the normalized document including samples of text belonging to spoken forms in the finite state grammar.

36. The device of claim 35 , wherein the concept comprises a semantic concept.

37. The device of claim 35 , wherein the concept comprises a syntactic concept.

38. The device of claim 35 , wherein all of the plurality of spoken forms have the same semantic meaning as each other.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: 3M INNOVATIVE PROPERTIES COMPANY
To: SOLVENTUM INTELLECTUAL PROPERTIES COMPANY
Reel/Frame 066435/0347 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2021
From: MMODAL IP LLC
To: 3M INNOVATIVE PROPERTIES COMPANY
Reel/Frame 057883/0129 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 22, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
To: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; MEDQUIST OF DELAWARE, INC.; MMODAL MQ INC.; MEDQUIST CM LLC
Reel/Frame 048411/0712 →
RELEASE OF SECURITY INTEREST Recorded Feb 1, 2019
From: CORTLAND CAPITAL MARKET SERVICES LLC, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 048211/0799 →
CHANGE OF ADDRESS Recorded Apr 14, 2017
From: MMODAL IP LLC
To: MMODAL IP LLC
Reel/Frame 042271/0858 →