IP Library Granted Patent US 8,918,317
Granted Patent B2
US 8,918,317 · App. 12/566,785 · Granted Dec 23, 2014

Decoding-time prediction of non-verbalized tokens

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,918,317
App. No.
12/566,785
Granted
Dec 23, 2014
Kind
B2
Abstract

Non-verbalized tokens, such as punctuation, are automatically predicted and inserted into a transcription of speech in which the tokens were not explicitly verbalized. Token prediction may be integrated with speech decoding, rather than performed as a post-process to speech decoding.

Claims (40)

1. A method performed by at least one computer processor executing computer program instructions stored on at least one non-transitory computer-readable medium, the method comprising:

(A) using a language model to decode, by a speech decoder executed by the at least one computer processor, a first portion of an audio signal into a first word in a token stream;

(B) after (A):

(B)(1) select a punctuation mark based on the language model and the first word; and

(B)(2) inserting, by the speech decoder, the punctuation mark into the token stream at a position after the first word, based on a lexical cue provided by the first word; and

(C) after (B), using the language model, the punctuation mark, and the first word to decode, by the speech decoder, a second portion of the audio signal into a second word in the token stream at a position after the punctuation mark.

2. The method of claim 1 , wherein the token stream comprises a document, wherein the first word comprises first text in the document, wherein the punctuation mark comprises second text in the document, and wherein the second word comprises third text in the document.

3. The method of claim 1 , wherein (B) comprises inserting the punctuation mark at a position immediately after the first word in the token stream, and wherein (C) comprises inserting the second word at a position immediately after the punctuation mark in the token stream.

4. The method of claim 1 , wherein the first portion and second portion of the audio signal are contiguous within the audio signal.

5. The method of claim 1 , wherein (B)(1) comprises selecting the punctuation mark based on the language model, the first word, and at least one additional word before the first word in the token stream.

6. The method of claim 1 , wherein (C) comprises using the language model, the punctuation mark, the first word, and at least one additional word before the first word in the text to decode the second portion of the audio signal into the second word.

7. The method of claim 1 , wherein (B)(1) comprises selecting the punctuation mark without using acoustic evidence from the audio signal.

8. The method of claim 1 , wherein (B)(1) comprises selecting the punctuation mark using the language model and without using an acoustic model.

9. The method of claim 1 , further comprising:

(D) before (A), training the language model using a document corpus, wherein the document corpus includes words and punctuation.

10. The method of claim 1 , further comprising:

(D) creating a data structure containing the first word at a first position and the second word at a second position that is after the first position in the data structure, wherein the punctuation mark is not between the first position and the second position in the data structure.

11. The method of claim 10 , wherein (D) comprises:

(D)(1) inserting the punctuation mark between the first position and the second position in the data structure; and

(D)(2) marking the punctuation mark as hidden within the data structure.

12. The method of claim 10 , wherein (D) comprises:

(D)(1) inserting the punctuation mark between the first position and the second position in the data structure; and

(D)(2) removing the punctuation mark from the data structure.

13. The method of claim 1 , wherein (B) comprises inserting the punctuation mark into the token stream at the position after the first word, based on the lexical cue provided by the first word and an acoustical cue.

14. The method of claim 1 , wherein (B) comprises inserting the punctuation mark into the token stream at the position after the first word, based on the lexical cue provided by the first word and a prosodic cue.

15. A non-transitory computer program product tangibly storing computer program instructions executable by a computer processor to perform a method, the method comprising:

(A) using a language model to decode a first portion of an audio signal into a first word in a token stream;

(B) after (A):

(B)(1) selecting a punctuation mark based on the language model and the first word;

(B)(2) inserting the punctuation mark into the token stream at a position after the first word, based on a lexical cue provided by the first word; and

(C) after (B), using the language model, the punctuation mark, and the first word to decode a second portion of the audio signal into a second word in the token stream at a position after the punctuation mark.

16. The computer program product of claim 15 , wherein the token stream comprises a document, wherein the first word comprises first text in the document, wherein the punctuation mark comprises second text in the document, and wherein the second word comprises third text in the document.

17. The computer program product of claim 15 , wherein the first portion and second portion of the audio signal are contiguous within the audio signal.

18. The computer program product of claim 15 :

wherein (A) comprises using a language model to decode the first portion of the audio signal into the first word;

wherein (B) comprises using the language model to select the punctuation mark; and

wherein (C) comprises using the language model to decode the second portion of the audio signal into the second word.

19. The computer program product of claim 15 , wherein (B)(1) comprises selecting the punctuation mark without using acoustic evidence from the audio signal.

20. The computer program product of claim 15 , wherein (B) comprises inserting the punctuation mark into the token stream at the position after the first word, based on the lexical cue provided by the first word and an acoustical cue.

21. The computer program product of claim 15 , wherein (B) comprises inserting the punctuation mark into the token stream at the position after the first word, based on the lexical cue provided by the first word and a prosodic cue.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 22, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
To: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; MEDQUIST OF DELAWARE, INC.; MMODAL MQ INC.; MEDQUIST CM LLC
Reel/Frame 048411/0712 →
RELEASE OF SECURITY INTEREST Recorded Feb 1, 2019
From: CORTLAND CAPITAL MARKET SERVICES LLC, AS ADMINISTRATIVE AGENT
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 048210/0792 →
PATENT SECURITY AGREEMENT Recorded Oct 10, 2014
From: MULTIMODAL TECHNOLOGIES, LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 033958/0511 →
SECURITY AGREEMENT Recorded Oct 8, 2014
From: MMODAL IP LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 034047/0527 →
RELEASE OF SECURITY INTEREST Recorded Aug 1, 2014
From: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 033459/0987 →
CHANGE OF NAME Recorded Oct 14, 2011
From: MULTIMODAL TECHNOLOGIES, INC.
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 027061/0492 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2009
From: FRITSCH, JUERGEN; DEORAS, ANOOP; KOLL, DETLEF
To: MULTIMODAL TECHNOLOGIES, INC.
Reel/Frame 023507/0547 →