IP Library Granted Patent US 7,996,223
Granted Patent B2
US 7,996,223 · App. 10/953,474 · Granted Aug 9, 2011

System and method for post processing speech recognition output

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,996,223
App. No.
10/953,474
Granted
Aug 9, 2011
Kind
B2
Abstract

A system and method may be disclosed for facilitating the conversion of dictation into usable and formatted documents by providing a method of post processing speech recognition output. In particular, the post processing system may be configured to implement rewrite rules and process raw speech recognition output or other raw data according to those rewrite rules. The application of the rewrite rules may format and/or normalize the raw speech recognition output into formatted or finalized documents and reports. The system may thereby reduce or eliminate the need for post processing by transcriptionists or dictation authors.

Claims (26)

1. A computer implemented method for altering output from a speech recognition engine, the method comprising:

inputting the output from a speech recognition engine into a post processor configured to convert the output, comprising speech-recognized dictation in the form of unformatted raw text or data, from the speech recognition engine to a formatted text document that is a transcription of the dictation, the post processor operatively configured to:

convert said output to a list of tokens;

identify a set of tokens from said list of tokens matching a first set of predetermined patterns; and

perform a set of rewrite rules based on said identified set of tokens, the rewrite rules used to transform the output of the speech recognition engine to the formatted text document that corresponds to the dictation, wherein the set of rewrite rules comprise:

identify section headings in the speech-recognized dictation;

format the section headings according to a standard practice of a specific site;

identify and transform number clusters adjacent to keywords in the speech-recognized dictation;

identify and transform number clusters not adjacent to keywords in the speech-recognized dictation;

perform consolidation or suppression of filled pauses and silences in the speech-recognized dictation;

perform attachment and format punctuation in the speech-recognized dictation; and

perform capitalization of words following punctuation in the speech-recognized dictation to thereby provide the formatted text document for writing to a storage device.

2. The method according to claim 1 , wherein a second set of rewrite rules may operate on the result of the first set of rewrite rules as well as the transformed number clusters.

3. The method according to claim 2 , further comprising performing at least a second rule interpretation performed to transform text that matches a second set of predetermined patterns.

4. The method according to claim 3 , further comprising converting punctuation tokens into symbols that cling to an adjacent word.

5. The method according to claim 4 , where the adjacent word is capitalized.

6. The method according to claim 5 , where filled pauses and silences are consolidated or suppressed and the boundaries between them and adjacent words are adjusted.

7. The method according to claim 6 , further comprising removing text used according to a set of post processor rules.

8. The method according to claim 7 , further comprising writing a file containing post-processed output to a target location.

9. The method according to claim 8 , where the target location is a buffer.

10. The method according to claim 9 , where the target location is a file.

11. A computer implemented method for converting output from a speech recognition engine, the method comprising:

inputting the output of the speech recognition engine to a post processor configured to convert the output, comprising speech-recognized dictation in the form of unformatted raw text or data, from the speech recognition engine to a formatted text document that is a transcription of the dictation;

converting, by the post processor, the output to a list of tokens;

identifying, by the post processor, a set of tokens from the list of tokens matching a first set of predetermined patterns; and

performing, by the post processor, a set of rewrite rules based on the identified set of tokens, the rewrite rules used to transform the output of the speech recognition engine to the formatted text document that corresponds to the dictation, wherein performing the set of rewrite rules comprises identifying and transforming number clusters adjacent to keywords in the speech-recognized dictation, identifying and transforming number clusters not adjacent to keywords in the speech-recognized dictation, performing consolidation or suppression of filled pauses and silences in the speech-recognized dictation, and performing attachment and formatting punctuation in the speech-recognized dictation to thereby provide the formatted text document for writing to a storage device.

Assignments (5)
PATENT RELEASE (REEL:017435/FRAME:0199) Recorded May 20, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
To: NUANCE COMMUNICATIONS, INC., AS GRANTOR; ART ADVANCED RECOGNITION TECHNOLOGIES, INC., A DELAWARE CORPORATION, AS GRANTOR; SPEECHWORKS INTERNATIONAL, INC., A DELAWARE CORPORATION, AS GRANTOR; TELELOGUE, INC., A DELAWARE CORPORATION, AS GRANTOR; DSP, INC., D/B/A DIAMOND EQUIPMENT, A MAINE CORPORATON, AS GRANTOR; SCANSOFT, INC., A DELAWARE CORPORATION, AS GRANTOR; DICTAPHONE CORPORATION, A DELAWARE CORPORATION, AS GRANTOR
Reel/Frame 038770/0824 →
PATENT RELEASE (REEL:018160/FRAME:0909) Recorded May 20, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
To: NUANCE COMMUNICATIONS, INC., AS GRANTOR; ART ADVANCED RECOGNITION TECHNOLOGIES, INC., A DELAWARE CORPORATION, AS GRANTOR; SPEECHWORKS INTERNATIONAL, INC., A DELAWARE CORPORATION, AS GRANTOR; TELELOGUE, INC., A DELAWARE CORPORATION, AS GRANTOR; DSP, INC., D/B/A DIAMOND EQUIPMENT, A MAINE CORPORATON, AS GRANTOR; HUMAN CAPITAL RESOURCES, INC., A DELAWARE CORPORATION, AS GRANTOR; INSTITIT KATALIZA IMENI G.K. BORESKOVA SIBIRSKOGO OTDELENIA ROSSIISKOI AKADEMII NAUK, AS GRANTOR; NOKIA CORPORATION, AS GRANTOR; MITSUBISH DENKI KABUSHIKI KAISHA, AS GRANTOR; STRYKER LEIBINGER GMBH & CO., KG, AS GRANTOR; NORTHROP GRUMMAN CORPORATION, A DELAWARE CORPORATION, AS GRANTOR; SCANSOFT, INC., A DELAWARE CORPORATION, AS GRANTOR; DICTAPHONE CORPORATION, A DELAWARE CORPORATION, AS GRANTOR
Reel/Frame 038770/0869 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2013
From: DICTAPHONE CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 029596/0836 →
MERGER Recorded Sep 13, 2012
From: DICTAPHONE CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028952/0397 →
SECURITY AGREEMENT Recorded Aug 24, 2006
From: NUANCE COMMUNICATIONS, INC.
To: USB AG. STAMFORD BRANCH
Reel/Frame 018160/0909 →