IP Library Granted Patent US 9,672,817
Granted Patent B2
US 9,672,817 · App. 14/748,474 · Granted Jun 6, 2017

Method and apparatus for optimizing a speech recognition result

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,672,817
App. No.
14/748,474
Granted
Jun 6, 2017
Kind
B2
Abstract

According to one embodiment, an apparatus for optimizing a speech recognition result comprising: a receiving unit configured to receive a speech recognition result; a calculating unit configured to calculate a pronunciation similarity between a segment of the speech recognition result and a key word in a key word list; and a replacing unit configured to replace the segment with the key word in a case that the pronunciation similarity is higher than a first threshold.

Claims (25)

1. An apparatus for optimizing a speech recognition result, comprising:

a computer that, by using computer executable instructions,

receives a speech recognition result from a speech recognition engine,

calculates a phoneme acoustic distance between a phoneme sequence of a segment of the speech recognition result and a phoneme sequence of a key word in a keyword list,

calculates a tone acoustic distance between a tone sequence of the segment and a tone sequence of the key word,

calculates a weighted average of the phoneme acoustic distance and the tone acoustic distance,

calculates an average acoustic distance which is obtained by dividing the weighted average by the number of characters, syllables or phonemes of the key word,

calculates a language model score of the segment, based on a language model score of each word in the segment,

replaces the segment with the key word in a case that the average acoustic distance is lower than a first threshold and the language model score is lower than a second threshold, and

outputs the speech recognition result of which the segment is replaced with the keyword to the speech recognition engine.

2. The apparatus according to claim 1 , wherein

the computer calculates the average acoustic distance between a segment of the speech recognition result, in which a language model score of the segment is lower than the second threshold, and a key word in the key word list.

3. The apparatus according to claim 1 , wherein

the computer calculates the phoneme acoustic distance between the phoneme sequence of the segment and the phoneme sequence of the key word by using a phoneme confusion matrix as a weight.

4. The apparatus according to claim 1 , wherein

the computer calculates the tone acoustic distance between the tone sequence of the segment and the tone sequence of the key word by using a tone confusion matrix as a weight.

5. A method for optimizing a speech recognition result, comprising steps of:

receiving a speech recognition result from a speech recognition engine;

calculating a phoneme acoustic distance between a phoneme sequence of a segment of the speech recognition result and a phoneme sequence of a key word in a keyword list;

calculating a tone acoustic distance between a tone sequence of the segment and a tone sequence of the key word;

calculating a weighted average of the phoneme acoustic distance and the tone acoustic distance;

calculating an average acoustic distance which is obtained by dividing the weighted average by the number of characters, syllables or phonemes of the key word;

calculating a language model score of the segment, based on a language model score of each word of the segment;

replacing the segment with the key word in a case that the average acoustic distance is lower than a first threshold and the language model score is lower than a second threshold; and

outputting the speech recognition result of which the segment is replaced with the keyword to the speech recognition engine.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2015
From: YONG, KUN; DING, PEI; ZHU, HUIFENG
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 035893/0663 →