IP Library Granted Patent US 10,535,339
Granted Patent B2
US 10,535,339 · App. 15/182,987 · Granted Jan 14, 2020

Recognition result output device, recognition result output method, and computer program product

Inventor: Hiroshi Fujimura (Kanagawa, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G10L15/183G10L15/01G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,535,339
App. No.
15/182,987
Granted
Jan 14, 2020
Kind
B2
Abstract

According to an embodiment, a speech recognition result output device includes a storage and processing circuitry. The storage is configured to store a language model for speech recognition. The processing circuitry is coupled to the storage and configured to acquire a phonetic sequence, convert the phonetic sequence into a phonetic sequence feature vector, convert the phonetic sequence feature vector into graphemes using the language model, and output the graphemes.

Claims (47)

1. A recognition result output device comprising:

first storage configured to store a language model for speech recognition;

second storage configured to store an acoustic model for speech recognition; and

processing circuitry coupled to the first storage and the second storage, and configured to function as a phonetic sequence acquirer, a first feature converter, a first grapheme converter, an input speech acquirer, a second feature converter, and a second graphene converter, wherein

the processing circuitry determines whether input is a phonetic sequence or a speech;

when the input is the phonetic sequence, the phonetic sequence acquirer acquires the phonetic sequence,

the first feature converter converts the phonetic sequence into a phonetic sequence feature vector, the phonetic sequence feature vector including a plurality of acoustic scores, each of the plurality of acoustic scores being an acoustic score of a phoneme included in the phonetic sequence,

the first grapheme converter converts the phonetic sequence feature vector into graphemes using the language model, and outputs the graphemes,

when the input is the speech, the input speech acquirer acquires the speech,

the second feature converter converts a speech waveform of the acquired speech into a speech feature vector for speech recognition,

the second grapheme converter converts the speech feature vector into graphemes using the language model and the acoustic model, and

the recognition result output device further comprises a display configured to display the output graphemes.

2. The device according to claim 1 , wherein the phonetic sequence feature vector is an acoustic score vector.

3. The device according to claim 1 , wherein the phonetic sequence feature vector is a phoneme state acoustic score vector in which an element of a phoneme state acoustic score corresponding to a phonetic sequence in a phoneme state acoustic score vector sequence is set to be higher than other phoneme state acoustic scores.

4. The device according to claim 1 , wherein

the acoustic model is a Gaussian distribution acoustic model, and

the phonetic sequence feature vector has average values of a plurality of dimensions of a mixture Gaussian acoustic model that represents a phonetic sequence state, as elements.

5. A recognition result output method employed in a recognition result output device comprising:

determining whether input is a phonetic sequence or a speech;

acquiring, when the input is the phonetic sequence, the phonetic sequence;

converting the acquired phonetic sequence into a phonetic sequence feature vector, the phonetic sequence feature vector including a plurality of acoustic scores, each of the plurality of acoustic scores being an acoustic score of a phoneme included in the phonetic sequence;

converting the phonetic sequence feature vector into graphemes using a language model that has language statistical information for speech recognition;

outputting the graphemes,

acquiring, when the input is the speech, the speech;

converting a speech waveform of the acquired speech into a speech feature vector for speech recognition;

converting the speech feature vector into graphemes using the language model and an acoustic model, and

displaying the output graphemes.

6. The method according to claim 5 , wherein the phonetic sequence feature vector is an acoustic score vector.

7. The method according to claim 5 , wherein the phonetic sequence feature vector is a phoneme state acoustic score vector in which an element of a phoneme state acoustic score corresponding to a phonetic sequence in a phoneme state acoustic score vector sequence is set to be higher than other phoneme state acoustic scores.

8. The method according to claim 5 , wherein

the acoustic model is a Gaussian distribution acoustic model, and

the phonetic sequence feature vector has average values of a plurality of dimensions of a mixture Gaussian acoustic model that represents a phonetic sequence state, as elements.

9. A computer program product comprising a non-transitory computer-readable medium containing a program executed by a computer, the program causing the computer to execute:

determining whether input is a phonetic sequence or a speech;

acquiring, when the input is the phonetic sequence, the phonetic sequence;

converting the acquired phonetic sequence into a phonetic sequence feature vector, the phonetic sequence feature vector including a plurality of acoustic scores, each of the plurality of acoustic scores being an acoustic score of a phoneme included in the phonetic sequence;

converting the phonetic sequence feature vector into graphemes using a language model that has language statistical information for speech recognition; and

outputting the graphemes,

acquiring, when the input is the speech, the speech;

converting a speech waveform of the acquired speech into a speech feature vector for speech recognition;

converting the speech feature vector into graphemes using the language model and an acoustic model, and

displaying the output graphemes.

10. The product according to claim 9 , wherein the phonetic sequence feature vector is an acoustic score vector.

11. The product according to claim 9 , wherein the phonetic sequence feature vector is a phoneme state acoustic score vector in which an element of a phoneme state acoustic score corresponding to a phonetic sequence in a phoneme state acoustic score vector sequence is set to be higher than other phoneme state acoustic scores.

12. The product according to claim 9 , wherein

the acoustic model is a Gaussian distribution acoustic model, and

the phonetic sequence feature vector has average values of a plurality of dimensions of a mixture Gaussian acoustic model that represents a phonetic sequence state, as elements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2016
From: FUJIMURA, HIROSHI
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 039892/0196 →
Priority Claims (1)
JP 2015-126246 · Jun 24, 2015 · national
Continuity (1)
Related Publication 20160379624A1 · Dec 29, 2016
Cited By (1)
US 12,451,124