IP Library Granted Patent US 8,762,142
Granted Patent B2
US 8,762,142 · App. 11/889,665 · Granted Jun 24, 2014

Multi-stage speech recognition apparatus and method

Inventors: So-young Jeong (Seoul, KR); Kwang-cheol Oh (Yongin-si, KR); Jae-hoon Jeong (Yongin-si, KR); Jeong-su Kim (Yongin-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/32G10L15/16G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,762,142
App. No.
11/889,665
Granted
Jun 24, 2014
Kind
B2
Abstract

Provided are a multi-stage speech recognition apparatus and method. The multi-stage speech recognition apparatus includes a first speech recognition unit performing initial speech recognition on a feature vector, which is extracted from an input speech signal, and generating a plurality of candidate words; and a second speech recognition unit rescoring the candidate words, which are provided by the first speech recognition unit, using a temporal posterior feature vector extracted from the speech signal.

Claims (24)

1. A multi-stage speech recognition apparatus comprising:

at least one processor that executes:

a first speech recognition unit performing initial speech recognition on a feature vector, which is extracted from an input speech signal, and generating a plurality of candidate words; and

a second speech recognition unit rescoring the candidate words, which are provided by the first speech recognition unit, using a temporal posterior feature vector reflected from time-varying voice characteristics and extracted from the feature vector, and outputting a word having a highest score as a final recognition result;

wherein the temporal posterior feature vector is a split-temporal context (STC)-TRAP feature vector comprising a left, center, and right context part.

2. The apparatus of claim 1 , wherein the first speech recognition unit comprises:

a first feature extractor extracting a spectrum feature vector from the speech signal; and

a recognizer performing the initial speech recognition using the spectrum feature vector.

3. The apparatus of claim 1 , wherein the second speech recognition unit comprises:

a second feature extractor extracting the temporal posterior feature vector from the feature vector; and

a rescorer performing forced alignment of the candidate words, which are provided by the first speech recognition unit, using the temporal posterior feature vector.

4. The apparatus of claim 1 , wherein the split-temporal context (STC)-TRAP feature vector is obtained by inputting feature vectors of adjacent frames placed before and after a current frame, which is to be converted, to a left context neural network, a center context neural network, and a right context neural network, for respective bands, and integrating outputs of the left context neural network, the center context neural network, and the right context neural network.

5. A multi-stage speech recognition method comprising:

performing, by at least one processor, initial speech recognition on a feature vector, which is extracted from an input speech signal, and generating a plurality of candidate words; and

rescoring, by the at least one processor, the candidate words, which are obtained from the initial speech recognition, using a temporal posterior feature vector reflected from time- varying voice characteristics and extracted from the speech signal, and outputting a word having a highest score as a final recognition result;

wherein the temporal posterior feature vector is a split-temporal context (STC)-TRAP feature vector comprising a left, center, and right context part.

6. The method of claim 5 , wherein the performing of the initial speech recognition comprises:

extracting a spectrum feature vector from the speech signal; and

performing the initial speech recognition using the spectrum feature vector.

7. The method of claim 5 , wherein the rescoring of the candidate words comprises:

extracting the temporal posterior feature vector from the feature vector; and

performing forced alignment of the candidate words, which are obtained from the initial speech recognition, using the temporal posterior feature vector.

8. The method of claim 5 , wherein the split-temporal context (STC)-TRAP feature vector is obtained by inputting feature vectors of adjacent frames placed before and after a current frame, which is to be converted, to a left context neural network, a center context neural network, and a right context neural network, for respective bands, and integrating outputs of the left context neural network, the center context neural network, and the right context neural network.

9. A non-transitory computer-readable recording medium storing a program to control at least one processing element to implement the method of claim 5 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2007
From: JEONG, SO-YOUNG; OH, KWANG-CHEOL; JEONG, JAE-HOON; KIM, JEONG-SU
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 019751/0786 →
Priority Claims (1)
KR 10-2007-0018666 · Feb 23, 2007 · national
Continuity (1)
Related Publication 20080208577A1 · Aug 28, 2008