IP Library › Granted Patent US 9,799,331
Granted Patent B2
US 9,799,331 · App. 15/074,579 · Granted Oct 24, 2017

Feature compensation apparatus and method for speech recognition in noisy environment

Inventors: Hyun Woo Kim (Daejeon-si, KR); Ho Young Jung (Daejeon-si, KR); Jeon Gue Park (Daejeon-si, KR); Yun Keun Lee (Daejeon-si, KR)
Assignee: Electronics and Telecommunications Research Institute
G10L15/20G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,799,331
App. No.
15/074,579
Filed
Mar 18, 2016
Granted
Oct 24, 2017
Kind
B2
Art Unit
2657
USPC
704/205
Abstract

A feature compensation apparatus includes a feature extractor configured to extract corrupt speech features from a corrupt speech signal with additive noise that consists of two or more frames; a noise estimator configured to estimate noise features based on the extracted corrupt speech features and compensated speech features; a probability calculator configured to calculate a correlation between adjacent frames of the corrupt speech signal; and a speech feature compensator configured to generate compensated speech features by eliminating noise features of the extracted corrupt speech features while taking into consideration the correlation between adjacent frames of the corrupt speech signal and the estimated noise features, and to transmit the generated compensated speech features to the noise estimator.

Claims (33)

1. A feature compensation apparatus for speech recognition in a noisy environment, the feature compensation apparatus comprising a computer readable medium storing computer readable code that implements:

a feature extractor configured to extract corrupt speech features from a corrupt speech signal with additive noise that consists of two or more frames;

a noise estimator configured to estimate noise features based on the extracted corrupt speech features and compensated speech features;

a linear model generator configured to approximate a Gaussian mixture model (GMM) probability distribution, the estimated noise features and the extracted corrupt speech features into a linear model;

a probability calculator configured to calculate a correlation between adjacent frames of the corrupt speech signal; and

a speech feature compensator configured to generate the compensated speech features by eliminating noise features of the extracted corrupt speech features while taking into consideration the correlation between adjacent frames of the corrupt speech signal and the estimated noise features, and to transmit the generated compensated speech features to the noise estimator,

wherein the noise estimator estimates an average and variance of the noise features based on a dynamics model of noise features of the extracted corrupt speech features and a nonlinear observation model of corrupt speech features, and

wherein the noise estimator reduces a Kalman gain of the average and variance of noise features that are to be updated in inverse proportion to a ratio of the extracted corrupt speech feature to the noise feature.

2. The feature compensation apparatus of claim 1 , wherein the probability calculator comprises

a probability distribution obtainer configured to obtain a GMM probability distribution of training speech features from training speech signals that consist of two or more frames,

a transition probability codebook obtainer configured to obtain a transition probability of a GMM mixture component between adjacent frames of the training speech features, and

a transition probability calculator configured to search transition probabilities of a GMM mixture component between adjacent frames of each of the training speech signals to calculate a transition probability of the GMM mixture component that corresponds to a transition probability of a mixture component between adjacent frames of the corrupt speech features extracted from the corrupt speech signal.

3. The feature compensation apparatus of claim 2 , wherein the speech feature compensator eliminates the noise features of the extracted corrupt speech features using the correlation between adjacent frames of the corrupt speech signal and the estimated noise features, wherein the correlation is based on the GMM probability distribution of the training speech features and the transition probability of a GMM mixture component.

4. The feature compensation apparatus of claim 1 , wherein the feature extractor converts each frame of the corrupt speech signal from time domain to frequency domain, and calculates a log energy value by taking a logarithm of energy which has been calculated by applying a Mel-scale filter bank to the converted corrupt speech signal, thereby extracting the corrupt speech features.

5. The feature compensation apparatus of claim 4 , wherein the feature extractor smooths the corrupt speech signal before taking a logarithm of the energy which has been calculated by applying the Mel-scale filter bank to the converted corrupt speech signal.

6. The feature compensation apparatus of claim 1 , wherein the probability calculator obtains a statistic model with a hidden Markov model (HMM) structure of training speech features from training speech signals that consist of two or more frames, decodes the training speech features into a HMM, and calculates HMM state probabilities.

7. The feature compensation apparatus of claim 6 , wherein the speech feature compensator eliminates the estimated noise features of the corrupt speech features using a statistical model of the training speech features, the estimated noise features, the extracted corrupt speech features, and the HMM state probabilities.

8. A feature compensation method for speech recognition in a noisy environment, the feature compensation method comprising:

extracting speech feature from a corrupt speech signal with additive noise that consists of two or more frames;

estimating noise features based on the extracted corrupt speech features and compensated speech features;

approximating a GMM probability distribution, the estimated noise features and the extracted corrupt speech features into a linear model;

calculating a correlation between adjacent frames of the corrupt speech signal; and

generating compensated speech features by eliminating noise features of the extracted corrupt speech features while taking into consideration the correlation between adjacent frames of the corrupt speech signal and the estimated noise features, and transmitting the generated compensated speech features,

wherein the estimation of the noise features comprises estimating an average and variance of noise features based on a dynamics model of noise features of the extracted corrupt speech features and a nonlinear observation model of corrupt speech features, and reducing a Kalman gain of the average and variance of noise features that are to be updated in inverse proportion to a ratio of the extracted corrupt speech feature to the noise feature.

9. The feature compensation method of claim 8 , wherein the calculation of the correlation comprises:

obtaining a Gaussian mixture model (GMM) probability distribution of training speech features from training speech signals that consist of two or more frames,

obtaining a transition probability of a GMM mixture component between adjacent frames of the training speech features, and

searching\transition probabilities of a GMM mixture component between adjacent frames of each of the training speech signals to calculate a transition probability of the GMM mixture component that corresponds to a transition probability of a mixture component between adjacent frames of the corrupt speech features extracted from the corrupt speech signal.

10. The feature compensation method of claim 9 , wherein the generation of the compensated speech features comprises eliminating the noise features of the extracted corrupt speech features using the correlation between adjacent frames of the corrupt speech signal and the estimated noise features, wherein the correlation is based on the GMM probability distribution of the training speech features and the transition probability of a GMM mixture component.

11. The feature compensation method of claim 8 , wherein the extraction of the corrupt speech features comprises converting each frame of the corrupt speech signal from time domain to frequency domain, and calculating a log energy value by taking a logarithm of energy which has been calculated by applying a Mel-scale filter bank to the converted corrupt speech signal, thereby extracting the corrupt speech features.

12. The feature compensation method of claim 11 , wherein the extraction of the corrupt speech features comprises smoothing the corrupt speech signal before taking a logarithm of the energy which has been calculated by applying the Mel-scale filter bank to the converted corrupt speech signal.

13. The feature compensation method of claim 8 , wherein the calculation of the correlation comprises obtaining a statistic model with a hidden Markov model (HMM) structure of training speech features from training speech signals that consist of two or more frames, decoding the training speech features into a HMM, and calculating HMM state probabilities.

14. The feature compensation method of claim 13 , wherein the generation of the compensated speech features comprises eliminating the estimated noise features of the corrupt speech features using a statistical model of the training speech features, the estimated noise features, the extracted corrupt speech features, and the HMM state probabilities.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2016
From: KIM, HYUN WOO; JUNG, HO YOUNG; PARK, JEON GUE; LEE, YUN KEUN
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 038224/0762 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2016
From: KIM, HYUN WOO; JUNG, HO YOUNG; PARK, JEON GUE
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 038035/0391 →
Priority Claims (1)
KR 10-2015-0039098 · Mar 20, 2015 · national
Continuity (1)
Related Publication 20160275964A1 · Sep 22, 2016