IP Library Granted Patent US 8,370,139
Granted Patent B2
US 8,370,139 · App. 11/723,410 · Granted Feb 5, 2013

Feature-vector compensating apparatus, feature-vector compensating method, and computer program product

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,370,139
App. No.
11/723,410
Granted
Feb 5, 2013
Kind
B2
Abstract

A noise-environment storing unit stores therein a compensation vector for compensating a feature vector of a speech. A feature-vector extracting unit extracts the feature vector of the speech in each of a plurality of frames. A noise-environment-series estimating unit estimates a noise-environment series based on a feature-vector series and a degree of similarity. A calculating unit obtains a compensation vector corresponding to each noise environment in estimated noise-environment series based on the compensation vector present in the noise-environment storing unit. A compensating unit compensates the extracted feature vector of the speech based on obtained compensation vector.

Claims (32)

1. A feature-vector compensating apparatus for compensating a feature vector of speech used in speech processing under a background noise environment, comprising:

a first storing unit that stores therein a compensation vector for compensating the feature vector of the speech for each of a plurality of noise environments;

a second storing unit that stores therein a noise-environment hidden Markov model that maintains each of the noise environments as a state and is obtained by modeling parameters of a Gaussian mixture model that is a probability model of the feature vector in each of the noise environments and a state transition probability between states;

a feature extracting unit that extracts the feature vector of the speech in each of a plurality of frames of input speech;

an estimating unit that estimates a noise-environment series based on the noise-environment hidden Markov model, a feature-vector series including a plurality of extracted feature vectors for the frames and a degree of similarity that indicates a certainty that a respective feature vector is generated under the noise environment in each current frame and at least one of an immediately previous frame and an immediately subsequent frame of the current frame, the noise-environment series being a series of noise environments which generates each of the plurality of extracted feature vectors in the feature-vector series;

a calculating unit that obtains a compensation vector corresponding to each noise environment in the estimated noise-environment series based on the compensation vectors present in the first storing unit, wherein the calculating unit obtains a first compensation vector from the compensation vector present in the first storing unit, and calculates a second compensation vector by performing a weighting addition of the obtained first compensation vector with an occupation probability of each state obtained from the noise-environment hidden Markov model as a weighting coefficient; and

a compensating unit that compensates the extracted feature vectors of the speech based on the obtained second compensation vectors.

2. The apparatus according to claim 1 , wherein

the feature extracting unit divides the input speech into a plurality of frames, and extracts the feature vector of the speech in each of the frames, and

the estimating unit estimates the noise-environment series based on the feature-vector series for the frames and the degree of similarity for the feature vectors in the frames.

3. The apparatus according to claim 1 , wherein

the estimating unit sequentially estimates the noise-environment series based on the feature-vector series for a plurality of frames from a predetermined frame to a current frame and the degree of similarity for the feature vector in the frames from the predetermined frame to the current frame.

4. The apparatus according to claim 1 , wherein

the compensating unit compensates the extracted feature vector of the speech by performing an addition of the second compensation vector to the feature vector.

5. The apparatus according to claim 1 , wherein

the first storing unit stores therein the compensation vector calculated from noisy speech that is speech under the noise environment and clean speech that is speech under an environment free from the noise, for each noise environment.

6. The apparatus according to claim 1 , wherein

the extracting unit extracts a Mel frequency cepstrum coefficient of the input speech as the feature vector.

7. A method executed by an apparatus for compensating a feature vector of speech used in speech processing under a background noise environment, the apparatus comprising:

a first storing unit that stores therein a compensation vector for compensating the feature vector of the speech for each of a plurality of noise environments; and

a second storing unit that stores therein a noise-environment hidden Markov model that maintains each of the noise environments as a state and is obtained by modeling parameters of a Gaussian mixture model that is a probability model of the feature vector in each of the noise environments and a state transition probability between states;

the method comprising:

extracting the feature vector of the speech in each of a plurality of frames of input speech;

estimating a noise-environment series based on the noise-environment hidden Markov model, a feature-vector series including a plurality of extracted feature vectors for the frames and a degree of similarity that indicates a certainty that a respective feature vector is generated under the noise environment in each current frame and at least one of an immediately previous frame and an immediately subsequent frame of the current frame, the noise-environment series being a series of the noise environments which generates each of the plurality of extracted feature vectors in the feature-vector series;

obtaining a compensation vector corresponding to each noise environment in the estimated noise-environment series based on a previously calculated compensation vector, the obtaining including obtaining a first compensation vector from the compensation vector present in the first storing unit, and calculates a second compensation vector by performing a weighting addition of the obtained first compensation vector with an occupation probability of each state obtained from the noise-environment hidden Markov model as a weighting coefficient; and

compensating the extracted feature vectors of the speech based on the obtained second compensation vectors.

8. A computer program product including a non-transitory computer readable medium storing program instructions,

wherein the instructions, when executed by a computer, cause the computer to perform operations comprising:

extracting a feature vector of speech in each of a plurality of frames of input speech;

estimating a noise-environment series based on a noise-environment hidden Markov model, a feature-vector series including a plurality of extracted feature vectors for the frames and a degree of similarity that indicates a certainty that a respective feature vector is generated under a noise environment in each current frame and at least one of an immediately previous frame and an immediately subsequent frame of the current frame, the noise-environment series being a series of a plurality of noise environments which generates each of the plurality of extracted feature vectors in the feature-vector series, the noise-environment hidden Markov model maintaining each of the noise environments as a state and being obtained by modeling parameters of a Gaussian mixture model that is a probability model of the feature vector in each of the noise environments and a state transition probability between states;

obtaining a compensation vector corresponding to each noise environment in the estimated noise-environment series based on a previously calculated compensation vector, the obtaining including obtaining a first compensation vector from the compensation vector present in a first storing unit that stores therein the compensation vector for compensating the feature vector of the speech for each of the plurality of noise environments, and calculates a second compensation vector by performing a weighting addition of the obtained first compensation vector with an occupation probability of each state obtained from the noise-environment hidden Markov model as a weighting coefficient; and

compensating the extracted feature vectors of the speech based on the obtained second compensation vectors.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2007
From: AKAMINE, MASAMI; MASUKO, TAKASHI; BARREDA, DANIEL; TEUNEN, REMCO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 019418/0134 →