IP Library Granted Patent US 9,224,392
Granted Patent B2
US 9,224,392 · App. 13/420,912 · Granted Dec 29, 2015

Audio signal processing apparatus and audio signal processing method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,224,392
App. No.
13/420,912
Granted
Dec 29, 2015
Kind
B2
Abstract

Likelihood calculation means extracts audio features expressing features of a voice signal and a non-voice signal from an acquired audio signal, and calculates likelihood expressing probability that the voice signal is included in the audio signal using the audio features. Spectral feature extraction means performs a frequency analysis to the audio signal to extract a spectral feature. Using the spectral feature, first basis matrix producing means produces a first basis matrix expressing the feature of the non-voice signal. Second basis matrix producing means specifies a component having a high association with the voice signal in the first basis matrix using the likelihood, and excludes the component to produce a second basis matrix. Spectral feature estimation means estimates a spectral feature of the voice signal or a spectral feature of the non-voice signal by performing nonnegative matrix factorization to the spectral feature using the second basis matrix.

Claims (56)

1. An audio signal processing apparatus utilized in pre-processing of speech recognition, the apparatus comprising:

a likelihood calculation unit, executed by a processor using a program stored in a storage, configured to extract audio features expressing features of a voice signal and a non-voice signal from an audio signal including the voice signal and the non-voice signal, and calculate a likelihood expressing a probability that the voice signal is included in the audio signal;

a spectral feature extraction unit, executed by the processor, configured to perform a frequency analysis to the audio signal to extract a spectral feature;

a first basis matrix producing unit, executed by the processor, configured to produce a first basis matrix expressing the feature of the non-voice signal using the spectral feature;

a second basis matrix producing unit, executed by the processor, configured to specify a component having a high association with the voice signal in the first basis matrix using the likelihood, and exclude the component from the first basis matrix to produce a second basis matrix; and

a spectral feature estimation unit, executed by the processor, configured to estimate a spectral feature of the voice signal or a spectral feature of the non-voice signal by performing nonnegative matrix factorization to the spectral feature of the audio signal using the second basis matrix,

wherein the spectral feature estimation unit produces a third basis matrix and a first coefficient matrix, which express the feature of the voice signal, by nonnegative matrix factorization in which the second basis matrix is used, and estimates the spectral feature of the voice signal included in the audio signal by a product of the third basis matrix and the first coefficient matrix, and wherein the product is used to separate the voice signal from the audio signal.

2. The apparatus according to claim 1 , wherein

the second basis matrix producing unit produces the second basis matrix by excluding a column vector having a high association with the voice signal from the first basis matrix.

3. The apparatus according to claim 1 , wherein

the second basis matrix producing unit produces the second basis matrix from the first basis matrix by replacing a value of a column vector having a high association with the voice signal with 0.

4. The apparatus according to claim 1 , wherein

the second basis matrix producing unit specifies a component having a high association with the voice signal in the first basis matrix by comparing the likelihood to a predetermined threshold.

5. The apparatus according to claim 1 , further comprising:

a voice/non-voice determination unit, executed by the processor, configured to extract the audio features expressing features of the voice signal and the non-voice signal from the audio signal, and determine whether the audio signal is the voice signal or the non-voice signal using the audio features,

wherein the first basis matrix producing unit produces the first basis matrix expressing the feature of the non-voice signal using the spectral feature of the audio signal in which the voice/non-voice determination unit determines that the audio signal is the non-voice signal.

6. The apparatus according to claim 1 , wherein

the spectral feature estimation unit produces a second coefficient matrix expressing the feature of the non-voice signal by the nonnegative matrix factorization in which the second basis matrix is used, and estimates the spectral feature of the non-voice signal included in the audio signal by a product of the second basis matrix and the second coefficient matrix.

7. The apparatus according to claim 1 , further comprising:

an inverse transformation unit, executed by the processor, configured to transform the spectral feature estimated by the spectral feature estimation unit into a temporal signal.

8. An audio signal processing apparatus utilized in pre-processing of speech recognition, the apparatus comprising:

a likelihood calculation unit, executed by a processor using a program stored in a storage, configured to extract audio features expressing features of a first audio signal and a second audio signal from a third audio signal including the first audio signal and the second audio signal, and calculate a likelihood expressing a probability that the first audio signal is included in the third audio signal;

a spectral feature extraction unit, executed by the processor, configured to perform a frequency analysis to the third audio signal to extract a spectral feature;

a first basis matrix producing unit, executed by the processor, configured to produce a first basis matrix expressing the feature of the second audio signal using the spectral feature;

a second basis matrix producing unit, executed by the processor, configured to specify a component having a high association with the first audio signal in the first basis matrix using the likelihood, and exclude the component from the first basis matrix to produce a second basis matrix; and

a spectral feature estimation unit, executed by the processor, configured to estimate a spectral feature of the first audio signal or a spectral feature of the second audio signal by performing nonnegative matrix factorization to the spectral feature of the third audio signal using the second basis matrix,

wherein the spectral feature estimation unit produces a third basis matrix and a first coefficient matrix, which express the feature of the first audio signal, by nonnegative matrix factorization in which the second basis matrix is used, and estimates the spectral feature of the first audio signal included in the third audio signal by a product of the third basis matrix and the first coefficient matrix and wherein the product is used to separate the voice signal from the audio signal.

9. The apparatus according to claim 8 , wherein

the second basis matrix producing unit produces the second basis matrix by excluding a column vector having a high association with the first audio signal from the first basis matrix.

10. The apparatus according to claim 8 , wherein

the second basis matrix producing unit produces the second basis matrix from the first basis matrix by replacing a value of a column vector having a high association with the first audio signal with 0.

11. The apparatus according to claim 8 , wherein

the second basis matrix producing unit specifies a component having a high association with the first audio signal in the first basis matrix by comparing the likelihood to a predetermined threshold.

12. The apparatus according to claim 8 , further comprising:

a first/second determination unit, executed by the processor, configured to extract the audio features expressing features of the first audio signal and the second audio signal from the third audio signal, and determine whether the third audio signal is the first audio signal or the second audio signal using the audio features,

wherein the first basis matrix producing unit produces the first basis matrix expressing the feature of the second audio signal using the spectral feature of the third audio signal in which the first/second determination unit determines that the third audio signal is the second audio signal.

13. The apparatus according to claim 8 , wherein

the spectral feature estimation unit produces a second coefficient matrix expressing the feature of the second audio signal by the nonnegative matrix factorization in which the second basis matrix is used, and estimates the spectral feature of the second audio signal included in the third audio signal by a product of the second basis matrix and the second coefficient matrix.

14. The apparatus according to claim 8 , further comprising:

an inverse transformation unit, executed by the processor, configured to transform the spectral feature estimated by the spectral feature estimation unit into a temporal signal.

15. The apparatus according to claim 8 , wherein the first audio signal is a voice signal and the second audio signal is a noise signal.

16. An audio signal processing method utilized in pre-processing of speech recognition, the method comprising:

extracting audio features expressing features of a first audio signal and a second audio signal from a third audio signal including the first audio signal and the second audio signal;

calculating a likelihood expressing a probability that the first audio signal is included in the third audio signal;

performing a frequency analysis to the third audio signal to extract a spectral feature;

producing a first basis matrix expressing the feature of the second audio signal using the spectral feature;

specifying a component having a high association with the first audio signal in the first basis matrix using the likelihood;

excluding the component from the first basis matrix to produce a second basis matrix; and

estimating a spectral feature of the first audio signal or a spectral feature of the second audio signal by performing nonnegative matrix factorization to the spectral feature of the third audio signal using the second basis matrix,

wherein

a third basis matrix and a first coefficient matrix, which express the feature of the first audio signal, are produced by nonnegative matrix factorization in which the second basis matrix is used, and

the spectral feature of the first audio signal included in the third audio signal, is estimated by a product of the third basis matrix and the first coefficient matrix and wherein the product is used to separate the voice signal from the audio signal.

17. The method according to claim 16 , further comprising:

extracting the audio features expressing features of the first audio signal and the second audio signal from the third audio signal,

determining whether the third audio signal is the first audio signal or the second audio signal using the audio features,

wherein the first basis matrix expressing the feature of the second audio signal using the spectral feature of the third audio signal in which the third audio signal is the second audio signal is determined, is produced.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2012
From: HIROHATA, MAKOTO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 027868/0430 →