IP Library Granted Patent US 9,754,603
Granted Patent B2
US 9,754,603 · App. 13/728,287 · Granted Sep 5, 2017

Speech feature extraction apparatus and speech feature extraction method

Inventors: Masanobu Nakamura (Tokyo, JP); Takashi Masuko (Kawasaki, JP)
Assignee: Kabushiki Kaisha Toshiba
G10L21/00G10L15/02G10L15/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,754,603
App. No.
13/728,287
Granted
Sep 5, 2017
Kind
B2
Abstract

According to one embodiment, a speech feature extraction apparatus includes an extraction unit and a calculation unit. The extraction unit extracts speech segments over a predetermined period at intervals of a unit time from either an input speech signal or a plurality of subband input speech signals obtained by extracting signal components of a plurality of frequency bands from the input speech signal, to generate either a unit speech signal or a plurality of subband unit speech signals. The calculation unit calculates either each average time of the unit speech signal in each of the plurality of frequency bands or each average time of each of the plurality of subband unit speech signals to obtain a speech feature.

Claims (63)

1. A speech feature extraction apparatus, comprising:

a computer programmed to comprise:

an extraction unit configured to extract speech segments over a predetermined period at intervals of a unit time from an input speech signal to generate a unit speech signal;

a first calculation unit configured to calculate each subband average time corresponding to time required to reach a center of energy gravity of the unit speech signal in each of a plurality of frequency bands obtained by dividing an overall frequency band into a number smaller than a bin number of frequency;

a generation unit configured to generate a speech feature in each of the frequency bands based on the subband average time; and

a decoder performing speech recognition processing to transform the input speech signal into words using the speech feature in the each of the frequency bands, wherein the speech feature is expressed in terms of time.

2. The apparatus according to claim 1 , wherein the computer is further programmed to comprise:

a second calculation unit configured to calculate a power spectrum of the unit speech signal, and

wherein the extraction unit extracts speech segments over the predetermined period from the input speech signal at intervals of the unit time to generate the unit speech signal, and

wherein the first calculation unit calculates the subband average time based on the power spectrum.

3. The apparatus according to claim 2 , wherein the computer is further programmed to comprise:

a third calculation unit configured to calculate a first product of a real part of a first spectrum of the unit speech signal and a real part of a second spectrum of a product of the unit speech signal and a time, to calculate a second product of an imaginary part of the first spectrum and an imaginary part of the second spectrum, and to add the first product and the second product together to obtain a third spectrum; and

wherein the first calculation unit calculates the subband average time based on the power spectrum and the third spectrum.

4. The apparatus according to claim 3 , wherein the computer is further programmed to comprise:

a first application unit configured to apply a first filter bank to the power spectrum to obtain a filtered power spectrum; and

a second application unit configured to apply a second filter bank to the third spectrum to obtain a filtered third spectrum, and

wherein the first calculation unit calculates the subband average time based on the filtered power spectrum and the filtered third spectrum.

5. The apparatus according to claim 3 , wherein the first calculation unit calculates the subband average time in a given frequency band of the frequency bands by dividing a summation of the third spectrum in the given frequency band by a summation of the power spectrum in the given frequency band.

6. The apparatus according to claim 2 , wherein the computer is further programmed to comprise:

a third calculation unit configured to calculate a group delay spectrum of the unit speech signal; and

a multiplication unit configured to multiply the power spectrum by the group delay spectrum to obtain a multiplication spectrum, and

wherein the first calculation unit calculates the subband average time based on the power spectrum and the multiplication spectrum.

7. The apparatus according to claim 6 , wherein the computer is further programmed to comprise:

a first application unit configured to apply a first filter bank to the power spectrum to obtain a filtered power spectrum; and

a second application unit configured to apply a second filter bank to the multiplication spectrum to obtain a filtered multiplication spectrum, and

wherein the first calculation unit calculates the subband average time based on the filtered power spectrum and the filtered multiplication spectrum.

8. The apparatus according to claim 2 , wherein the computer is further programmed to comprise:

an application unit configured to apply a filter bank to the power spectrum to obtain a filtered power spectrum, and

wherein the first calculation unit calculates the subband average time based on the filtered power spectrum.

9. The apparatus according to claim 1 , wherein the generation unit generates the speech feature by applying an axis transformation process on the subband average time.

10. A non-transitory computer readable storage medium storing instructions of a computer program which when executed by a computer results in performance of steps comprising:

extracting speech segments over a predetermined period at intervals of a unit time from an input speech signal to generate a unit speech signal;

calculating each subband average time corresponding to time required to reach a center of energy gravity of the unit speech signal in each of a plurality of frequency bands obtained by dividing an overall frequency band into a number smaller than a bin number of frequency;

generating a speech feature in each of the frequency bands based on the subband average time; and

transforming, by a decoder that performs speech recognition processing, the input speech signal into words using the speech feature in the each of the frequency bands, wherein the speech feature is expressed in terms of time.

11. A speech feature extraction method, comprising:

controlling a computer to:

extract speech segments over a predetermined period at intervals of a unit time from an input speech signal to generate a unit speech signal;

calculate each subband average time corresponding to time required to reach a center of energy gravity of the unit speech signal in each of a plurality of frequency bands obtained by dividing an overall frequency band into a number smaller than a bin number of frequency;

generate a speech feature in each of the frequency bands based on the subband average time; and

transform, by a decoder that performs speech recognition processing, the input speech signal into words using the speech feature in the each of the frequency bands, wherein the speech feature is expressed in terms of time.

12. A speech feature extraction method, comprising:

controlling a computer to:

extract speech segments over a predetermined period at intervals of a unit time from the plurality of subband input speech signals obtained by extracting signal components of a plurality of frequency bands from the input speech signal to generate a plurality of subband unit speech signals;

calculate each subband average time corresponding to a center of energy gravity of power of each of the plurality of subband unit speech signals within a predetermined interval;

generate a speech feature in each of the frequency bands based on the subband average time; and

transform, by a decoder that performs speech recognition processing, the input speech signal into words using the speech feature in the each of the frequency bands, wherein the speech feature is expressed in terms of time.

13. A non-transitory computer readable storage medium storing instructions of a computer program which when executed by a computer results in performance of steps comprising:

extracting speech segments over a predetermined period at intervals of a unit time from a plurality of subband input speech signals obtained by extracting signal components of a plurality of frequency bands from an input speech signal to generate a plurality of subband unit speech signals, wherein the plurality of subband input speech signals is obtained from the input speech signal by a band-pass filter;

calculating each subband average time corresponding to a center of energy gravity of power of each of the plurality of subband unit speech signals within a predetermined interval;

generating a speech feature in each of the frequency bands based on the subband average time; and

transforming, by a decoder that performs speech recognition processing, the input speech signal into words using the speech feature in the each of the frequency bands, wherein the speech feature is expressed in terms of time.

14. A speech feature extraction apparatus, comprising:

a computer programmed to comprise:

an extraction unit configured to extract speech segments over a predetermined period at intervals of a unit time from the plurality of subband input speech signals obtained by extracting signal components of a plurality of frequency bands from the input speech signal to generate a plurality of subband unit speech signals;

a calculation unit configured to calculate each subband average time corresponding to a center of energy gravity of power of each of the plurality of subband unit speech signals within a predetermined interval;

a generation unit configured to generate a speech feature in each of the frequency bands based on the subband average time; and

a decoder performing speech recognition processing to transform the input speech signal into words using the speech feature in the each of the frequency bands, wherein the speech feature is expressed in terms of time.

15. The apparatus according to claim 14 , wherein the generation unit generates the speech feature by applying an axis transformation process on the subband average time.

16. The apparatus according to claim 14 , wherein the computer is further programmed to comprise:

an application unit configured to apply a plurality of band-pass filters to the input speech signal to obtain the plurality of subband input speech signals, and

wherein the extraction unit extracts speech segments over the predetermined period from the plurality of subband input speech signals at intervals of the unit time to generate the plurality of subband unit speech signals, and

wherein the calculation unit calculates the subband average time.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2012
From: NAKAMURA, MASANOBU; MASUKO, TAKASHI
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 029534/0269 →
Priority Claims (2)
JP 2012-002133 · Jan 10, 2012 · national
JP 2012-053506 · Mar 9, 2012 · national
Continuity (1)
Related Publication 20130179158A1 · Jul 11, 2013