IP Library Granted Patent US 9,330,682
Granted Patent B2
US 9,330,682 · App. 13/232,469 · Granted May 3, 2016

Apparatus and method for discriminating speech, and computer readable medium

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,330,682
App. No.
13/232,469
Granted
May 3, 2016
Kind
B2
Abstract

According to one embodiment, an apparatus for discriminating speech/non-speech of a first acoustic signal includes a weight assignment unit, a feature extraction unit, and a speech/non-speech discrimination unit. The first acoustic signal includes a user's speech and a reproduced sound. The reproduced sound is a system sound having a plurality of channels reproduced from a plurality of speakers. The weight assignment unit is configured to assign a weight to each frequency band based on the system sound. The feature extraction unit is configured to extract a feature from a second acoustic signal based on the weight of each frequency band. The second acoustic signal is the first acoustic signal in which the reproduced sound is suppressed. The speech/non-speech discrimination unit is configured to discriminate speech/non-speech of the first acoustic signal based on the feature.

Claims (81)

1. A speech recognition system, comprising:

a plurality of speakers that reproduce a system sound having a plurality of channels;

a microphone that generates a first acoustic signal including a user's speech and an echo of the system sound;

an echo canceller that generates a second acoustic signal by suppressing the echo included in the first acoustic signal;

a monophonization unit that generates a third acoustic signal by monophonizing the system sound and outputs the third acoustic signal to the echo canceller,

the third acoustic signal of a single channel reflecting a characteristic of the plurality of channels by channel-weighted sum if a correlation among the plurality of channels is smaller than a threshold or by channel-selection if the correlation is not smaller than the threshold;

a weight assignment unit that assigns a weight to each frequency band of the system sound, based on an amplitude of a frequency spectrum of the third acoustic signal;

a feature extraction unit that extracts a feature from the second acoustic signal based on the weight of each frequency band; and

a speech/non-speech discrimination unit that discriminates speech/non-speech of the second acoustic signal based on the feature; and

a speech recognition unit that recognizes the second acoustic signal, based on a discrimination result of the speech/non-speech;

wherein the echo canceller suppresses the echo included in the first acoustic signal, based on the third acoustic signal.

2. The system according to claim 1 , wherein

the weight assignment unit assigns a predetermined weight to a frequency band in which a frequency spectrum of the third acoustic signal is larger than a first threshold, and

the feature extraction unit extracts the feature by excluding frequency spectrums of the frequency band to which the predetermined weight is assigned.

3. The system according to claim 1 , wherein

the weight assignment unit assigns a predetermined weight to a frequency band in which a frequency spectrum of the third acoustic signal is larger than a first threshold and a frequency spectrum of the second acoustic signal is smaller than a second threshold, and

the feature extraction unit extracts the feature by excluding frequency spectrums of the frequency band to which the predetermined weight is assigned.

4. The system according to claim 1 , wherein

the weight assignment unit assigns a weight to each frequency band so that the weight is smaller when a frequency spectrum of the third acoustic signal is larger, and

the feature extraction unit makes a degree which a frequency spectrum of a frequency band contributes to the feature be smaller when the weight assigned to the frequency band is smaller.

5. A speech recognition system, comprising:

a plurality of speakers that reproduce a system sound having a plurality of channels;

a microphone that generates a first acoustic signal including a user's speech and an echo of the system sound;

an echo canceller that generates a second acoustic signal by suppressing the echo included in the first acoustic signal and generates a third acoustic signal by monophonizing the system sound,

the third acoustic signal of a single channel reflecting a characteristic of the plurality of channels by channel-weighted sum if a correlation among the plurality of channels is smaller than a threshold or by channel-selection if the correlation is not smaller than the threshold;

a weight assignment unit that assigns a weight to each frequency band of the system sound, based on an amplitude of a frequency spectrum of each channel composing the system sound;

a feature extraction unit that extracts a feature from the second acoustic signal based on the weight of each frequency band;

a speech/non-speech discrimination unit that discriminates speech/non-speech of the second acoustic signal based on the features; and

a speech recognition unit that recognizes the second acoustic signal, based on a discrimination result of the speech/non-speech;

wherein the echo canceller suppresses the echo included in the first acoustic signal, based on the third acoustic signal.

6. The system according to claim 5 , wherein

the weight assignment unit assigns a predetermined weight to a frequency band in which a frequency spectrum of one of the plurality of channels is larger than a first threshold and a frequency spectrum of the second acoustic signal is smaller than a second threshold, and

the feature extraction unit extracts the feature by excluding frequency spectrums of the frequency band to which the predetermined weight is assigned.

7. A method for controlling a speech recognition system

the method, comprising:

reproducing by a plurality of speakers, a system sound having a plurality of channels;

generating by a microphone, a first acoustic signal including a user's speech and an echo of the system sound;

generating by an echo canceller, a second acoustic signal by suppressing the echo included in the first acoustic signal;

generating by a monophonization unit in a speech discrimination apparatus, a third acoustic signal by monophonizing the system sound,

the third acoustic signal of a single channel reflecting a characteristic of the plurality of channels by channel-weighted sum if a correlation among the plurality of channels is smaller than a threshold or channel selection if the correlation is not smaller than the threshold;

assigning by a weight assignment unit in the speech discrimination apparatus, a weight to each frequency band of the system sound, based on an amplitude of a frequency spectrum of the third acoustic signal;

extracting by a feature extraction unit in the speech discrimination apparatus, a feature from a second acoustic signal based on the weight of each frequency band;

discriminating by a speech/non-speech discrimination unit in the speech discrimination apparatus, speech/non-speech of the second acoustic signal based on the feature; and

outputting by the speech discrimination apparatus, a discrimination result of the speech/non-speech to a speech recognition unit to recognize the second acoustic signal; and

outputting by the speech discrimination apparatus, the third acoustic signal to the echo canceller to suppress the echo included in the first acoustic signal.

8. A non-transitory computer readable medium for causing a computer to perform a method for controlling a speech recognition system

the method comprising:

reproducing by a plurality of speakers, a system sound having a plurality of channels;

generating by a microphone, a first acoustic signal including a user's speech and an echo of the system sound;

generating by an echo canceller, a second acoustic signal by suppressing the echo included in the first acoustic signal;

generating by a monophonization unit in a speech discrimination apparatus, a third acoustic signal by monophonizing the system sound,

the third acoustic signal of a single channel reflecting a characteristic of the plurality of channels by channel-weighted sum if a correlation among the plurality of channels is smaller than a threshold or by channel selection if the correlation is not smaller than the threshold;

assigning by a weight assignment unit in the speech discrimination apparatus, a weight to each frequency band of the system sound, based on an amplitude of a frequency spectrum of the third acoustic signal;

extracting by a feature extraction unit in the speech discrimination apparatus, a feature from the second acoustic signal based on the weight of each frequency band;

discriminating by a speech/non-speech discrimination unit in the speech discrimination apparatus, speech/non-speech of the second acoustic signal based on the feature;

outputting by the speech discrimination apparatus, a discrimination result of the speech/non-speech to a speech recognition unit to recognize the second acoustic signal; and

outputting by the speech discrimination apparatus, the third acoustic signal to the echo canceller to suppress the echo included in the first acoustic signal.

9. A method for controlling a speech recognition system

the method comprising:

reproducing by a plurality of speakers, a system sound having the plurality of channels;

generating by a microphone, a first acoustic signal including a user's speech and an echo of the system sound;

generating by an echo canceller, a second acoustic signal by suppressing the echo included in the first acoustic signal;

generating by a monophonization unit in the echo canceller, a third acoustic signal by monophonizing the system sound,

the third acoustic signal of a single channel reflecting a characteristic of the plurality of channels by channel-weighted sum if a correlation among the plurality of channels is smaller than a threshold or by channel-selection if the correlation is not smaller than the threshold;

assigning by a weight assignment unit in the speech discrimination apparatus, a weight to each frequency band of the system sound, based on an amplitude of a frequency spectrum of each channel composing the system sound;

extracting by a feature extraction unit in the speech discrimination apparatus, a feature from the second acoustic signal based on the weight of each frequency band;

discriminating by a speech/non-speech discrimination unit in the speech discrimination apparatus, speech/non-speech of the second acoustic signal based on the feature;

outputting by the speech discrimination apparatus a discrimination result of the speech/non-speech to a speech recognition unit to recognize the second acoustic signal; and

suppressing by the echo canceller, the echo included in the first acoustic signal, based on the third acoustic signal.

10. A non-transitory computer readable medium for causing a computer to perform a method for controlling a speech recognition system

wherein the method comprising:

reproducing by a plurality of speakers, a system sound having a plurality of channels;

generating by a microphone, a first acoustic signal including a user's speech and an echo of the system sound;

generating by an echo canceller, a second acoustic signal by suppressing the echo included in the first acoustic signal;

generating by a monophonization unit in the echo canceller, a third acoustic signal by monophonizing the system sound;

the third acoustic signal of a single channel reflecting a characteristic of the plurality of channels by channel-weighted sum if a correlation among the plurality of channels is smaller than a threshold or by channel-selection if the correlation is not smaller than the threshold;

assigning by a weight assignment unit in a speech discrimination apparatus, a weight to each frequency band of the system sound, based on an amplitude of a frequency spectrum of each channel composing the system sound;

extracting by a feature extraction unit in the speech discrimination apparatus, a feature from the second acoustic signal based on the weight of each frequency band;

discriminating by a speech/non-speech discrimination unit in the speech discrimination apparatus, speech/non-speech of the second acoustic signal based on the feature;

outputting by the speech discrimination apparatus, a discrimination result of the speech/non-speech to the speech recognition unit to recognize the second acoustic signal; and

suppressing by the echo canceller, the echo included in the first acoustic signal, based on the third acoustic signal.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2011
From: SUZUKI, KAORU; SAKAI, MASARU; KIDA, YUSUKE
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 027103/0811 →