IP Library Granted Patent US 9,330,683
Granted Patent B2
US 9,330,683 · App. 13/232,491 · Granted May 3, 2016

Apparatus and method for discriminating speech of acoustic signal with exclusion of disturbance sound, and non-transitory computer readable medium

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,330,683
App. No.
13/232,491
Granted
May 3, 2016
Kind
B2
Abstract

According to one embodiment, an apparatus for discriminating speech/non-speech of a first acoustic signal includes a weight assignment unit, a feature extraction unit, and a speech/non-speech discrimination unit. The weight assignment unit is configured to assign a weight to each frequency band, based on a frequency spectrum of the first acoustic signal including a user's speech and a frequency spectrum of a second acoustic signal including a disturbance sound. The feature extraction unit is configured to extract a feature from the frequency spectrum of the first acoustic signal, based on the weight of each frequency band. The speech/non-speech discrimination unit is configured to discriminate speech/non-speech of the first acoustic signal, based on the feature.

Claims (41)

1. An apparatus for discriminating speech/non-speech of a first acoustic signal, comprising:

a memory to store computer executable instructions;

a processor configured to execute the computer executable instructions to perform operations comprising:

assigning a weight to each, frequency band, based on both a frequency spectrum of the first acoustic signal including a user's speech and a frequency spectrum of a second acoustic signal including a disturbance sound,

wherein the first acoustic signal is acquired via a main microphone, and the second acoustic signal is acquired via a sub microphone located at a position farther than the main microphone from the user;

extracting a feature from the frequency spectrum of the first acoustic signal, based on an updated weight of each frequency band; and

discriminating speech/non-speech of the first acoustic signal, based on the feature, wherein,

the assigning assigns a first weight to a frequency band in which the frequency spectrum of the first acoustic signal is smaller than a first threshold, assigns a second weight larger than the first weight to frequency bands in which the frequency spectrum of the first acoustic signal is not smaller than the first threshold, and updates the first weight already assigned to the frequency band in which the frequency spectrum of the second acoustic signal is not larger than a second threshold, to the second weight,

the extracting extracts the feature by excluding frequency spectrums of the frequency band to which the first weight is assigned.

2. The apparatus according to claim 1 , the operations further comprising:

suppressing a noise included in the first acoustic signal, based on the second acoustic signal;

wherein the assigning utilizes the frequency spectrum of the first acoustic signal in which the noise is suppressed.

3. The apparatus according to claim 2 , the operations further comprising:

extracting the first acoustic signal in which the user's sound is emphasized by processing acoustic signals of a plurality of channels; and

extracting the second acoustic signal in which the disturbance sound is emphasized by processing at least two of the acoustic signals;

wherein the suppressing suppresses the noise included in the first acoustic signal extracted, based on the second acoustic signal extracted.

4. The apparatus according to claim 1 , the operations further comprising:

extracting the first acoustic signal in which the user's sound is emphasized by processing acoustic signals of a plurality of channels; and

extracting the second acoustic signal in which the disturbance sound is emphasized by processing at least two of the acoustic signals;

wherein the assigning utilizes the frequency spectrum of the first acoustic signal extracted and the frequency spectrum of the second acoustic signal extracted.

5. The apparatus according to claim 1 , the operations further comprising:

mixing a system sound into the second acoustic signal;

wherein the assigning utilizes the frequency spectrum of the second acoustic signal in which the system sound is mixed.

6. A method for discriminating speech/non-speech of a first acoustic signal, comprising:

assigning a weight to each frequency band, based on both a frequency spectrum of the first acoustic signal including a user's speech and a frequency spectrum of a second acoustic signal including a disturbance sound,

wherein the first acoustic signal is acquired via a main microphone, and the second acoustic signal is acquired via a sub microphone located at a position farther than the main microphone from the user;

extracting a feature from the frequency spectrum of the first acoustic signal, based on an updated weight of each frequency band; and

discriminating speech/non-speech of the first acoustic signal, based on the feature, wherein,

the assigning includes assigning a first weight to a frequency band in which the frequency spectrum of the first acoustic signal is smaller than a first threshold,

assigning a second weight larger than the first weight to frequency bands in which the frequency spectrum of the first acoustic signal is not smaller than the first threshold, and

updating the first weight already assigned to the frequency band in which the frequency spectrum of the second acoustic signal is not larger than a second threshold, to the second weight,

the extracting includes extracting the feature by excluding frequency spectrums of the frequency band to which the first weight is assigned.

7. A non-transitory computer readable medium storing instructions thereon, that when executed by a processor, perform operations for discriminating speech/non-speech of a first acoustic signal, the operations comprising:

assigning a weight to each frequency band, based on both a frequency spectrum of the first acoustic signal including a user's speech and a frequency spectrum of a second acoustic signal including a disturbance sound,

wherein the first acoustic signal is acquired via a main microphone, and the second acoustic signal is acquired via a sub microphone located at a position farther than the main microphone from the user;

extracting a feature from the frequency spectrum of the first acoustic signal, based on an updated weight of each frequency band; and

discriminating speech/non-speech of the first acoustic signal, based on the feature, wherein,

the assigning includes assigning a first weight to a frequency band in which the frequency spectrum of the first acoustic signal is smaller than a first threshold,

assigning a second weight larger than the first weight to frequency bands in which the frequency spectrum of the first acoustic signal is not smaller than the first threshold, and

updating the first weight already assigned to the frequency band in which the frequency spectrum of the second acoustic signal is not larger than a second threshold, to the second weight,

the extracting includes extracting the feature by excluding frequency spectrums of the frequency band to which the first weight is assigned.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2011
From: SUZUKI, KAORU; SAKAI, MASARU; KIDA, YUSUKE
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 027103/0814 →