IP Library Granted Patent US 9,626,987
Granted Patent B2
US 9,626,987 · App. 14/072,937 · Granted Apr 18, 2017

Speech enhancement apparatus and speech enhancement method

Inventor: Naoshi Matsuo (Yokohama, JP)
Assignee: FUJITSU LIMITED
G10L21/0232G10L21/0316
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,626,987
App. No.
14/072,937
Granted
Apr 18, 2017
Kind
B2
Abstract

A speech enhancement apparatus includes: a noise estimating unit which estimates a noise component contained in a speech signal for each frequency band; a signal-to-noise ratio computing unit which computes, for each frequency band, a signal-to-noise ratio; a gain computing unit which selects a frequency band whose computed signal-to-noise ratio indicates that the signal component contained in the speech signal for the frequency band is recognizable, and which determines a gain indicating the degree of enhancement to be applied to the speech signal in accordance with the signal-to-noise ratio of the selected frequency band; and an enhancing unit which amplifies an amplitude component of a frequency domain signal in each frequency band in accordance with the gain, and which corrects the amplitude component of the frequency domain signal by subtracting the noise component from the amplitude component in each frequency band.

Claims (40)

1. A speech enhancement apparatus comprising:

a processor configured to:

compute a frequency domain signal for each of a plurality of frequency bands by transforming a speech signal containing a signal component and a noise component into a frequency domain;

estimate the noise component based on the frequency domain signal for each frequency band;

compute, for each frequency band, a signal-to-noise ratio representing the ratio of the signal component to the noise component;

select each frequency band whose computed signal-to-noise ratio is not smaller than a predetermined threshold value among the plurality of frequency bands;

determine a gain indicating the degree of enhancement to be applied to the speech signal in accordance with the signal-to-noise ratio of the selected frequency band;

amplify an amplitude component of the frequency domain signal in each frequency band in accordance with the gain, and which corrects the amplitude component of the frequency domain signal by subtracting the noise component from the amplitude component in each frequency band; and

compute a corrected speech signal by transforming the frequency domain signal having the corrected amplitude component in each frequency band into a time domain, wherein

the determining of the gain sets the gain larger as the number of selected frequency bands is larger.

2. The speech enhancement apparatus according to claim 1 , wherein the determining the gain sets the gain larger as an average value of the signal-to-noise ratio of the selected frequency band is higher.

3. The speech enhancement apparatus according to claim 1 , wherein the processor is further configured to adjust the gain for each of the plurality of frequency bands so that the gain decreases as the signal-to-noise ratio of the frequency band increases, and wherein

for each of the plurality of frequency bands, the amplifying the amplitude component amplifies the amplitude component in accordance with the gain adjusted for the frequency band.

4. The speech enhancement apparatus according to claim 3 , wherein when an average value of the signal-to-noise ratio of the selected frequency band is higher than or equal to a predetermined value, the gain computing unit sets the gain to a first value, and

for any frequency band in which the signal-to-noise ratio is higher than the predetermined value, the adjusting the gain for each of the plurality of frequency bands adjusts the gain so that the gain decreases as the signal-to-noise ratio of the frequency band increases.

5. The speech enhancement apparatus according to claim 1 , wherein for each of the plurality of frequency bands, the amplifying the amplitude component computes the corrected amplitude component by subtracting the noise component from the amplified amplitude component.

6. A speech enhancement method comprising:

computing a frequency domain signal for each of a plurality of frequency bands by transforming a speech signal containing a signal component and a noise component into a frequency domain;

estimating the noise component based on the frequency domain signal for each frequency band;

computing, for each frequency band, a signal-to-noise ratio representing the ratio of the signal component to the noise component;

selecting each frequency band whose computed signal-to-noise ratio is not smaller than a predetermined threshold value among the plurality of frequency bands;

determining a gain indicating the degree of enhancement to be applied to the speech signal in accordance with the signal-to-noise ratio of the selected frequency band;

amplifying an amplitude component of the frequency domain signal in each frequency band in accordance with the gain, and correcting the amplitude component of the frequency domain signal by subtracting the noise component from the amplitude component in each frequency band; and

computing a corrected speech signal by transforming the frequency domain signal having the corrected amplitude component in each frequency band into a time domain, wherein

the determining of the gain sets the gain lamer as the number of selected frequency bands is lamer.

7. The speech enhancement method according to claim 6 , wherein the determining the gain sets the gain larger as an average value of the signal-to-noise ratio of the selected frequency band is higher.

8. The speech enhancement method according to claim 6 , further comprising adjusting the gain for each of the plurality of frequency bands so that the gain decreases as the signal-to-noise ratio of the frequency band increases, and wherein

for each of the plurality of frequency bands, the amplifying the amplitude component amplifies the amplitude component in accordance with the gain adjusted for the frequency band.

9. The speech enhancement method according to claim 8 , wherein when an average value of the signal-to-noise ratio of the selected frequency band is higher than or equal to a predetermined value, the determining the gain sets the gain to a first value, and

for any frequency band in which the signal-to-noise ratio is higher than the predetermined value, the adjusting the gain for each of the plurality of frequency bands adjusts the gain so that the gain decreases as the signal-to-noise ratio of the frequency band increases.

10. The speech enhancement method according to claim 6 , wherein for each of the plurality of frequency bands, the amplifying the amplitude component computes the corrected amplitude component by subtracting the noise component from the amplified amplitude component.

11. A non-transitory computer-readable recording medium having recorded thereon a speech enhancement computer program that causes a computer to execute a process comprising:

computing a frequency domain signal for each of a plurality of frequency bands by transforming a speech signal containing a signal component and a noise component into a frequency domain;

estimating the noise component based on the frequency domain signal for each frequency band;

computing, for each frequency band, a signal-to-noise ratio representing the ratio of the signal component to the noise component;

selecting each frequency band whose computed signal-to-noise ratio is not smaller than a predetermined threshold value among the plurality of frequency bands;

determining a gain indicating the degree of enhancement to be applied to the speech signal in accordance with the signal-to-noise ratio of the selected frequency band;

amplifying an amplitude component of the frequency domain signal in each frequency band in accordance with the gain, and correcting the amplitude component of the frequency domain signal by subtracting the noise component from the amplitude component in each frequency band; and

computing a corrected speech signal by transforming the frequency domain signal having the corrected amplitude component in each frequency band into a time domain, wherein

the determining of the gain sets the gain larger as the number of selected frequency bands is lamer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 26, 2013
From: MATSUO, NAOSHI
To: FUJITSU LIMITED
Reel/Frame 031677/0018 →
Priority Claims (1)
JP 2012-261704 · Nov 29, 2012 · national
Continuity (1)
Related Publication 20140149111A1 · May 29, 2014