IP Library Granted Patent US 11,804,234
Granted Patent B2
US 11,804,234 · App. 17/124,794 · Granted Oct 31, 2023

Method for enhancing telephone speech signals based on Convolutional Neural Networks

Inventors: Javier Gallart Mauri (Saragossa, ES); Iñigo Garcia Morte (Saragossa, ES); Dayana Ribas Gonzalez (Saragossa, ES); Antonio Miguel Artiaga (Saragossa, ES); Alfonso Ortega Gimenez (Saragossa, ES); Eduardo Lleida Solano (Saragossa, ES)
Assignee: SYSTEM ONE NOC & DEVELOPMENT SOLUTIONS, S.A.
G10L21/0232G06N3/04G06N3/08G10L15/16G10L25/18G10L25/21G10L25/24G10L25/30G10L25/60G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,804,234
App. No.
17/124,794
Granted
Oct 31, 2023
Kind
B2
Abstract

A method for enhancing telephone speech signals based on Deep Convolutional Neural Network (CNN) is disclosed. The method is able to reduce the effect of acoustic distortions in daily scenarios during a telephone call. It is a single-channel, speech-oriented method with causal design and low latency. The novelty lies in the noise reduction method which, based on the classical gain method, uses a CNN to learn the Wiener estimator. Then, it computes the gain of the filter to enhance the speech power over the noise power for each time-frequency component of the signal. The selection of the Wiener gain estimator as an essential element of the method, decreases the vulnerability to estimation errors since the characteristics of this measure make it very appropriate to be estimated by deep learning approaches.

Claims (166)

1. A method for enhancing telephone speech signals based on convolutional neural networks, wherein the method comprises:

extracting, during a pre processing stage, a magnitude and a phase of a spectral representation of a telephone speech signal;

applying, during a noise reduction stage, the following steps to the magnitude of the spectral representation of the telephone speech signal:

applying a spectral estimator;

computing a perceptual representation;

applying a Convolutional Neural Network which, with a plurality of inputs corresponding to a spectral estimate and the perceptual representation, generates as an output a Wiener gain estimate consisting of a matrix/vector dependent on a frequency and which varies in time;

using the Wiener gain estimate within an enhancement filter of a function f1:

G

filtro

(

t

,

f

)

=

[

G

^

Wiener

(

t

,

f

)

exp

(

1

2

v

(

t

,

f

)

e

-

t

t

dt

)

]

p

(

t

,

f

)

G

min

1

-

p

(

t

,

f

)

wherein t is a time segment, f is a frequency bin, Ĝ Wiener =DNN(x t , x t-1 , . . . ) with x t as a vector of a plurality of spectral and perceptual parameters, G min is a constant, p(t, f) is a speech presence probability and

v

(

t

,

f

)

=

G

^

Wiener

1

-

G

^

Wiener

;

using the Wiener gain estimate as a probability of a presence of speech;

applying the function f1 as a speech enhancement filter; and

merging, during a post processing state, an initial phase with the magnitude enhanced in the noise reduction stage.

2. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 1 , wherein the Convolutional Neural Network is trained with a cost function which is a mean squared error between an optimal Wiener gain estimate and an output of the Convolutional Neural Network defined by:

F

coste

=

1

T

t

=

1

T

f

=

1

F

(

G

Wiener

(

t

,

f

)

-

G

^

Wiener

(

t

,

f

)

)

2

wherein

G

Wiener

(

t

,

f

)

=

S

X

(

t

,

f

)

S

X

(

t

,

f

)

+

S

N

(

t

,

f

)

 is obtained in a supervised manner, S X(t,f) and S N(t,f) respectively being a plurality of estimates of a plurality of power spectral densities of a clean speech signal and a noise.

3. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 1 , wherein extracting the magnitude and the phase of the spectral representation of the telephone speech signal further comprises dividing the telephone speech signal into a plurality of overlapping segments of tens of milliseconds to which a Hamming or Hanning window is applied, and subsequently a Fourier transform.

4. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 1 , wherein the spectral estimator is calculated by Welch's method.

5. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 1 , wherein the perceptual representation is calculated by applying a Mel scale filter bank.

6. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 1 , wherein the perceptual representation is performed with Mel-frequency cepstral coefficients “MFCC”.

7. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 6 , wherein merging the initial phase with the magnitude enhanced in the noise reduction stage further comprises applying an inverse Fourier transform, and subsequently, a temporal reconstruction algorithm.

8. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 2 , wherein the Convolutional Neural Network comprises at least one convolutional layer which is causal and has low latency.

9. The method for enhancing telephone speech signals based on convolutional neural networks, according to claim 1 , further comprising objectively evaluating, during the pre-processing stage, a quality of the telephone speech signal by using an acoustic quality measure selected from SNR, distortion and POLQA.

Assignments (2)
CHANGE OF NAME Recorded Aug 26, 2025
From: SYSTEM ONE NOC & DEVELOPMENT SOLUTIONS, S.A.
To: BTS TECHNOLOGY SERVICES, S.A.
Reel/Frame 072552/0944 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2020
From: GALLART MAURI, JAVIER; GARCIA MORTE, IÑIGO; RIBAS GONZALEZ, DAYANA; MIGUEL ARTIAGA, ANTONIO; ORTEGA GIMENEZ, ALFONSO; LLEIDA SOLANO, EDUARDO
To: SYSTEM ONE NOC & DEVELOPMENT SOLUTIONS, S.A.
Reel/Frame 054680/0048 →
Priority Claims (1)
EP 20382110 · Feb 14, 2020 · regional
Continuity (1)
Related Publication 20210256988A1 · Aug 19, 2021
Cited By (1)
US 12,231,851