IP Library Granted Patent US 11,756,564
Granted Patent B2
US 11,756,564 · App. 16/442,279 · Granted Sep 12, 2023

Deep neural network based speech enhancement

Inventors: Ganesh Sivaraman (Atlanta, GA); Elie Khoury (Atlanta, GA)
Assignee: PINDROP SECURITY, INC.
G10L21/0232G06N3/048G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,756,564
App. No.
16/442,279
Granted
Sep 12, 2023
Kind
B2
Abstract

A computer may segment a noisy audio signal into audio frames and execute a deep neural network (DNN) to estimate an instantaneous function of clean speech spectrum and noisy audio spectrum in the audio frame. This instantaneous function may correspond to a ratio of an a-priori signal to noise ratio (SNR) and an a-posteriori SNR of the audio frame. The computer may add estimated instantaneous function to the original noisy audio frame to output an enhanced speech audio frame.

Claims (36)

1. A computer-implemented method comprising:

segmenting, by a computer, an audio signal into a plurality of audio frames;

generating, by the computer, a feature vector for an audio frame of the plurality of audio frames, the feature vector including values of a predetermined number of frequency channels of the audio frame and values of frequency channels of a predetermined number of audio frames on each side of the audio frame;

executing, by the computer, a deep neural network (DNN) on the feature vector to estimate an instantaneous function of a clean audio spectrum and a noisy audio spectrum of the audio frame; and

generating, by the computer, an enhanced speech audio frame corresponding to the audio frame based on noisy audio spectrum of the audio frame and the estimated instantaneous function of the clean audio spectrum and the noisy audio spectrum of the audio frame.

2. The computer-implemented method of claim 1 , further comprising:

outputting, by the computer, an enhanced speech audio signal corresponding to the audio signal and containing enhanced speech audio frame.

3. The computer-implemented method of claim 1 , further comprising:

estimating, by the computer, a frame-wise voice activity in association with estimating the instantaneous function of the clean audio spectrum and the noisy audio spectrum.

4. The computer-implemented method of claim 1 , wherein the instantaneous function of the clean audio spectrum and the noisy audio spectrum of the audio frame corresponds to a ratio of an a-priori signal-to-noise ratio (SNR) and an a-posteriori SNR of the audio frame.

5. The computer-implemented method of claim 1 , wherein the output layer of the DNN has a sigmoid activation function to constrain values in an output vector of the DNN between 0 and 1.

6. The computer-implemented method of claim 5 , further comprising:

estimating, by the computer, the instantaneous function by applying an inverse of the sigmoid activation function to the output vector of the DNN.

7. The computer-implemented method of claim 1 , wherein the step of generating the feature vector of the audio frame further comprises:

performing, by the computer, mean and variance normalization of corresponding values in the frequency channels of the audio frame.

8. The computer-implemented method of claim 1 , wherein the instantaneous function is a logarithmic ratio of the clean audio spectrum and the noisy audio spectrum of the audio frame.

9. The computer-implemented method of claim 1 , wherein the DNN is trained with a binary cross entropy loss function.

10. A system comprising:

a non-transitory storage medium storing a plurality of computer program instructions and a trained deep neural network (DNN); and

a processor electrically coupled to the non-transitory storage medium and configured to execute the plurality of computer program instructions to:

segment an audio signal into a plurality of audio frames;

generate a feature vector for an audio frame of the plurality of audio frames, the feature vector including a predetermined number of frequency channels of the audio frame and frequency channels of a predetermined number of audio frames on each side of the audio frame;

feed the feature vector to the DNN to estimate an instantaneous function of a clean audio spectrum and a noisy audio spectrum of the audio frame; and

generate an enhanced speech audio frame corresponding to the audio frame based on noisy audio spectrum of the audio frame and the estimated instantaneous function of the clean audio spectrum and the noisy audio spectrum of the audio frame.

11. The system of claim 10 , wherein the processor is configured to further execute the plurality of computer program instructions to:

output an enhanced speech audio signal corresponding to the audio signal and containing enhanced speech audio frame.

12. The system of claim 10 , wherein the processor is configured to further execute the plurality of computer program instructions to:

estimate a frame-wise voice activity in association with estimating the instantaneous function of the clean audio spectrum and the noisy audio spectrum.

13. The system of claim 10 , wherein the instantaneous function of the clean audio spectrum and the noisy audio spectrum of the audio frame correspond to a ratio of an a-priori signal-to-noise ratio (SNR) and an a-posteriori SNR of the audio frame.

14. The system of claim 10 , wherein the output layer of the DNN has a sigmoid activation function to constrain values in an output vector of the DNN between 0 and 1.

15. The system of claim 14 , wherein the processor is configured to further execute the computer program instructions to:

estimate the instantaneous function by applying an inverse of the sigmoid activation function to the output vector of the DNN.

16. The system of claim 10 , wherein the processor is configured to further execute the computer program instructions to:

perform mean and variance normalization of corresponding values in the frequency channels of the audio frame.

17. The system of claim 10 , wherein the instantaneous function is a logarithmic ratio of the clean audio spectrum and the noisy audio spectrum of the audio frame.

18. The system of claim 10 , wherein the DNN is trained with a binary cross entropy loss function.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Dec 18, 2024
From: HERCULES CAPITAL, INC., AS AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069629/0384 →
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2019
From: SIVARAMAN, GANESH; KHOURY, ELIE
To: PINDROP SECURITY, INC.
Reel/Frame 049478/0418 →
Continuity (2)
Provisional Application 62685146 · Jun 14, 2018
Related Publication 20190385630A1 · Dec 19, 2019