IP Library Granted Patent US 9,805,738
Granted Patent B2
US 9,805,738 · App. 14/423,543 · Granted Oct 31, 2017

Formant dependent speech signal enhancement

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,805,738
App. No.
14/423,543
Granted
Oct 31, 2017
Kind
B2
Abstract

An arrangement is described for speech signal processing. An input microphone signal is received that includes a speech signal component and a noise component. The microphone signal is transformed into a frequency domain set of short-term spectra signals. Then speech formant components within the spectra signals are estimated based on detecting regions of high energy density in the spectra signals. One or more dynamically adjusted gain factors are applied to the spectra signals to enhance the speech formant components.

Claims (34)

1. A computer-implemented method employing at least one hardware implemented computer processor for speech signal processing comprising:

receiving an input microphone signal having a speech signal component and a noise component;

transforming the microphone signal into a frequency domain set of short term spectra signals;

estimating speech formant components within the spectra signals based on detecting regions of high energy density in the spectra signals;

applying one or more dynamically adjusted gain factors to the spectra signals to enhance the speech formant components only during voiced speech phonemes and on the speech formant components having signal-to-noise ratio above a threshold;

adjusting the gain factors around a center frequency of the speech formant components based upon a presumed reliability of the estimation of the speech formant components, including adjusting the gain factors to boost the speech formant components more for higher reliability formant estimations than lower reliability formant estimations; and

requiring a minimum clearance between ones of the speech formant components.

2. The method according to claim 1 , wherein the speech formant components are estimated based on finding spectral peaks using a linear predictive coding filter.

3. The method according to claim 1 , wherein the speech formant components are estimated based on infinite impulse response smoothing of the spectral signals using a plurality of different smoothing constants.

4. The method according to claim 1 , wherein the gain factors are based on shaped windows concentrated on frequency regions corresponding to the speech formant components.

5. The method according to claim 4 , wherein the shaped windows are dynamically adjusted as a function of a corresponding phoneme associated with the speech signal component.

6. The method according to claim 4 , wherein the shaped windows are dynamically adjusted as a function of a signal to noise ratio of the microphone signal.

7. The method according to claim 1 , wherein the gain factors are applied to underestimate the noise component so as to reduce speech distortion in formant regions of the spectra signals.

8. The method according to claim 1 , further comprising:

combining the gain factors with one or more noise suppression coefficients to increase broadband signal to noise ratio.

9. The method according to claim 1 , further comprising:

outputting the formant enhanced spectra signals to at least one of a mobile telephony application and a speech recognition application.

10. The method according to claim 1 , wherein local maxima are determined by finding zeros of a derivative of the spectra signals after smoothing.

11. The method according to claim 1 , further including applying the one or more dynamically adjusted gain factors at a substantial center of the respective speech formant components.

12. The method according to claim 1 , wherein the speech signal component comprises non-whispered speech.

13. A speech signal processing system comprising:

a speech signal input for receiving a microphone signal having a speech signal component and a noise component;

a signal pre-processor for transforming the microphone signal into a frequency domain set of short term spectra signals;

a formant estimating module for estimating speech formant components within the spectra signals based on detecting regions of high energy density in the spectra signals; and

a formant enhancement module for applying one or more dynamically adjusted gain factors to the spectra signals to enhance the speech formant components only during voiced speech phonemes and on the speech formant components having signal-to-noise ratio above a threshold and for adjusting the gain factors around a center frequency of the speech formant components based upon a presumed reliability of the estimation of the speech formant components, wherein the gain factors are adjusted to boost the speech formant components more for higher reliability formant estimations than lower reliability formant estimations, and wherein there is a minimum clearance between ones of the speech formant components.

14. The system according to claim 13 , wherein the formant estimating module estimates the speech formant components based on finding spectral peaks in a linear predictive coding filter.

15. The system according to claim 13 , wherein the formant estimating module estimates the speech formant components based on infinite impulse response smoothing of the spectral signals using a plurality of different smoothing constants.

16. The system according to claim 13 , wherein the gain factors are based on shaped windows concentrated on frequency regions corresponding to the speech formant components.

17. The system according to claim 16 , the formant enhancement module dynamically adjusts the shaped windows as a function of a corresponding phoneme associated with the speech signal component.

18. The system according to claim 16 , wherein the formant enhancement module dynamically adjusts the shaped windows as a function of a signal to noise ratio of the microphone signal.

19. The system according to claim 13 , wherein the formant enhancement module applies the gain factors to underestimate the noise component so as to reduce speech distortion in formant regions of the spectra signals.

20. The system according to claim 13 , wherein the formant enhancement module further combines the gain factors with one or more noise suppression coefficients to increase broadband signal to noise ratio.

21. The system according to claim 13 , further comprising:

a processing output for providing the formant enhanced spectra signals to at least one of a mobile telephony application and a speech recognition application.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2015
From: KRINI, MOHAMED; SCHALK-SCHUPP, INGO; BUCK, MARKUS
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 035201/0138 →