IP Library Granted Patent US 9,373,342
Granted Patent B2
US 9,373,342 · App. 14/312,074 · Granted Jun 21, 2016

System and method for speech enhancement on compressed speech

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,373,342
App. No.
14/312,074
Granted
Jun 21, 2016
Kind
B2
Abstract

The present disclosure is directed towards a method for speech intelligibility. The method may include receiving, at one or more computing devices, a first speech input from a first user and performing voice activity detection upon the first speech input. The method may also include analyzing a spectral tilt associated with the first speech input, wherein analyzing includes computing an impulse response of a linear predictive coding (“LPC”) synthesis filter in a linear pulse code modulation (“PCM”) domain and wherein the one or more computing devices includes an adaptive high pass filter configured to recalculate one or more linear prediction coefficients.

Claims (39)

1. A method for speech intelligibility comprising:

receiving, at one or more computing devices, a first speech input from a first user;

performing voice activity detection upon the first speech input;

calculating one or more linear prediction coefficients; and

analyzing a spectral tilt associated with the first speech input, wherein analyzing includes computing an impulse response of a linear predictive coding (“LPC”) synthesis filter in a linear pulse code modulation (“PCM”) domain and wherein the one or more computing devices includes an adaptive high pass filter configured to recalculate the one or more linear prediction coefficients.

2. The method of claim 1 , wherein the one or more recalculated linear prediction coefficients includes at least one of a line spectral frequency (“LSF”) and a linear prediction coefficient (“LPC”).

3. The method of claim 2 , further comprising:

partially decoding a bit stream associated with the first speech input based upon, at least in part, at least one of the line spectral frequency (“LSF”) and the linear prediction coefficient (“LPC”).

4. The method of claim 1 , wherein the spectral tilt includes a ratio of frame energies between a low-pass and high-pass version of a portion of the first speech input.

5. The method of claim 1 , wherein the adaptive high pass filter is a two-tap finite impulse response (“FIR”) filter.

6. The method of claim 1 , further comprising:

determining if the first speech signal is a voiced speech signal using an unvoiced speech detection module.

7. The method of claim 1 further comprising:

performing an input power estimation analysis and a gain calculation analysis to determine an input power level and an output power level.

8. The method of claim 7 , further comprising:

determining a final speech output based upon, at least in part, a weighted average of an output of the adaptive high-pass filter and the gain calculation analysis.

9. A system for speech intelligibility comprising:

one or more computing devices configured to receive a first speech input from a first user and to perform voice activity detection upon the first speech input and to calculate one or more linear prediction coefficients, the one or more computing devices further configured to analyze a spectral tilt associated with the first speech input, wherein analyzing includes computing an impulse response of a linear predictive coding (“LPC”) synthesis filter in a linear pulse code modulation (“PCM”) domain and wherein the one or more computing devices includes an adaptive high pass filter configured to recalculate the one or more linear prediction coefficients.

10. The system of claim 9 , wherein the one or more recalculated linear prediction coefficients includes at least one of a line spectral frequency (“LSF”) and a linear prediction coefficient (“LPC”).

11. The system of claim 10 , further comprising:

partially decoding a bit stream associated with the first speech input based upon, at least in part, at least one of the line spectral frequency (“LSF”) and the linear prediction coefficient (“LPC”).

12. The system of claim 9 , wherein the spectral tilt includes a ratio of frame energies between a low-pass and high-pass version of a portion of the first speech input.

13. The system of claim 9 , wherein the adaptive high pass filter is a two-tap finite impulse response (“FIR”) filter.

14. The system of claim 9 , further comprising:

determining if the first speech signal is a voiced speech signal using an unvoiced speech detection module.

15. The system of claim 9 , further comprising:

performing an input power estimation analysis and a gain calculation analysis to determine an input power level and an output power level.

16. The system of claim 15 , further comprising:

determining a final speech output based upon, at least in part, a weighted average of an output of the adaptive high-pass filter and the gain calculation analysis.

17. A method comprising:

receiving, at one or more computing devices, a first speech input from a first user;

decoding the first speech input;

performing speech enhancement on the first speech input to generate an enhanced speech signal;

receiving the enhanced speech signal at an analysis filter configured to generate an excitation vector;

comparing the excitation vector to an original excitation vector obtained from an original bitstream to determine a final bitstream value; and

updating a partial encoder based upon, at least in part, the final bitstream value.

18. The method of claim 17 , wherein comparing includes comparing at least one of an original fixed codebook gain, a fixed codebook index, an adaptive codebook gain, and an adaptive codebook index.

19. The method of claim 17 , wherein the analysis filter is computed from the original bitstream line spectral frequency (“LSF”).

20. The method of claim 17 , wherein if the excitation vector and the original excitation vector are within a certain threshold then the original bitstream is the final bitstream value and if the excitation vector and the original excitation vector are outside of the certain threshold then a new gain is computed prior to generating the final bitstream value.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2014
From: PILLI, SRIDHAR; GODAVARTI, MAHESH; TANG, QIAN-YU; LAINEZ, JOSE; BALAM, JAGADEESH
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 033353/0181 →