IP Library Granted Patent US 10,650,806
Granted Patent B2
US 10,650,806 · App. 15/959,606 · Granted May 12, 2020

System and method for discriminative training of regression deep neural networks

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,650,806
App. No.
15/959,606
Granted
May 12, 2020
Kind
B2
Abstract

A method, computer program product, and computer system for transforming, by a computing device, a speech signal into a speech signal representation. A regression deep neural network may be trained with a cost function to minimize a mean squared error between actual values of the speech signal representation and estimated values of the speech signal representation, wherein the cost function may include one or more discriminative terms. Bandwidth of the speech signal may be extended by extending the speech signal representation of the speech signal using the regression deep neural network trained with the cost function that includes the one or more discriminative terms.

Claims (23)

1. A computer-implemented method comprising:

transforming, by a computing device, a speech signal into a speech signal representation;

training, with a cost function, a regression deep neural network to minimize a mean squared error between actual values of the speech signal representation and estimated values of the speech signal representation, wherein the cost function includes one or more discriminative terms; and

extending bandwidth of the speech signal by extending the speech signal representation of the speech signal using the regression deep neural network trained with the cost function that includes the one or more discriminative terms;

wherein at least one of: (i) the one or more discriminative terms include at least one of a fricative-to-vowel power ratio and a function thereof; (ii) the one or more discriminative terms preserve relations of statistics between different phoneme classes in the actual values of the speech signal representation and the estimated values of the speech signal representation; and (iii) the method further comprising reproducing an average power ratio at an output of the regression deep neural network.

2. The computer-implemented method of claim 1 wherein the speech signal representation is obtained by decomposing the speech signal into a spectral envelope and an excitation signal, and wherein the spectral envelope is extended using the regression deep neural network trained with the cost function.

3. The computer-implemented method of claim 1 wherein the cost function preserves a power ratio between the different phoneme classes in the actual values of the speech signal representation and the estimated values of the speech signal representation.

4. The computer-implemented method of claim 1 wherein the cost function preserves the power ratio between the different phoneme classes using a weighted sum of K power ratio errors between the different phoneme classes.

5. A computer program product residing on a non-transitory computer readable storage medium having a plurality of instructions stored thereon which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising:

transforming a speech signal into a speech signal representation;

training, with a cost function, a regression deep neural network to minimize a mean squared error between actual values of the speech signal representation and estimated values of the speech signal representation, wherein the cost function includes one or more discriminative terms; and

extending bandwidth of the speech signal by extending the speech signal representation of the speech signal using the regression deep neural network trained with the cost function that includes the one or more discriminative terms;

wherein at least one of: (i) the one or more discriminative terms include at least one of a fricative-to-vowel power ratio and a function thereof; (ii) the one or more discriminative terms preserve relations of statistics between different phoneme classes in the actual values of the speech signal representation and the estimated values of the speech signal representation; and (iii) the operations further comprising reproducing an average power ratio at an output of the regression deep neural network.

6. The computer program product of claim 5 wherein the speech signal representation is obtained by decomposing the speech signal into a spectral envelope and an excitation signal, and wherein the spectral envelope is extended using the regression deep neural network trained with the cost function.

7. The computer program product of claim 5 wherein the cost function preserves a power ratio between the different phoneme classes in the actual values of the speech signal representation and the estimated values of the speech signal representation.

8. The computer program product of claim 5 wherein the cost function preserves the power ratio between the different phoneme classes using a weighted sum of K power ratio errors between the different phoneme classes.

9. A computing system including one or more processors and one or more memories configured to perform operations comprising:

transforming a speech signal into a speech signal representation;

training, with a cost function, a regression deep neural network to minimize a mean squared error between actual values of the speech signal representation and estimated values of the speech signal representation, wherein the cost function includes one or more discriminative terms; and

extending bandwidth of the speech signal by extending the speech signal representation of the speech signal using the regression deep neural network trained with the cost function that includes the one or more discriminative terms;

wherein at least one of: (i) the one or more discriminative terms include at least one of a fricative-to-vowel power ratio and a function thereof; (ii) the one or more discriminative terms preserve relations of statistics between different phoneme classes in the actual values of the speech signal representation and the estimated values of the speech signal representation; and (iii) the operations further comprising reproducing an average power ratio at an output of the regression deep neural network.

10. The computing system of claim 9 wherein the speech signal representation is obtained by decomposing the speech signal into a spectral envelope and an excitation signal, and wherein the spectral envelope is extended using the regression deep neural network trained with the cost function.

11. The computing system of claim 9 wherein the cost function preserves a power ratio between the different phoneme classes in the actual values of the speech signal representation and the estimated values of the speech signal representation, and wherein the cost function preserves the power ratio between the different phoneme classes using a weighted sum of K power ratio errors between the different phoneme classes.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2018
From: FAUBEL, FRIEDRICH; SAUTTER, JONAS
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 045610/0041 →