IP Library Granted Patent US 10,283,112
Granted Patent B2
US 10,283,112 · App. 15/905,789 · Granted May 7, 2019

System and method for neural network based feature extraction for acoustic model development

Inventors: Srinath Cheluvaraja (Carmel, IN); Ananth Nagaraja Iyer (Carmel, IN)
G10L15/144G06N3/02G10L15/14G10L15/16G10L25/24G10L25/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,283,112
App. No.
15/905,789
Granted
May 7, 2019
Kind
B2
Abstract

A system and method are presented for neural network based feature extraction for acoustic model development. A neural network may be used to extract acoustic features from raw MFCCs or the spectrum, which are then used for training acoustic models for speech recognition systems. Feature extraction may be performed by optimizing a cost function used in linear discriminant analysis. General non-linear functions generated by the neural network are used for feature extraction. The transformation may be performed using a cost function from linear discriminant analysis methods which perform linear operations on the MFCCs and generate lower dimensional features for speech recognition. The extracted acoustic features may then be used for training acoustic models for speech recognition systems.

Claims (26)

1. A method for training acoustic models in speech recognition systems, the method comprising:

extracting high dimensional acoustic features from one or more input speech signals;

extracting, using a neural network, linear discriminant analysis features from the high dimensional acoustic features, the neural network comprising a plurality of neurons and a plurality of activation functions, wherein the neural network is trained by computing parameters of the activation functions and weights of connections between neurons of the neural network using stochastic gradient descent on a cost function based on training data; and

training an acoustic model using the linear discriminant analysis features.

2. The method of claim 1 , wherein the activation functions comprise linear discriminant analysis functions.

3. The method of claim 1 , wherein the training data comprises comprising prealigned feature data and linear discriminant analysis class labels.

4. The method of claim 3 , wherein the cost function represents a measure of inter-class separation between the class labels.

5. The method of claim 1 , wherein the cost function is a linear discriminant analysis cost function.

6. The method of claim 1 , wherein the high dimensional acoustic features comprise mel-frequency cepstral coefficients.

7. The method of claim 1 , wherein the linear discriminant analysis features have a lower dimension than the high dimensional acoustic features.

8. The method of claim 1 , wherein the neural network is a recurrent network.

9. The method of claim 1 , wherein the training comprises using at least one of maximum likelihood method and an expectation maximization method.

10. A system for training acoustic models in speech recognition systems, the system comprising:

a processor; and

a memory in communication with the processor, the memory storing instructions therein that, when executed by the processor, cause the processor to:

extract high dimensional acoustic features from one or more input speech signals;

extract linear discriminant analysis features from the high dimensional acoustic features using a neural network, the neural network comprising a plurality of neurons and a plurality of activation functions, wherein the neural network is trained by computing parameters of the activation functions and weights of connections between neurons of the neural network using stochastic gradient descent on a cost function based on training data; and

train an acoustic model using the linear discriminant analysis features.

11. The system of claim 10 , wherein the activation functions comprise linear discriminant analysis functions.

12. The system of claim 10 , wherein the training data comprises prealigned feature data and linear discriminant analysis class labels.

13. The system of claim 12 , wherein the cost function represents a measure of inter-class separation between the class labels.

14. The system of claim 10 , wherein the cost function is a linear discriminant analysis cost function.

15. The system of claim 10 , wherein the high dimensional acoustic features comprise mel-frequency cepstral coefficients.

16. The system of claim 10 , wherein the linear discriminant analysis features have a lower dimension than the high dimensional acoustic features.

17. The system of claim 10 , wherein the neural network is a recurrent network.

18. The system of claim 10 , wherein the memory further stores instructions that, when executed by the processor, cause the processor to train the acoustic model by using at least one of a maximum likelihood method and an expectation maximization method.

Assignments (6)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 050860/0227 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070096/0452 →
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067646/0448 →
CORRECTIVE ASSIGNMENT TO CORRECT THE TO ADD PAGE 2 OF THE SECURITY AGREEMENT WHICH WAS INADVERTENTLY OMITTED PREVIOUSLY RECORDED ON REEL 049916 FRAME 0454. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY AGREEMENT. Recorded Oct 29, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 050860/0227 →
SECURITY AGREEMENT Recorded Jul 31, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 049916/0454 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2018
From: CHELUVARAJA, SRINATH; IYER, ANANTH NAGARAJA
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 045793/0529 →
Continuity (2)
Continuation 14985560 · Dec 31, 2015
Related Publication 20180190267A1 · Jul 5, 2018