IP Library Granted Patent US 9,972,310
Granted Patent B2
US 9,972,310 · App. 14/985,560 · Granted May 15, 2018

System and method for neural network based feature extraction for acoustic model development

Inventors: Srinath Cheluvaraja (Carmel, IN); Ananth Nagaraja Iyer (Carmel, IN)
Assignee: Interactive Intelligence Group, Inc.
G10L15/144G06N3/02G10L15/14G10L15/16G10L25/24G10L25/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,972,310
App. No.
14/985,560
Granted
May 15, 2018
Kind
B2
Abstract

A system and method are presented for neural network based feature extraction for acoustic model development. A neural network may be used to extract acoustic features from raw MFCCs or the spectrum, which are then used for training acoustic models for speech recognition systems. Feature extraction may be performed by optimizing a cost function used in linear discriminant analysis. General non-linear functions generated by the neural network are used for feature extraction. The transformation may be performed using a cost function from linear discriminant analysis methods which perform linear operations on the MFCCs and generate lower dimensional features for speech recognition. The extracted acoustic features may then be used for training acoustic models for speech recognition systems.

Claims (42)

1. A method for training acoustic models in speech recognition systems, wherein the speech recognition system comprises a neural network, the method comprising the steps of:

a. extracting acoustic features from a speech signal using the neural network; and

b. processing the acoustic features into an acoustic model by the speech recognition system,

wherein the neural network comprises at least one of: activation functions with parameters, prealigned feature data, and training,

wherein the training is performed using a stochastic gradient descent method on a cost function, and

wherein the cost function is a linear discriminant analysis cost function.

2. The method of claim 1 , wherein the acoustic features are extracted from Mel Frequency Cepstral Coefficients.

3. The method of claim 1 , wherein the features are extracted from a speech signal spectrum.

4. The method of claim 1 , wherein the processing comprises using at least one of a maximum likelihood method and an expectation maximization method.

5. A method for training acoustic models in speech recognition systems, wherein the speech recognition system comprises a neural network, the method comprising the steps of:

a. extracting acoustic features from a speech signal using the neural network; and

b. processing the acoustic features into an acoustic model by the speech recognition system,

wherein the extracting of step (a) further comprises the step of optimizing a cost function,

wherein the cost function is capable of transforming general non-linear functions generated by the neural network.

6. The method of claim 5 , wherein the transforming comprises:

a. performing non-linear operations on the features; and

b. generating lower dimensional features for speech recognition.

7. The method of claim 6 , wherein the transforming further comprises the generation of Linear Discriminant Analysis transforms, wherein the transforms are generated by optimizing the cost function with linear activation functions.

8. The method of claim 5 , wherein the neural network carries activation functions with variable parameters.

9. The method of claim 8 , wherein the neural network is recurrent.

10. The method of claim 8 , wherein the parameters are determined during training of the neural network.

11. A method for training acoustic models in speech recognition systems, wherein the speech recognition system comprises a neural network, the method comprising the steps of:

a. extracting trainable features from an incoming audio signal using the neural network; and

b. processing the trainable features into an acoustic model by the speech recognition system,

wherein the neural network comprises at least one of: activation functions

with parameters, prealigned feature data, and training,

wherein the training is performed using a stochastic gradient descent

method on a cost function, and

wherein the cost function is a linear discriminant analysis cost function.

12. The method of claim 11 , wherein the trainable features are extracted from Mel Frequency Cepstral Coefficients, wherein the Mel Frequency Cepstral Coefficients are extracted from the audio signal.

13. The method of claim 11 , wherein the processing comprises using at least one of a maximum likelihood method and an expectation maximization method.

14. A method for training acoustic models in speech recognition systems, wherein the speech recognition system comprises a neural network, the method comprising the steps of:

a. extracting trainable features from an incoming audio signal using the neural network; and

b. processing the trainable features into an acoustic model by the speech recognition system,

wherein the extracting of step (a) further comprises the step of optimizing a cost function, wherein the cost function is capable of transforming general non-linear functions generated by the neural network.

15. The method of claim 14 , wherein the neural network carries activation functions with variable parameters.

16. The method of claim 15 , wherein the neural network is recurrent.

17. The method of claim 15 , wherein the parameters are determined during training of the neural network.

18. The method of claim 14 , wherein the transforming comprises:

a. performing non-linear operations on the features; and

b. generating lower dimensional features for speech recognition.

19. The method of claim 18 , wherein the transforming further comprises the generation of Linear Discriminant Analysis transforms, wherein the transforms are generated by optimizing the cost function with linear activation functions.

Assignments (7)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 04814/0387 Recorded Feb 5, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070115/0445 →
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 040815/0001 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070498/0001 →
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067644/0877 →
SECURITY AGREEMENT Recorded Feb 22, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.; ECHOPASS CORPORATION; GREENEDEN U.S. HOLDINGS II, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 048414/0387 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
SECURITY AGREEMENT Recorded Dec 5, 2016
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC., AS GRANTOR; ECHOPASS CORPORATION; INTERACTIVE INTELLIGENCE GROUP, INC.; BAY BRIDGE DECISION TECHNOLOGIES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 040815/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2016
From: CHELUVARAJA, SRINATH; IYER, ANANTH NAGARAJA
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 037821/0701 →
Continuity (1)
Related Publication 20170193988A1 · Jul 6, 2017