IP Library Granted Patent US 10,089,979
Granted Patent B2
US 10,089,979 · App. 14/737,907 · Granted Oct 2, 2018

Signal processing algorithm-integrated deep neural network-based speech recognition apparatus and learning method thereof

Inventors: Hoon Chung (Daejeon, KR); Jeon Gue Park (Daejeon, KR); Sung Joo Lee (Daejeon, KR); Yun Keun Lee (Daejeon, KR)
Assignee: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,089,979
App. No.
14/737,907
Filed
Jun 12, 2015
Granted
Oct 2, 2018
Kind
B2
Examiner
KY, KEVIN
Art Unit
2669
USPC
704/232
Abstract

Provided are a signal processing algorithm-integrated deep neural network (DNN)-based speech recognition apparatus and a learning method thereof. A model parameter learning method in a deep neural network (DNN)-based speech recognition apparatus implementable by a computer includes converting a signal processing algorithm for extracting a feature parameter from a speech input signal of a time domain into signal processing deep neural network (DNN), fusing the signal processing DNN and a classification DNN, and learning a model parameter in a deep learning model in which the signal processing DNN and the classification DNN are fused.

Claims (39)

1. A speech recognition method performed by a deep neural network (DNN)-based speech recognition apparatus implementable by a computer, the speech recognition method comprising:

performing a model parameter learning method comprising:

converting a signal processing algorithm into a signal processing deep neural network (DNN), the signal processing algorithm being for extracting a feature parameter from a speech input signal of a time domain, converting the signal processing algorithm including:

converting pre-emphasis, windowing, and discrete Fourier transform operations of the signal processing algorithm into a first single linear operation (W 1 ),

converting a mel-frequency filter bank and a logarithm operation of the signal processing algorithm into a single linear and log operation (W 2 ), and

converting a discrete cosine transform operation of the signal processing algorithm into a second single linear operation (W 3 );

initializing the signal processing DNN with a signal processing algorithm coefficient;

converting a speech input signal (X) to a first value (X 1 ) using the first single linear operation (W 1 );

converting the first value (X 1 ) to a second value (X 2 ) using the single linear and log operation (W 2 );

converting the second value (X 2 ) to a third value (X 3 ) using the second single linear operation (W 3 );

outputting the third value (X 3 );

using the third value (X 3 ) as an input with respect to a classification DNN;

fusing the signal processing DNN and the classification DNN to create a deep learning model;

receiving the speech input signal; and

learning, using the speech input signal of the time domain, a model parameter in the deep learning model; and

performing speech recognition using the deep learning model having the learned model parameter to output a word recognized by the deep learning model.

2. The speech recognition method of claim 1 , wherein the converting a signal processing algorithm comprises: converting a plurality of linear operations forming the signal processing algorithm into a single linear operation by using a matrix inner product or outer product.

3. The speech recognition method of claim 1 , wherein the fusing the signal processing DNN and a classification DNN comprises: inputting a feature parameter output from the signal processing DNN to the classification DNN.

4. The speech recognition method of claim 1 , wherein the learning a model parameter comprises: adjusting a model parameter to generate a target output value with respect to the speech input signal of the time domain.

5. The speech recognition method of claim 4 , wherein the learning a model parameter comprises: determining a model parameter with which an error of an output value with respect to the speech input signal of the time domain is minimized by using a back-propagation learning algorithm.

6. A signal processing algorithm-integrated deep neural network-based speech recognition apparatus including at least one processor and a non-volatile memory storing a code executed by the at least one processor, wherein the processor:

converts a signal processing algorithm into a signal processing deep neural network (DNN), the signal processing algorithm being for extracting a feature parameter from a speech input signal of a time domain, converting the signal processing algorithm including:

converting pre-emphasis, windowing, and discrete Fourier transform operations of the signal processing algorithm into a first single linear operation (W 1 ),

converting a mel-frequency filter bank and a logarithm operation of the signal processing algorithm into a single linear and log operation (W 2 ), and

converting a discrete cosine transform operation of the signal processing algorithm into a second single linear operation (W 3 ),

initializes the signal processing DNN with a signal processing algorithm coefficient;

converts a speech input signal (X) to a first value (X 1 ) using the first single linear operation (W 1 );

converts the first value (X 1 ) to a second value (X 2 ) using the single linear and log operation (W 2 );

converts the second value (X 2 ) to a third value (X 3 ) using the second single linear operation (W 3 );

outputs the third value (X 3 ); and

uses the third value (X 3 ) as an input with respect to a classification DNN

fuses the signal processing DNN and the classification DNN to form a deep learning model,

receives a speech input signal of the time domain,

learns, using the speech input signal of the time domain, a model parameter in the deep learning model, and

performs speech recognition using the deep learning model having the learned model parameter to output a word recognized by the deep learning model.

7. The signal processing algorithm-integrated deep neural network-based speech recognition apparatus of claim 6 , wherein the processor converts a plurality of linear operations included in the signal processing algorithm into a single linear operation by using a matrix inner product or outer product.

8. The signal processing algorithm-integrated deep neural network-based speech recognition apparatus of claim 6 , wherein the processor inputs the feature parameter output from the signal processing DNN to the classification DNN.

9. The signal processing algorithm-integrated deep neural network-based speech recognition apparatus of claim 6 , wherein the processor adjusts a model parameter to generate a target output value with respect to the speech input signal of the time domain.

10. The signal processing algorithm-integrated deep neural network-based speech recognition apparatus of claim 9 , wherein the processor determines a model parameter with which an error of an output value with respect to the speech input signal of the time domain is minimized by using a back-propagation learning algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2015
From: CHUNG, HOON; PARK, JEON GUE; LEE, SUNG JOO; LEE, YUN KEUN
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 035905/0866 →
Priority Claims (1)
KR 10-2014-0122803 · Sep 16, 2014 · national
Continuity (1)
Related Publication 20160078863A1 · Mar 17, 2016
Cited By (1)
US 12,413,930