IP Library › Granted Patent US 10,621,972
Granted Patent B2
US 10,621,972 · App. 15/914,066 · Granted Apr 14, 2020

Method and device extracting acoustic feature based on convolution neural network and terminal device

Inventors: Chao Li (Beijing, CN); Xiangang Li (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G10L15/02G10L15/16G10L15/22G10L25/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,621,972
App. No.
15/914,066
Granted
Apr 14, 2020
Kind
B2
Abstract

The present disclosure provides a method and a device for extracting an acoustic feature based on a convolution neural network and a terminal device. The method includes: arranging speech to be recognized into a speech spectrogram with a predetermined dimension number; and recognizing the speech spectrogram with the predetermined dimension number by the convolution neural network to obtain the acoustic feature of the speech to be recognized.

Claims (45)

1. A method for extracting an acoustic feature based on a convolution neural network, comprising:

arranging speech to be recognized into a speech spectrogram with a predetermined dimension number; and

recognizing the speech spectrogram with the predetermined dimension number by the convolution neural network, to obtain the acoustic feature of the speech to be recognized,

wherein the convolution neural network comprises a structure of one of the following:

a residual network structure; and

a jump link structure;

wherein before recognizing the speech spectrogram with the predetermined dimension number by the convolution neural network, further comprising:

configuring a model of the structure of the convolution neural network;

wherein configuring the model of the structure of the convolution neural network comprises:

for a 64-channel filter block consisting of a convolution directed acycline graph of a 64-channel filter bank, performing a down-sampling by a pooling layer both in a time domain and in a frequency domain;

for a 128-channel filter block consisting of a convolution directed acycline graph of a 128-channel filter bank, performing the down-sampling by the pooling layer both in the time domain and in the frequency domain;

for a 256-channel filter block consisting of a convolution directed acycline graph of a 256-channel filter bank, performing the down-sampling by the pooling layer in the frequency domain; and

for a 512-channel filter block consisting of a convolution directed acycline graph of a 512-channel filter bank, performing the down-sampling by the pooling layer in the frequency domain.

2. The method according to claim 1 , wherein arranging speech to be recognized into a speech spectrogram with a predetermined dimension number comprises:

extracting a predetermined multidimensional feature vector from the speech to be recognized every predetermined time interval to arrange the speech to be recognized into the speech spectrogram with the predetermined dimension number.

3. A device for extracting an acoustic feature based on a convolution neural network, comprising:

one or more processors;

a memory, configured to store one or more program modules executable by the one or more processors, wherein the one or more program modules comprise:

a generating module, configured to arrange speech to be recognized into a speech spectrogram with a predetermined dimension number; and

a recognizing module, configured to recognize the speech spectrogram with the predetermined dimension number by the convolution neural network to obtain the acoustic feature of the speech to be recognized,

wherein the convolution neural network comprises a structure of one of the following:

a residual network structure; and

a jump link structure;

wherein the one or more program modules further comprise:

a configuring module, configured to configure a model of the structure of the convolution neural network before the recognizing module recognizes the speech spectrogram with the predetermined dimension number;

wherein, the configuring module is configured to:

for a 64-channel filter block consisting of a convolution directed acycline graph of a 64-channel filter bank, perform a down-sampling by a pooling layer both in a time domain and in a frequency domain;

for a 128-channel filter block consisting of a convolution directed acycline graph of a 128-channel filter bank, perform the down-sampling by the pooling layer both in the time domain and in the frequency domain;

for a 256-channel filter block consisting of a convolution directed acycline graph of a 256-channel filter bank, perform the down-sampling by the pooling layer in the frequency domain; and

for a 512-channel filter block consisting of a convolution directed acycline graph of a 512-channel filter bank, perform the down-sampling by the pooling layer in the frequency domain.

4. The device according to claim 3 , wherein,

the generating module is configured to extract a predetermined multidimensional feature vector from the speech to be recognized every predetermined time interval to arrange the speech to be recognized into the speech spectrogram with the predetermined dimension number.

5. A non-transitory computer readable storage medium comprising computer executable instructions configured to perform a method for extracting an acoustic feature based on a convolution neural network when executed by a computer processor, the method comprising:

arranging speech to be recognized into a speech spectrogram with a predetermined dimension number; and

recognizing the speech spectrogram with the predetermined dimension number by the convolution neural network, to obtain the acoustic feature of the speech to be recognized,

wherein the convolution neural network comprises a structure of one of the following:

a residual network structure; and

a jump link structure;

wherein before recognizing the speech spectrogram with the predetermined dimension number by the convolution neural network, further comprising:

configuring a model of the structure of the convolution neural network;

wherein configuring the model of the structure of the convolution neural network comprises:

for a 64-channel filter block consisting of a convolution directed acycline graph of a 64-channel filter bank, performing a down-sampling by a pooling layer both in a time domain and in a frequency domain;

for a 128-channel filter block consisting of a convolution directed acycline graph of a 128-channel filter bank, performing the down-sampling by the pooling layer both in the time domain and in the frequency domain;

for a 256-channel filter block consisting of a convolution directed acycline graph of a 256-channel filter bank, performing the down-sampling by the pooling layer in the frequency domain; and

for a 512-channel filter block consisting of a convolution directed acycline graph of a 512-channel filter bank, performing the down-sampling by the pooling layer in the frequency domain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2018
From: LI, CHAO
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 045645/0869 →
Priority Claims (1)
CN 2017 1 0172622 · Mar 21, 2017 · national
Continuity (1)
Related Publication 20180277097A1 · Sep 27, 2018