IP Library Granted Patent US 11,264,044
Granted Patent B2
US 11,264,044 · App. 16/074,367 · Granted Mar 1, 2022

Acoustic model training method, speech recognition method, acoustic model training apparatus, speech recognition apparatus, acoustic model training program, and speech recognition program

Inventors: Marc Delcroix (Soraku-gun, JP); Keisuke Kinoshita (Soraku-gun, JP); Atsunori Ogawa (Soraku-gun, JP); Takuya Yoshioka (Soraku-gun, JP); Tomohiro Nakatani (Soraku-gun, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L21/0208G10L15/063G10L15/16G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,264,044
App. No.
16/074,367
Granted
Mar 1, 2022
Kind
B2
Abstract

To begin with, an acoustic model training apparatus extracts speech features representing speech characteristics, and calculates an acoustic-condition feature representing a feature of an acoustic condition of the speech data using an acoustic-condition calculation model that is represented as a neural network, based on an acoustic-condition calculation model parameter characterizing the acoustic-condition calculation model. The acoustic model training apparatus then generates an adjusted parameter that is an acoustic model parameter adjusted based on the acoustic-condition feature, the acoustic model parameter characterizing an acoustic model represented as a neural network to which an output layer of the acoustic-condition calculation model is coupled. The acoustic model training apparatus then updates the acoustic model parameter based on the adjusted parameter and the speech features, and updates the acoustic-condition calculation model parameters based on the adjusted parameter and the speech features.

Claims (44)

1. A speech recognition apparatus comprising:

a memory; and

a processor coupled to the memory and programmed to execute a process comprising:

transforming speech data to be recognized into information identifying a symbol sequence by a neural network; and

adjusting at least a part of parameters of the neural network based on input acoustic-condition features, wherein

the transforming transforms the speech data into the information identifying the symbol sequence by the neural network in which the at least a part of parameters is adjusted by the adjusting,

wherein the neural network includes a same number of divided hidden layers as the number of the acoustic-condition features, and a layer for acquiring information identifying the symbol sequence by using an intermediate state output from each of the divided hidden layers, and

the adjusting adjusts parameters of each of the hidden layers based on the acoustic-condition features corresponding to each hidden layer.

2. The speech recognition apparatus according to claim 1 , wherein

the adjusting calculates a weighted sum as adjusted parameters by multiplying each parameters of each of the hidden layers by the acoustic-condition features adjusts parameters.

3. An acoustic model training apparatus configured to learn parameters of a neural network which transforms input speech data into information identifying a symbol sequence corresponding to the speech data, the apparatus comprising:

a memory; and

a processor coupled to the memory and programmed to execute a process comprising:

transforming training speech data into information identifying a symbol sequence by the neural network;

adjusting at least a part of parameters of the neural network based on input acoustic-condition features; and

updating each parameter of the neural network based on a result of a comparison between the information identifying the symbol sequence corresponding to training speech data obtained by the transformation of the training speech data using the neural network, in which the at least a part of parameters is adjusted by the adjusting, and a reference corresponding to the information identifying the symbol sequence,

wherein the neural network includes a same number of divided hidden layers as the number of the acoustic-condition features, and a layer for acquiring information identifying the symbol sequence by using an intermediate state output from each of the divided hidden layers, and

the adjusting adjusts parameters of each of the hidden layers based on the acoustic-condition features corresponding to each hidden layer.

4. The acoustic model training apparatus according to claim 3 , wherein

the adjusting calculates a weighted sum as adjusted parameters by multiplying each parameters of each of the hidden layers by the acoustic-condition features.

5. A speech recognition method executed by a speech recognition apparatus, the method comprising:

transforming speech data to be recognized into information identifying a symbol sequence by a neural network; and

adjusting at least a part of parameters of the neural network based on input acoustic-condition features, wherein

the transforming transforms the speech data into the information identifying the symbol sequence by the neural network in which the at least a part of parameters is adjusted by the adjusting,

wherein the neural network includes a same number of divided hidden layers as the number of the acoustic-condition features, and a layer for acquiring information identifying the symbol sequence by using an intermediate state output from each of the divided hidden layers, and

the adjusting adjusts parameters of each of the hidden layers based on the acoustic-condition features corresponding to each hidden layer.

6. An acoustic model training method executed by an acoustic model training apparatus configured to learn parameters of a neural network which transforms input speech data into information identifying a symbol sequence corresponding to the speech data, the apparatus, the method comprising:

transforming training speech data into information identifying a symbol sequence by the neural network;

adjusting at least a part of parameters of the neural network based on input acoustic-condition features; and

updating each parameter of the neural network based on a result of a comparison between the information identifying the symbol sequence corresponding to training speech data obtained by the transformation of the training speech data using the neural network, in which the at least a part of parameters is adjusted by the adjusting, and a reference corresponding to the information identifying the symbol sequence,

wherein the neural network includes a same number of divided hidden layers as the number of the acoustic-condition features, and a layer for acquiring information identifying the symbol sequence by using an intermediate state output from each of the divided hidden layers, and

the adjusting adjusts parameters of each of the hidden layers based on the acoustic-condition features corresponding to each hidden layer.

7. A non-transitory computer-readable recording medium having stored a program for speech recognition that causes a computer to execute a process comprising:

transforming speech data to be recognized into information identifying a symbol sequence by a neural network, and

adjusting at least a part of parameters of the neural network based on input acoustic-condition features, wherein

the transforming transforms the speech data into the information identifying the symbol sequence by the neural network in which the at least a part of parameters is adjusted by the adjusting,

wherein the neural network includes a same number of divided hidden layers as the number of the acoustic-condition features, and a layer for acquiring information identifying the symbol sequence by using an intermediate state output from each of the divided hidden layers, and

the adjusting adjusts parameters of each of the hidden layers based on the acoustic-condition features corresponding to each hidden layer.

8. A non-transitory computer-readable recording medium having stored a program for training an acoustic model that causes a computer to execute a process comprising:

transforming the training speech data into information identifying a symbol sequence by the neural network;

adjusting at least a part of parameters of the neural network based on the input acoustic-condition features; and

updating each parameter of the neural network based on a result of a comparison between the information identifying the symbol sequence corresponding to training speech data obtained by the transformation of the training speech data using the neural network, in which the at least a part of parameters is adjusted by the adjusting, and a reference corresponding to the information identifying the symbol sequence,

wherein the neural network includes a same number of divided hidden layers as the number of the acoustic-condition features, and a layer for acquiring information identifying the symbol sequence by using an intermediate state output from each of the divided hidden layers, and

the adjusting adjusts parameters of each of the hidden layers based on the acoustic-condition features corresponding to each hidden layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2018
From: DELCROIX, MARC; KINOSHITA, KEISUKE; OGAWA, ATSUNORI; YOSHIOKA, TAKUYA; NAKATANI, TOMOHIRO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 046516/0773 →
Priority Claims (1)
JP JP2016-018016 · Feb 2, 2016 · national
Continuity (1)
Related Publication 20210193161A1 · Jun 24, 2021
Cited By (2)
US 12,418,270 US 12,499,875