IP Library › Granted Patent US 12,346,804
Granted Patent B2
US 12,346,804 · App. 17/288,848 · Granted Jul 1, 2025

Acoustic model learning apparatus, model learning apparatus, method and program for the same

Inventors: Takafumi Moriya (Tokyo, JP); Yusuke Shinohara (Tokyo, JP); Yoshikazu Yamaguchi (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06N3/08G06N3/045G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,804
App. No.
17/288,848
Granted
Jul 1, 2025
Kind
B2
Abstract

An acoustic model learning apparatus includes a parameter updating part configured to update a parameter of a second acoustic model on the basis of a first loss for a feature amount for training, based on output probability distribution of the second acoustic model which is a neural network acoustic model to be trained, and a second loss for a feature amount for training, based on an intermediate feature amount of a first acoustic model which is a trained neural network acoustic model and an intermediate feature amount of the second acoustic model.

Claims (32)

1. A model learning apparatus for training an acoustic model comprising:

a hardware processor that:

retrieves an intermediate feature amount of a first model, wherein the first model represents a trained acoustic model comprising a combination of an initial layer, a plurality of intermediate layers, and an output layer, and

the trained acoustic model converts a feature amount extracted from an acoustic signal for training into a first output probability distribution of phoneme state in voice recognition;

updates a parameter of a second model on a basis of a first loss for a feature amount for training and a second loss for the feature amount for training,

wherein the first loss is based on output probability distribution of an output layer of the second model which is a neural network model as a voice recognition model to be trained,

the second loss is based on the intermediate feature amount of an m-th intermediate layer of the first model and an intermediate feature amount of an n-th intermediate layer of the second model,

a number of dimensions of the m-th intermediate layer of the first model is same as a number of dimensions of the n-th intermediate layer of the second model, and

wherein a trained second model is used in voice recognition where an input speech signal is converted into output text.

2. The model learning apparatus according to claim 1 ,

wherein the second model is a neural network acoustic model and the first model is a neural network acoustic model.

3. The model learning apparatus according to claim 1 ,

wherein, in a case where the number of dimensions of an intermediate layer of the second model is smaller than the number of dimensions of an intermediate layer of a model to be utilized as the first model, a bottleneck layer having the same number of dimensions as the number of dimensions of the intermediate layer of the second model is provided between an output layer of the model to be utilized as the first model and a layer one layer before the output layer, and a model which is trained again on a basis of the feature amount for training and a correct solution to the feature amount for training is used as the first model.

4. The model learning apparatus according to claim 1 ,

wherein the hardware processor calculates a first A loss from output probability distribution of the first model for the feature amount for training and the output probability distribution of the second model for the feature amount for training; and

the hardware processor calculates a first B loss from a correct solution to the feature amount for training and the output probability distribution of the second model for the feature amount for training, and

the first loss includes the first A loss and the first B loss.

5. A model learning method, implemented by a model learning apparatus that includes a hardware processor, comprising:

a retrieving step in which the hardware processor retrieves an intermediate feature amount of a first model, wherein the first model represents a trained acoustic model comprising a combination of an initial layer, a plurality of intermediate layers, and an output layer, and

the trained acoustic model converts a feature amount extracted from an acoustic signal for training into a first output probability distribution of phoneme state in voice recognition;

a parameter updating step in which the hardware processor updates a parameter of a second model on a basis of a first loss for a feature amount for training and a second loss for the feature amount for training,

wherein the first loss is based on output probability distribution of an output layer of the second model which is a neural network model as a voice recognition model to be trained,

the second loss is based on the intermediate feature amount of an m-th intermediate layer of the first model and an intermediate feature amount of an n-th intermediate layer of the second model,

a number of dimensions of the m-th intermediate layer of the first model is same as a number of dimensions of the n-th intermediate layer of the second model, and

wherein a trained second model is used in voice recognition where an input speech signal is converted into output text.

6. The model learning method according to claim 5 ,

wherein the second model is a neural network acoustic model and the first model is a neural network acoustic model.

7. The model learning method according to claim 5 ,

wherein the number of dimensions of an m-th layer of the first model is the same as the number of dimensions of an n-th layer of the second model.

8. The model learning method according to claim 5 ,

wherein, in a case where the number of dimensions of an intermediate layer of the second model is smaller than the number of dimensions of an intermediate layer of a model to be utilized as the first model, a bottleneck layer having the same number of dimensions as the number of dimensions of the intermediate layer of the second model is provided between an output layer of the model to be utilized as the first model and a layer one layer before the output layer, and a model which is trained again on a basis of the feature amount for training and a correct solution to the feature amount for training is used as the first model.

9. A non-transitory computer-readable recording medium having recorded thereon a program for causing a computer to function as the model learning apparatus according to claim 1 .

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0641 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2021
From: MORIYA, TAKAFUMI; SHINOHARA, YUSUKE; YAMAGUCHI, YOSHIKAZU
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 056043/0863 →
Priority Claims (1)
JP 2018-203094 · Oct 29, 2018 · national
Continuity (1)
Related Publication 20220004868A1 · Jan 6, 2022
References Cited (4)
US 20170169313A1 · Choi · 2017 [cited by examiner]
Tsukimoto, et al., “Extracting rules from trained neural networks”, IEEE Transactions on Neural Networks, vol. 11, No. 2, Mar. 2000 (Year: 2000). [cited by examiner]
Geoffrey Hinton et al. (2012) “Deep Neural Networks for Acoustic Modeling in Speech Recognition,” IEEE Signal Processing Magazine, vol. 29, No. 6, pp. 82-97. [cited by applicant]
Geoffrey Hinton et al. (2015) “Distilling the Knowledge in a Neural Network.” arXiv: 1503.02531v1, Mar. 9, 2015. [cited by applicant]