IP Library › Granted Patent US 10,950,225
Granted Patent B2
US 10,950,225 · App. 16/337,081 · Granted Mar 16, 2021

Acoustic model learning apparatus, method of the same and program

Inventors: Taichi Asami (Yokosuka, JP); Takashi Nakamura (Yokosuka, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L15/16G06F17/18G06N3/0472G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,950,225
App. No.
16/337,081
Granted
Mar 16, 2021
Kind
B2
Abstract

An acoustic model learning apparatus includes a first output probability distribution calculating part that calculates a first output probability distribution including a distribution of output probabilities of respective units of an output layer using a feature amount obtained from an acoustic signal for learning and a learned first acoustic model including a neural network, and the first output probability distribution calculating part obtains the first output probability distribution using a smoothing parameter made up of a real value greater than 0 as input so that the first output probability distribution approaches a uniform distribution as the smoothing parameter is greater, and calculates the first output probability distribution by obtaining logits of respective units of an output layer using the feature amount obtained from the acoustic signal for learning and the first acoustic model and setting a value of the smoothing parameter greater in the case where an output unit number with the greatest logit value is different from a correct unit number than in the case where the output unit number with the greatest logit value matches the correct unit number.

Claims (33)

1. An acoustic model learning apparatus comprising:

processing circuitry configured to:

execute a first output probability distribution calculating processing in which the processing circuitry calculates a first output probability distribution including a distribution of output probabilities of respective units of an output layer using a feature amount obtained from an acoustic signal for learning and a learned first acoustic model including a neural network;

execute a second output probability distribution calculating processing in which the processing circuitry calculates a second output probability distribution including a distribution of output probabilities of respective units of an output layer using the feature amount obtained from the acoustic signal for learning, and a second acoustic model which is different from the first acoustic model and which includes a neural network; and

execute a corrected model updating processing in which the processing circuitry calculates a second loss function from a correct unit number corresponding to the acoustic signal for learning and the second output probability distribution, calculates a cross entropy of the first output probability distribution and the second output probability distribution, obtains a weighted sum of the second loss function and the cross entropy, and updates a parameter of the second acoustic model so that the weighted sum decreases,

wherein in the first output probability distribution calculating processing the processing circuitry obtains the first output probability distribution using a smoothing parameter made up of a real value greater than 0 as input so that the first output probability distribution approaches a uniform distribution as the smoothing parameter is greater, and calculates the first output probability distribution by obtaining logits of respective units of an output layer using the feature amount obtained from the acoustic signal for learning and the first acoustic model, and setting a value of the smoothing parameter greater in a case where an output unit number with the greatest logit value is different from the correct unit number, than in a case where the output unit number with the greatest logit value matches the correct unit number.

2. The acoustic model learning apparatus according to claim 1 ,

the processing circuitry being further configured to:

execute a first output probability distribution correcting processing in which the processing circuitry sets a distribution of output probabilities in which, in a case where, among logits of respective units of an output layer obtained using the feature amount obtained from the acoustic signal for learning and the first acoustic model, an output unit number with the greatest logit value is different from the correct unit number, an output probability of an output unit corresponding to the output unit number with the greatest logit value is replaced with an output probability of an output unit corresponding to the correct unit number in the first output probability distribution, as a corrected first output probability distribution,

wherein in the corrected model updating processing the processing circuitry uses the corrected first output probability distribution.

3. An acoustic model learning apparatus comprising:

processing circuitry configured to:

execute a first output probability distribution calculating processing in which the processing circuitry calculates a first output probability distribution including a distribution of output probabilities of respective units of an output layer using a feature amount obtained from an acoustic signal for learning and a learned first acoustic model including a neural network;

execute a second output probability distribution calculating processing in which the processing circuitry calculates a second output probability distribution including a distribution of output probabilities of respective units of an output layer using the feature amount obtained from the acoustic signal for learning, and a second acoustic model which is different from the first acoustic model and which includes a neural network;

execute a first output probability distribution correcting processing in which the processing circuitry sets a distribution of output probabilities in which, in a case where, among logits of respective units of an output layer obtained using the feature amount obtained from the acoustic signal for learning and the first acoustic model, an output unit number with the greatest logit value is different from a correct unit number corresponding to the acoustic signal for learning, an output probability of an output unit corresponding to the output unit number with the greatest logit value is replaced with an output probability of an output unit corresponding to the correct unit number in the first output probability distribution, as a corrected first output probability distribution; and

execute a corrected model updating processing in which the processing circuitry calculates a second loss function from the correct unit number and the second output probability distribution, calculates a cross entropy of the corrected first output probability distribution and the second output probability distribution, obtains a weighted sum of the second loss function and the cross entropy, and updates a parameter of the second acoustic model so that the weighted sum decreases.

4. The acoustic model learning apparatus according to claim 3 ,

wherein in the first output probability distribution calculating processing the processing circuitry obtains the first output probability distribution using a smoothing parameter made up of a real value greater than 0 as input so that the first output probability distribution approaches a uniform distribution as the smoothing parameter is greater.

5. The acoustic model learning apparatus according to any one of claims 1 to 4 ,

the processing circuitry being further configured to:

execute an initial value setting processing in which the processing circuitry sets a parameter of the second acoustic model using a parameter of the first acoustic model, the first acoustic model and the second acoustic model including neural networks having similar structures.

6. An acoustic model learning method comprising:

a first output probability distribution calculating step of calculating a first output probability distribution including a distribution of output probabilities of respective units of an output layer using a feature amount obtained from an acoustic signal for learning and a learned first acoustic model including a neural network;

a second output probability distribution calculating step of calculating a second output probability distribution including a distribution of output probabilities of respective units of an output layer using the feature amount obtained from the acoustic signal for learning, and a second acoustic model which is different from the first acoustic model and which includes a neural network; and

a corrected model updating step of calculating a second loss function from a correct unit number corresponding to the acoustic signal for learning and the second output probability distribution, calculating a cross entropy of the first output probability distribution and the second output probability distribution, obtaining a weighted sum of the second loss function and the cross entropy, and updating a parameter of the second acoustic model so that the weighted sum decreases,

wherein, in the first output probability distribution calculating step, the first output probability distribution is obtained using a smoothing parameter made up of a real value greater than 0 as input so that the first output probability distribution approaches a uniform distribution as the smoothing parameter is greater, and the first output probability distribution is calculated by obtaining logits of respective units of an output layer using the feature amount obtained from the acoustic signal for learning and the first acoustic model, and setting a value of the smoothing parameter greater in a case where an output unit number with the greatest logit value is different from the correct unit number, than in a case where the output unit number with the greatest logit value matches the correct unit number.

7. An acoustic model learning method comprising:

a first output probability distribution calculating step of calculating a first output probability distribution including a distribution of output probabilities of respective units of an output layer using a feature amount obtained from an acoustic signal for learning and a learned first acoustic model including a neural network;

a second output probability distribution calculating step of calculating a second output probability distribution including a distribution of output probabilities of respective units of an output layer using the feature amount obtained from the acoustic signal for learning, and a second acoustic model which is different from the first acoustic model and which includes a neural network;

a first output probability distribution correcting step of setting a distribution of output probabilities in which, in a case where, among logits of respective units of an output layer obtained using the feature amount obtained from the acoustic signal for learning and the first acoustic model, an output unit number with the greatest logit value is different from a correct unit number corresponding to the acoustic signal for learning, an output probability of an output unit corresponding to the output unit number with the greatest logit value is replaced with an output probability of an output unit corresponding to the correct unit number in the first output probability distribution, as a corrected first output probability distribution; and

a corrected model updating step of calculating a second loss function from the correct unit number and the second output probability distribution, calculating a cross entropy of the corrected first output probability distribution and the second output probability distribution, obtaining a weighted sum of the second loss function and the cross entropy, and updating a parameter of the second acoustic model so that the weighted sum decreases.

8. A non-transitory computer-readable recording medium that records a program for causing a computer to function as the acoustic model learning apparatus according to any one of claims 1 to 4 .

9. A non-transitory computer-readable recording medium that records a program for causing a computer to function as the acoustic model learning apparatus according to claim 5 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2019
From: ASAMI, TAICHI; NAKAMURA, TAKASHI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 048713/0211 →
Priority Claims (1)
JP JP2016-193533 · Sep 30, 2016 · national
Continuity (1)
Related Publication 20200035223A1 · Jan 30, 2020