Speaker embedding device, speaker embedding method, and speaker embedding program
A speaker embedding apparatus includes processing circuitry configured to accept input of voice data, generate utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data, and use a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and train a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input.
1 . A speaker embedding apparatus comprising:
processing circuitry configured to:
accept input of voice data;
generate utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data; and
use a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and train a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input,
the processing circuitry is further configured to:
accept input of voice data to be converted into a speaker vector,
generate the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and
input a duration length for each utterance indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.
2 . The speaker embedding apparatus according to claim 1 , wherein the processing circuitry is further configured to:
use utterances and a duration length for each of the utterances indicated in the utterance unit segmentation information as training data, and when utterances of a speaker and a duration length for each of the utterances are input, train a speaker identification model for outputting an identification result of the speaker.
3 . The speaker embedding apparatus according to claim 2 , wherein the processing circuitry is further configured to:
accept input of voice data to be converted into a speaker vector,
generate the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and
input utterances and a duration length for each of the utterances indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.
4 . The speaker embedding apparatus according to claim 1 , wherein the processing circuitry is further configured to:
convert a duration length for each utterance indicated in the generated utterance unit segmentation information into a one-dimensional numerical expression, and
train the speaker identification model using the converted one-dimensional numerical expression of the duration length for each utterance as training data.
5 . The speaker embedding apparatus according to claim 4 , wherein the processing circuitry is further configured to:
convert utterances and the duration length for each of the utterances indicated in the generated utterance unit segmentation information into a one-dimensional numerical expression, and
train the speaker identification model using the converted one-dimensional numerical expression of the utterances and the duration length for each of the utterances as training data.
6 . A speaker embedding method executed by a speaker embedding apparatus, comprising:
accepting input of voice data;
generating utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data; and
using a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and training a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input,
the speaker embedding method is further configured to:
accepting input of voice data to be converted into a speaker vector,
generating the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and
inputting a duration length for each utterance indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.
7 . A non-transitory computer-readable recording medium storing therein a speaker embedding program that causes a computer to execute a process comprising:
accepting input of voice data;
generating utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data; and
using a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and training a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input,
the process is further configured to:
accepting input of voice data to be converted into a speaker vector,
generating the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and
inputting a duration length for each utterance indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.