DIALOGUE LEARNING APPARATUS, RESPONSE AUDIO GENERATION APPARATUS, DIALOGUE LEARNING METHOD, RESPONSE AUDIO GENERATION METHOD AND PROGRAM
A dialogue learning device comprising: a dialogue data acquisition unit configured to acquire dialogue data including a dialogue context indicating text of a dialogue and speech data of a response sentence of the dialogue; an acoustic feature value calculation unit configured to calculate an acoustic feature value on the basis of the speech data; and a dialogue learning unit configured to learn a dialogue generation model for generating a dialogue on the basis of the dialogue context and data indicating the calculated acoustic feature value.
1 . A dialogue learning device comprising:
a dialogue data acquisition unit configured to acquire dialogue data including a dialogue context indicating text of a dialogue and speech data of a response sentence of the dialogue;
an acoustic feature value calculation unit configured to calculate an acoustic feature value on the basis of the speech data; and
a dialogue learning unit configured to learn a dialogue generation model for generating a dialogue on the basis of the dialogue context and data indicating the calculated acoustic feature value.
2 . The dialogue learning device according to claim 1 further comprising
a quantized acoustic feature value calculation unit configured to calculate a quantized acoustic feature value by clustering on the basis of the calculated acoustic feature value, wherein
the dialogue learning unit is configured to learn the dialogue generation model on the basis of the dialogue context and data indicating the quantized acoustic feature value.
3 . The dialogue learning device according to claim 1 , wherein
the dialogue data acquisition unit is configured to acquire dialogue data including the dialogue context and text data of the response sentence of the dialog, and
the acoustic feature value calculation unit is configured to calculate an acoustic feature value on the basis of the text data.
4 . The dialogue learning device according to claim 1 , wherein
the dialogue learning unit is configured to learn the learned dialogue generation model learned based on the text data based on the basis of the dialogue context and the data indicating the calculated acoustic feature value.
5 . A response speech generation device comprising:
a dialogue context acquisition unit configured to acquire a dialogue context indicating text of a dialogue;
a quantized acoustic feature value calculation unit configured to calculate a quantized acoustic feature value on the basis of the dialogue context using a dialogue generation model for generating a dialog, and
a response sentence speech data generation unit configured to generate speech data indicating a response sentence on the basis of the quantized data indicating the acoustic feature value.
6 . A dialogue learning method executed by a dialogue learning device, comprising:
acquiring dialogue data including a dialogue context indicating text of a dialogue and speech data of a response sentence of the dialogue;
calculating an acoustic feature value on the basis of the speech data; and
learning a dialogue generation model for generating a dialogue on the basis of the dialogue context and data indicating the calculated acoustic feature value.
7 . A response speech generation method executed by a response speech generation device, comprising:
acquiring a dialogue context indicating text of a dialogue;
calculating a quantized acoustic feature value on the basis of the dialogue context using a dialogue generation model for generating a dialogue; and
generating speech data indicating a response sentence on the basis of data indicating the quantized acoustic feature value.
8 . (canceled)
9 . The dialogue learning method according to claim 6 further comprising
calculating a quantized acoustic feature value by clustering on the basis of the calculated acoustic feature value, wherein
the dialogue generation model is learnt on the basis of the dialogue context and data indicating the quantized acoustic feature value.
10 . The dialogue learning method according to claim 6 , wherein
acquiring dialogue data including the dialogue context and text data of the response sentence of the dialog, and
calculating an acoustic feature value on the basis of the text data.
11 . The dialogue learning method according to claim 6 , wherein
the learned dialogue generation model is learnt based on the text data of the dialogue context and the data indicating the calculated acoustic feature value.
12 . The response speech generation device according to claim 5 , wherein a quantized acoustic feature quantity is calculated from a discretized dialogue context, wherein the quantized acoustic feature quantity indicates speech of an appropriate response sentence.
13 . The response speech generation device according to claim 5 , wherein a plurality of clusters is generated based on collected speech vectors.
14 . The response speech generation device according to claim 13 , wherein a representative point representing average of the speech vectors and a cluster number of the plurality of clusters are paired and stored as a codebook.
15 . The response speech generation method according to claim 7 , wherein acoustic feature data corresponding to the dialogue context is quantized without generating text data.
16 . The response speech generation method according to claim 7 , comprising:
calculating a quantized acoustic feature quantity from a discretized dialogue context, wherein the quantized acoustic feature quantity indicates speech of an appropriate response sentence.
17 . The response speech generation method according to claim 7 , wherein a plurality of clusters is generated based on collected speech vectors.
18 . The response speech generation method according to claim 17 , wherein a representative point representing average of the speech vectors and a cluster number of the plurality of clusters are paired and stored as a codebook.
19 . The response speech generation method according to claim 7 , wherein acoustic feature data corresponding to the dialogue context is quantized without generating text data.
20 . The response speech generation method according to claim 7 , wherein a generation model is trained based on a text-speech pair.
21 . The response speech generation method according to claim 7 , wherein a response sentence text data is extracted from a dialogue data and the acoustic feature amount is calculated from the response sentence text data.