IMAGE CAPTION GENERATION MODEL LEARNING APPARATUS, IMAGE CAPTION GENERATION APPARATUS, IMAGE CAPTION GENERATION MODEL LEARNING METHOD, IMAGE CAPTION GENERATION METHOD, AND PROGRAM
An image caption generation model learning apparatus uses, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data; and learns an image parameter that is a model parameter for image hidden information generation, a text parameter that is a model parameter for text hidden information generation, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.
1 . An image caption generation model learning apparatus comprising:
processing circuitry configured to
use, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data; and
learn an image parameter that is a model parameter for image hidden information generation, a text parameter that is a model parameter for text hidden information generation, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.
2 . The image caption generation model learning apparatus according to claim 1 ,
the processing circuitry configured to
generate image hidden information from the image and the image parameter;
generate text hidden information from the first language text and the text parameter;
generate inter-crossmodal invariant information from the image hidden information or the text hidden information and the crossmodal parameter;
generate a text generation probability from the inter-crossmodal invariant information and the output parameter; and
estimate various model parameters such that a sum of a text generation probability of a text corresponding to a caption and a text generation probability of a text corresponding to a translation result of machine translation becomes maximum.
3 . An image caption generation apparatus comprising:
processing circuitry configured to
generate a caption describing an input image based on an image parameter that is model parameter for image hidden information generation learned by using, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.
4 . The image caption generation apparatus according to claim 3 ,
the processing circuitry configured to
generate image hidden information from an image and the image parameter;
generate inter-crossmodal invariant information from the image hidden information and the crossmodal parameter; and
generate a text generation probability from the inter-crossmodal invariant information and the output parameter and generates a text serving as the caption of the image.
5 . An image caption generation model learning method executed by an image caption generation model learning apparatus, the image caption generation model learning method comprising:
using, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data; and
learning an image parameter that is a model parameter for image hidden information generation, a text parameter that is a model parameter for text hidden information generation, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.
6 . An image caption generation method executed by an image caption generation apparatus, the image caption generation method comprising:
generating a caption describing an input image based on an image parameter that is model parameter for image hidden information generation learned by using, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.
7 . A non-transitory computer readable medium storing a computer program for causing a computer to function as the image caption generation model learning apparatus according to claim 1 .
8 . A non-transitory computer readable medium storing a computer program for causing a computer to function as the image caption generation apparatus according to claim 3 .