SUMMARY GENERATION APPARATUS, SUMMARY MODEL LEARNING APPARATUS, SUMMARY GENERATION METHOD, SUMMARY MODEL LEARNING METHOD, AND PROGRAM
A summary generation device includes: an image processing unit that receives an input of an image related to a moving image, and extracts at least a text from the image; a sound processing unit that receives an input of sound in the moving image, and extracts at least a text from the sound; and a summary generation unit that generates a summary text of the moving image from information extracted from the image and information extracted from the sound, using a trained summary model.
1 . A device comprising a processor configured to execute operations comprising:
receiving an input of an image related to a moving image, and extracts at least a text from the image;
receiving an input of sound in the moving image, and extracts at least a text from the sound; and
generating a summary text of the moving image from information extracted from the image and information extracted from the sound, using a trained summary model.
2 . A device comprising a processor configured to execute operations comprising:
receiving an input of an image of a moving image;
extracting at least a first training text from the image;
receiving an input of sound in the moving image;
extracting at least a second training text from the sound; and
training a summary model, using the first training text, the second training text, and a correct summary text of the moving image.
3 . The device according to claim 2 , the processor further configured to execute operations comprising:
acquiring the moving image and the correct summary text from a server in a network.
4 . The device according to claim 2 , the processor further configured to execute operations comprising:
performing pre-training on the summary model, using a pre-training text in a field associated with the moving image and a correct summary text of the pre-training text.
5 . The device according to claim 2 , the processor further configured to execute operations comprising:
generating at least one further training data set from a training data set,
wherein the training data set comprises first information extracted from the image, second information extracted from the sound, and the correct summary text of the training moving image,
the first information extracted from the image comprises the first training text, and
the second information extracted from the sound comprises the second training text.
6 . A method implemented by a computer, comprising:
a first image processing step of extracting at least a text from an image related to a moving image;
a first sound processing step of extracting at least a text from sound in the moving image; and
a summary generating step of generating a summary text of the moving image from information extracted from the image related to the moving image and information extracted from the sound, using a trained summary model.
7 . The method according to claim 6 , comprising:
a second image processing step of extracting at least a first training text from a training image related to a training moving image;
a second sound processing step of extracting at least a second training text from sound in the training moving image; and
a summary model training step of training a summary model, using the first training text, the second training text, and a correct summary text of the training moving image.
8 . (canceled)
9 . The device according to claim 1 , wherein the image related to the moving image is based on a frame of the moving image.
10 . The device according to claim 1 , the processor further configured to execute operations comprising:
performing pre-training on a summary model, using a pre-training text in a field associated with a pre-training moving image and a correct summary text of the pre-training text.
11 . The device according to claim 10 , the processor further configured to execute operations comprising:
receiving an input of a training image of a training moving image;
extracting at least a first training text from the training image;
receiving an input of training sound in the training moving image;
extracting at least a second training text from the training sound; and
training the summary model, using the first training text, the second training text, and a correct summary text of the training moving image.
12 . The device according to claim 11 , wherein a first amount of training data based on a first combination comprising the pre-training text and the correct summary text of the pre-training text is larger than a second amount of training data based on a second combination comprising the first training text, the second training text, and the correct summary text of the training moving image, and the training of the summary model represents a fine-tuning of the summary model.
13 . The device according to claim 11 , the processor further configured to execute operations comprising:
acquiring the training moving image and the correct summary text from a server in a network.
14 . The method according to claim 6 , further comprising:
performing pre-training on a summary model, using a pre-training text in a field associated with a pre-training moving image and a correct summary text of the pre-training text.
15 . The method according to claim 6 , further comprising:
wherein the image related to the moving image is based on a frame of the moving image.
16 . The method according to claim 7 , further comprising:
acquiring the training moving image and the correct summary text from a server in a network.
17 . The method according to claim 7 , further comprising:
performing pre-training on the summary model, using a pre-training text in a field associated with a pre-training moving image and a correct summary text of the pre-training text.
18 . The method according to claim 17 ,
wherein a first amount of training data comprising a first combination of the pre-training text and the correct summary text of the pre-training text is larger than a second amount of training data comprising a second combination of the first training text, the second training text, and the correct summary text of the training moving image, and
the training of the summary model represents a fine-tuning of the summary model.