Information processing device, computer program product, and information processing method
According to one embodiment, an information processing device includes a memory and one or more processors coupled to the memory. The one or more processors are configured to: receive input of a prompt including a first text and an expected value of an answer related to the first text; predict, upon input of the first text and at least one image, the answer for each of the at least one image by using an AI model that outputs the answer; compute accuracy of the answer from the expected value and the answer; and display, on a display device, display information including at least the prompt, the answer, and the accuracy.
1 . An information processing device comprising:
a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group; and
at least one hardware processor operably coupled to the memory and configured to execute processes comprising:
receiving input of a prompt including a first text and an expected value of an answer related to the first text;
predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;
calculating an accuracy of the answer from the expected value and the answer;
extracting K samples, from a sample image dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;
generating a suggested text based on the second text included in each of the K samples; and
controlling a display device to display information including at least the prompt, the answer, the accuracy, and the suggested text.
2 . The device according to claim 1 , wherein:
the processes further comprise receiving input of a plurality of images including at least one positive example image indicating a positive example and at least one negative example image indicating a negative example, and
the display information includes the answer for each of the at least one positive example image and the answer for each of the at least one negative example image.
3 . The device according to claim 2 , wherein;
the AI model is configured to carry out a visual question answering (VQA) task,
the first text includes a question for the plurality of images, and
the expected value of the answer includes a correct answer to the question.
4 . The device according to claim 2 , wherein:
the AI model is configured to carry out an image retrieval task retrieving an image having a specific feature,
the first text includes a query for retrieving the image having the specific feature, and
the expected value of the answer means that a first similarity to the image having the specific feature is greater than a threshold for the positive example image and is equal to or less than the threshold for the negative example image.
5 . The device according to claim 1 , wherein:
the AI model is configured to carry out a visual grounding task identifying a specific region,
the first text includes a query for indicating the specific region, and
the expected value of the answer includes a coordinate indicating a position of the specific region.
6 . The device according to claim 1 , wherein;
the processes further comprise visualizing a region of interest noted by the AI model for the at least one image according to a word contained in the first text, and
the display information further includes information indicating the region of interest noted according to selection of the word, for each of the at least one image.
7 . The device according to claim 1 , wherein:
the processes further comprise, upon processing of the at least one image, visualizing a word of interest noted by the AI model among words contained in the first text, and
the display information further includes information indicating the word of interest for each of the at least one image.
8 . The device according to claim 1 , wherein:
the processes further comprise retrieving the at least one image from a network based on the first text, and
the display information further includes a button and the at least one image retrieved, the button being configured to command retrieval of the at least one image.
9 . The device according to claim 1 , wherein;
the processes further comprise generating a caption describing the at least one image, and
the display information further includes the at least one image to which the caption is added.
10 . The device according to claim 1 , wherein the processes further comprise calculating a loss with a preset loss function from the answer output from the AI model and the expected value included in the prompt, and updating the AI model by error-back-propagating the loss.
11 . An information processing device comprising:
a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group; and
at least one hardware processor operably coupled to the memory and configured to execute processes comprising:
receiving input of a prompt including a first text and an expected value of an answer related to the first text;
predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;
calculating an accuracy of the answer from the expected value and the answer;
extracting K samples, from a prompt dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;
calculating a third similarity between the first text and the prompt dataset based on the second similarity of each of the K samples; and
controlling a display device to display information including at least the prompt, the answer, the accuracy, and the third similarity.
12 . A non-transitory computer-readable storage medium storing a program thereon, the program being executable by at least one hardware processor operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the program being executable to control the at least one hardware processor to execute processes comprising:
receiving input of a prompt including a first text and an expected value of an answer related to the first text;
predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;
calculating an accuracy of the answer from the expected value and the answer;
extracting K samples, from a sample image dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;
generating a suggested text based on the second text included in each of the K samples; and
controlling a display device to display information including at least the prompt, the answer, the accuracy, and the suggested text.
13 . An information processing method executed under control of at least one hardware processor of an information processing device, the at least one hardware processor being operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the method comprising:
receiving input of a prompt including a first text and an expected value of an answer related to the first text;
predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;
calculating an accuracy of the answer from the expected value and the answer;
extracting K samples, from a sample image dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;
generating a suggested text based on the second text included in each of the K samples; and
controlling a display device to display information including at least the prompt, the answer, the accuracy, and the suggested text.
14 . A non-transitory computer-readable storage medium storing a program thereon, the program being executable by at least one hardware processor operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the program being executable to control the at least one hardware processor to execute processes comprising:
receiving input of a prompt including a first text and an expected value of an answer related to the first text;
predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;
calculating an accuracy of the answer from the expected value and the answer;
extracting K samples, from a prompt dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;
calculating a third similarity between the first text and the prompt dataset based on the second similarity of each of the K samples; and
controlling a display device to display information including at least the prompt, the answer, the accuracy, and the third similarity.
15 . An information processing method executed under control of at least one hardware processor of an information processing device, the at least one hardware processor being operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the method comprising:
receiving input of a prompt including a first text and an expected value of an answer related to the first text;
predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;
calculating an accuracy of the answer from the expected value and the answer;
extracting K samples, from a prompt dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;
calculating a third similarity between the first text and the prompt dataset based on the second similarity of each of the K samples; and
controlling a display device to display information including at least the prompt, the answer, the accuracy, and the third similarity.