IP Library › Granted Patent US 12,657,219
Granted Patent B2
US 12,657,219 · App. 18/586,639 · Granted Jun 16, 2026

Information processing device, computer program product, and information processing method

Inventors: Nao Mishima (Tokyo, JP); Reiko Noda (Kawasaki, JP); Tatsuo Kozakaya (Kawasaki, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G06F16/3329G06F16/5846G06V10/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,219
App. No.
18/586,639
Granted
Jun 16, 2026
Kind
B2
Abstract

According to one embodiment, an information processing device includes a memory and one or more processors coupled to the memory. The one or more processors are configured to: receive input of a prompt including a first text and an expected value of an answer related to the first text; predict, upon input of the first text and at least one image, the answer for each of the at least one image by using an AI model that outputs the answer; compute accuracy of the answer from the expected value and the answer; and display, on a display device, display information including at least the prompt, the answer, and the accuracy.

Claims (74)

1 . An information processing device comprising:

a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group; and

at least one hardware processor operably coupled to the memory and configured to execute processes comprising:

receiving input of a prompt including a first text and an expected value of an answer related to the first text;

predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;

calculating an accuracy of the answer from the expected value and the answer;

extracting K samples, from a sample image dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;

generating a suggested text based on the second text included in each of the K samples; and

controlling a display device to display information including at least the prompt, the answer, the accuracy, and the suggested text.

2 . The device according to claim 1 , wherein:

the processes further comprise receiving input of a plurality of images including at least one positive example image indicating a positive example and at least one negative example image indicating a negative example, and

the display information includes the answer for each of the at least one positive example image and the answer for each of the at least one negative example image.

3 . The device according to claim 2 , wherein;

the AI model is configured to carry out a visual question answering (VQA) task,

the first text includes a question for the plurality of images, and

the expected value of the answer includes a correct answer to the question.

4 . The device according to claim 2 , wherein:

the AI model is configured to carry out an image retrieval task retrieving an image having a specific feature,

the first text includes a query for retrieving the image having the specific feature, and

the expected value of the answer means that a first similarity to the image having the specific feature is greater than a threshold for the positive example image and is equal to or less than the threshold for the negative example image.

5 . The device according to claim 1 , wherein:

the AI model is configured to carry out a visual grounding task identifying a specific region,

the first text includes a query for indicating the specific region, and

the expected value of the answer includes a coordinate indicating a position of the specific region.

6 . The device according to claim 1 , wherein;

the processes further comprise visualizing a region of interest noted by the AI model for the at least one image according to a word contained in the first text, and

the display information further includes information indicating the region of interest noted according to selection of the word, for each of the at least one image.

7 . The device according to claim 1 , wherein:

the processes further comprise, upon processing of the at least one image, visualizing a word of interest noted by the AI model among words contained in the first text, and

the display information further includes information indicating the word of interest for each of the at least one image.

8 . The device according to claim 1 , wherein:

the processes further comprise retrieving the at least one image from a network based on the first text, and

the display information further includes a button and the at least one image retrieved, the button being configured to command retrieval of the at least one image.

9 . The device according to claim 1 , wherein;

the processes further comprise generating a caption describing the at least one image, and

the display information further includes the at least one image to which the caption is added.

10 . The device according to claim 1 , wherein the processes further comprise calculating a loss with a preset loss function from the answer output from the AI model and the expected value included in the prompt, and updating the AI model by error-back-propagating the loss.

11 . An information processing device comprising:

a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group; and

at least one hardware processor operably coupled to the memory and configured to execute processes comprising:

receiving input of a prompt including a first text and an expected value of an answer related to the first text;

predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;

calculating an accuracy of the answer from the expected value and the answer;

extracting K samples, from a prompt dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;

calculating a third similarity between the first text and the prompt dataset based on the second similarity of each of the K samples; and

controlling a display device to display information including at least the prompt, the answer, the accuracy, and the third similarity.

12 . A non-transitory computer-readable storage medium storing a program thereon, the program being executable by at least one hardware processor operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the program being executable to control the at least one hardware processor to execute processes comprising:

receiving input of a prompt including a first text and an expected value of an answer related to the first text;

predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;

calculating an accuracy of the answer from the expected value and the answer;

extracting K samples, from a sample image dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;

generating a suggested text based on the second text included in each of the K samples; and

controlling a display device to display information including at least the prompt, the answer, the accuracy, and the suggested text.

13 . An information processing method executed under control of at least one hardware processor of an information processing device, the at least one hardware processor being operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the method comprising:

receiving input of a prompt including a first text and an expected value of an answer related to the first text;

predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;

calculating an accuracy of the answer from the expected value and the answer;

extracting K samples, from a sample image dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;

generating a suggested text based on the second text included in each of the K samples; and

controlling a display device to display information including at least the prompt, the answer, the accuracy, and the suggested text.

14 . A non-transitory computer-readable storage medium storing a program thereon, the program being executable by at least one hardware processor operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the program being executable to control the at least one hardware processor to execute processes comprising:

receiving input of a prompt including a first text and an expected value of an answer related to the first text;

predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;

calculating an accuracy of the answer from the expected value and the answer;

extracting K samples, from a prompt dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;

calculating a third similarity between the first text and the prompt dataset based on the second similarity of each of the K samples; and

controlling a display device to display information including at least the prompt, the answer, the accuracy, and the third similarity.

15 . An information processing method executed under control of at least one hardware processor of an information processing device, the at least one hardware processor being operably coupled to a memory storing an artificial intelligence (AI) model configured to output an answer based on input of (i) a text comprising a question and (ii) an image, the AI model being trained using a dataset including an image group and a question group associated with the image group, and the method comprising:

receiving input of a prompt including a first text and an expected value of an answer related to the first text;

predicting, upon input of the first text and at least one image, the answer for each of the at least one image by using the AI model;

calculating an accuracy of the answer from the expected value and the answer;

extracting K samples, from a prompt dataset storing therein, as a sample, the at least one image and a second text associated with the at least one image, in descending order of second similarity between the first text and the second text;

calculating a third similarity between the first text and the prompt dataset based on the second similarity of each of the K samples; and

controlling a display device to display information including at least the prompt, the answer, the accuracy, and the third similarity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2024
From: MISHIMA, NAO; NODA, REIKO; KOZAKAYA, TATSUO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 067001/0764 →
Priority Claims (1)
JP 2023-089841 · May 31, 2023 · national
Continuity (1)
Related Publication 20240403337A1 · Dec 5, 2024
References Cited (22)
US 20080030488A1 · Kobayashi · 2008 [cited by examiner]
US 20220067081A1 · Saito et al. · 2022 [cited by applicant]
US 20220129693A1 · Pham · 2022 [cited by examiner]
US 20220383037A1 · Pham · 2022 [cited by examiner]
US 20230077031A1 · Hosoya · 2023 [cited by examiner]
US 20230252344A1 · Mishima · 2023 [cited by examiner]
US 20230419652A1 · Tiong · 2023 [cited by examiner]
US 20240119257A1 · Guo · 2024 [cited by examiner]
US 20240193200A1 · Watanabe et al. · 2024 [cited by applicant]
US 20240242487A1 · Conde et al. · 2024 [cited by applicant]
JP 2022038941A · 2022 [cited by applicant]
JP 2022071675A · 2022 [cited by applicant]
JP 2022190985A · 2022 [cited by applicant]
JP 2023039656A · 2023 [cited by applicant]
JP 2023117248A · 2023 [cited by applicant]
JP 2024074523A · 2024 [cited by applicant]
JP 2024521118A · 2024 [cited by applicant]
JP 2024082634A · 2024 [cited by applicant]
Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, ICML2021, 2021. [cited by applicant]
Selvaraju, et al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization”, ICCV2017, 2017. [cited by applicant]
Japanese Office Action (and an English language translation thereof) dated Mar. 3, 2026, issued in corresponding Japanese Application No. 2023-089841. [cited by applicant]
“What is ChatGPT?”, [Online], Windows Forest, Mar. 15, 2023, <Retrieved from the Internet, URL: https://forest.watch.impress.co.jp/docs/review/1499657.html>. [cited by applicant]