IP Library › Granted Patent US 12,086,210
Granted Patent B2
US 12,086,210 · App. 17/460,387 · Granted Sep 10, 2024

State determination apparatus and image analysis apparatus

Inventors: Quoc Viet Pham (Yokohama Kanagawa, JP); Toshiaki Nakasu (Chofu Tokyo, JP); Nao Mishima (Inagi Tokyo, JP); Shojun Nakayama (Yokohama Kanagawa, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G06F18/22G06T7/11G06V10/22G06V30/40G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,086,210
App. No.
17/460,387
Granted
Sep 10, 2024
Kind
B2
Abstract

According to one embodiment, a state determination apparatus includes a processor. The processor acquires a targeted image. The processor acquires a question concerning the targeted image and an expected answer to the question. The processor generates an estimated answer estimated with respect to the question concerning the targeted image using a trained model trained to estimate an answer based on a question concerning an image. The processor determines a state of a target for determination in accordance with a similarity between the expected answer and the estimated answer.

Claims (47)

1. A state determination apparatus comprising a processor configured to:

acquire a targeted image;

acquire, from among a plurality of stored questions and associated expected answers, a question concerning the targeted image and an expected answer to the question;

generate an estimated answer estimated with respect to the question concerning the targeted image using a trained model trained to estimate an answer based on a question concerning an image; and

determine whether a state of a determination target is a normal state or an anomalous state, based on a similarity between the expected answer and the estimated answer,

wherein the expected answer assumes the normal state, and the processor determines that the determination target is in the anomalous state when the similarity is smaller than a threshold.

2. The apparatus according to claim 1 , wherein when the determination target is in the anomalous state, the processor determines that the determination target is in a dangerous state.

3. The apparatus according to claim 1 , wherein the processor is further configured to:

refer to a database in which the question is associated with a solution when the determination target is determined to be in the anomalous state; and

present the solution.

4. The apparatus according to claim 1 , wherein the processor is further configured to extract and generate the question and the expected answer to the question from a manual.

5. The apparatus according to claim 1 , wherein:

the processor is further configured to generate a plurality of sets of questions and associated expected answers with respect to one determination item assuming a normal state in a manual, and

the processor determines that the determination item is in an anomalous state when a number of sets is smaller than a second threshold value, the number of sets being a number of times that a similarity between the expected answer and the estimated answer generated by using the trained model with respect to each of the questions is equal to or greater than a first threshold value.

6. The apparatus according to claim 1 , wherein the trained model is a model relating to visual question answering (VQA).

7. The apparatus according to claim 1 , wherein when a situation in which the similarity is smaller than a threshold value lasts for a predetermined period or longer or occurs a predetermined number of times or more, the processor determines that the determination target is in the anomalous state.

8. A state determination apparatus comprising a processor configured to:

acquire a targeted image;

acquire, from among a plurality of stored questions and associated expected answers, a question concerning the targeted image and an expected answer to the question;

generate an estimated answer estimated with respect to the question concerning the targeted image using a trained model trained to estimate an answer based on a question concerning an image; and

determine a state of a determination target based on a similarity between the expected answer and the estimated answer,

wherein the expected answer assumes an anomalous state, and the processor determines that the determination target is in the anomalous state when the similarity is equal to or greater than a threshold value.

9. An image analysis apparatus comprising a processor configured to:

acquire an image;

acquire a question pertaining to the image from among a plurality of stored questions;

detect a region of interest (ROI) in the image;

calculate a first image feature amount relating to the detected ROI;

divide the image into image regions through a semantic segmentation process;

calculate a second image feature amount with respect to each divided image region;

calculate an image feature amount from the image by combining the first image feature amount and the second image feature amount;

calculate a text feature amount representing a semantic content of the question; and

estimate an answer to the question based on the image feature amount and the text feature amount.

10. The apparatus according to claim 9 , wherein the processor is configured to:

generate a combined ROI by combining the detected ROI and a divided image region; and

calculate the image feature amount with respect to the combined ROI.

11. The apparatus according to claim 10 , wherein the processor calculates, as the combined ROI, a sum of the detected ROI and the divided image region.

12. The apparatus according to claim 10 , wherein the processor calculates, as the combined ROI, an ROI in which an overlap region between the detected ROI and the divided image region is equal to or greater than a threshold value.

13. The apparatus according to claim 9 , wherein:

the first image feature amount and the second image feature amount are represented as vectors; and

the processor combines a vector of the first image feature amount and a vector of the second image feature amount.

14. The apparatus according to claim 9 , wherein the processor combines the image feature amount with a feature amount based on a label applied to a divided image region obtained through the semantic segmentation process.

15. The apparatus according to claim 9 , wherein the processor extracts information on a scene graph representing a positional relation between objects and a semantic relation between the objects, and calculates the image feature amount by combining the information on the scene graph and the second image feature amount.

16. A state determination apparatus comprising a processor configured to:

acquire a targeted image;

acquire, from among a plurality of stored questions and associated expected answers, a question concerning the targeted image and an expected answer to the question;

generate an estimated answer estimated for the question concerning the targeted image using the image analysis apparatus according to claim 9 ; and

determine a state of a determination target based on a similarity between the expected answer and the estimated answer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2021
From: PHAM, QUOC VIET; NAKASU, TOSHIAKI; MISHIMA, NAO; NAKAYAMA, SHOJUN
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 057322/0131 →
Priority Claims (1)
JP 2020-180756 · Oct 28, 2020 · national
Continuity (1)
Related Publication 20220129693A1 · Apr 28, 2022