IP Library › Granted Patent US 12,183,450
Granted Patent B2
US 12,183,450 · App. 17/884,525 · Granted Dec 31, 2024

Constructing trained models to associate object in image with description in sentence where feature amount for sentence is derived from structured information

Inventor: Akimichi Ichinose (Tokyo, JP)
Assignee: FUJIFILM Corporation
G16H30/40G06T7/0012G06T7/60G06T7/74G06T7/75G06T2207/10072G06T2207/20081G06T2207/20084G06T2207/30004G06T2207/30064G06T2207/30096
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,183,450
App. No.
17/884,525
Granted
Dec 31, 2024
Kind
B2
Abstract

A processor derives a first feature amount for an object included in an image by a first neural network, structures a sentence including description of the object included in the image to derive structured information for the sentence, and derives a second feature amount for the sentence from the structured information by a second neural network. The processor trains the first neural network and the second neural network such that, in a feature space to which the first feature amount and the second feature amount belong, a distance between the derived first feature amount and second feature amount is reduced in a case in which the object included in the image and the object described in the sentence correspond to each other.

Claims (76)

1. A learning device comprising:

at least one processor,

wherein the processor

derives a first feature amount for an object included in an image by a first neural network,

structures a sentence including description of the object included in the image to derive structured information for the sentence,

derives a second feature amount for the sentence from the structured information by a second neural network, and

constructs a first derivation model that derives a feature amount for the object included in the image and a second derivation model that derives a feature amount for the sentence including the description of the object by training the first neural network and the second neural network such that, in a feature space to which the first feature amount and the second feature amount belong, a distance between the derived first feature amount and second feature amount is smaller in a case in which the object included in the image and the object described in the sentence correspond to each other than a case in which the object included in the image and the object described in the sentence do not correspond to each other.

2. The learning device according to claim 1 ,

wherein the processor trains the first neural network and the second neural network such that, in the feature space, the distance between the derived first feature amount and second feature amount is larger in a case in which the object included in the image and the object described in the sentence do not correspond to each other than a case in which the object included in the image and the object described in the sentence correspond to each other.

3. The learning device according to claim 1 ,

wherein the processor extracts one or more unique expressions for the object from the sentence and determines factuality for the unique expression to derive the unique expression and a determination result of the factuality as the structured information.

4. The learning device according to claim 3 ,

wherein the unique expression represents at least one of a position, an opinion, or a size of the object, and

the determination result of the factuality represents any of positivity, negativity, or suspicion for the opinion.

5. The learning device according to claim 3 ,

wherein, in a case in which a plurality of the unique expressions are extracted, the processor further derives a relationship between the unique expressions as the structured information.

6. The learning device according to claim 5 ,

wherein the relationship represents whether or not the plurality of unique expressions are related to each other.

7. The learning device according to claim 3 ,

wherein the processor normalizes the unique expression and the factuality to derive normalized structured information.

8. The learning device according to claim 1 ,

wherein the image is a medical image,

the object included in the image is a lesion included in the medical image, and

the sentence is an opinion sentence in which an opinion about the lesion is described.

9. An information processing apparatus comprising:

at least one processor,

wherein the processor

derives a first feature amount for one or more objects included in a target image by the first derivation model constructed by the learning device according to claim 1 ,

structures one or more target sentences including description of the object to derive structured information for the target sentence,

derives a second feature amount for the target sentence from the structured information for the target sentence by the second derivation model constructed by the learning device according to claim 1 ,

specifies the first feature amount corresponding to the second feature amount based on a distance between the derived first feature amount and second feature amount in a feature space, and

displays the object from which the specified first feature amount is derived, in distinction from other regions in the target image.

10. An information processing apparatus comprising:

at least one processor,

wherein the processor

receives input of a target sentence including description of an object,

structures the target sentence to derive structured information for the target sentence,

derives a second feature amount for the input target sentence from the structured information for the target sentence by the second derivation model constructed by the learning device according to claim 1 ,

refers to a database in which a first feature amount for one or more objects included in a plurality of reference images, which is derived by the first derivation model constructed by the learning device according to claim 1 , is associated with each of the reference images, to specify at least one first feature amount corresponding to the second feature amount based on a distance between the first feature amounts for the plurality of reference images and the derived second feature amount in a feature space, and

specifies the reference image associated with the specified first feature amount.

11. The information processing apparatus according to claim 9 ,

wherein the processor gives a notification of a unique expression that contributes to association with the first feature amount.

12. A learning method comprising:

deriving a first feature amount for an object included in an image by a first neural network;

structuring a sentence including description of the object included in the image to derive structured information for the sentence;

deriving a second feature amount for the sentence from the structured information by a second neural network; and

constructing a first derivation model that derives a feature amount for the object included in the image and a second derivation model that derives a feature amount for the sentence including the description of the object by training the first neural network and the second neural network such that, in a feature space to which the first feature amount and the second feature amount belong, a distance between the derived first feature amount and second feature amount is smaller in a case in which the object included in the image and the object described in the sentence correspond to each other than a case in which the object included in the image and the object described in the sentence do not correspond to each other.

13. An information processing method comprising:

deriving a first feature amount for one or more objects included in a target image by the first derivation model constructed by the learning device according to claim 1 ;

structuring one or more target sentences including description of the object to derive structured information for the target sentence;

deriving a second feature amount for the target sentence from the structured information for the target sentence by the second derivation model constructed by the learning device according to claim 1 ;

specifying the first feature amount corresponding to the second feature amount based on a distance between the derived first feature amount and second feature amount in a feature space; and

displaying the object from which the specified first feature amount is derived, in distinction from other regions in the target image.

14. An information processing method comprising:

receiving input of a target sentence including description of an object;

structuring the target sentence to derive structured information for the target sentence;

deriving a second feature amount for the input target sentence from the structured information for the target sentence by the second derivation model constructed by the learning device according to claim 1 ;

referring to a database in which a first feature amount for one or more objects included in a plurality of reference images, which is derived by the first derivation model constructed by the learning device according to claim 1 , is associated with each of the reference images, to specify at least one first feature amount corresponding to the second feature amount based on a distance between the first feature amounts for the plurality of reference images and the derived second feature amount in a feature space; and

specifying the reference image associated with the specified first feature amount.

15. A non-transitory computer-readable storage medium that stores a learning program causing a computer to execute:

a procedure of deriving a first feature amount for an object included in an image by a first neural network;

a procedure of structuring a sentence including description of the object included in the image to derive structured information for the sentence;

a procedure of deriving a second feature amount for the sentence from the structured information by a second neural network; and

a procedure of constructing a first derivation model that derives a feature amount for the object included in the image and a second derivation model that derives a feature amount for the sentence including the description of the object by training the first neural network and the second neural network such that, in a feature space to which the first feature amount and the second feature amount belong, a distance between the derived first feature amount and second feature amount is smaller in a case in which the object included in the image and the object described in the sentence correspond to each other than a case in which the object included in the image and the object described in the sentence do not correspond to each other.

16. A non-transitory computer-readable storage medium that stores an information processing program causing a computer to execute:

a procedure of deriving a first feature amount for one or more objects included in a target image by the first derivation model constructed by the learning device according to claim 1 ;

a procedure of structuring one or more target sentences including description of the object to derive structured information for the target sentence;

a procedure of deriving a second feature amount for the target sentence from the structured information for the target sentence by the second derivation model constructed by the learning device according to claim 1 ;

a procedure of specifying the first feature amount corresponding to the second feature amount based on a distance between the derived first feature amount and second feature amount in a feature space; and

a procedure of displaying the object from which the specified first feature amount is derived, in distinction from other regions in the target image.

17. A non-transitory computer-readable storage medium that stores an information processing program causing a computer to execute:

a procedure of receiving input of a target sentence including description of an object;

a procedure of structuring the target sentence to derive structured information for the target sentence;

a procedure of deriving a second feature amount for the input target sentence from the structured information for the target sentence by the second derivation model constructed by the learning device according to claim 1 ;

a procedure of referring to a database in which a first feature amount for one or more objects included in a plurality of reference images, which is derived by the first derivation model constructed by the learning device according to claim 1 , is associated with each of the reference images, to specify at least one first feature amount corresponding to the second feature amount based on a distance between the first feature amounts for the plurality of reference images and the derived second feature amount in a feature space; and

a procedure of specifying the reference image associated with the specified first feature amount.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2022
From: ICHINOSE, AKIMICHI
To: FUJIFILM CORPORATION
Reel/Frame 060776/0432 →
Priority Claims (1)
JP 2021-132919 · Aug 17, 2021 · national
Continuity (1)
Related Publication 20230054096A1 · Feb 23, 2023
Cited By (1)
US 12,431,236