IP Library › Granted Patent US 11,810,312
Granted Patent B2
US 11,810,312 · App. 17/236,191 · Granted Nov 7, 2023

Multiple instance learning method

Inventors: Sang Hyun Park (Daegu, KR); Philip Chikontwe (Daegu, KR)
Assignee: DAEGU GYEONGBUK INSTITUTE OF SCIENCE AND TECHNOLOGY
G06T7/593G06N20/00G06T2207/10081G06T2207/20081G06T2207/30061G16H30/40G16H50/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,810,312
App. No.
17/236,191
Granted
Nov 7, 2023
Kind
B2
Abstract

A multiple instance learning device for analyzing 3D images, comprises a memory in which a multiple instance learning model is stored and at least one processor electrically connected to the memory, wherein the multiple instance learning model comprises a convolution block configured to derive a feature map for each of 2D instances of a 3D image inputted to the multiple instance learning model, a spatial attention block configured to derive spatial attention maps of the instances from the feature maps derived from the convolution block, an instance attention block configured to receive a result of combining the feature maps and the spatial attention maps and derive an attention score for each instance, and derive an aggregated embedding for the 3D image by aggregating embeddings of the instances according to the attention scores.

Claims (254)

1. A multiple instance learning device for analyzing 3D images, comprising:

a memory in which a multiple instance learning model is stored; and

at least one processor electrically connected to the memory,

wherein the multiple instance learning model comprises:

a convolution block configured to derive a feature map for each of 2D instances of a 3D image inputted to the multiple instance learning model;

a spatial attention block configured to derive spatial attention maps of the instances from the feature maps derived from the convolution block;

an instance attention block configured to receive a result of combining the feature maps and the spatial attention maps and derive an attention score for each instance, and derive an aggregated embedding for the 3D image by aggregating embeddings of the instances according to the attention scores; and

an output block configured to output an analysis result for the 3D image based on the aggregated embedding.

2. The multiple instance learning device of claim 1 , wherein the at least one processor is configured to perform an operation of, in a training phase of the multiple instance learning model, training the multiple instance learning model such that a total loss function (L) of the multiple instance learning model has a minimum value, using training data including 3D images labeled with ground-truths for the analysis result,

wherein the total loss function is a combination of a bag-level loss function (L B ) for a result of the output block and a contrastive loss function (L F ) between the instance embeddings and the aggregated embedding.

3. The multiple instance learning device of claim 2 , wherein the operation of training the multiple instance learning model comprises an operation of adjusting parameters of the convolution block, the spatial attention block, and the instance attention block such that the total loss function (L) of the multiple instance learning model has a minimum value.

4. The multiple instance learning device of claim 2 , wherein the total loss function (L) is expressed by Equation 1 below,

=λ B +(1−λ) F   Equation 1:

wherein λ is a value between 0 and 1 and is a parameter representing a weight of the bag-level loss function.

5. The multiple instance learning device of claim 2 , wherein the bag-level loss function (L B ) is expressed by Equation 2 below, and the contrastive loss function (L F ) is expressed by Equation 3 below,

B =−Σy i log ŷ.   Equation 2:

wherein y i is an instance label; and ŷ is the probability of the 3D image being labelled as y i ,

ℒ

F

⁡

(

z

′

,

z

,

τ

)

=

-

log

⁢

⁢

∑

i

,

j

=

1

N

⁢

exp

⁢

⁢

(

sim

⁡

(

z

i

′

,

z

j

)

/

τ

)

∑

k

=

1

2

⁢

N

⁢

ℚ

[

k

≠

i

]

⁢

exp

⁡

(

sim

⁡

(

z

i

′

,

z

k

)

/

τ

)

Equation

⁢

⁢

3

wherein z′ is an instance-level feature; z is a bag-level feature; [k≠i] ϵ{0, 1} has a value of 1 if k is not equal to i and a value of 0 if k=i; τ is a temperature parameter; sim(⋅, ⋅) is a similarity function; and N is the number of instances of the 3D image.

6. The multiple instance learning device of claim 1 , wherein the at least one processor is configured to perform an operation of, in a training phase of the multiple instance learning model, training the multiple instance learning model such that a final loss function (L*) of the multiple instance learning model has a minimum value, using training data including 3D images labeled with ground-truths for the analysis result,

wherein the final loss function is a combination of a bag-level loss function (L B ) for a result of the output block, ins and an instance-level loss function (L I ), and a center loss function (L c ).

7. The multiple instance learning device of claim 6 , wherein the final loss is expressed by Equation 4 as follows:

*=α I +λ C +β B   Equation 4:

wherein L is the final loss, L I is the instance loss, L B is the bag loss, L C is the center loss, and α, β, and λ are each loss balance parameters set to arbitrary positive decimal values.

8. A multiple instance learning device for analyzing 3D images, comprising:

a memory in which at least one instruction is stored; and

at least one processor operating in conjunction with the memory and configured to execute the at least one instruction,

wherein the at least one instruction is configured to, when executed by the at least one processor, cause the at least one processor to perform operations of:

deriving a feature map for each of 2D instances of a 3D image inputted to a multiple instance learning model;

deriving spatial attention maps of the instances from the feature maps;

receiving a result of combining the feature maps and the spatial attention maps, and deriving an attention score for each instance;

deriving an aggregated embedding for the 3D image by aggregating embeddings of the instances according to the attention scores; and

outputting an analysis result for the 3D image based on the aggregated embedding.

9. The multiple instance learning device of claim 8 , wherein the at least one processor is configured to perform an operation of, in a training phase of the multiple instance learning model, training the multiple instance learning model such that a total loss function (L) of the multiple instance learning model has a minimum value, using training data including 3D images labeled as ground-truths for the analysis result,

wherein the total loss function is a combination of a bag-level loss function (L B ) for a result of the output block and a contrastive loss function (L F ) between the instance embeddings and the aggregated embedding.

10. The multiple instance learning device of claim 9 , wherein the operation of training the multiple instance learning model comprises an operation of adjusting parameters used in the operation of deriving the feature maps, the operation of deriving the spatial attention maps, the operation of deriving the attention scores, and the operation of deriving the aggregated embedding, such that the total loss function (L) of the multiple instance learning model has a minimum value.

11. The multiple instance learning device of claim 9 , wherein the total loss function (L) is expressed by Equation 1 below,

=λ B +(1−λ) F   Equation 1:

wherein λ is a value between 0 and 1 and is a parameter representing a weight of the bag-level loss function.

12. The multiple instance learning device of claim 9 , wherein the bag-level loss function (L B ) is expressed by Equation 2 below, and the contrastive loss function (L F ) is expressed by Equation 3 below,

B =−Σy i log ŷ   Equation 2:

wherein y i is an instance label; and ŷ is the probability of the 3D image being labelled as y i ,

ℒ

F

=

-

log

⁢

⁢

exp

⁢

⁢

(

sim

⁡

(

z

i

′

,

z

j

)

/

τ

)

∑

k

=

1

2

⁢

N

⁢

ℚ

[

k

≠

i

]

⁢

exp

⁡

(

sim

⁡

(

z

i

′

,

z

k

)

/

τ

)

Equation

⁢

⁢

3

wherein z′ is an instance-level feature; z is a bag-level feature; [k≠i] ϵ{0, 1} has a value of 1 if k is not equal to i and a value of 0 if k=i; τ is a temperature parameter; sim(⋅, ⋅) is a similarity function; and N is the number of instances of the 3D image.

13. A multiple instance learning method for analyzing 3D images, the method performed by at least one processor in a computing device or a computing network, and the method comprising:

deriving a feature map for each of 2D instances of a 3D image inputted to a multiple instance learning model;

deriving spatial attention maps of the instances from the feature maps derived;

receiving a result of combining the feature maps and the spatial attention maps, and deriving an attention score for each instance;

deriving an aggregated embedding for the 3D image by aggregating embeddings of the instances according to the attention scores; and

outputting an analysis result for the 3D image based on the aggregated embedding.

14. The multiple instance learning method of claim 13 ,

wherein the 3D image is labelled as ground-truth for the analysis result, and

wherein the multiple instance learning method further comprises, in a training phase of the multiple instance learning model, training the multiple instance learning model such that a total loss function (L) of the multiple instance learning model has a minimum value, using training data including the 3D image labeled as a ground-truth for the analysis result,

wherein the total loss function is a combination of a bag-level loss function (L B ) for a result of the output block and a contrastive loss function (L F ) between the instance embeddings and the aggregated embedding.

15. The multiple instance learning method of claim 14 , wherein the training of the multiple instance learning model comprises adjusting parameters used in the deriving of the feature maps, the deriving of the spatial attention maps, the deriving of the attention scores, and the deriving of the aggregated embedding, such that the total loss function (L) of the multiple instance learning model has a minimum value.

16. The multiple instance learning method of claim 14 , wherein the total loss function (L) is expressed by Equation 1 below,

=λ B +(1−λ) F   Equation 1:

wherein λ is a value between 0 and 1 and is a parameter representing a weight of the bag-level loss function.

17. The multiple instance learning method of claim 14 , wherein the bag-level loss function (L B ) is expressed by Equation 2 below, and the contrastive loss function (L F ) is expressed by Equation 3 below,

B =−Σy i log ŷ.   Equation 2:

wherein y i is an instance label; and ŷ is the probability of the 3D image being labelled as y i ,

ℒ

F

=

-

log

⁢

⁢

exp

⁢

⁢

(

sim

⁡

(

z

i

′

,

z

j

)

/

τ

)

∑

k

=

1

2

⁢

N

⁢

ℚ

[

k

≠

i

]

⁢

exp

⁡

(

sim

⁡

(

z

i

′

,

z

k

)

/

τ

)

Equation

⁢

⁢

3

wherein z′ is an instance-level feature; z is a bag-level feature; [k≠i] ϵ{0, 1} has a value of 1 if k is not equal to i and a value of 0 if k=i; τ is a temperature parameter; sim (⋅, ⋅) is a similarity function; and N is the number of instances of the 3D image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2021
From: PARK, SANG HYUN; CHIKONTWE, PHILIP
To: DAEGU GYEONGBUK INSTITUTE OF SCIENCE AND TECHNOLOGY
Reel/Frame 055987/0099 →
Priority Claims (2)
KR 10-2020-0047888 · Apr 21, 2020 · national
KR 10-2021-0051331 · Apr 20, 2021 · national
Continuity (1)
Related Publication 20210334994A1 · Oct 28, 2021
Cited By (1)
US 12,541,959