IP Library › Granted Patent US 12,536,784
Granted Patent B2
US 12,536,784 · App. 18/018,310 · Granted Jan 27, 2026

Prediction of labels for digital images, especially medical ones, and supply of explanations associated with these labels

Inventor: Gwenolė Quellec (Brest, FR)
Assignees: UNIVERSITE BRESTBRETAGNE OCCIDENTALE; INSTITUT NATIONAL DE LA SANTÉ ET DE LA RECHERCHE MÉDICALE (INSERM)
G06V10/82G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,784
App. No.
18/018,310
Granted
Jan 27, 2026
Kind
B2
Abstract

Method for the prediction of labels associated with a digital image, comprising a prediction phase consisting of: supplying the image to a segmentation neural network configured to predict a classification of the pixels of the image into a first set of classes; and supplying at least part of this classification to a classification neural network configured to predict a set of labels for said image, based on the classification P of the pixels; said segmentation and classification neural networks being determined by a learning phase comprising, for each image of a training set, the first and second steps; and determining a location of the background of said image, based on the classification of the pixels, and optimizing the weights of the neural networks according to a set of cost functions configured, by iteration, to maximize the quality of the set of labels as a function of labels previously established and associated with the image, and to maximize the probability of not predicting any label for the background.

Claims (68)

1 . A method for the prediction of labels associated with a digital image, comprising a prediction phase consisting of:

supplying, in a first step S 1 , said image to a segmentation neural network configured to predict a classification P of the pixels of said image into a first set of classes; and

supplying, in a second step S 2 , said classification to a classification neural network configured to predict a set of labels p(I) for said image, based on said classification P of the pixels, except for a segment corresponding to a background of said image;

said segmentation and classification neural networks being determined by a learning phase comprising, for each image of a training set;

performing said first step S 1 and second steps S 2 ;

determining a location of said background of said image, based on the classification of the pixels; and

optimizing the weights of said neural networks, according to a set of cost functions configured, by iteration, to maximize the quality of said set of labels p(I) as a function of labels previously established and associated with said image, and to maximize the probability of not predicting any label for said background.

2 . The method according to claim 1 , wherein the determination of a location of the background of said image comprises the determination of an occluded image Î, based on the classification P 1 of the pixels, corresponding to the background of the image, said occluded image being defined by Î=I×P 1 ={I x,y ·P 1,x,y ,∀(x,y)},

(x, y) defining a pixel of said image.

3 . The method according to claim 2 , wherein, during the learning phase, an auxiliary classification neural network (CN′) is trained in order to optimize the classification of said occluded image.

4 . The method according to claim 3 , wherein said set of cost functions includes a total cost function total Which is expressed as:

total = +α· ′+β· occlusion +γ· sparsity

where

L is a cost function which allows maximizing the quality of said set of labels p(I) on the basis of previously established labels;

L′ is a cost function which allows maximizing the quality of the predictions of said auxiliary classification neural network (CN′) on the basis of said previously established labels;

L occlusion is a cost function which allows maximizing the probability of not predicting any label for said occluded image; and

L sparsity is a cost function which allows maximizing a surface area of said classification P 1 of the pixels; and

α, β and γ are parameters.

5 . The method according to claim 1 , wherein said segmentation neural network is an encoder-decoder network formed of an encoder neural network (EN) and a decoder neural network (DN), arranged in cascade.

6 . The method according to claim 1 , wherein said classification neural network (CN) is composed of summary layers Π and classification Δ layers.

7 . The method according to claim 6 , wherein the output from said classification layer Δ can be expressed as a function of an input vector z m , Δ being expressed as:

Δ

⁡

(

z

)

=

{

σ

⁡

(

∑

m

=

2

M

z

m

⁢

w

m

,

n

2

+

b

n

)

,

∀

n

∈

{

1

,

2

,

…

,

N

}

}

with w m,n representing positive synaptic weights, σ representing the activation function for the neurons of said classification layer, N representing the number of said labels, and b n representing biases.

8 . The method according to claim 7 , wherein an explanation associated with each label of said set of labels is provided during said prediction phase.

9 . The method according to claim 8 , wherein said explanation is based on said synaptic weights w m,n of said classification layer Δ and on the outputs from said summary layers Π.

10 . The method according to claim 7 , wherein, at the end of the learning phase, names are associated with the probability maps P m , and said names are provided with said explanations during the prediction phase.

11 . A device for predicting labels associated with a digital image, comprising a computer configured for implementing the method according to claim 1 .

12 . A non-transitory computer readable storage medium having stored thereon code instructions which, when executed by a computer, cause said computer to carry out the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2023
From: QUELLEC, GWENOLÉ
To: UNIVERSITE BREST BRETAGNE OCCIDENTALE; INSTITUT NATIONAL DE LA SANTÉ ET DE LA RECHERCHE MÉDICALE (INSERM)
Reel/Frame 063217/0481 →
Priority Claims (1)
FR 2007935 · Jul 28, 2020 · national
Continuity (1)
Related Publication 20230306731A1 · Sep 28, 2023
References Cited (8)
US 10810543B2 · Hsieh · 2020 [cited by examiner]
WO WO2019233394A1 · 2019 [cited by examiner]
Chen, et al. (Computer English translation of WO-2019233394, pp. 1-17. (Year: 2019). [cited by examiner]
International Search Report related to Application No. PCT/FR2021/051363; reported on Oct. 27, 2021. [cited by applicant]
Suo Qiu, et al., “Global Weighted Average Pooling Bridges Pixel-level Localization and Image-level Classification”, arixiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 21, … [cited by applicant]
Florian Dubost, et al., “Weakly Supervised Objection Detection with 2D and 3D Regression Neural Networks”, arixiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 5, 2019, XP08… [cited by applicant]
Shahriyar Shaikh Akib, et al., “An Approach for Multi Label Image Classification Using Single Label Convolutional Neural Network”, 2018 21st International Conference of Computer and Information Technology (ICCIT), IEEE,… [cited by applicant]
Gwenole' Quellec, et al., “ExplAIn: Explanatory Artificial Intelligence for Diabetic Retinopathy Diagnosis”, arixiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 1, 2020, XP… [cited by applicant]