IP Library › Granted Patent US 12,657,894
Granted Patent B2
US 12,657,894 · App. 18/528,692 · Granted Jun 16, 2026

Apparatus and method for classifying domain non-specific images using text

Inventors: Jin Kyu Kim (Seoul, KR); No Kyung Park (Seoul, KR)
Assignee: KOREA UNIVERSITY RESEARCH AND BUSINESS FOUNDATION
G06V10/811G06F40/284G06V10/764G06V10/7715G06V10/774G06V10/776G06V10/806G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,894
App. No.
18/528,692
Granted
Jun 16, 2026
Kind
B2
Abstract

Disclosed is an apparatus for classifying domain non-specific images using text according to one embodiment of the present invention. According to the present invention, the learning process is performed not only using the images but also text information together, and thus even during training on images from only a few specific domains, it is possible to effectively classify the images from different domains with high accuracy by applying the human inference process.

Claims (151)

1 . An apparatus for classifying domain non-specific images using text, the apparatus comprising:

an image classification unit configured to receive a learning image to generate a visual feature from the received learning image, to generate and output a classification result of the learning image using the generated visual feature, and to perform a learning process using a first loss function (L task );

an image-text fusion unit configured to receive a learning text to generate a textual feature from the received learning text, to align the visual feature with the textual feature, and to perform the learning process using a second loss function (L align ); and

a text description generation unit configured to receive the visual feature and the classification result of the learning image to generate a text describing the learning image based on the inputs and to perform the learning process using a third loss function (L expl ),

wherein the learning image comprises one or more images represented in one or more domains for one or more classes, and the learning text comprises one or more texts described in one or more texts for the one or more classes,

wherein the learning image and the learning text are included in one learning dataset, and

wherein the image classification unit is configured to perform the learning process using the first loss function (L task ) according to Equation 1 below:

L

t

⁢

a

⁢

s

⁢

k

=

-

∑

i

⁢

y

i

⁢

log

⁡

(

y

ˆ

i

)

Equation

⁢

1

where y refers to a one-hot vector in which a ground-truth label is 1 and others are 0, and ŷ refers to a softmax distribution that is the classification result of the learning image generated using the visual feature.

2 . The apparatus of claim 1 , wherein the image classification unit comprises a Residual Neural Network (ResNet) as an image encoder configured to generate the visual feature from the received learning image.

3 . The apparatus of claim 1 , wherein the image-text fusion unit comprises a text encoder based on a contrastive language-image pre-training (CLIP) model configured to generate the textual feature from the received learning text.

4 . The apparatus of claim 1 , wherein the image-text fusion unit is configured to use the text description generation unit as a text encoder configured to generate the textual feature from the received learning text.

5 . The apparatus of claim 1 , wherein the image-text fusion unit is configured to perform the learning process using the second loss function (L align ) according to Equation 2 below:

L

align

=

f

proj

(

v

)

-

g

prog

(

x

)

2

-

∑

i

⁢

y

i

⁢

log

⁡

(

y

˜

i

)

Equation

⁢

2

where f proj refers to the projection layer for textual features, g proj refers to the projection layer for visual features, y refers to a one-hot vector in which a ground-truth label is 1 and others are 0, and {tilde over (y)} refers to a softmax distribution obtained from g prog (x).

6 . The apparatus of claim 1 , wherein the text description generation unit comprises a first long short-term memory (LSTM) layer and a second LSTM layer,

wherein the second LSTM layer is configured to receive an output token from the first LSTM layer, the visual features and the classification result of the learning image, to generate a softmax probability value for each word held by the text description generation unit, and to output a word token with the highest probability value, and

wherein the output of the word token is repeated until the word token output by the second LSTM layer satisfies an-end of-sentence token (EOS) or a predetermined number of sampled word tokens.

7 . The apparatus of claim 1 , wherein the text description generation unit is configured to perform the learning process using a third loss function (L expl ) according to Equation 3 below:

L

expl

=

-

∑

t

log

⁢

P

⁡

(

o

t

+

1

|

o

0

:

t

,

I

,

C

)

-

E

o

~

∼

P

⁡

(

o

|

I

,

c

)

[

R

⁡

(

O

~

)

]

Equation

⁢

3

where, Õ˜p(o|I, c) refers to a sentence sampled to generate a text describing the learning image, p(O|I, C) refers to an estimated distribution for the description (o) when the learning image (I) and the class (C) are given as conditions, and R (O) refers to a softmax probability per class when there is a sentence (Õ) generated from P(c|õ).

8 . The apparatus of claim 1 , wherein the apparatus is configured to perform the learning process in a direction of minimizing a total loss function (L total ) that is equal to a weighted sum of the first loss function (L task ), the second loss function (L align ), and the third loss function (L expl ).

9 . A method for classifying domain non-specific images using text, performed by an apparatus comprising a process and a memory, the method comprising:

(a) a first step of receiving a learning dataset as an input to perform a learning process in a direction of minimizing a total loss function (L total ); and

(b) a second step of receiving an executable image represented in a specific domain to output a classification result of the executable image,

wherein the learning dataset contains one or more learning images represented in one or more domains for one or more classes and one or more learning texts described in one or more texts for the one or more classes,

wherein the first step comprises the steps of:

(1-1) receiving the learning image to generate a visual feature from the received learning image, generating and outputting a classification result of the learning image using the generated visual feature, and performing the learning process using a first loss function (L task );

(1-2) receiving the learning text to generate a textual feature from the received learning text, aligning the visual feature with the textual feature, and performing the learning process using a second loss function (L align ); and

(1-3) receiving the visual feature and the classification result of the learning image to generate a text describing the learning image based on the inputs and performing the learning process using a third loss function (L expl ), and

wherein the learning process is performed using the first loss function (L task ) according to Equation 1 below:

L task =−Σ i y i log( ŷ i )  Equation 1:

where y refers to a one-hot vector in which a ground-truth label is 1 and others are 0, and ŷ refers to a softmax distribution that is the classification result of the learning image generated using the visual feature.

10 . A computer program stored on a non-transitory computer-readable medium, when executed on a computing device performing:

(AA) a first step of receiving a learning dataset as an input to perform a learning process in a direction of minimizing a total loss function (L total ); and

(BB) a second step of receiving an executable image represented in a specific domain to output a classification result of the executable image,

wherein the learning dataset contains one or more learning images represented in one or more domains for one or more classes and one or more learning texts described in one or more texts for the one or more classes,

wherein the first step comprises the steps of:

(1-1) receiving the learning image to generate a visual feature from the received learning image, generating and outputting a classification result of the learning image using the generated visual feature, and performing the learning process using a first loss function (L task );

(1-2) receiving the learning text to generate a textual feature from the received learning text, aligning the visual feature with the textual feature, and performing the learning process using a second loss function (L align ); and

(1-3) receiving the visual feature and the classification result of the learning image to generate a text describing the learning image based on the inputs and performing the learning process using a third loss function (L expl ), and

wherein the learning process is performed using the first loss function (L task ) according to Equation 1 below:

L task =−Σ i y i log( ŷ i )  Equation 1:

where y refers to a one-hot vector in which a ground-truth label is 1 and others are 0, and ŷ refers to a softmax distribution that is the classification result of the learning image generated using the visual feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2023
From: KIM, JIN KYU; PARK, NO KYUNG
To: KOREA UNIVERSITY RESEARCH AND BUSINESS FOUNDATION
Reel/Frame 065756/0924 →
Priority Claims (1)
KR 10-2023-0047711 · Apr 11, 2023 · national
Continuity (1)
Related Publication 20240346812A1 · Oct 17, 2024
References Cited (12)
US 20220391755A1 · Li · 2022 [cited by examiner]
US 20230104127A1 · Babagholami et al. · 2023 [cited by applicant]
US 20240153258A1 · Mangla · 2024 [cited by examiner]
CN 110443293A · 2019 [cited by applicant]
CN 113569932A · 2021 [cited by examiner]
CN 114821196A · 2022 [cited by examiner]
CN 115205592A · 2022 [cited by applicant]
CN 115410031A · 2022 [cited by applicant]
KR 102279797B1 · 2021 [cited by applicant]
KR 1020220138696A · 2022 [cited by applicant]
KR 1020230014034A · 2023 [cited by applicant]
Korean Notice of Allowance issued on Dec. 26, 2023, in connection with the Korean Patent Application No. 10-20230047711, with its English translation, 10 pages. [cited by applicant]