IP Library Granted Patent US 12,437,565
Granted Patent B2
US 12,437,565 · App. 17/925,354 · Granted Oct 7, 2025

Apparatus and method for automatically generating image caption by applying deep learning algorithm to an image

Inventors: Seung Ho Han (Daejeon, KR); Ho Jin Choi (Daejeon, KR)
Assignee: KOREA ADVANCED INSTITUTE OF SCIENCE AND TECHNOLOGY
G06V20/70G06F40/44G06T7/90G06V10/44G06N3/02G06T2207/10024G06T2207/20081G06T2207/20084G06V10/768G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,565
App. No.
17/925,354
Filed
Nov 15, 2022
Granted
Oct 7, 2025
Kind
B2
Examiner
SABAH, HARIS
Art Unit
2682
USPC
382/190
Abstract

An apparatus and method for automatically generating an image caption is provided capable of giving an explanation by using Bayesian inference and an image area-word mapping module on the basis of a deep learning algorithm. An apparatus for automatically generating an image caption, according to one embodiment of the present invention, includes: an automatic caption generation module for creating a caption by applying a deep learning algorithm to an image received from a client; a caption basis generation module for creating a basis for the caption by mapping a partial area in the image received from the client with respect to important words in the caption received from the automatic caption generation module; and a visualization module for visualizing the caption received from the automatic caption generation module and the basis for the caption received from the caption basis generation module to return same to the client.

Claims (25)

1. An apparatus for automatically generating an image caption, the apparatus comprising:

an automatic caption generation module configured to generate a caption by applying a deep learning algorithm to an image received from a client;

a caption basis generation module configured to generate a basis for the caption by mapping a partial area in the image received from the client with respect to important words in the caption received from the automatic caption generation module; and

a visualization module configured to visualize the caption received from the automatic caption generation module and the basis for the caption received from the caption basis generation module to return the visualized caption and basis to the client,

wherein the caption basis generation module includes:

an object recognition module configured to recognize one or more objects included in the image received from the client and extract one or more object areas;

an image area-word mapping module configured to train a relevance between words in the caption generated by the automatic caption generation module and each of the object areas extracted by the object recognition module using a deep learning algorithm, and output a weight matrix as a result of the training; and

an interpretation reinforcement module configured to extract a word having a highest weight for each object area from the weight matrix received from the image area-word mapping module, and calculate a posterior probability for each word.

2. The apparatus of claim 1 , wherein the automatic caption generation module includes:

an image feature extraction module configured to extract an image feature vector from the image received from the client using a convolutional neural network algorithm; and

a language generation module configured to pre-train a predefined image feature vector and an actual caption (ground truth) for the predefined image feature vector and generate the caption corresponding to a result of the training with respect to the image feature vector extracted by the image feature extraction module.

3. The apparatus of claim 1 , wherein the visualization module displays the one or more object areas and the caption on the image received from the client, displays the words in the caption corresponding to the object areas as the basis for the caption with the same color as the object areas, and returns an output image indicating a relevance value between the words having the same color as the object areas to the client.

4. A method of automatically generating an image caption, the method comprising:

generating, by an automatic caption generation module, a caption by applying a deep learning algorithm to an image received from a client;

generating, by a caption basis generation module, a basis for the caption by mapping a partial area in the image received from the client with respect to important words in the caption received from the automatic caption generation module; and

visualizing, by a visualization module, the caption received from the automatic caption generation module and the basis for the caption received from the caption basis generation module to return the visualized caption and basis to the client,

wherein the generating of the basis includes:

recognizing, by an object recognition module, one or more objects included in the image received from the client and extracting one or more object areas;

training, by an image area-word mapping module, a relevance between words in the caption generated by the automatic caption generation module and each of the object areas extracted by the object recognition module using a deep learning algorithm, and outputting a weight matrix as a result of the training; and

extracting, by an interpretation reinforcement module, a word having a highest weight for each object area from the weight matrix received from the image area-word mapping module, and calculating a posterior probability for each word.

5. The method of claim 4 , wherein the generating of the caption includes:

extracting, by an image feature extraction module, an image feature vector from the image received from the client using a convolutional neural network algorithm; and

pre-training, by a language generation module, a predefined image feature vector and an actual caption (ground truth) for the predefined image feature vector and generating the caption corresponding to a result of the training with respect to the image feature vector extracted by the image feature extraction module.

6. The method of claim 4 , wherein the returning of the visualized caption and basis includes displaying, by the visualization module, the one or more object areas and the caption on the image received from the client, and by the visualization module, displaying the words in the caption corresponding to the object areas as the basis for the caption with the same color as the object areas, and returning an output image indicating a relevance value between the words having the same color as the object areas to the client.

7. A computer program product comprising a non-transitory computer readable storage medium having instructions, which when executed by a processor, executes the method of claim 4 using a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: HAN, SEUNG HO; CHOI, HO JIN
To: KOREA ADVANCED INSTITUTE OF SCIENCE AND TECHNOLOGY
Reel/Frame 061768/0909 →
Continuity (1)
Related Publication 20230177854A1 · Jun 8, 2023
References Cited (15)
US 9836671B2 · Gao · 2017 [cited by examiner]
US 10198671B1 · Yang et al. · 2019 [cited by applicant]
US 20170200065A1 · Wang · 2017 [cited by examiner]
US 20170200066A1 · Wang · 2017 [cited by examiner]
US 20180181281A1 · Suki · 2018 [cited by examiner]
US 20190227988A1 · Kwon · 2019 [cited by examiner]
KR 1020110033179 · 2011 [cited by applicant]
KR 1020130127458 · 2013 [cited by applicant]
KR 101930940B1 · 2018 [cited by applicant]
KR 101996371B1 · 2019 [cited by applicant]
KR 1020190108378A · 2019 [cited by applicant]
KR 1020200065832A · 2020 [cited by applicant]
KR 1020200106115A · 2020 [cited by applicant]
Machine Translation in English of Korean Prior art Pub (KR 101930940) to Lim et al. [cited by examiner]
PCT International Search Report for PCT Application No. PCT/KR2020/007755, dated Mar. 5, 2021. [cited by applicant]