IP Library Granted Patent US 12,614,401
Granted Patent B2
US 12,614,401 · App. 18/499,717 · Granted Apr 28, 2026

Apparatus and method for generating text from image and method of training model for generating text from image

Inventors: Dong Hyuck Im (Daejeon, KR); Jung Hyun Kim (Daejeon, KR); Hye Mi Kim (Daejeon, KR); Jee Hyun Park (Daejeon, KR); Yong Seok Seo (Daejeon, KR); Won Young Yoo (Daejeon, KR)
Assignee: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
G06V20/70G06F40/40G06T11/00G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,401
App. No.
18/499,717
Granted
Apr 28, 2026
Kind
B2
Abstract

An apparatus for generating text from an image may comprise: a memory configured to store at least one instruction; and a processor configured to execute the at least one instruction, wherein the processor is further configured to generate encoding information for an image based on the image and extract text information related to content of the image based on a degree of association with the encoding information.

Claims (41)

1 . An apparatus for generating text from an image, the apparatus comprising:

a memory configured to store at least one instruction; and

a processor configured to execute the at least one instruction,

wherein the processor is further configured to:

convert the image into an image token sequence using a learned codebook;

generate encoding information comprising the image token sequence; and

generate text information describing semantic content of the image by applying an image-to-text translation model that receives the encoding information as an input sequence, and

wherein the image-to-text translation model is configured to generate the text information based on learned degrees of association between image tokens in the encoding information and text tokens.

2 . The apparatus of claim 1 , wherein each token of the image token sequence corresponds to a unit block of the image, the unit block having a size of N pixels×M pixels, N and M being positive integers greater than one.

3 . The apparatus of claim 2 , wherein N and M are equivalent.

4 . The apparatus of claim 1 , wherein the image-to-text translation model comprises a transformer-based encoder-decoder architecture configured to generate the text information based on the learned degrees of association between the image tokens and the text tokens.

5 . The apparatus of claim 4 , wherein

at least one encoder layer of the transformer-based encoder-decoder architecture comprises a self-attention neural network layer and a feed-forward neural network layer, and

at least one decoder layer of the transformer-based encoder-decoder architecture comprises a self-attention neural network layer, an encoder-decoder attention neural network layer, and a feed-forward neural network layer.

6 . The apparatus of claim 1 , wherein the processor is further configured to generate image-text synthetic data for the image by combining the image and the text information.

7 . The apparatus of claim 6 , wherein the processor is further configured to generate an image based on the text using the image-text synthetic data.

8 . A method of generating text from an image performed by a processor executing at least one instruction stored in a memory, the method comprising:

converting the image into an image token sequence using a learned codebook;

generating encoding information comprising the image token sequence; and

generating text information describing semantic content of the image by applying an image-to-text translation model that receives the encoding information as an input sequence,

wherein the image-to-text translation model is configured to generate the text information based on learned degrees of association between image tokens in the encoding information and text tokens.

9 . The method of claim 8 , wherein each token of the image token sequence corresponds to a unit block of the image, the unit block having a size of N pixels×M pixels, N and M being positive integers greater than one.

10 . The method of claim 9 , wherein N and M are equivalent.

11 . The method of claim 8 , wherein the image-to-text translation model comprises a transformer-based encoder-decoder architecture configured to generate the text information based on the learned degrees of association between the image tokens and the text tokens.

12 . The method of claim 11 , wherein

at least one encoder layer of the transformer-based encoder-decoder architecture comprises a self-attention neural network layer and a feed-forward neural network layer, and

at least one decoder layer of the transformer-based encoder-decoder architecture comprises a self-attention neural network layer, an encoder-decoder attention neural network layer, and a feed-forward neural network layer.

13 . The method of claim 8 , further comprising combining the image and the text information to generate image-text synthetic data for the image.

14 . The method of claim 13 , further comprising generating an image based on the text using the image-text synthetic data.

15 . A method of training an image-to-text translation model for generating text describing semantic content of an image, performed by a processor executing at least one instruction stored in a memory, the method comprising:

converting an image into an image token sequence using a learned codebook;

generating encoding information comprising the image token sequence; and

training the image-to-text translation model to learn a function for generating text information describing semantic content of the image using the encoding information as an input sequence,

wherein the function enables the image-to-text translation model to generate the text information based on learned degrees of association between image tokens in the encoding information and text tokens.

16 . The method of claim 15 , wherein each token of the image token sequence corresponds to a unit block of the image, the unit block having a size of N pixels×M pixels, N and M being positive integers greater than one.

17 . The method of claim 15 , further comprising:

inputting a training image to an encoder-decoder model; and

training the encoder-decoder model to learn a function for generating the learned codebook from the training image.

18 . The method of claim 15 , wherein the image-to-text translation model comprises a transformer-based encoder-decoder architecture configured to generate, using the function, the text information based on the learned degrees of association between the image tokens and the text tokens.

19 . The method of claim 15 , further comprising combining the image and the text information to generate image-text synthetic data for the image.

20 . The method of claim 19 , further comprising training a text-based image generation model to learn a function for generating an image based on text using the image-text synthetic data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2023
From: IM, DONG HYUCK; KIM, JUNG HYUN; KIM, HYE MI; PARK, JEE HYUN; SEO, YONG SEOK; YOO, WON YOUNG
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 065423/0712 →
Priority Claims (1)
KR 10-2022-0158863 · Nov 24, 2022 · national
Continuity (1)
Related Publication 20240177507A1 · May 30, 2024
References Cited (13)
US 9298682B2 · Makadia et al. · 2016 [cited by applicant]
US 9858524B2 · Bengio et al. · 2018 [cited by applicant]
US 10949744B2 · Lin et al. · 2021 [cited by applicant]
US 11195048B2 · Bui et al. · 2021 [cited by applicant]
US 11281709B2 · Zheng et al. · 2022 [cited by applicant]
US 20220188703A1 · Hong et al. · 2022 [cited by applicant]
US 20230206661A1 · Choi et al. · 2023 [cited by applicant]
US 20240087083A1 · Wang · 2024 [cited by examiner]
US 20240127005A1 · Panda · 2024 [cited by examiner]
CN 101620615A · 2010 [cited by examiner]
KR 1020210130980 · 2021 [cited by applicant]
KR 1020220082618 · 2022 [cited by applicant]
KR 1020220098956 · 2022 [cited by applicant]