IP Library › Patent Application 18997002
Patent Application
App. No. 18/997,002

IMAGE CAPTION GENERATION MODEL LEARNING APPARATUS, IMAGE CAPTION GENERATION APPARATUS, IMAGE CAPTION GENERATION MODEL LEARNING METHOD, IMAGE CAPTION GENERATION METHOD, AND PROGRAM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/997,002
Abstract

An image caption generation model learning apparatus uses, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data; and learns an image parameter that is a model parameter for image hidden information generation, a text parameter that is a model parameter for text hidden information generation, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.

Claims (26)

1 . An image caption generation model learning apparatus comprising:

processing circuitry configured to

use, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data; and

learn an image parameter that is a model parameter for image hidden information generation, a text parameter that is a model parameter for text hidden information generation, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.

2 . The image caption generation model learning apparatus according to claim 1 ,

the processing circuitry configured to

generate image hidden information from the image and the image parameter;

generate text hidden information from the first language text and the text parameter;

generate inter-crossmodal invariant information from the image hidden information or the text hidden information and the crossmodal parameter;

generate a text generation probability from the inter-crossmodal invariant information and the output parameter; and

estimate various model parameters such that a sum of a text generation probability of a text corresponding to a caption and a text generation probability of a text corresponding to a translation result of machine translation becomes maximum.

3 . An image caption generation apparatus comprising:

processing circuitry configured to

generate a caption describing an input image based on an image parameter that is model parameter for image hidden information generation learned by using, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.

4 . The image caption generation apparatus according to claim 3 ,

the processing circuitry configured to

generate image hidden information from an image and the image parameter;

generate inter-crossmodal invariant information from the image hidden information and the crossmodal parameter; and

generate a text generation probability from the inter-crossmodal invariant information and the output parameter and generates a text serving as the caption of the image.

5 . An image caption generation model learning method executed by an image caption generation model learning apparatus, the image caption generation model learning method comprising:

using, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data; and

learning an image parameter that is a model parameter for image hidden information generation, a text parameter that is a model parameter for text hidden information generation, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.

6 . An image caption generation method executed by an image caption generation apparatus, the image caption generation method comprising:

generating a caption describing an input image based on an image parameter that is model parameter for image hidden information generation learned by using, as inputs, pair data of an image that is learning data for image caption generation and text data that is a caption describing the image and pair data of a first language text and a second language text that are machine translation data, a crossmodal parameter that is a model parameter for crossmodal invariant information embedment, and an output parameter that is a model parameter for text generation.

7 . A non-transitory computer readable medium storing a computer program for causing a computer to function as the image caption generation model learning apparatus according to claim 1 .

8 . A non-transitory computer readable medium storing a computer program for causing a computer to function as the image caption generation apparatus according to claim 3 .

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0641 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2025
From: MASUMURA, RYO; TAKASHIMA, AKIHIKO; IHORI, MANA
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 069925/0128 →