IP Library › Granted Patent US 12,167,100
Granted Patent B2
US 12,167,100 · App. 17/292,627 · Granted Dec 10, 2024

Method, apparatus, device and medium for generating captioning information of multimedia data

Inventors: Ke Lin (Beijing, CN); Zhuoxin Gan (Beijing, CN); Yingying Jiang (Beijing, CN)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H04N21/4884G06N3/04G06V10/454G06V10/764G06V10/82G06V20/41G06V20/635H04N21/234336
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,167,100
App. No.
17/292,627
Granted
Dec 10, 2024
Kind
B2
Abstract

Embodiments of the present disclosure provide a method, an apparatus, a device, and a medium for generating captioning information of multimedia data. The method includes extracting characteristic information of multimedia data to be processed, wherein the multimedia data comprises a video or an image; and generating a text caption of the multimedia data based on the extracted characteristic information. According to the method provided in the embodiments of the present disclosure, the accuracy of the generated text caption of the multimedia data can be effectively improved.

Claims (53)

1. A method for generating captioning information of multimedia data, comprising:

extracting characteristic information of multimedia data to be processed, wherein the multimedia data comprises a video or an image; and

generating a text caption of the multimedia data based on the extracted characteristic information, wherein the generating the text caption of the multimedia data based on the extracted characteristic information comprises:

determining a first application scenario for the multimedia data by analyzing the multimedia data;

obtaining length information of the text caption to be generated for the first application scenario, wherein the length information indicates at least one length of a plurality of lengths that respectively correspond to a plurality of application scenarios including the first application scenario; and

generating the text caption based on the length information and the extracted characteristic information.

2. The method of claim 1 , wherein the extracting characteristic information of the multimedia data to be processed comprises at least one of the following:

extracting local visual features of targets contained in respective target regions of each image in the multimedia data;

extracting semantic features of the multimedia data;

extracting spatial-temporal visual features of the multimedia data when the multimedia data is a video;

extracting global visual features of the multimedia data;

extracting attribute features of the targets contained in the respective target regions of each image in the multimedia data; and

extracting global attribute features of each image in the multimedia data.

3. The method of claim 2 , wherein the characteristic information comprises the local visual features of the targets contained in respective target regions in each image of the multimedia data, and the generating the text caption of the multimedia data based on the extracted characteristic information, comprising:

obtaining relationship features between the targets based on the local visual features of each target in the image;

constructing a scene graph of the image based on the local visual features and the relationship features;

obtaining graph convolution features of the image based on the scene graph of the image; and

generating the text caption of the multimedia data based on the graph convolution features of each image of the multimedia data.

4. The method of claim 3 , wherein the scene graph comprises a plurality of nodes and a plurality of edges, wherein one node represents a local visual feature of one target, and each of the plurality of edges represents the relationship feature between two connected nodes.

5. The method of claim 3 , wherein the characteristic information comprises the attribute features of the targets contained in respective target regions of each image in the multimedia data;

the constructing of the scene graph of the image based on the local visual features and the relationship features comprises:

constructing the scene graph of the image based on the local visual features of each target, the relationship features between the targets, and the attribute features of each target, wherein one node in the scene graph represents the local visual features or attribute features of one target.

6. The method of claim 3 , wherein, when the multimedia data is the video, the images of the multimedia data are a plurality of frames selected from the video, and when the target regions of two adjacent frames comprise the same targets, the scene graphs of the two adjacent frames have temporal edges between the nodes corresponding to the same target.

7. The method of claim 3 , wherein the obtaining the graph convolution features of the image based on the scene graph of the image comprises:

obtaining a target dimension of feature vector by encoding nodes and edges in the scene graph; and

obtaining the graph convolution features by using a graph convolution network based on the obtained feature vector.

8. The method of claim 2 , wherein when the characteristic information of the multimedia data comprises at least two of the local visual feature, the semantic feature, the spatial-temporal visual feature, and the global feature, the generating the text caption of the multimedia data based on the extracted characteristic information comprises:

determining weights of each characteristic information;

weighting each characteristic information based on the weights of each characteristic information; and

generating the text caption of the multimedia data based on the weighted characteristic information.

9. The method of claim 2 , wherein the generating the text caption of the multimedia data based on the extracted characteristic information comprises:

encoding the obtained characteristic information by using self-attention-based encoder;

inputting the encoded characteristic information to a decoder to generate the text caption of the multimedia data;

wherein when the multimedia data is an image, the self-attention-based encoder is a self-attention-based intra-frame encoder; when the multimedia data is a video, the self-attention-based encoder comprises a self-attention-based intra-frame encoder and/or a self-attention-based inter-frame encoder.

10. The method of claim 1 , wherein the generating the text caption of the multimedia data based on the extracted characteristic information comprises:

inputting the extracted characteristic information into a plurality of decoders, respectively; and

generating the text caption of the multimedia data based on decoding results of the decoders.

11. The method of claim 1 , wherein the text caption of the multimedia data is generated through a multimedia data captioning model, wherein the multimedia data captioning model is obtained by training in the following manner:

obtaining training samples, wherein the training samples comprise a first sample multimedia data with captioning labels;

training an initial captioning model based on the first sample multimedia data until a model loss function converges; and taking the trained captioning model as the multimedia data captioning model.

12. The method of claim 11 , wherein the training samples further comprise a second sample multimedia data without the captioning labels, and the model loss function comprises a first loss function and a second loss function;

the training the initial captioning model based on the first sample multimedia data until the model loss function converges comprises:

training a preset captioning model based on the first sample multimedia data to obtain a value of the first loss function, and training the captioning model based on the second sample multimedia data to obtain a value of the second loss function;

obtaining a value of the final loss function based on the value of the first loss function and the value of the second loss function; and

training the captioning model based on the value of the final loss function until the final loss function converges.

13. An apparatus for generating captioning information of multimedia data, comprising:

a memory storing one or more instructions; and

a processor configured to execute the one or more instructions stored in the memory to:

extract characteristic information of multimedia data to be processed, wherein the multimedia data comprises a video or an image;

determine a first application scenario for the multimedia data by analyzing the multimedia data;

obtain length information of the text caption to be generated for the first application scenario, wherein the length information indicates at least one length of a plurality of lengths that respectively correspond to a plurality of application scenarios including the first application scenario; and

generate a text caption of the multimedia data based on the length information and the extracted characteristic information.

14. A computer program product including a non-transitory computer-readable storage medium, wherein the storage medium stores a computer program that, when executed by a processor, performs the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2021
From: LIN, KE; GAN, ZHUOXIN; JIANG, YINGYING
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 056189/0960 →
Priority Claims (4)
CN 201910219009.4 · Mar 21, 2019 · national
CN 201910270450.5 · Apr 4, 2019 · national
CN 201911115147.4 · Nov 14, 2019 · national
CN 202010152713.5 · Mar 6, 2020 · national
Continuity (1)
Related Publication 20220014807A1 · Jan 13, 2022
Cited By (3)
US 12,437,528 US 12,541,999 US 12,567,269