IP Library › Granted Patent US 12,125,271
Granted Patent B2
US 12,125,271 · App. 17/626,171 · Granted Oct 22, 2024

Image paragraph description generating method and apparatus, medium and electronic device

Inventors: Yingwei Pan (Beijing, CN); Ting Yao (Beijing, CN); Tao Mei (Beijing, CN)
Assignees: BEIJING JINGDONG SHANGKE INFORMATION TECHNOLOGY CO., LTD.; BEIJING JINGDONG CENTURY TRADING CO., LTD.
G06V10/82G06N3/08G06V10/40G06V10/771G06V10/776G06V10/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,125,271
App. No.
17/626,171
Granted
Oct 22, 2024
Kind
B2
Abstract

An image paragraph description generating method and apparatus, a medium and an electronic device. The method comprises: obtaining image features of an image (S 101 ); determining the topic of the image according to the image features by using a convolutional automatic coding method (S 102 ); and determining image description information of the image according to the topic by using a long short-term memory (LSTM)-based paragraph coding method (S 103 ), wherein the LSTM comprises a sentence-level LSTM and a paragraph-level LSTM.

Claims (64)

1. An image paragraph description generating method, comprising:

acquiring image features of an image;

determining a topic of the image according to the image features by using a convolutional automatic encoding method; and

determining image description information of the image according to the topic by using a long short-term memory (LSTM) based paragraph encoding method; wherein the LSTM comprises a sentence-level LSTM and a paragraph-level LSTM;

wherein the image features comprise initial regional features of the image, and determining the topic of the image according to the image feature by using the convolutional automatic encoding method comprises:

constructing an initial regional feature vector by concatenating the initial regional features;

obtaining a topic vector by convolving the initial regional feature vector using the convolutional coding method; and

determining the topic of the image based on the topic vector; and

wherein the method further comprises:

reconstructing the topic vector by using a de-convolution decoding method to obtain a reconstructed regional feature vector; and

determining a reconstruction loss of the topic by calculating a distance between the initial regional feature vector and the reconstructed regional feature vector.

2. The method according to claim 1 , further comprising:

determining a number of sentences contained in the image description information according to the topic vector.

3. The method according to claim 1 , further comprising:

obtaining a fused image feature by averagely fusing the initial regional features.

4. The method according to claim 3 , wherein the determining image description information of the image according to the topic by using the LSTM based paragraph encoding method comprises:

determining, according to the fused image feature, an inter-sentences dependency in the image description information and an output vector of the paragraph-level LSTM using the paragraph-level LSTM;

determining an attention distribution of the fused image feature according to the output vector of the paragraph-level LSTM and the topic vector;

obtaining a noticed image feature by performing a weighting processing on the fused image feature based on the attention distribution;

obtaining a sentence generation condition of the topic and words describing the topic by inputting the noticed image feature, the topic vector, and the output vector of the paragraph-level LSTM into the sentence-level LSTM; and

determining the image description information according to the sentence generation condition and the words describing the topic.

5. The method according to claim 1 , further comprising:

obtaining a sequence-level reward of the image by evaluating a coverage range of the image description information using a self-criticism method;

determining a coverage rate of high-frequency objects of the image description information with respect to ground-truth paragraph information; and

obtaining a final reward for the image description information by weighting the coverage rate and then adding the weighted coverage rate with the sequence-level reward.

6. A non-transitory computer-readable medium with a computer program stored thereon, wherein the program is executed by a processor to implement an image paragraph description generating method,

wherein the image paragraph description generating method comprises:

acquiring image features of an image;

determining a topic of the image according to the image features by using a convolutional automatic encoding method; and

determining image description information of the image according to the topic by using a long short-term memory (LSTM) based paragraph encoding method; wherein the LSTM comprises a sentence-level LSTM and a paragraph-level LSTM;

wherein the image features comprise initial regional features of the image, and determining the topic of the image according to the image feature by using the convolutional automatic encoding method comprises:

constructing an initial regional feature vector by concatenating the initial regional features;

obtaining a topic vector by convolving the initial regional feature vector using the convolutional coding method; and

determining the topic of the image based on the topic vector; and

wherein the method further comprises:

reconstructing the topic vector by using a de-convolution decoding method to obtain a reconstructed regional feature vector; and

determining a reconstruction loss of the topic by calculating a distance between the initial regional feature vector and the reconstructed regional feature vector.

7. An electronic device, comprising:

one or more processors;

storage device configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are configured to:

acquire image features of an image;

determine a topic of the image according to the image features by using a convolutional automatic encoding method; and

determine image description information of the image according to the topic by using a long short-term memory (LSTM) based paragraph encoding method; wherein the LSTM comprises a sentence-level LSTM and a paragraph-level LSTM;

wherein the image features comprise initial regional features of the image, and the one or more processors are configured to:

construct an initial regional feature vector by concatenating the initial regional features;

obtain a topic vector by convolving the initial regional feature vector using the convolutional coding method; and

determine the topic of the image based on the topic vector; and

wherein the one or more processors are further configured to:

reconstruct the topic vector by using a de-convolution decoding method to obtain a reconstructed regional feature vector; and

determine a reconstruction loss of the topic by calculating a distance between the initial regional feature vector and the reconstructed regional feature vector.

8. The electronic device according to claim 7 , wherein the processors are further configured to:

determine a number of sentences contained in the image description information according to the topic vector.

9. The electronic device according to claim 7 , wherein the processors are further configured to:

obtain a fused image feature by averagely fusing the initial regional features.

10. The electronic device according to claim 9 , wherein the processors are configured to:

determine, according to the fused image feature, an inter-sentences dependency in the image description information and an output vector of the paragraph-level LSTM using the paragraph-level LSTM;

determine an attention distribution of the fused image feature according to the output vector of the paragraph-level LSTM and the topic vector;

obtain a noticed image feature by performing a weighting processing on the fused image feature based on the attention distribution;

obtain a sentence generation condition of the topic and words describing the topic by inputting the noticed image feature, the topic vector, and the output vector of the paragraph-level LSTM into the sentence-level LSTM; and

determine the image description information according to the sentence generation condition and the words describing the topic.

11. The electronic device according to claim 7 , wherein the processors are configured to:

obtain a sequence-level reward of the image by evaluating a coverage range of the image description information using a self-criticism method;

determine a coverage rate of high-frequency objects of the image description information with respect to ground-truth paragraph information; and

obtain a final reward for the image description information by weighting the coverage rate and then adding the weighted coverage rate with the sequence-level reward.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2022
From: PAN, YINGWEI; YAO, TING; MEI, TAO
To: BEIJING JINGDONG SHANGKE INFORMATION TECHNOLOGY CO., LTD.; BEIJING JINGDONG CENTURY TRADING CO., LTD.
Reel/Frame 058615/0633 →
Priority Claims (1)
CN 201910629398.8 · Jul 12, 2019 · national
Continuity (1)
Related Publication 20220270359A1 · Aug 25, 2022