IP Library › Granted Patent US 11,790,650
Granted Patent B2
US 11,790,650 · App. 16/998,876 · Granted Oct 17, 2023

Contrastive captioning for image groups

Inventors: Quan Hung Tran (San Jose, CA); Long Thanh Mai (San Jose, CA); Zhe Lin (Fremont, CA); Zhuowan Li (Baltimore, MD)
G06V20/30G06F16/535G06F16/55G06F18/214G06F40/205G06V10/751G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,650
App. No.
16/998,876
Filed
Aug 20, 2020
Granted
Oct 17, 2023
Kind
B2
Art Unit
2661
USPC
382/159
Abstract

A group captioning system includes computing hardware, software, and/or firmware components in support of the enhanced group captioning contemplated herein. In operation, the system generates a target embedding for a group of target images, as well as a reference embedding for a group of reference images. The system identifies information in-common between the group of target images and the group of reference images and removes the joint information from the target embedding and the reference embedding. The result is a contrastive group embedding that includes a contrastive target embedding and a contrastive reference embedding with which to construct a contrastive group embedding, which is then input to a model to obtain a group caption for the target group of images.

Claims (84)

1. A method of generating captions for image groups, the method comprising:

in an embedding module:

generating a target embedding for a group of target images; and

generating a reference embedding for a group of reference images;

in a contrast module:

identifying joint information in-common between the group of target images and the group of reference images;

removing the joint information from the target embedding and the reference embedding, resulting in a contrastive target embedding and a contrastive reference embedding;

generating a contrastive group embedding based at least on the contrastive target embedding and the contrastive reference embedding; and

submitting the contrastive group embedding to a model to obtain a group caption for the group of target images.

2. The method of claim 1 wherein:

the model includes a self-attention layer configured to identify and promote prominent features of image groups;

generating the target embedding for the group of target images comprises submitting the group of target images to the model to obtain the target embedding; and

generating the reference embedding for the group of reference images comprises submitting the group of reference images to the model to obtain the reference embedding.

3. The method of claim 1 further comprising training the model on a dataset comprising target groups and reference groups, wherein:

each of the target groups comprises a corresponding group of target images and a corresponding group caption for the group of target images; and

each of the reference groups corresponds to a different one of the target groups and comprises a corresponding group of reference images.

4. The method of claim 3 further comprising generating the dataset on which to train the model by at least:

parsing each image caption, in a set of image captions corresponding to a set of training images, into a scene graph;

identifying the target groups from within the set of training images, wherein each of the target groups comprises a subset of the set of training images having a shared scene graph;

identifying the reference groups from within the set of training images, wherein each of the reference groups corresponds to a different one of the target groups and comprises a different subset of the set of training images having scene graphs that only partially overlap with the shared scene graph of a corresponding one of the target groups; and

generating the group caption for each of the target groups based at least on the shared scene graph for a given target group.

5. The method of claim 1 wherein:

identifying the joint information in-common between the group of target images and the group of reference images comprises generating a joint embedding for a joint group of images that includes the group of target images and the group of reference images; and

removing the joint information from the target embedding and the reference embedding comprises subtracting the joint embedding from the target embedding, to produce the contrastive target embedding, and subtracting the joint embedding from the reference embedding, to produce the contrastive reference embedding.

6. The method of claim 1 wherein:

the group of target images comprises a subset of images returned by a query;

the group of reference images comprises a different subset of the images returned by the query; and

the method further comprises generating a refinement to the query based on the group caption generated for the group of target images.

7. The method of claim 1 wherein the contrastive group embedding comprises a concatenation of the contrastive target embedding and the contrastive reference embedding.

8. A computing apparatus comprising:

one or more computer readable storage media;

one or more processors operatively coupled to the one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media and comprising:

an embedding module that, when executed by the one or more processors, direct the computing apparatus to:

generate a target embedding for a group of target images; and

generate a reference embedding for a group of reference images; and

a contrast module that, when executed by the one or more processors, directs the computing apparatus to:

identify joint information in-common between the group of target images and the group of reference images;

remove the joint information from the target embedding and the reference embedding, to produce a contrastive target embedding and a contrastive reference embedding;

generate a contrastive group embedding based at least on the contrastive target embedding and the contrastive reference embedding; and

generate a group caption for the group of target images based at least on processing the contrastive group embedding using a model.

9. The computing apparatus of claim 8 wherein, to generate the group caption for the group of target images, the program instructions when executed by the one or more processors direct the computing apparatus to input the contrastive group embedding to the model to obtain the group caption.

10. The computing apparatus of claim 9 wherein:

the model includes a self-attention layer configured to identify and promote prominent features of image groups;

to generate the target embedding for the group of target images, the program instructions when executed by the one or more processors direct the computing apparatus to input the group of target images to the model to obtain the target embedding; and

to generate the reference embedding for the group of reference images, the program instructions when executed by the one or more processors direct the computing apparatus to input the group of reference images to the model to obtain the reference embedding.

11. The computing apparatus of claim 9 wherein:

the model comprises an artificial neural network trained on a dataset comprising target groups and reference groups:

each of the target groups comprises a corresponding group of target images and a corresponding group caption for the group of target images; and

each of the reference groups corresponds to a different one of the target groups and comprises a corresponding group of reference images.

12. The computing apparatus of claim 8 wherein:

to identify the joint information in-common between the group of target images and the group of reference images, the program instructions when executed by the one or more processors direct the computing apparatus to generate a joint embedding for a joint group of images that includes the group of target images and the group of reference images; and

to remove the joint information from the target embedding and the reference embedding, the program instructions when executed by the one or more processors direct the computing apparatus to subtract the joint embedding from the target embedding, and to subtract the joint embedding from the reference embedding, to produce the contrastive target embedding and the contrastive reference embedding, respectively.

13. The computing apparatus of claim 8 wherein:

the group of target images comprises a subset of images returned by a query;

the group of reference images comprises a different subset of the images returned by the query; and

the program instructions when executed by the one or more processors further direct the computing apparatus to generate a refinement to the query based on the group caption generated for the group of target images.

14. The computing apparatus of claim 8 wherein the contrastive group embedding comprises a concatenation of the contrastive target embedding and the contrastive reference embedding.

15. A method comprising:

generating a target embedding for a group of target images;

generating a reference embedding for a group of reference images;

identifying joint information in-common between the group of target images and the group of reference images;

removing the joint information from the target embedding and the reference embedding, resulting in a contrastive target embedding and a contrastive reference embedding;

generating a contrastive group embedding based at least on the contrastive target embedding and the contrastive reference embedding; and

submitting the contrastive group embedding to a model to obtain a group caption for the group of target images.

16. The method of claim 15 wherein:

the model includes a self-attention layer configured to identify and promote prominent features of image groups;

generating the target embedding for the group of target images comprises submitting the group of target images to the model to obtain the target embedding; and

generating the reference embedding for the group of reference images comprises submitting the group of reference images to the model to obtain the reference embedding.

17. The method of claim 15 further comprising training the model on a dataset comprising target groups and reference groups, wherein:

each of the target groups comprises a corresponding group of target images and a corresponding group caption for the group of target images; and

each of the reference groups corresponds to a different one of the target groups and comprises a group of reference images.

18. The method of claim 17 further comprising generating the dataset on which to train the model by at least:

parsing each image caption, in a set of image captions corresponding to a set of training images, into a scene graph;

identifying the target groups from within the set of training images, wherein each of the target groups comprises a subset of the set of training images having a shared scene graph;

identifying the reference groups from within the set of training images, wherein each of the reference groups corresponds to a different one of the target groups and comprises a different subset of the set of training images having scene graphs that only partially overlap with the shared scene graph of a corresponding one of the target groups; and

generating the group caption for each of the target groups based at least on the shared scene graph for a given target group.

19. The method of claim 15 wherein:

identifying the joint information in-common between the group of target images and the group of reference images comprises generating a joint embedding for a joint group of images that includes the group of target images and the group of reference images; and

removing the joint information from the target embedding and the reference embedding comprises subtracting the joint embedding from the target embedding, to produce the contrastive target embedding, and subtracting the joint embedding from the reference embedding, to produce the contrastive reference embedding.

20. The method of claim 15 wherein:

the group of target images comprises a subset of images returned by a query;

the group of reference images comprises a different subset of images returned by the query; and

the method further comprises generating a refinement to the query based on the group caption generated for the group of target images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2020
From: TRAN, QUAN HUNG; LI, ZHUOWAN; LIN, ZHE; MAI, LONG THANH
To: ADOBE INC.
Reel/Frame 053556/0401 →
Continuity (1)
Related Publication 20220058390A1 · Feb 24, 2022