IP Library › Granted Patent US 12,586,392
Granted Patent B2
US 12,586,392 · App. 18/179,177 · Granted Mar 24, 2026

Perturbation robust metric for evaluating image captions

Inventors: Seunghyun Yoon (San Jose, CA); Trung Bui (San Jose, CA)
Assignee: Adobe Inc.
G06V20/70G06F40/58G06T1/0021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,392
App. No.
18/179,177
Granted
Mar 24, 2026
Kind
B2
Abstract

Embodiments are disclosed for training an image caption evaluation system to perform evaluations of image captions. In particular, in one or more embodiments, the disclosed systems and methods comprise receiving a training image, a ground truth image caption for the training image, and a perturbed image caption for the training image, where the perturbed image caption includes modifications to the ground truth image caption. The disclosed systems and methods further comprise generating, by a visual encoder, a visual embedding representation of the training image and generating, by a perturbation-aware text encoder, a first text embedding for the ground truth image caption and a second text embedding for the perturbed image caption. The disclosed systems and methods further comprise computing losses between the visual embedding, the first text embedding, and the second text embedding and training the perturbation-aware text encoder based on the computed losses.

Claims (73)

1 . A computer-implemented method, comprising:

receiving a training image and a ground truth image caption for the training image;

generating a perturbed image caption for the training image by performing modifications to text elements of the ground truth image caption;

generating, by a visual encoder, a visual embedding representation of the training image;

generating, by a perturbation-aware text encoder, a first text embedding for the ground truth image caption and a second text embedding for the perturbed image caption;

computing losses between the visual embedding representation of the training image, the first text embedding for the ground truth image caption, and the second text embedding for the perturbed image caption; and

training the perturbation-aware text encoder based on the computed losses.

2 . The computer-implemented method of claim 1 , further comprising:

performing an initial training phase of the perturbation-aware text encoder using a set of ground truth image captions in a first language and the set of ground truth image captions translated into a second language.

3 . The computer-implemented method of claim 2 , wherein performing the initial training phase of the perturbation-aware text encoder further comprises:

for each ground truth image caption of the set of ground truth image captions:

translating the ground truth image caption from the first language to the second language;

generating, by a text encoder, a first text embedding representation of the ground truth image caption in the first language;

generating, by the perturbation-aware text encoder, a second text embedding representation of the ground truth image caption in the second language;

computing a loss between the first text embedding representation and the second text embedding representation; and

backpropagating the loss to train the perturbation-aware text encoder.

4 . The computer-implemented method of claim 1 , further comprising:

generating the perturbed image caption by replacing first text elements in the ground truth image caption with second text elements from a different image caption.

5 . The computer-implemented method of claim 1 , further comprising:

generating the perturbed image caption by swapping text elements within the ground truth image caption.

6 . The computer-implemented method of claim 1 , further comprising:

generating the perturbed image caption by removing text elements within the ground truth image caption.

7 . The computer-implemented method of claim 1 , wherein computing the losses between the visual embedding representation of the training image, the first text embedding for the ground truth image caption, and the second text embedding for the perturbed image caption further comprises:

computing a first loss between the visual embedding representation of the training image and the first text embedding for the ground truth image caption;

computing a second loss between the visual embedding representation of the training image and the second text embedding for the perturbed image caption;

computing a third loss between the first text embedding for the ground truth image caption and the second text embedding for the perturbed image caption; and

aggregating the first loss, the second loss, and the third loss to generate an overall loss.

8 . The computer-implemented method of claim 1 , wherein the training image is part of a training dataset, and wherein each image in the training dataset is associated with a ground truth image caption for the image, a description of an image background, a description of objects in the image, and a description of a relationship between the objects in the image.

9 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a training image and a ground truth image caption for the training image;

generating a perturbed image caption for the training image by performing modifications to text elements of the ground truth image caption;

generating, by a visual encoder, a visual embedding representation of the training image;

generating, by a perturbation-aware text encoder, a first text embedding for the ground truth image caption and a second text embedding for the perturbed image caption;

computing losses between the visual embedding representation of the training image, the first text embedding for the ground truth image caption, and the second text embedding for the perturbed image caption; and

training the perturbation-aware text encoder based on the computed losses.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the processing device to perform operations comprising:

performing an initial training phase of the perturbation-aware text encoder using a set of ground truth image captions in a first language and the set of ground truth image captions translated into a second language.

11 . The non-transitory computer-readable storage medium of claim 10 , wherein to perform the initial training phase of the perturbation-aware text encoder the instructions further cause the processing device to perform operations comprising:

for each ground truth image caption of the set of ground truth image captions:

translating the ground truth image caption from the first language to the second language;

generating, by a text encoder, a first text embedding representation of the ground truth image caption in the first language;

generating, by the perturbation-aware text encoder, a second text embedding representation of the ground truth image caption in the second language;

computing a loss between the first text embedding representation and the second text embedding representation; and

backpropagating the loss to train the perturbation-aware text encoder.

12 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the processing device to perform operations comprising:

generating the perturbed image caption by replacing first text elements in the ground truth image caption with second text elements from a different image caption.

13 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the processing device to perform operations comprising:

generating the perturbed image caption by swapping text elements within the ground truth image caption.

14 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the processing device to perform operations comprising:

generating the perturbed image caption by removing text elements within the ground truth image caption.

15 . The non-transitory computer-readable storage medium of claim 9 , wherein to computer the losses between the visual embedding representation of the training image, the first text embedding for the ground truth image caption, and the second text embedding for the perturbed image caption the instructions further cause the processing device to perform operations comprising:

computing a first loss between the visual embedding representation of the training image and the first text embedding for the ground truth image caption;

computing a second loss between the visual embedding representation of the training image and the second text embedding for the perturbed image caption;

computing a third loss between the first text embedding for the ground truth image caption and the second text embedding for the perturbed image caption; and

aggregating the first loss, the second loss, and the third loss to generate an overall loss.

16 . The non-transitory computer-readable storage medium of claim 9 , wherein the training image is part of a training dataset, and wherein each image in the training dataset is associated with a ground truth image caption for the image, a description of an image background, a description of objects in the image, and a description of a relationship between the objects in the image.

17 . A computer-implemented method, comprising:

receiving an image and an image caption, the image caption describing content in the image;

generating, by a visual encoder, a visual embedding representation of the image;

generating, by a perturbation-aware text encoder, a text embedding for the image caption, wherein the perturbation-aware text encoder is trained to recognize perturbations in image captions; and

computing an image caption score for the image caption using the visual embedding representation of the image and text embedding for the image caption, wherein the image caption score is a metric indicating an accuracy of the image caption to the content in the image.

18 . The computer-implemented method of claim 17 , wherein the perturbation-aware text encoder is trained by:

receiving a training image and a ground truth image caption for the training image;

generating a perturbed image caption for the training image by performing modifications to text elements of the ground truth image caption;

generating, by the visual encoder, a visual embedding representation of the training image;

generating, by the perturbation-aware text encoder, a first text embedding for the ground truth image caption and a second text embedding for the perturbed image caption;

computing losses between the visual embedding representation of the training image, the first text embedding for the ground truth image caption, and the second text embedding for the perturbed image caption; and

training the perturbation-aware text encoder to recognize the perturbations in the perturbed image caption using the computed losses.

19 . The computer-implemented method of claim 17 , wherein computing an image caption score using the visual embedding representation of the image and text embedding for the image caption comprises:

computing a cosine similarity of the visual embedding representation of the image and the text embedding for the image caption; and

applying a weighting to the computed cosine similarity to compute the image caption score for the image caption.

20 . The computer-implemented method of claim 17 , wherein the perturbation-aware text encoder is trained by:

performing an initial training phase of the perturbation-aware text encoder using a set of ground truth image captions in a first language and the set of ground truth image captions translated into a second language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2023
From: YOON, SEUNGHYUN; BUI, TRUNG
To: ADOBE INC.
Reel/Frame 062897/0158 →
Continuity (1)
Related Publication 20240304009A1 · Sep 12, 2024
References Cited (18)
US 20220058390A1 · Tran · 2022 [cited by examiner]
Anderson et al., “SPICE: Semantic Propositional Image Caption Evaluation”, Jul. 29, 2016, pp. 1-17. [cited by applicant]
Banerjee et al., “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments”, Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and… [cited by applicant]
Chen et al., “UNITER: UNiversal Image-TExt Representation Learning”, Jul. 17, 2020, pp. 1-26. [cited by applicant]
Chin-Yew Lin, “Rouge: A Package for Automatic Evaluation of Summaries”, 2004, 8 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, May 24, 2019, 16 pages. [cited by applicant]
Github, “Multilingual-CLIP”, available online at <https://github.com/FreddeFrallan/Multilingual-CLIP>, Oct. 11, 2024, 5 pages. [cited by applicant]
Hessel et al., CLIPScore: A Reference-free Evaluation Metric for Image Captioning, Mar. 23, 2022, 15 pages. [cited by applicant]
Kusner et al., “From Word Embeddings to Document Distances”, 2015, 10 pages. [cited by applicant]
Lee et al., “UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning”, Jun. 26, 2021, 8 pages. [cited by applicant]
Lee et al., “VILBERTScore: Evaluating Image Caption Using Vision-and-Language BERT”, Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP), Nov. 20, 2020, pp. 34-39. [cited by applicant]
Madhyastha et al., “VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions”, Jul. 22 2019, 12 pages. [cited by applicant]
Papineni et al., “BLEU: a Method for Automatic Evaluation of Machine Translation”, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Jul. 2002, pp. 311-318. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Feb. 26, 2021, 48 pages. [cited by applicant]
Sai et al., “Perturbation CheckLists for Evaluating NLG Evaluation Metrics”, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Nov. 7-11, 2021, pp. 7219-7234. [cited by applicant]
Vedantam et al., “CIDEr: Consensus-based Image Description Evaluation”, Jun. 3, 2015, pp. 1-17. [cited by applicant]
Yi et al., “Improving Image Captioning Evaluation by Considering Inter References Variance”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5-10, 2020, pp. 985-994. [cited by applicant]
Zhang et al., “BERTScore: Evaluating Text Generation With BERT”, ICLR, Feb. 24, 2020, pp. 1-43. [cited by applicant]