IP Library Granted Patent US 12,645,877
Granted Patent B2
US 12,645,877 · App. 18/315,391 · Granted Jun 2, 2026

Systems and methods for generating text descriptive of digital images

Inventors: Pranav Aggarwal (San Jose, CA); Di Pu (San Jose, CA); Daniel ReMine (San Jose, CA); Ajinkya Kale (San Jose, CA)
Assignee: Adobe Inc.
G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,877
App. No.
18/315,391
Granted
Jun 2, 2026
Kind
B2
Abstract

In implementation of techniques for generating text descriptive of digital images, a computing device implements a captioning system to generate a text description of an image using a machine learning model. The captioning system identifies a negative word in the text description of the image based on a quantitative risk level for the negative word. Based on the text description of the image and the quantitative risk level, the captioning system determines an amount of suppression for use of the negative word. The captioning system then generates an updated text description of the image by applying the amount of suppression to the machine learning model to change a likelihood of using the negative word.

Claims (40)

1 . A method comprising:

generating, by a processing device, a text description of an image using a machine learning model;

identifying, by the processing device, a negative word describing an entity in the text description of the image and a category corresponding to the negative word;

determining, by the processing device, a quantitative risk level for the negative word based on use of the negative word in the text description;

determining, by the processing device, a threshold risk level corresponding to the category from multiple different threshold risk levels;

determining, by the processing device, an amount of suppression for use of the negative word based on the text description of the image and the quantitative risk level meeting the threshold risk level; and

generating, by the processing device, an updated text description of the image including replacing the negative word with an alternative word by applying the amount of suppression to the machine learning model.

2 . The method of claim 1 , wherein the machine learning model is biased based on the amount of suppression.

3 . The method of claim 1 , wherein the text description of the image is formatted for audible output.

4 . The method of claim 1 , wherein the text description of the image is used to generate a database of images.

5 . The method of claim 1 , further comprising generating a caption using the text description of the image for display associated with the image.

6 . The method of claim 1 , wherein the negative word is identified from a list of negative words.

7 . The method of claim 6 , wherein the list of negative words is organized by the quantitative risk level.

8 . The method of claim 1 , wherein the negative word is a geographic location and the text description of the image is adjusted to replace the geographic location with a generalized geographic location.

9 . The method of claim 1 , wherein the amount of suppression is based on a confidence of use of the negative word in the text description.

10 . The method of claim 1 , wherein the threshold risk level depends on a type of the negative word.

11 . The method of claim 1 , wherein the machine learning model is biased based on the amount of suppression by multiplying post-softmax likelihoods or pre-softmax logits by a factor inversely proportional to the quantitative risk level for the negative word.

12 . The method of claim 1 , further comprising training the machine learning model to reduce selection of the negative word.

13 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

generating a text description of an image using a machine learning model;

identifying a negative word describing an entity in the text description of the image and a category corresponding to the negative word;

determining a quantitative risk level for use of the negative word based on use of the negative word in the text description;

determining a threshold risk level corresponding to the category from multiple different threshold risk levels;

determining an amount of suppression for use of the negative word based on the text description of the image and the quantitative risk level meeting the threshold risk level; and

generating an updated text description of the image by replacing the negative word with an alternative word by applying the amount of suppression to the machine learning model.

14 . The system of claim 13 , further comprising mapping the quantitative risk level to the amount of suppression.

15 . The system of claim 13 , wherein the text description of the image is formatted for audible output.

16 . The system of claim 13 , wherein the text description of the image is used to generate a database of images.

17 . The system of claim 13 , further comprising generating a caption using the text description of the image for display associated with the image.

18 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

generating a text description of an image using a machine learning model;

identifying a negative word describing an entity in the text description of the image and a category corresponding to the negative word;

determining a quantitative risk level for use of the negative word based on use of the negative word in the text description;

determining a threshold risk level corresponding to the category from multiple different threshold risk levels;

determining an amount of suppression for use of the negative word based on the text description of the image and the quantitative risk level meeting the threshold risk level; and

generating an updated text description of the image including replacing the negative word with an alternative word by applying the amount of suppression to the machine learning model.

19 . The non-transitory computer-readable storage medium of claim 18 , wherein the text description of the image is used to generate a database of images.

20 . The non-transitory computer-readable storage medium of claim 18 , wherein the negative word is identified from a list of negative words that is organized by the quantitative risk level.

Continuity (2)
Continuation 16998730 · Aug 20, 2020
Related Publication 20230315988A1 · Oct 5, 2023
References Cited (43)
US 5796948A · Cohen · 1998 [cited by examiner]
US 6972802B2 · Bray · 2005 [cited by examiner]
US 8965752B2 · Chalmers · 2015 [cited by examiner]
US 9069794B1 · Bandukwala · 2015 [cited by examiner]
US 9575959B2 · Takeuchi · 2017 [cited by examiner]
US 10467339B1 · Shen · 2019 [cited by examiner]
US 10795944B2 · Brown et al. · 2020 [cited by applicant]
US 10803247B2 · Bhatt · 2020 [cited by examiner]
US 11687714B2 · Aggarwal et al. · 2023 [cited by applicant]
US 12182525B2 · Beshara · 2024 [cited by examiner]
US 20040193870A1 · Redlich · 2004 [cited by examiner]
US 20110107205A1 · Chow · 2011 [cited by examiner]
US 20110191097A1 · Spears · 2011 [cited by examiner]
US 20130010062A1 · Redmann · 2013 [cited by applicant]
US 20130339437A1 · De Armas · 2013 [cited by applicant]
US 20130339907A1 · Matas et al. · 2013 [cited by applicant]
US 20150067867A1 · Blake · 2015 [cited by examiner]
US 20160189414A1 · Baker et al. · 2016 [cited by applicant]
US 20170061250A1 · Gao et al. · 2017 [cited by applicant]
US 20170098153A1 · Mao et al. · 2017 [cited by applicant]
US 20170200066A1 · Wang et al. · 2017 [cited by applicant]
US 20170300472A1 · Parikh · 2017 [cited by examiner]
US 20180025292A1 · Senci · 2018 [cited by examiner]
US 20180144747A1 · Skarbovsky · 2018 [cited by examiner]
US 20180329892A1 · Lubbers et al. · 2018 [cited by applicant]
US 20180341637A1 · Gaur · 2018 [cited by examiner]
US 20180373979A1 · Wang · 2018 [cited by examiner]
US 20190116136A1 · Baudart · 2019 [cited by examiner]
US 20190130221A1 · Bose et al. · 2019 [cited by applicant]
US 20190266325A1 · Scherman · 2019 [cited by examiner]
US 20190286931A1 · Kim et al. · 2019 [cited by applicant]
US 20190325023A1 · Awadallah et al. · 2019 [cited by applicant]
US 20190377987A1 · Price · 2019 [cited by examiner]
US 20190379943A1 · Ayala · 2019 [cited by applicant]
US 20200117951A1 · Li et al. · 2020 [cited by applicant]
US 20200320353A1 · Yang et al. · 2020 [cited by applicant]
US 20210064879A1 · Gupta · 2021 [cited by examiner]
US 20210295093A1 · Pan · 2021 [cited by examiner]
US 20220058340A1 · Aggarwal et al. · 2022 [cited by applicant]
US 20220108222A1 · Brannon · 2022 [cited by examiner]
U.S. Appl. No. 16/998,730, filed Nov. 7, 2022 , “Non-Final Office Action”, U.S. Appl. No. 16/998,730, filed Nov. 7, 2022, 20 pages. [cited by applicant]
U.S. Appl. No. 16/998,730, filed Feb. 10, 2023 , “Notice of Allowance”, U.S. Appl. No. 16/998,730, filed Feb. 10, 2023, 14 pages. [cited by applicant]
Zhu, et al., “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books”, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, 9 pages. [cited by applicant]