IP Library Granted Patent US 11,687,714
Granted Patent B2
US 11,687,714 · App. 16/998,730 · Granted Jun 27, 2023

Systems and methods for generating text descriptive of digital images

Inventors: Pranav Aggarwal (San Jose, CA); Di Pu (Bothell, WA); Daniel ReMine (Seattle, WA); Ajinkya Kale (San Jose, CA)
Assignee: Adobe Inc.
G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,687,714
App. No.
16/998,730
Granted
Jun 27, 2023
Kind
B2
Abstract

Disclosed are computer-implemented methods and systems for generating text descriptive of digital images, comprising using a machine learning model to pre-process an image to generate initial text descriptive of the image; adjusting one or more inferences of the machine learning model, the inferences biasing the machine learning model away from associating negative words with the image; using the machine learning model comprising the adjusted inferences to post-process the image to generate updated text descriptive of the image; and processing the generated updated text descriptive of the image outputted by the machine learning model to fine-tune the updated text descriptive of the image.

Claims (58)

1. A method comprising:

pre-processing an image, using a machine learning model, to generate an initial text description of the image;

adjusting inferences of the machine learning model to bias the machine learning model from associating negative words with the image by:

identifying the negative words and a risk level for each of the negative words in the initial text description of the image; and

mapping the risk level to a suppression level to reduce a likelihood of the machine learning model selecting the negative words;

post-processing the image using the machine learning model and the adjusted inferences to generate an updated text description of the image; and

adjusting the initial text description of the image based on the updated text description of the image.

2. The method of claim 1 , further comprising:

adjusting the inferences of the machine learning model during a beam search of the machine learning model to adjust a posterior probability of selected words of the initial text description of the image; and

processing the updated text description of the image using pure natural language processing or text-based rules overlaid on part-of-speech (POS) libraries.

3. The method of claim 1 , wherein mapping the risk level further comprises multiplying post-softmax likelihoods by a factor inversely proportional to the risk level for each of the negative words in the initial text description of the image.

4. The method of claim 1 , wherein adjusting the initial text description of the image comprises applying one or more of:

a gender mitigation mechanism;

an offensive adjective mitigation mechanism;

a low confidence adjective mitigation mechanism;

a geo-location generalizing mechanism; and

an image with text templatizing mechanism.

5. The method of claim 4 , wherein applying the gender mitigation mechanism further comprises, for each triplet of a gender-triplet-reference collection biasing selection of a gender neutral term over a male version of a word or a female version of a word.

6. The method of claim 1 , wherein the initial text description of the image is a title of the image.

7. The method of claim 1 , wherein the machine learning model includes a sentence encoder neural network that encodes the initial text description of the image as a vector.

8. The method of claim 7 , wherein post-processing the image further includes decoding the vector using a sentence decoder neural network to generate the updated text description of the image.

9. The method of claim 4 , wherein the offensive adjective mitigation mechanism detects a noun in the initial text description of the image and infers whether an adjective proceeding the noun is present in a list of offensive adjectives.

10. A method comprising:

generating, using a first machine learning model, an initial text description of an image to pre-process the image;

adjusting one or more inferences of the first machine learning model, the inferences biasing the first machine learning model from associating negative words with the image;

generating an updated text description of the image to post-process the image using a second machine learning model and the adjusted inferences; and

processing the updated text description of the image to adjust the updated text description of the image by:

identifying negative words and a risk level for each of the negative words; and

mapping the risk level to a suppression level to reduce a likelihood of the machine learning model selecting the negative words.

11. The method of claim 10 , further comprising:

adjusting the one or more inferences of the first machine learning model during a beam search of the first machine learning model to adjust a posterior probability of selected words of the text; and

processing the updated text description of the image using pure natural language processing.

12. The method of claim 10 , further comprising multiplying post-softmax likelihoods by a factor inversely proportional to the risk level for each of the negative words.

13. The method of claim 10 , wherein processing the updated text description of the image comprises applying one or more of:

a gender mitigation mechanism;

an offensive adjective mitigation mechanism;

a low confidence adjective mitigation mechanism;

a geo-location generalizing mechanism; and

an image with text templatizing mechanism.

14. The method of claim 10 , wherein applying the gender mitigation mechanism further comprises, for each triplet of a gender-triplet-reference collection biasing selection of a gender neutral term over a male version of a word or a female version of a word.

15. The method of claim 10 , wherein the initial text description of the image is a caption of the image.

16. A system comprising:

a first machine learning model for pre-processing an image to generate an initial text description of the image;

an inference adjustment module for adjusting inferences of the first machine learning model, the inferences biasing the first machine learning model from associating negative words with the image, including:

a second machine learning model for identifying negative words and a risk level for each of the negative words and configured for mapping the risk level to a suppression level to reduce a likelihood of the first machine learning model selecting the negative words;

a third machine learning model for post-processing the image using the adjusted inferences to generate an updated text description of the image; and

a post-processing module for adjusting the initial text description of the image based on the updated text description of the image.

17. The system of claim 16 , wherein:

the inference adjustment module adjusts the inferences of the first machine learning model during a beam search to adjust a posterior probability of selected words of the text; and

the third machine learning model post-processes the updated text description of the image using pure natural language processing or text-based rules overlaid on part-of-speech (POS) libraries.

18. The system of claim 16 , wherein the second machine learning model is further configured to multiply post-softmax likelihoods by a factor inversely proportional to the risk level for each of the negative words.

19. The system of claim 16 , wherein the post-processing module is further configured to apply one or more of:

a gender mitigation mechanism;

an offensive adjective mitigation mechanism;

a low confidence adjective mitigation mechanism;

a geo-location generalizing mechanism; and

an image with text templatizing mechanism.

20. The system of claim 16 , wherein the initial text description of the image is a title of the image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2020
From: AGGARWAL, PRANAV; PU, DI; REMINE, DANIEL; KALE, AJINKYA
To: ADOBE INC.
Reel/Frame 053658/0989 →
Continuity (1)
Related Publication 20220058340A1 · Feb 24, 2022
Cited By (2)
US 12,206,962 US 12,645,877