IP Library › Granted Patent US 12,249,116
Granted Patent B2
US 12,249,116 · App. 17/656,147 · Granted Mar 11, 2025

Concept disambiguation using multimodal embeddings

Inventors: Venkata Naveen Kumar Yadav Marri (Fremont, CA); Ajinkya Gorakhnath Kale (San Jose, CA)
Assignee: ADOBE INC.
G06V10/761G06N3/088G06V10/771G06V10/7715G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,249,116
App. No.
17/656,147
Granted
Mar 11, 2025
Kind
B2
Abstract

Systems and methods for image processing are described. Embodiments of the present disclosure identify a plurality of candidate concepts in a knowledge graph (KG) that correspond to an image tag of an image; generate an image embedding of the image using a multi-modal encoder; generate a concept embedding for each of the plurality of candidate concepts using the multi-modal encoder; select a matching concept from the plurality of candidate concepts based on the image embedding and the concept embedding; and generate association data between the image and the matching concept.

Claims (65)

1. A method for image processing, comprising:

identifying a plurality of candidate concepts in a knowledge graph (KG) that correspond to an image tag of an image, wherein the knowledge graph comprises a plurality of nodes corresponding to the plurality of candidate concepts;

generating an image embedding of the image using a multi-modal encoder;

generating a text embedding for each of the plurality of candidate concepts using the multi-modal encoder used to generate the image embedding, wherein the image embedding and the text embedding are located in a same embedding space;

selecting a matching concept from the plurality of candidate concepts based on the image embedding and the text embedding;

generating association data between the image and the matching concept; and

transmitting information from the knowledge graph corresponding to the image based on the association data between the image and the matching concept.

2. The method of claim 1 , further comprising:

extracting a plurality of image tags based on the image using an image tagger neural network, wherein the plurality of image tags comprises the image tag.

3. The method of claim 2 , further comprising:

identifying a plurality of similar images based on the image embedding;

identifying a plurality of additional image tags associated with the plurality of similar images; and

computing a tag similarity score for each of the plurality of additional image tags, wherein the image tag is selected based on the tag similarity score.

4. The method of claim 1 , further comprising:

identifying a description from the KG for each of the plurality of candidate concepts; and

applying the multi-modal encoder to the description, wherein the text embedding is based on the description.

5. The method of claim 1 , further comprising:

identifying a name for each of the plurality of candidate concepts in the KG; and

determining that the name corresponds to the image tag, wherein the plurality of candidate concepts are identified based on the determination.

6. The method of claim 1 , further comprising:

computing a similarity score between the image embedding and the text embedding; and

comparing the similarity score of each of the plurality of candidate concepts, wherein the matching concept is selected based on the comparison.

7. The method of claim 1 , further comprising:

identifying a plurality of images;

associating each of the plurality of images with a concept in the KG; and

augmenting the KG with the plurality of images based on the association.

8. The method of claim 1 , further comprising:

receiving a query including the image;

identifying a description of the matching concept; and

transmitting the description in response to the query.

9. The method of claim 1 , further comprising:

receiving a plurality of training images and a plurality of captions corresponding to each of the plurality of training images as input to the multi-modal encoder; and

training the multi-modal encoder based on the input using contrastive self-supervised learning.

10. The method of claim 9 , further comprising:

maximizing a similarity between an anchor image of the plurality of training images and a corresponding caption of the anchor image;

minimizing a similarity between the anchor image of the plurality of training images and the plurality of captions excluding the corresponding caption; and

updating parameters of the multi-modal encoder based on the maximization and the minimization.

11. A method for image processing, comprising:

identifying a plurality of candidate concepts in a knowledge graph (KG) that correspond to an image tag of an image, wherein the knowledge graph comprises a plurality of nodes corresponding to the plurality of candidate concepts;

generating an image embedding of the image using a multi-modal encoder;

generating a text embedding for each of the plurality of candidate concepts using the multi-modal encoder used to generate the image embedding, wherein the image embedding and the text embedding are located in a same embedding space;

computing a similarity score between the image embedding and the text embedding;

comparing the similarity score for each of the plurality of candidate concepts; and

selecting a matching concept from the plurality of candidate concepts based on the comparison.

12. The method of claim 11 , further comprising:

extracting a plurality of image tags based on the image using an image tagger, wherein the plurality of image tags comprises the image tag.

13. The method of claim 11 , further comprising:

identifying a description for each of the plurality of candidate concepts, wherein the text embedding is based on the description.

14. The method of claim 11 , further comprising:

identifying a name for each of the plurality of candidate concepts in the KG; and

determining that the name corresponds to the image tag.

15. An apparatus for image processing, comprising:

a knowledge graph (KG) component configured to identify a plurality of candidate concepts in a knowledge graph that correspond to an image tag of an image, wherein the knowledge graph comprises a plurality of nodes corresponding to the plurality of candidate concepts;

a multi-modal encoder configured to generate an image embedding of the image, and to generate a text embedding for each of the plurality of candidate concepts, wherein the image embedding and the text embedding are located in a same embedding space; and

a matching component configured to select a matching concept from the plurality of candidate concepts based on the image embedding and the text embedding.

16. The apparatus of claim 15 , further comprising:

an image tagger configured to extract a plurality of image tags based on the image, wherein the plurality of image tags comprises the image tag.

17. The apparatus of claim 15 , further comprising:

an association manager configured to generate association data between the image and the matching concept.

18. The apparatus of claim 15 , wherein:

the matching component computes a similarity score between the image embedding and the text embedding, and compares the similarity score for each of the plurality of candidate concepts, wherein the matching concept is selected based on the comparison.

19. The apparatus of claim 18 , wherein:

the similarity score comprises a cosine similarity score or a L2 norm.

20. The apparatus of claim 15 , wherein:

the multi-modal encoder comprises a text-to-visual embedding model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: MARRI, VENKATA NAVEEN KUMAR YADAV; KALE, AJINKYA GORAKHNATH
To: ADOBE INC.
Reel/Frame 059380/0135 →
Continuity (1)
Related Publication 20230326178A1 · Oct 12, 2023
References Cited (9)
US 9830526B1 · Lin et al. · 2017 [cited by applicant]
US 20160259862A1 · Navanageri · 2016 [cited by examiner]
US 20180204111A1 · Zadeh · 2018 [cited by examiner]
US 20180232443A1 · Delgo · 2018 [cited by examiner]
US 20190327331A1 · Natarajan · 2019 [cited by examiner]
US 20200379787A1 · Martin · 2020 [cited by examiner]
US 20210067684A1 · Kim · 2021 [cited by examiner]
US 20210365727A1 · Aggarwal et al. · 2021 [cited by applicant]
Aggarwal, et al., Towards Zero-shot Cross-lingual Image retrieval. arXiv preprint arXiv:2012.05107v1 [cs.CL] Nov. 24, 2020, 7 pages. [cited by applicant]