IP Library › Granted Patent US 12,682,612
Granted Patent B2
US 12,682,612 · App. 18/313,642 · Granted Jul 14, 2026

Weak supervised training data for image tagging models

Inventors: Venkata Naveen Kumar Yadav Marri (Newark, CA); Ajinkya Gorakhnath Kale (San Jose, CA)
Assignee: ADOBE INC.
G06V10/774G06T9/00G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,612
App. No.
18/313,642
Filed
May 8, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2663
USPC
382/159
Abstract

Systems and methods for image tagging are provided. One aspect of the systems and methods includes encoding an image and a tag of the image using a multimodal encoder to obtain an image embedding and a text embedding, respectively. Another aspect of the systems and methods includes generating training data for a machine learning model by filtering a plurality of image-tag pairs based on a similarity between the image embedding and the text embedding. Another aspect of the systems and methods includes training the machine learning model using the training data.

Claims (48)

1 . A computer-implemented method of training a neural network for image processing, the method comprising:

obtaining a plurality of image-tag pairs including an image-tag pair comprising an image and a tag associated with the image;

encoding the image and the tag associated with the image using a multimodal encoder to obtain an image embedding and a text embedding, respectively, in a common embedding space;

generating training data for a machine learning model comprising an artificial neural network (ANN) by computing a similarity score between the image embedding and the text embedding, comparing the similarity score to a threshold, and filtering the plurality of image-tag pairs based on the comparison, wherein the similarity score represents a semantic similarity between the image and the tag of the image-tag pair based on a similarity between the image embedding and the text embedding in the common embedding space; and

training the machine learning model to perform image processing by updating parameters of the ANN using the training data.

2 . The method of claim 1 , further comprising:

obtaining user-interaction data for the image and the tag, wherein the filtering is further based on the user-interaction data.

3 . The method of claim 2 , further comprising:

computing, based on the user-interaction data, a behavioral score indicating user interactions with the image in response to searches corresponding to the tag, wherein the filtering is based on the behavioral score.

4 . The method of claim 1 , further comprising:

receiving the image and the tag of the image from a user upload.

5 . The method of claim 1 , further comprising:

identifying a set of frequently-used tags; and

selecting the plurality of image-tag pairs based on the set of frequently-used tags.

6 . The method of claim 1 , further comprising:

computing an offensiveness value for the image or the tag; and

selecting the plurality of image-tag pairs based on the offensiveness value.

7 . The method of claim 1 , further comprising:

computing an editorial value for the image or the tag; and

selecting the plurality of image-tag pairs based on the editorial value.

8 . The method of claim 1 , further comprising:

obtaining an additional image; and

generating an additional tag for the additional image using the machine learning model.

9 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device configured to perform image processing using a machine learning model, wherein the machine learning model is trained by:

obtaining a plurality of image-tag pairs including an image-tag pair comprising an image and a tag associated with the image;

encoding the image and the tag associated with the image using a multimodal encoder to obtain an image embedding and a text embedding, respectively, in a common embedding space;

generating training data for a machine learning model comprising an artificial neural network (ANN) by computing a similarity score between the image embedding and the text embedding, comparing the similarity score to a threshold, and filtering the plurality of image-tag pairs based on the comparison, wherein the similarity score represents a semantic similarity between the image and the tag of the image-tag pair based on a similarity between the image embedding and the text embedding in the common embedding space; and

training the machine learning model to perform image processing by updating parameters of the ANN using the training data.

10 . The system of claim 9 , further comprising:

obtaining user-interaction data for the image and the tag, wherein the filtering is further based on the user-interaction data.

11 . The system of claim 10 , further comprising:

computing, based on the user-interaction data, a behavioral score indicating user interactions with the image in response to searches corresponding to the tag, wherein the filtering is based on the behavioral score.

12 . The system of claim 9 , further comprising:

receiving the image and the tag of the image from a user upload.

13 . The system of claim 9 , further comprising:

identifying a set of frequently-used tags; and

selecting the plurality of image-tag pairs based on the set of frequently-used tags.

14 . The system of claim 9 , further comprising:

computing an offensiveness value for the image or the tag; and

selecting the plurality of image-tag pairs based on the offensiveness value.

15 . The system of claim 9 , further comprising:

computing an editorial value for the image or the tag; and

selecting the plurality of image-tag pairs based on the editorial value.

16 . The system of claim 9 , further comprising:

obtaining an additional image; and

generating an additional tag for the additional image using the machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2023
From: MARRI, VENKATA NAVEEN KUMAR YADAV; KALE, AJINKYA GORAKHNATH
To: ADOBE INC.
Reel/Frame 063565/0987 →
Continuity (1)
Related Publication 20240378863A1 · Nov 14, 2024
References Cited (7)
US 10803231B1 · Wang · 2020 [cited by examiner]
US 20200380403A1 · Aggarwal · 2020 [cited by examiner]
US 20210271707A1 · Lin · 2021 [cited by examiner]
US 20220121702A1 · Kale · 2022 [cited by examiner]
US 20230154213A1 · Gao · 2023 [cited by examiner]
Interactive Weak Supervision: Learning Useful Heuristics for Data Labeling, by Boecking et al., arXiv: 2012.06046v2 [cs. LG] Jan. 25, 2021 (Year: 2021). [cited by examiner]
1Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv preprint arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021, pp. 1-48. [cited by applicant]