IP Library › Granted Patent US 12,169,520
Granted Patent B2
US 12,169,520 · App. 18/103,476 · Granted Dec 17, 2024

Machine learning selection of images

Inventor: Jessica Zhang (Los Angeles, CA)
Assignee: Intuit Inc.
G06F16/583G06F16/951G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,169,520
App. No.
18/103,476
Filed
Jan 30, 2023
Granted
Dec 17, 2024
Kind
B2
Art Unit
2165
USPC
707/706
Abstract

A method including receiving an input and embedding the input into a first data structure that defines first relationships among images and texts. The method also includes comparing the first data structure to an index including a second data structure that defines second relationships among pre-determined texts and pre-determined images. The method also includes returning a subset of images from the pre-determined images. The subset includes those images in the pre-determined images for which matches exist between the first relationships and the second relationships.

Claims (61)

1. A method comprising:

receiving an input comprising a plurality of texts from a source of text and a plurality of images from a source of images, wherein the plurality of texts is separate from the plurality of images;

embedding the input into a first data structure that defines first relationships among the plurality of images from the source of images and the plurality of texts from the source of text;

comparing the first data structure to an index comprising a second data structure that defines second relationships among a plurality of pre-determined texts and a plurality of pre-determined images, wherein:

the plurality of pre-determined texts have known relationships to the plurality of pre-determined images,

each pre-determined image in the plurality of pre-determined images is related to one or more instances of the plurality of pre-determined texts, and

the plurality of pre-determined texts is separate from the plurality of texts; and

returning a subset of images from the plurality of pre-determined images, wherein the subset comprises those images in the plurality of pre-determined images for which matches exist between the first relationships and the second relationships.

2. The method of claim 1 , further comprising:

receiving a user selection of a selected image from the subset of images; and

inserting the selected image into an electronic document.

3. The method of claim 2 , wherein the electronic document is selected from the group consisting of: an email document, a word processing document, a presentation building document, a portable document format (PDF) document, and an image manipulation document.

4. The method of claim 1 , further comprising, prior to receiving the input:

scraping a website to generate a text associated with a scraped image; and

generating the input from the text and the scraped image.

5. The method of claim 1 , further comprising, prior to receiving the input:

inputting an email to a natural language processing machine learning model;

outputting, from the natural language processing machine learning model, at least one keyword; and

generating the input from the at least one keyword.

6. The method of claim 1 , wherein the input comprises at least one of a text and an image.

7. The method of claim 1 , wherein comparing comprises performing a nearest neighbor comparison between the first data structure and the second data structure.

8. The method of claim 1 , wherein returning the subset of images comprises:

transmitting the subset of images to a remote user device; and

displaying the subset of images on a display device of the remote user device.

9. The method of claim 1 , further comprising:

ranking the subset of images according to a ranking criterion.

10. A system comprising:

a processor;

a data repository, in communication with the processor, and storing:

an input comprising a plurality of texts from a source of text and a plurality of images from a source of images, wherein the plurality of texts is separate from the plurality of images;

a first data structure that embeds the input, wherein the first data structure defines first relationships among the plurality of images from the source of images and the plurality of texts from the source of text;

an index comprising a second data structure that defines second relationships among a plurality of pre-determined texts and a plurality of pre-determined images, wherein:

the plurality of pre-determined texts have known relationships to the plurality of pre-determined images,

each pre-determined image in the plurality of pre-determined images is related to one or more instances of the plurality of pre-determined texts, and

the plurality of pre-determined texts is separate from the plurality of texts, and

a subset of images, wherein the subset comprises those images in the plurality of pre-determined images for which matches exist between the first relationships and the second relationships; and

a server controller which, when executed by the processor:

embed the input into the first data structure;

compares the first data structure to the second data structure; and

returns the subset of images from the plurality of pre-determined images.

11. The system of claim 10 , further comprising:

a nearest neighbor machine learning model which, when executed by the processor, takes as input the plurality of pre-determined texts and the plurality of pre-determined images, and which generates the index as output.

12. The system of claim 10 , further comprising:

a natural language machine learning model which, when executed by the processor, takes the input and generates the plurality of texts as output, wherein the plurality of texts comprises keywords that represent at least a portion of the input, and wherein the plurality of texts is represented as a vector.

13. The system of claim 10 , wherein the server controller is further configured to:

generate the input by scraping a third party website;

receive a user selection of a selected image from the subset of images; and

insert the selected image into an electronic document.

14. The system of claim 13 , wherein the electronic document comprises an email, and wherein the server controller is further configured to:

format the selected image together with pre-determined email text into a plurality of different formatted emails;

receive, from a user, a selected email format from the plurality of different formatted emails; and

return the selected email format.

15. The system of claim 10 , wherein the server controller further comprises a Contrastive Language-Image Pre-training (CLIP) machine learning model, and wherein the server controller embeds the input into the first data structure by being programmed to:

receive, as input to the CLIP machine learning model, the plurality of images and the plurality of texts, and to generate, as output, the first data structure.

16. The system of claim 15 , further comprising:

a training controller which, when executed by the processor, is configured to:

receive a plurality of emails containing a plurality of raw images and a plurality of associated raw text associated with the plurality of raw images;

extract a training portion of the plurality of raw images and the plurality of associated raw text, wherein, after extracting, a remaining portion of raw images and associated raw text remains;

generate, for the training portion, labeled data by labeling the plurality of raw images as being associated with the associated raw text;

embed the labeled data into a known vector; and

train the CLIP machine learning model using the remaining portion as input and the known vector as a known result.

Continuity (1)
Related Publication 20240256597A1 · Aug 1, 2024