IP Library Granted Patent US 11,755,641
Granted Patent B2
US 11,755,641 · App. 17/575,209 · Granted Sep 12, 2023

Image searches based on word vectors and image vectors

Inventor: Jenhao Hsiao (Palo Alto, CA)
G06F16/5846G06F16/53G06F16/538G06F16/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,755,641
App. No.
17/575,209
Granted
Sep 12, 2023
Kind
B2
Abstract

A mobile device hosts an artificial intelligence model trained for text-based image searches. Images associated with an image album of the mobile device are indexed by generating, based on the artificial intelligence model, image vectors from the images and word vectors from the image vectors. In response to user input that includes text representing a keyword search, a word vector is generated from the text based on the artificial model. A match is determined between the word vector and one or more of the word vectors to generate a search result that identifies one or more images corresponding to the one or more word vectors. The mobile device displays the search result on a user interface.

Claims (70)

1. A system comprising:

one or more processors; and

one or more memories storing instructions that, upon execution by at least one of the one or more processors, cause the system to perform operations including:

generating an image vector from an image based on an artificial model;

generating a first word vector from the image vector based on the artificial model;

receiving a query associated with an image search;

generating, based on the artificial model, a second word vector from text associated with the query;

determining a match between the first word vector and the second word vector; and

generating, based on the match, a search result that identifies the image;

wherein the operations further include training the artificial model by at least:

generating, based on the artificial model, a third word vector from a label associated with a training image;

generating, based on the artificial model, a second image vector from the training image;

generating, based on the artificial model, a first predicted word vector from the second image vector;

computing a loss of the artificial model based on the third word vector and the first predicted word vector; and

updating a parameter of the artificial model based on the loss.

2. The system of claim 1 , wherein the system is a mobile device, and wherein the operations further include:

storing the artificial model in the one or more memories of the mobile device; and

displaying the search result on a user interface of the mobile device.

3. The system of claim 2 , wherein the operations further include:

storing the image in association with an image album; and

storing word vectors for images associated with the image album,

wherein the match is determined based on Euclidean distances between the first word vector and the word vectors.

4. The system of claim 3 , wherein displaying the search result includes displaying a subset of the images in an order of smallest Euclidean distance to largest Euclidean distance.

5. The system of claim 1 , wherein the training further includes:

generating, based on the artificial model, a second predicted word vector from the second image vector; and

generating a triplet that includes the third word vector, the first predicted word vector, and the second predicted word vector, wherein the loss is computed based on the triplet.

6. The system of claim 5 , wherein computing the loss includes computing a total distance of the triplet based on third word vector, the first predicted word vector, and the second predicted word vector, and wherein the loss is based on the total distance.

7. The system of claim 1 , wherein the artificial model includes a language model, a visual model, and a visual-semantic model.

8. The system of claim 7 , wherein the image vector is an output of the visual model, wherein the first word vector is an output of the visual-semantic model, and wherein the second word vector is an output of the language model.

9. The system of claim 7 , wherein the operations further include training the visual-semantic model based on word vectors that are output of the language model and on image vectors that are output of the visual model.

10. A method implemented by a system, the method including:

generating an image vector from an image based on an artificial model;

generating a first word vector from the image vector based on the artificial model;

receiving a query associated with an image search;

generating, based on the artificial model, a second word vector from text associated with the query;

determining a match between the first word vector and the second word vector; and

generating, based on the match, a search result that identifies the image;

wherein the method further includes training the artificial model by at least:

generating, based on the artificial model, a third word vector from a label associated with a training image;

generating, based on the artificial model, a second image vector from the training image;

generating, based on the artificial model, a first predicted word vector from the second image vector;

computing a loss of the artificial model based on the third word vector and the first predicted word vector; and

updating a parameter of the artificial model based on the loss.

11. The method of claim 10 , further including:

storing the artificial model in a one or more memories of the system; and

displaying the search result on a user interface of the system.

12. The method of claim 11 , further including:

storing the image in association with an image album; and

storing word vectors for images associated with the image album,

wherein the match is determined based on Euclidean distances between the first word vector and the word vectors.

13. The method of claim 12 , wherein displaying the search result includes displaying a subset of the images in an order of smallest Euclidean distance to largest Euclidean distance.

14. The method of claim 10 , wherein the training further includes:

generating, based on the artificial model, a second predicted word vector from the second image vector; and

generating a triplet that includes the third word vector, the first predicted word vector, and the second predicted word vector, wherein the loss is computed based on the triplet.

15. The method of claim 14 , wherein computing the loss includes computing a total distance of the triplet based on third word vector, the first predicted word vector, and the second predicted word vector, and wherein the loss is based on the total distance.

16. A non-transitory computer readable media storing instructions that, upon execution on a system, cause the system to perform operations including:

generating an image vector from an image based on an artificial model;

generating a first word vector from the image vector based on the artificial model;

receiving a query associated with an image search;

generating, based on the artificial model, a second word vector from text associated with the query;

determining a match between the first word vector and the second word vector; and

generating, based on the match, a search result that identifies the image;

wherein the operations further include training the artificial model by at least:

generating, based on the artificial model, a third word vector from a label associated with a training image;

generating, based on the artificial model, a second image vector from the training image;

generating, based on the artificial model, a first predicted word vector from the second image vector;

computing a loss of the artificial model based on the third word vector and the first predicted word vector; and

updating a parameter of the artificial model based on the loss.

17. The non-transitory computer readable media of claim 16 , wherein the operations further include storing the artificial model as including a language model, a visual model, and a visual-semantic model, wherein the image vector is an output of the visual model, wherein the first word vector is an output of the visual-semantic model, and wherein the second word vector is an output of the language model.

18. The non-transitory computer readable media of claim 17 , wherein the operations further include training the visual-semantic model based on word vectors that are output of the language model and on image vectors that are output of the visual model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2022
From: HSIAO, JENHAO
To: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
Reel/Frame 058734/0871 →
Continuity (3)
Continuation PCTCN2020091055 · May 19, 2020
Provisional Application 62895309 · Sep 3, 2019
Related Publication 20220138252A1 · May 5, 2022
Cited By (3)
US 12,541,544 US 12,591,559 US 12,682,179