IP Library Granted Patent US 11,620,331
Granted Patent B2
US 11,620,331 · App. 17/191,449 · Granted Apr 4, 2023

Textual and image based search

Inventors: Dmitry Olegovich Kislyuk (San Francisco, CA); Jeffrey Harris (Oakland, CA); Anton Herasymenko (San Francisco, CA); Eric Kim (San Francisco, CA); Yiming Jen (San Jose, CA)
Assignee: Pinterest, Inc.
G06F16/5838G06F16/248G06F16/24578G06F16/56G06F16/583G06F16/5866G06F16/738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,620,331
App. No.
17/191,449
Granted
Apr 4, 2023
Kind
B2
Abstract

Described is a system and method for enabling visual search for information. With each selection of an object included in an image, additional images that include visually similar objects are determined and presented to the user.

Claims (91)

1. A computing system, comprising:

one or more processors; and

a memory storing program instructions that when executed by the one or more processors, cause the one or more processors to at least:

receive a text query;

provide a visual refinement option;

determine and return a plurality of results corresponding to the text query;

receive an object as part of the visual refinement option;

generate an object label for the object;

generate an object feature vector representative of the object;

determine, based at least in part on the object label for the object, a first plurality of stored feature vectors from a plurality of stored feature vectors corresponding to segments of visual items included in the plurality of results;

compare the object feature vector with the first plurality of stored feature vectors to determine a plurality of respective similarity scores, each of the respective similarity scores representative of a similarity between the object feature vector and a respective one of the plurality of stored feature vectors;

generate a ranked list of a plurality of results based at least in part on the plurality of respective similarity scores; and

present at least some results of the ranked list in response to the receipt of the object.

2. The computing system of claim 1 , wherein execution of the programming instructions further cause the one or more processors to, at least:

identify a plurality of items represented in the object;

determine an item of interest of the plurality of items; and

generate a feature vector for the item of interest as the object feature vector representative of the object.

3. The computing system of claim 2 , wherein execution of the programming instructions further causes the one or more processors to, at least:

determine a defined category corresponding to the text query;

determine a label for each item of the plurality of items represented in the object; and

identify a first item of the plurality of items whose label is associated with the defined category, the first item being the item of interest.

4. The computing system of claim 3 , wherein execution of the programming instructions further causes the one or more processors to, at least:

for items whose label is associated with the defined category:

present the labels as user-selectable controls in proximity to the corresponding items represented in the object;

receive a selection of a first label; and

identify the item corresponding to the first label as the first item.

5. The computing system of claim 2 , wherein execution of the programming instructions further causes the one or more processors to, at least:

identify a plurality of segments of the object, each segment comprising less than the entire object; and

identify the plurality of items represented in the object from items represented in the plurality of segments.

6. The computing system of claim 5 , wherein the object is one of a video content or an image content, and wherein the plurality of items represented in the object are identified from items represented in the plurality of segments of the object according to an edge detection algorithm.

7. The computing system of claim 5 , further comprising:

for each segment of the plurality of segments of the object:

identifying a background of a segment; and

removing the background of the segment; and

processing the plurality of segments with the background removed to identify the plurality of items represented in the object.

8. A computer-implemented method, comprising:

under the control of one or more processors executing instructions on a computer system, and in response to a presentation of a first image on a display of the computer system:

determining a first object of a plurality of objects represented in the first image;

generating a first label for the first object;

generating a first feature vector representative of the first object;

determining, based at least in part on the first label for the first object, a first plurality of stored feature vectors from a plurality of stored feature vectors, each stored feature vector of the plurality of stored feature vectors corresponding to an image maintained in a data store;

comparing the first feature vector with the first plurality of stored feature vectors to generate a plurality of similarity scores between the first feature vector and the first plurality of stored feature vectors;

selecting a result set of feature vectors of the first plurality of stored feature vectors having highest similarity scores from the plurality of similarity scores;

identifying a result set of images corresponding to the feature vectors of the results set of feature vectors; and

presenting at least some images of the result set of images on the display of the computer system.

9. The computer-implemented method of claim 8 , further comprising:

segmenting the first image into a plurality of segments; and

identifying the plurality of objects from the plurality of segments of the first image.

10. The computer-implemented method of claim 9 , further comprising:

for each segment of the plurality of segments:

identifying a background of a segment; and

removing the background of the segment; and

identifying the plurality of objects from the plurality of segments having the background removed.

11. The computer-implemented method of claim 8 , wherein the first object is determined from the plurality of objects represented in the first image according to a foreground and centered position of the first object in the first image.

12. The computer-implemented method of claim 8 , wherein determining the first object of the plurality of objects comprises:

presenting user-selectable selectors on the display corresponding to at least some objects of the plurality of objects;

receiving an indication of a selection of a first user-selectable selector corresponding to a selected object of the at least some objects; and

identifying the selected object as the first object.

13. The computer-implemented method of claim 12 , further comprising:

determining a label for each object of the at least some objects; and

wherein each user-selectable selector includes a presentation of the label of the corresponding object.

14. A computer-implemented method, comprising:

under a control of one or more processors executing instructions on a computer system:

receiving a text query for items of content maintained in a data store by the computer system;

receiving a visual content as a refinement to the text query;

processing the visual content to detect a plurality of objects represented within the visual content;

determining a text label for each object of the plurality of objects;

presenting selectors corresponding to the text labels of at least some objects of the plurality of objects, wherein presenting the selectors includes presenting each selector in proximity to a corresponding object of the at least some objects and concurrently with the visual content on a display of the computer system, and wherein each selector is user-selectable;

receiving a selection of a first selector corresponding to a first text label of a first object of the at least some objects;

determining, based at least in part on the first text label corresponding to the selected first selector, a first plurality of items of content from the items of content;

identifying a set of items of content from the first plurality of items of content according to the text query as refined by the first text label; and

presenting the set of items of content on the display of the computer system.

15. The computer-implemented method of claim 14 , further comprising:

determining a defined category corresponding to the text query; and

identifying a subset of objects of the plurality of objects, wherein the text label of each object of the subset of objects is associated with the defined category; and

wherein the at least some objects of the plurality of objects are included in the subset of objects.

16. The computer-implemented method of claim 15 , further comprising:

generating an embedding vector for the first object;

comparing the embedding vector for the first object with embedding vectors associated with the objects of the subset of objects to generate a similarity score for each object of the subset of objects, each similarity score representative of a similarity between a corresponding object of the subset of objects and the first object; and

identifying the at least some objects of the subset of objects as objects having highest similarity scores.

17. The computer-implemented method of claim 14 , further comprising and concurrent with the presentation of the at least some items of content on the display of the computer:

displaying the text query in an input field on the display, and further displaying an image representative of the first object in the input field with the text query.

18. The computer-implemented method of claim 14 , further comprising:

processing the visual content into a plurality of segments; and

processing the plurality of segments to identify the plurality of objects represented within the visual content.

19. The computer-implemented method of claim 18 , further comprising:

for each segment of the plurality of segments:

identifying a background of a segment; and

removing the background of the segment; and

processing the plurality of segments with the background removed to identify the plurality of objects represented within the visual content.

20. The computer-implemented method of claim 18 , wherein the plurality of segments are processed according to an edge detection algorithm to identify the plurality of objects represented within the visual content.

Assignments (2)
SECURITY INTEREST Recorded Oct 25, 2022
From: PINTEREST, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 061767/0853 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2021
From: KISLYUK, DMITRY OLEGOVICH; HARRIS, JEFFREY; HERASYMENKO, ANTON; KIM, ERIC; JEN, YIMING
To: PINTEREST, INC.
Reel/Frame 055484/0241 →
Continuity (2)
Continuation 15713567 · Sep 22, 2017
Related Publication 20210256054A1 · Aug 19, 2021
Cited By (3)
US 12,242,491 US 12,314,347 US 12,326,867