IP Library › Granted Patent US 11,804,035
Granted Patent B2
US 11,804,035 · App. 17/512,389 · Granted Oct 31, 2023

Intelligent online personal assistant with offline visual search database

Inventors: Ajinkya Gorakhnath Kale (San Jose, CA); Fan Yang (San Jose, CA); Qiaosong Wang (San Francisco, CA); Mohammadhadi Kiapour (San Francisco, CA); Robinson Piramuthu (Oakland, CA)
Assignee: eBay Inc.
G06V10/82G06F16/248G06F16/24578G06F16/3344G06F16/51G06F16/5838G06F18/2411G06Q30/0643G06V10/758G06N3/044G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,804,035
App. No.
17/512,389
Granted
Oct 31, 2023
Kind
B2
Abstract

Systems, methods, and computer program products for identifying a candidate product in an electronic marketplace based on a visual comparison between candidate product image visual content and input query image visual content. Embodiments generate and store descriptive image signatures from candidate product images or selected portions of such images. A subsequently calculated visual similarity measure serves as a visual search result score for the candidate product in comparison to an input query image. Any number of images of any number of candidate products may be analyzed, such as for items available for sale in an online marketplace. Image analysis results are stored in a database and made available for subsequent automated on-demand visual comparisons to an input query image. The embodiments enable substantially real time visual based product searching of a potentially vast catalog of items.

Claims (71)

1. A method, comprising:

receiving a natural language utterance associated with an input query image;

generating, in real-time using a neural network, an image signature for the input query image, the image signature numerically describing image content;

calculating, in real-time, a visual similarity measure between a plurality of candidate product images and the input query image based on a plurality of image signatures corresponding to the plurality of candidate product images, the visual similarity measure indicating a correlation between a candidate product image and the input query image;

determining a list of candidate products based on the visual similarity measure and knowledge graph data, the knowledge graph data used to confirm the correlation of a candidate product image of the list of candidate products generated based on the visual similarity measure;

analyzing the natural language utterance to determine an intent;

determining whether the list of candidate products satisfies the intent; and

when the list of candidate products does not satisfy the intent:

identifying a second list of candidate products based upon the intent; and

generating a response to the natural language utterance, the response comprising a natural language prompt and one or more products from the second list of candidate products.

2. The method of claim 1 , wherein generating the image signature comprises:

generating metadata associated with the input query image; and

retrieving metadata associated with the plurality of candidate product images, wherein calculating the visual similarity measure is based at least in part on the generated metadata associated with the input query image and the retrieved metadata associated with the plurality of candidate product images.

3. The method of claim 1 , wherein generating the image signature comprises:

analyzing the input query image, based at least in part on object localization, object recognition, optical character recognition, inventory matching, or any combination thereof, to generate the image signature.

4. The method of claim 1 , further comprising:

collecting user data associated with a user that provided the natural language utterance, the user data comprising identity data, behavioral data, or both; and

extrapolating the collected user data to generate extrapolated user data, wherein the intent is determined based at least in part on the user data, the extrapolated user data, or both.

5. The method of claim 1 , wherein analyzing the natural language utterance comprises:

decoding the natural language utterance with a speech to text decoder that comprises a feature extraction component, an acoustical model component, a language model component, or any combination thereof, to generate text, wherein the intent is determined based at least in part on the text.

6. The method of claim 1 , wherein analyzing the natural language utterance comprises:

retrieving one or more language samples associated with a target domain; and

analyzing the natural language utterance to determine the intent based at least in part on the retrieved one or more language samples.

7. The method of claim 1 , wherein analyzing the natural language utterance to determine the intent is based at least in part on one or more statistical models of speech units.

8. The method of claim 7 , wherein the one or more statistical models of speech units comprise a Gaussian mixture model or a deep neural network.

9. An apparatus, comprising:

a processor;

memory coupled with the processor; and

instructions stored in the memory and executable by the processor to cause the apparatus to perform operations comprising:

receiving a natural language utterance associated with an input query image;

generating, in real-time using a neural network, an image signature for the input query image, the image signature numerically describing image content;

calculating, in real-time, a visual similarity measure between a plurality of candidate product images and the input query image based on a plurality of image signatures corresponding to the plurality of candidate product images, the visual similarity measure indicating a correlation between a candidate product image and the input query image;

determining a list of candidate products based on the visual similarity measure and knowledge graph data, the knowledge graph data used to confirm the correlation of a candidate product image of the list of candidate products generated based on the visual similarity measure;

analyzing the natural language utterance to determine an intent;

determining whether the list of candidate products satisfies the intent; and

when the list of candidate products do not satisfy the intent:

identifying a second list of candidate products based upon the intent; and

generating a response to the natural language utterance, the response comprising a natural language prompt and one or more products from the second list of candidate products.

10. The apparatus of claim 9 , the operations further comprising:

generating metadata associated with the input query image; and

retrieving metadata associated with the plurality of candidate product images, wherein calculating the visual similarity measure is based at least in part on the generated metadata associated with the input query image and the retrieved metadata associated with the plurality of candidate product images.

11. The apparatus of claim 9 , the operations further comprising:

analyzing the input query image, based at least in part on object localization, object recognition, optical character recognition, inventory matching, or any combination thereof, to generate the image signature.

12. The apparatus of claim 9 , the operations further comprising:

collecting user data associated with a user that provided the natural language utterance, the user data comprising identity data, behavioral data, or both; and

extrapolating the collected user data to generate extrapolated user data, wherein the intent is determined based at least in part on the user data, the extrapolated user data, or both.

13. The apparatus of claim 9 , wherein analyzing the natural language utterance further comprises:

decoding the natural language utterance with a speech to text decoder that comprises a feature extraction component, an acoustical model component, a language model component, or any combination thereof, to generate text, wherein the intent is determined based at least in part on the text.

14. The apparatus of claim 9 , wherein analyzing the natural language utterance further comprises:

retrieving one or more language samples associated with a target domain; and

analyzing the natural language utterance to determine the intent based at least in part on the retrieved one or more language samples.

15. The apparatus of claim 9 , wherein analyzing the natural language utterance to determine the intent is based at least in part on one or more statistical models of speech units.

16. The apparatus of claim 15 , wherein the one or more statistical models of speech units comprise a Gaussian mixture model or a deep neural network.

17. A non-transitory computer-readable medium storing code, the code comprising instructions executable by a processor to perform operations comprising:

receiving a natural language utterance associated with an input query image;

generating, in real-time using a neural network, an image signature for the input query image, the image signature numerically describing image content;

calculating, in real-time, a visual similarity measure between a plurality of candidate product images and the input query image based on a plurality of image signatures corresponding to the plurality of candidate product images, the visual similarity measure indicating a correlation between a candidate product image and the input query image;

determining a list of candidate products based on the visual similarity measure and knowledge graph data, the knowledge graph data used to confirm the correlation of a candidate product image of the list of candidate products generated based on the visual similarity measure;

analyzing the natural language utterance to determine an intent;

determining whether the list of candidate products satisfies the intent; and

when the list of candidate products do not satisfy the intent:

identifying a second list of candidate products based upon the intent; and

generating a response to the natural language utterance, the response comprising a natural language prompt and one or more products from the second list of candidate products.

18. The non-transitory computer-readable medium of claim 17 , the operations further comprising:

generating metadata associated with the input query image; and

retrieving metadata associated with the plurality of candidate product images, wherein calculating the visual similarity measure is based at least in part on the generated metadata associated with the input query image and the retrieved metadata associated with the plurality of candidate product images.

19. The non-transitory computer-readable medium of claim 17 , the operations further comprising:

collecting user data associated with a user that provided the natural language utterance, the user data comprising identity data, behavioral data, or both; and

extrapolating the collected user data to generate extrapolated user data, wherein the intent is determined based at least in part on the user data, the extrapolated user data, or both.

20. The non-transitory computer-readable medium of claim 17 , wherein analyzing the natural language utterance further comprises:

decoding the natural language utterance with a speech to text decoder that comprises a feature extraction component, an acoustical model component, a language model component, or any combination thereof, to generate text, wherein the intent is determined based at least in part on the text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2021
From: KALE, AJINKYA GORAKHNATH; YANG, FAN; WANG, QIAOSONG; KIAPOUR, MOHAMMADHADI; PIRAMUTHU, ROBINSON
To: EBAY INC.
Reel/Frame 057937/0604 →
Continuity (2)
Continuation 15294767 · Oct 16, 2016
Related Publication 20220050870A1 · Feb 17, 2022
Cited By (3)
US 12,223,533 US 12,272,130 US 12,548,059