IP Library › Granted Patent US 12,279,017
Granted Patent B1
US 12,279,017 · App. 17/841,418 · Granted Apr 15, 2025

Enhanced shopping based on recognition of objects presented in video

Inventors: Lingyun Wang (Bothell, WA); Qipin Chen (Bellevue, WA); Xing Ju (Bellevue, WA); Gordon Snow Zhang (Seattle, WA); Jack Patrick Copeland (Clyde Hill, WA)
Assignee: Amazon Technologies, Inc.
H04N21/47815G06Q30/0623G06V20/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,279,017
App. No.
17/841,418
Granted
Apr 15, 2025
Kind
B1
Abstract

Devices, systems, and methods are provided for smart shopping based on recognition of objects presented in video. A method may include identifying, by a first device, using a machine learning model using a computer vision technique, objects represented in video content; determining that an object of the objects is available for purchase using an online retail system; causing concurrent presentation of the video content and a first indication that the object is available for purchase using the online retail system; receiving, from a second device, a second indication of a user selection of the first indication, wherein the user selection is indicative of a request to present additional information associated with the object; generating, based on the request, presentation data including the additional information and an option to purchase the object using the online retail system; and causing, based on the request, presentation of the presentation data.

Claims (73)

1. A method for facilitating shopping for items presented in video content, the method comprising:

identifying, by at least one processor of a first device, using a machine learning model using a computer vision technique to classify objects represented in video content based on comparisons of the video content to object images, the objects represented in the video content;

determining, by the at least one processor, based on a comparison of first image data from the video content to a first object image in an online retail system, that a first object of the objects is available for purchase using the online retail system;

determining, by the at least one processor, based on a comparison of second image data from the video content to a second object image in the online retail system, that a second object of the objects is available for purchase using the online retail system;

causing, by the at least one processor, concurrent presentation, at a second device, of the video content and a first indication that the first object is available for purchase using the online retail system;

receiving, by the at least one processor, from the second device, a second indication of a user selection of the first indication, wherein the user selection is indicative of a request to present, at a third device, additional information associated with the first object;

generating, by the at least one processor, based on the request, presentation data comprising the additional information and an option to purchase the first object using the online retail system; and

causing, by the at least one processor, based on the request, presentation of the presentation data at the third device.

2. The method of claim 1 , further comprising:

determining, based on the user selection, a video frame of the video content from which the user selection was made;

determining, a location within the video frame from where the user selection was made; and

determining that the location is indicative of the first object,

wherein generating the presentation data is based on determining that the location is indicative of the first object.

3. The method of claim 1 , further comprising:

generating a metric indicative of the user selection, the video content, and the first object;

generating a ranking of video titles base on the metric; and

causing presentation, at the second device, of second video content based on the ranking.

4. The method of claim 1 , further comprising:

identifying, in a catalog of products available for purchase using the online retail system, the first object image,

wherein determining that the first object of the objects is available for purchase using the online retail system is based on the first object image in the catalog compared to the first image data.

5. The method of claim 1 , further comprising:

providing, to the online retail system a search query comprising a text string indicative of the first object; and

identifying search results of the search query; and

identifying the first object image in the search results.

6. The method of claim 1 , wherein the concurrent presentation further comprises a second indication that the second object is available for purchase using the online retail system.

7. A method for facilitating shopping for items presented in video content, the method comprising:

identifying, by at least one processor of a first device, using a machine learning model using a computer vision technique, objects represented in video content;

determining, by the at least one processor, based on a comparison of first image data from the video content to a first object image in an online retail system, that an object of the objects is associated with a product of the online retail system;

causing, by the at least one processor, concurrent presentation, at a second device, of the video content and a first indication that the object is associated with the product of the online retail system;

receiving, by the at least one processor, from the second device, a second indication of a user selection of the first indication, wherein the user selection is indicative of a request to present additional information associated with the product;

generating, by the at least one processor, based on the request, presentation data comprising the additional information and an option to purchase the product using the online retail system; and

causing, by the at least one processor, based on the request, presentation of the presentation data.

8. The method of claim 7 , further comprising:

generating, using the machine learning model, location coordinates of the object within a video frame of the video content.

9. The method of claim 8 , further comprising:

determining, a location within the video frame from where the user selection was made; and

determining that the location comprises the location coordinates,

wherein generating the presentation data is based on determining that the location comprises the location coordinates.

10. The method of claim 7 , further comprising:

selecting the video content from among multiple video titles based on a first ranking of video titles;

generating a metric indicative of the user selection, the video content, and the object;

generating, using a reinforced learning model, a second ranking of video titles based on the metric as feedback to the reinforced learning model; and

causing presentation of second video content based on the second ranking.

11. The method of claim 7 , further comprising:

providing, to the online retail system a search query comprising a text string indicative of the object; and

identifying search results of the search query; and

identifying the first object image in the search results.

12. The method of claim 7 , further comprising:

determining, by the at least one processor, that a second object of the objects is associated with a second product of the online retail system,

wherein the concurrent presentation further comprises a second indication that the second object is associated with a second product of the online retail system.

13. The method of claim 7 , further comprising:

determining a maximum number of objects permitted to be concurrently presented with the video content,

wherein generating the presentation data is based on the maximum number.

14. The method of claim 7 , wherein the first device is associated with a back-end of a video application, wherein the second device is associated with a front-end of the video application, and wherein the concurrent presentation occurs at the second device.

15. The method of claim 7 , wherein the presentation data are presented using a product page for the object.

16. The method of claim 7 , wherein the presentation data are presented using a message with a uniform resource locator associated with a product page for the object.

17. A system for facilitating shopping for items presented in video content, the system comprising memory coupled to at least one processor, the at least one processor configured to:

identify, using a machine learning model using a computer vision technique, objects represented in video content;

determine, based on a comparison of first image data from the video content to a first object image in an online retail system, that an object of the objects is associated with a product of the online retail system;

cause concurrent presentation, at a first device, of the video content and a first indication that the object is associated with the product of the online retail system;

receive, from the first device, a second indication of a user selection of the first indication, wherein the user selection is indicative of a request to present additional information associated with the product;

generate, based on the request, presentation data comprising the additional information and an option to purchase the product using the online retail system; and

cause, based on the request, presentation of the presentation data.

18. The system of claim 17 , wherein the at least one processor is further configured to:

generate, using the machine learning model, location coordinates of the object within a video frame of the video content.

19. The system of claim 18 , wherein the at least one processor is further configured to:

determine a location within the video frame from where the user selection was made; and

determine that the location comprises the location coordinates,

wherein to generate the presentation data is based on determining that the location comprises the location coordinates.

20. The system of claim 17 , wherein the at least one processor is further configured to:

provide, to the online retail system a search query comprising a text string indicative of the object; and

identify search results of the search query; and

identify the first object image in the search results.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2022
From: WANG, LINGYUN; CHEN, QIPIN; JU, XING; ZHANG, GORDON SNOW; COPELAND, JACK PATRICK
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 060237/0440 →
References Cited (12)
US 8682739B1 · Feinstein · 2014 [cited by applicant]
US 9973819B1 · Taylor et al. · 2018 [cited by applicant]
US 10055783B1 · Feinstein · 2018 [cited by applicant]
US 10440435B1 · Erdmann et al. · 2019 [cited by applicant]
US 10440436B1 · Taylor et al. · 2019 [cited by applicant]
US 20140259056A1 · Grusd · 2014 [cited by examiner]
US 20140282677A1 · Mantell · 2014 [cited by examiner]
US 20150100989A1 · Gellman · 2015 [cited by examiner]
US 20220182725A1 · Song · 2022 [cited by examiner]
US 20220309553A1 · Gibbon · 2022 [cited by examiner]
US 20230065762A1 · Gupta · 2023 [cited by examiner]
US 20230283839A1 · Vella · 2023 [cited by examiner]
Cited By (1)
US 12,597,218