IP Library Granted Patent US 12,380,484
Granted Patent B1
US 12,380,484 · App. 17/385,406 · Granted Aug 5, 2025

Contextually relevant user-based product recommendations based on scene information

Inventors: Yash Chaturvedi (Issaquah, WA); Mohamed Kamal Omar (Seattle, WA); Alexander Ratnikov (Redmond, WA); Ahmed Aly Saad Ahmed (Bothell, WA); Steven James Cox (Mill Creek, WA); Prasanth Saraswatula (Bellevue, WA); Jingxiang Chen (Bellevue, WA)
Assignee: AMAZON TECHNOLOGIES, INC.
G06Q30/0631G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,484
App. No.
17/385,406
Granted
Aug 5, 2025
Kind
B1
Abstract

Systems, devices, and methods are provided for determining contextually relevant user-based product recommendations based on scene information. In at least one embodiment, techniques described herein may be used to determine, using a first machine-learning model, first information associated with a first object within a first image of digital content, determine, using a second machine-learning model, similarity scores between the first object and a first plurality of products of an online purchasing system, detect, in association with the first image of the digital content, performance of a first computer-based action by a user, determine, using a third machine-learning model and based on contextual data of the user, one or more affinity scores for the user, select a first product based on the one or more affinity scores, and present a recommendation to the user to perform a second computer-based action in association with the first product.

Claims (58)

1. A system, comprising:

one or more processors; and

memory storing executable instructions that, as a result of execution by the one or more processors, cause the system to:

parse a scene in a video content to identify a first object in multiple frames of the scene, the multiple frames including a first frame;

determine first information associated with the first object;

determine, using a second machine-learning model and based on the first information, similarity scores between the first object and one or more products of an online purchasing system, wherein the second machine-learning model is trained on a database of product images including images for the one or more products;

detect, in association with the first frame, performance of a first computer-based action by a user;

determine, based on contextual data of the user, one or more affinity scores for the user, wherein a first affinity score associated with a first product of the one or more products indicates a likelihood the user is to perform a second computer-based action in association with the first product;

select the first product based on the one or more affinity scores;

encode the video content to include an annotation associated with the first object, the annotation including third information associated with the first product; and

transmit the encoded video content to a user device.

2. The system of claim 1 , wherein:

the first information comprises bounding box information for the first object; and

the instructions include further instructions that, as a result of execution by the one or more processors, further cause the system to:

extract a patch of the first frame corresponding to the bounding box; and

compute, using the second machine-learning model, cosine similarity scores between the patch and images of the one or more products.

3. The system of claim 1 , wherein the instructions include further instructions that, as a result of execution by the one or more processors, further cause the system to:

determine second information associated with a second object within the first frame of video content.

4. The system of claim 3 , wherein

determine, using the second machine-learning model, second similarity scores between the second object and a second one or more products of the online purchasing system;

wherein the one or more affinity scores comprises a second affinity score associated with a second product of the second one or more products; and

the first product is selected further based on the first affinity score being greater than the second affinity score.

5. The system of claim 1 , wherein:

the contextual data comprises text-based data; and

the instructions include further instructions that, as a result of execution by the one or more processors, further cause the system to use a natural language processing (NLP) engine to extract embeddings from the text-based data.

6. The system of claim 5 , wherein the text-based data comprises a review, a comment, or a post by the user.

7. The system of claim 1 , wherein

a first similarity score for the first object is determined, using the second machine-learning model, further based on a second frame of the video content, further wherein:

the first object is depicted in the first frame from a first perspective; and

the first object is depicted in the second frame from a second perspective different from the first perspective.

8. The system of claim 1 , wherein the recommendation comprises a user-interactable object to perform the second computer-based action.

9. A non-transitory computer-readable storage medium storing executable instructions that, as a result of being executed by one or more processors of a computer system, cause the computer system to at least:

parse a scene in a video content to identify a first object in multiple frames of the scene, the multiple frames including a first frame;

determine first information associated with the first object within the first frame of video content;

determine, using a second machine-learning model and the first information, similarity scores between the first object and one or more products of an online purchasing system, wherein the second machine-learning model is trained on a database of product images including images for the one or more products;

detect, in association with the first frame of the video content, performance of a first computer-based action by a user;

determine, based on contextual data of the user, one or more affinity scores for the user, wherein a first affinity score associated with a first product of the one or more products indicates a likelihood the user is to perform a second computer-based action in association with the first product;

select the first product based on the one or more affinity scores;

encode the video content to include an annotation associated with the first object, the annotation including third information associated with the first product; and

transmit the encoded video content to a user device.

10. The non-transitory computer-readable storage medium of claim 9 , wherein:

the first information comprises a classification of the first object and

the instructions, as a result of being executed by the one or more processors of the computer system, further cause the system to:

track the first object over a plurality of frames of the video content;

determine confidence of the classification of the first object exceeds a threshold over the plurality of frames; and

determine the similarity scores between the first object and the one or more products of the online purchasing system based on the confidence of the classification of the first object exceeding the threshold over the plurality of frames.

11. The non-transitory computer-readable storage medium of claim 9 , wherein a similarity score between the first object and the first product is computed based on:

a first cosine similarity score determined based on:

a first portion of the first frame corresponding to the first object from a first perspective; and

a second image, of the first product; and

a second cosine similarity score determined based on:

a second portion of a third frame corresponding to the first object from a second perspective; and

a fourth image, of the first product.

12. The non-transitory computer-readable storage medium of claim 9 , wherein the contextual data comprises a budget, shopping history, browsing history, wish list, or shopping cart of the user.

13. The non-transitory computer-readable storage medium of claim 9 , wherein the performance of the first computer-based action by the user comprises a command to pause the video content at the first frame.

14. The non-transitory computer-readable storage medium of claim 9 , wherein the contextual data comprises a product review written by the user.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the recommendation comprises a user-interactable object to perform the second computer-based action.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the recommendation further comprises a product description generated based on the product review or video content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2021
From: CHATURVEDI, YASH; OMAR, MOHAMED KAMAL; RATNIKOV, ALEXANDER; AHMED, AHMED ALY SAAD; COX, STEVEN JAMES; SARASWATULA, PRASANTH; CHEN, JINGXIANG
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 057003/0564 →
References Cited (3)
US 20190287139A1 · Yadav · 2019 [cited by examiner]
US 20190362154A1 · Moore · 2019 [cited by examiner]
Gharibshah, Zhabiz, et al., “User Response Prediction in Online Advertising”. ACM Computing Surveys (CSUR), 2021, preprint. https://doi.org/10.48550/arXiv.2101.02342 (Year: 2021). [cited by examiner]
Cited By (6)
US 12,586,351 US 12,632,896 US 12,651,085 US 12,689,787 US 12,694,430 US 12,707,114