IP Library › Granted Patent US 11,804,044
Granted Patent B1
US 11,804,044 · App. 17/934,033 · Granted Oct 31, 2023

Intelligent correlation of video and text content using machine learning

Inventor: Dylan James Rose (Medforn, MA)
Assignee: Amazon Technologies, Inc.
G06V20/49G06V20/46H04N21/4884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,804,044
App. No.
17/934,033
Granted
Oct 31, 2023
Kind
B1
Abstract

Systems, methods, and computer-readable media are disclosed for systems and methods for intelligent correlation of video and text content using machine learning. Example methods include determining first video content having a first video segment and first text corresponding to the first video segment, generating a first embedded vector using the first video segment and the first text, and determining, using a first machine learning model, a first image relevance score for the first video content based at least in part on the first embedded vector. Example methods may include determining a first text score for the first video content based at least in part on the first embedded vector, determining, using the first image relevance score and the first text score, that the first video segment depicts content of a first undesired category, and causing presentation of a notification indicating presence of content of the first undesired category.

Claims (65)

1. A method comprising:

determining, by one or more computer processors coupled to memory, a video file for a movie, the video file comprising first video segment and first text corresponding to the first video segment;

generating a first embedded vector using the first video segment and the first text;

determining, using a first machine learning model, a first image relevance score for the first video content based at least in part on the first embedded vector;

determining, using the first machine learning model, a first text score for the first video content based at least in part on the first embedded vector;

determining a first audio segment corresponding to the first video content; and

determining, using the first machine learning model, a first audio score for the first video content based at least in part on the first embedded vector;

determining, based at least in part on the first image relevance score, the first text score, and the first audio segment, that the first video segment depicts content of a first undesired category;

causing presentation of a notification indicating presence of content of the first undesired category at a display during playback of the first video content; and

causing presentation of an option to skip the first video segment.

2. The method of claim 1 , further comprising:

determining that a number of instances content of the first undesired category appear in the first video content is greater than a smoothing threshold; and

determining an increased confidence threshold for the first video content to reduce a number of notifications associated with the first video content.

3. The method of claim 1 , wherein the first text score represents a likelihood the first video segment depicts content of the first undesired category, and the first image relevance score is indicative of a match between the first text and content depicted in the first frame.

4. The method of claim 1 , further comprising:

receiving feedback data, wherein the feedback data comprises an indication that presence of content of the first undesired category is inaccurate; and

using the feedback data associated with the presentation of the notification to retrain the machine learning model.

5. A method comprising:

determining, by one or more computer processors coupled to memory, first video content comprising a first video segment and first text corresponding to the first video segment;

generating a first embedded vector using the first video segment and the first text;

determining, using a first machine learning model, a first image relevance score for the first video content based at least in part on the first embedded vector;

determining, using the first machine learning model, a first text score for the first video content based at least in part on the first embedded vector;

determining, based at least in part on the first image relevance score and the first text score, that the first video segment depicts content of a first undesired category; and

causing presentation of a notification indicating presence of content of the first undesired category at a display.

6. The method of claim 5 , further comprising:

determining a first audio segment corresponding to the first video content; and

determining, using the first machine learning model, a first audio score for the first video content based at least in part on the first embedded vector;

wherein generating the first embedded vector comprises generating the first embedded vector using the first video segment, the first text, and the first audio segment.

7. The method of claim 5 , further comprising:

determining a confidence score based at least in part on the first image relevance score and the first text score; and

determining that the confidence score satisfies a confidence threshold.

8. The method of claim 7 , wherein the confidence score is customizable, the method further comprising:

determining user preferences associated with a user account;

wherein the user preferences comprise the confidence threshold, and wherein the confidence score is associated with the first undesired category.

9. The method of claim 5 , wherein determining that the first video segment depicts content of a first undesired category is performed at a time the first video content is selected for playback.

10. The method of claim 5 , wherein the notification is presented during playback of the first video content, and within a predetermined time interval prior to presentation of the first video segment.

11. The method of claim 10 , further comprising:

during playback of the first video content, causing presentation of an option to skip the first video segment.

12. The method of claim 5 , wherein the first text score represents a likelihood the first video segment depicts content of the first undesired category; and

wherein the first video segment comprises a first frame, and the first image relevance score is indicative of a match between the first text and content depicted in the first frame.

13. The method of claim 5 , further comprising:

determining that a number of instances content of the first undesired category appear in the first video content is greater than a smoothing threshold; and

determining an increased confidence threshold for the first video content to reduce a number of notifications associated with the first video content.

14. The method of claim 5 , wherein the first text is at least one of: subtitle text or closed caption text.

15. The method of claim 5 , further comprising:

using feedback data associated with the presentation of the notification to retrain the machine learning model.

16. The method of claim 15 , further comprising:

receiving the feedback data, wherein the feedback data comprises an indication that presence of content of the first undesired category is inaccurate.

17. A device comprising:

memory that stores computer-executable instructions; and

at least one processor configured to access the memory and execute the computer-executable instructions to:

determine first video content comprising a first video segment and first text corresponding to the first video segment;

generate a first embedded vector using the first video segment and the first text;

determine, using a first machine learning model, a first image relevance score for the first video content based at least in part on the first embedded vector;

determine, using the first machine learning model, a first text score for the first video content based at least in part on the first embedded vector;

determine, based at least in part on the first image relevance score and the first text score, that the first video segment depicts content of a first undesired category; and

cause presentation of a notification indicating presence of content of the first undesired category at a display.

18. The device of claim 17 , wherein the at least one processor is further configured to access the memory and execute the computer-executable instructions to:

determine a confidence score based at least in part on the first image relevance score and the first text score; and

determine that the confidence score satisfies a confidence threshold.

19. The device of claim 17 , wherein the notification is presented during playback of the first video content within a predetermine time interval prior to presentation of the first video segment, and wherein the at least one processor is further configured to access the memory and execute the computer-executable instructions to:

during playback of the first video content, cause presentation of an option to skip the first video segment.

20. The device of claim 17 , wherein the at least one processor is further configured to access the memory and execute the computer-executable instructions to:

determine that a number of instances content of the first undesired category appear in the first video content is greater than a smoothing threshold; and

determine an increased confidence threshold for the first video content to reduce a number of notifications associated with the first video content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: ROSE, DYLAN JAMES
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 061170/0660 →
Cited By (2)
US 12,340,551 US 12,556,781