IP Library Granted Patent US 12,382,115
Granted Patent B2
US 12,382,115 · App. 18/618,080 · Granted Aug 5, 2025

Machine learning based media content annotation

Inventors: Jonathan Bennett-James (Wales, GB); Craig Holbrook (Wales, GB)
Assignee: NAGRAVISION SARL
H04N21/2353G06N20/00G06V20/47H04N21/26603
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,382,115
App. No.
18/618,080
Granted
Aug 5, 2025
Kind
B2
Abstract

Systems and techniques are described herein for annotating media content. For example, a process can include obtaining media content and generate, use one or more machine learning models, a metadata file for at least a portion of the media content. The metadata file includes one or more metadata descriptions. The process can include generating a text description of the media content based on the one or more metadata descriptions of the metadata file. The process can include annotating the media content use the text description.

Claims (62)

1. A method of annotating media content, the method comprising:

obtaining media content;

generating, using one or more machine learning models, a metadata file for at least a portion of the media content, the metadata file including one or more metadata descriptions;

generating a plurality of sentences using the one or more metadata descriptions;

determining, using a machine learning model, a corresponding sentiment associated with each sentence of the plurality of sentences;

comparing the corresponding sentiment associated with each sentence of the plurality of sentences to a sentiment associated with at least the portion of the media content;

determining a subset of sentences from among the plurality of sentences to generate a text description of a scene, wherein the subset of sentences are within a sentiment threshold of the sentiment associated with at least the portion of the media content; and

annotating the media content using the text description of the scene of the media content.

2. The method of claim 1 , wherein each metadata description of the one or more metadata descriptions is associated with at least one of a character depicted in at least the portion of the media content, a facial expression of the character depicted in at least the portion of the media content, an object depicted in at least the portion of the media content, and an action occurring in at least the portion of the media content.

3. The method of claim 2 , wherein generating the metadata file for at least the portion of the media content includes:

determining, using the one or more machine learning models, at least one of the character depicted in at least the portion of the media content, the facial expression of the character depicted in at least the portion of the media content, the object depicted in at least the portion of the media content, and the action occurring in at least the portion of the media content; and

generating the one or more metadata descriptions for at least one of the character depicted in at least the portion of the media content, the facial expression of the character depicted in at least the portion of the media content, the object depicted in at least the portion of the media content, and the action occurring in at least the portion of the media content.

4. The method of claim 1 , further comprising:

determining, using a machine learning model, a first character and a second character depicted in at least the portion of the media content;

determining a first priority score for the first character and a second priority score for the second character; and

adding the first priority score and the second priority score to the metadata file.

5. The method of claim 4 , further comprising:

determining the first priority score is higher than the second priority score; and

based on determining the first priority score being higher than the second priority score, generating the text description of the scene of the media content using audio data associated with the first character.

6. The method of claim 1 , further comprising:

generating one or more metadata files for a plurality of portions of the media content, each metadata file of the one or more metadata files being associated with a corresponding timestamp within the media content.

7. The method of claim 1 , wherein generating the plurality of sentences using the one or more metadata descriptions includes:

determining a subset of metadata descriptions from the one or more metadata descriptions having confidence scores greater than a confidence threshold; and

generating the plurality of sentences using the subset of metadata descriptions having confidence scores greater than the confidence threshold.

8. The method of claim 7 , further comprising:

discarding one or more metadata descriptions from the one or more metadata descriptions having a confidence score greater than the confidence threshold.

9. The method of claim 7 , wherein generating the plurality of sentences using the subset of metadata descriptions includes:

obtaining a plurality of template sentences, each template sentence of the plurality of template sentences including one or more placeholder metadata tags; and

replacing placeholder metadata tags of the plurality of template sentences with the subset of metadata descriptions having confidence scores greater than the confidence threshold.

10. The method of claim 1 , wherein annotating the media content using the text description of the scene of the media content includes:

generating an audio file using the text description.

11. The method of claim 10 , wherein generating the audio file includes converting the text description to an audio description, and further comprising embedding the audio file into a file of the media content.

12. The method of claim 1 , wherein annotating the media content using the text description of the scene of the media content includes:

generating a media summary of the media content using the text description.

13. A system for annotating media content, including:

a memory; and

one or more processors coupled to the memory and configured to:

obtain media content;

generate, using one or more machine learning models, a metadata file for at least a portion of the media content, the metadata file including one or more metadata descriptions;

generate a plurality of sentences using the one or more metadata descriptions;

determine, using a machine learning model, a corresponding sentiment associated with each sentence of the plurality of sentences;

compare the corresponding sentiment associated with each sentence of the plurality of sentences to a sentiment associated with at least the portion of the media content;

determining a subset of sentences from among the plurality of sentences to generate a text description of a scene, wherein the subset of sentences are within a sentiment threshold of the sentiment associated with at least the portion of the media content; and

annotate the media content use the text description of the scene of the media content.

14. The system of claim 13 , wherein each metadata description of the one or more metadata descriptions is associated with at least one of a character depicted in at least the portion of the media content, a facial expression of the character depicted in at least the portion of the media content, an object depicted in at least the portion of the media content, and an action occurring in at least the portion of the media content.

15. The system of claim 14 , wherein the one or more processors are configured to:

determine, use the one or more machine learning models, at least one of the character depicted in at least the portion of the media content, the facial expression of the character depicted in at least the portion of the media content, the object depicted in at least the portion of the media content, and the action occurring in at least the portion of the media content; and

generate the one or more metadata descriptions for at least one of the character depicted in at least the portion of the media content, the facial expression of the character depicted in at least the portion of the media content, the object depicted in at least the portion of the media content, and the action occurring in at least the portion of the media content.

16. The system of claim 13 , wherein the one or more processors are configured to:

determine, use a machine learning model, a first character and a second character depicted in at least the portion of the media content;

determine a first priority score for the first character and a second priority score for the second character; and

add the first priority score and the second priority score to the metadata file.

17. The system of claim 16 , wherein the one or more processors are configured to:

determine the first priority score is higher than the second priority score; and

based on determining the first priority score is higher than the second priority score, generate the text description of the scene of the media content using audio data associated with the first character.

18. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:

obtain media content;

generate, using one or more machine learning models, a metadata file for at least a portion of the media content, the metadata file including one or more metadata descriptions;

generate a plurality of sentences using the one or more metadata descriptions;

determine, using a machine learning model, a subset of sentences from the plurality of sentences to generate a text description of a scene of the media content; and

annotate the media content use the text description of the scene of the media content.

19. The non-transitory computer-readable storage medium of claim 18 , wherein each metadata description of the one or more metadata descriptions is associated with at least one of a character depicted in at least the portion of the media content, a facial expression of the character depicted in at least the portion of the media content, an object depicted in at least the portion of the media content, and an action occurring in at least the portion of the media content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: BENNETT-JAMES, JONATHAN; HOLBROOK, CRAIG
To: NAGRAVISION MEDIA UK LTD.
Reel/Frame 071349/0477 →
Continuity (3)
Continuation 17510722 · Oct 26, 2021
Provisional Application 63106784 · Oct 28, 2020
Related Publication 20240314372A1 · Sep 19, 2024
References Cited (4)
US 11151386B1 · Aggarwal · 2021 [cited by examiner]
US 20180277093A1 · Carr · 2018 [cited by examiner]
US 20210117685A1 · Sureshkumar · 2021 [cited by examiner]
US 20210151038A1 · Manjunath · 2021 [cited by examiner]