IP Library Granted Patent US 12671878
Granted Patent B2
US 12671878 · App. 18/427,443 · Granted Jun 30, 2026

Systems and methods for automated metadata generation for multimedia content using multimodal data

Inventors: Adwait Ashish Murudkar (Somerville, NJ); Vidhya Seran (Irving, TX); Sergey Virodov (San Diego, CA)
Assignee: Verizon Patent and Licensing Inc.
H04N21/84G06F16/783G06V20/41G06V20/49G06V20/70H04N21/23418H04N21/44008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12671878
App. No.
18/427,443
Filed
Jan 30, 2024
Granted
Jun 30, 2026
Kind
B2
Examiner
TRAN, LOI H
Art Unit
2484
USPC
386/239
Abstract

In some implementations, a device may receive a multimedia content file. The device may divide the multimedia content file into a set of logical entities. The device may generate a set of embeddings for each logical entity of the set of logical entities. The device may compare groups of logical entities, of the set of logical entities, to generate a similarity metric. The device may selectively merge, based on the similarity metric satisfying a threshold, a pair of logical entities, in a group of logical entities of the groups of logical entities, to generate one or more logical entity sets. The device may process the one or more logical entity sets to generate one or more metadata tags for the one or more logical entity sets. The device may store a metadata file including the one or more metadata tags.

Claims (47)

1 . A device, comprising:

one or more processors configured to:

receive a video file;

process the video file to divide the video file into a set of logical camera shots;

process the set of logical camera shots using a machine learning model to generate a set of embeddings for each logical camera shot, of the set of logical camera shots, wherein a quantity of embeddings generated for each logical camera shot of the set of logical camera shots is based on a length of each logical camera shot;

compare groups of logical camera shots, of the set of logical camera shots, to generate a similarity metric;

selectively merge, based on whether the similarity metric satisfies a threshold, a pair of logical camera shots, in a group of logical camera shots of the groups of logical camera shots, to generate one or more logical camera scenes, wherein a logical camera scene, of the one or more logical camera scenes, includes one or more logical camera shots;

process the one or more logical camera scenes to generate one or more metadata tags for the one or more logical camera scenes;

transmit the one or more metadata tags concurrently with the generation of the one or more logical camera scenes, wherein the video file is associated with a live stream and the one or more metadata tags are included into the live stream; and

store a metadata file including the one or more metadata tags in connection with the video file;

wherein the one or more instructions, that cause the device to generate the one or more metadata tags, cause the device to: generate a metadata tag for a logical entity set, of the one or more logical entity sets, based on a position of the logical entity set within the one or more logical entity sets.

2 . The device of claim 1 , wherein the one or more processors are further configured to: receive a request for the video file; and transmit a message including the video file and the metadata file.

3 . The device of claim 1 , wherein the one or more processors, to receive the video file, are configured to: obtain the video file from a data repository; and wherein the one or more processors, to store the metadata file, are configured to: store the metadata file in the data repository in association with the video file.

4 . The device of claim 1 , wherein the one or more processors are further configured to: receive a search query; search, using the search query, a data repository storing a set of metadata files that includes the metadata file for the video file; determine a match between the search query and a metadata tag of the metadata file; and return, based on the match between the search query and the metadata tag, the video file.

5 . The device of claim 1 , wherein the one or more processors are further configured to: receive a search query; search, using the search query, a data repository storing a set of metadata files that includes the metadata file for the video file; determine a match between the search query and a metadata tag of the metadata file; and return, based on the match between the search query and the metadata tag, a new video file including a logical entity set, of the one or more logical camera scenes, associated with the match between the search query and the metadata tag.

6 . A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

receive a multimedia content file;

process the multimedia content file to divide the multimedia content file into a set of logical entities;

process the set of logical entities using a machine learning model to generate a set of embeddings for each logical entity, of the set of logical entities, wherein a quantity of embeddings generated for each logical entity of the set of logical entities is based on a length of each logical entity;

compare groups of logical entities, of the set of logical entities, to generate a similarity metric;

selectively merge, based on the similarity metric satisfying a threshold, a pair of logical entities, in a group of logical entities of the groups of logical entities, to generate one or more logical entity sets, wherein a logical entity set, of the one or more logical entity sets, includes one or more logical entities;

generate one or more metadata tags for the one or more logical entity sets; and

transmit the one or more metadata tags concurrently with the generation of the one or more logical entities, wherein the multimedia content file is associated with a live stream and the one or more metadata tags are included into the live stream;

wherein the one or more instructions, that cause the device to generate the one or more metadata tags, cause the device to: generate a metadata tag for a logical entity set, of the one or more logical entity sets, based on a position of the logical entity set within the one or more logical entity sets.

7 . The non-transitory computer-readable medium of claim 6 , wherein the one or more instructions, that cause the device to generate the one or more metadata tags, cause the device to: generate a metadata tag for a logical entity set, of the one or more logical entity sets, based on a computer vision analysis of one or more objects detected within the logical entity set.

8 . The non-transitory computer-readable medium of claim 6 , wherein the one or more instructions, that cause the device to generate the one or more metadata tags, cause the device to: generate a metadata tag for a logical entity set, of the one or more logical entity sets, based on a natural language processing analysis of audio associated with the logical entity set.

9 . The non-transitory computer-readable medium of claim 6 , wherein a first logical entity set, of the one or more logical entity sets, includes a first logical entity and a second logical entity, and a second logical entity set, of the one or more logical entity sets, includes a third logical entity, and wherein the third logical entity is between the first logical entity and the second logical entity in an order of logical entities within the multimedia content file.

10 . The non-transitory computer-readable medium of claim 6 , wherein a logical entity set, of the one or more logical entity sets, includes a first logical entity, a second logical entity, and a third logical entity in consecutive order, and wherein the first logical entity and the third logical entity are associated with the similarity metric satisfying the threshold, and wherein the second logical entity is included in the logical entity set based at least on being between the first logical entity and the third logical entity in consecutive order.

11 . The non-transitory computer-readable medium of claim 6 , wherein the machine learning model includes at least one of a zero-shot video transformer model or a modality transformer model.

12 . The non-transitory computer-readable medium of claim 6 , wherein the similarity metric is a cosine similarity metric.

13 . The non-transitory computer-readable medium of claim 6 , wherein the multimedia content file includes video content, wherein a logical entity, of the set of logical entities, is a logical camera shot, and wherein a logical entity set, of the one or more logical entity sets, is a logical camera scene.

14 . The non-transitory computer-readable medium of claim 6 , wherein the multimedia content file includes audio content, wherein a logical entity, of the set of logical entities, is a logical audio recording, and wherein a logical entity set, of the one or more logical entity sets, is a logical camera scene.

15 . A method, comprising:

receiving, by a device, a multimedia content file;

dividing, by the device, the multimedia content file into a set of logical entities;

generating, by the device and using a machine learning model, a set of embeddings for each logical entity of the set of logical entities, wherein a quantity of embeddings generated for each logical entity of the set of logical entities is based on a length of each logical entity;

comparing, by the device, groups of logical entities, of the set of logical entities, to generate a similarity metric;

selectively merging, by the device and based on the similarity metric satisfying a threshold, a pair of logical entities, in a group of logical entities of the groups of logical entities, to generate one or more logical entity sets, wherein a logical entity set, of the one or more logical entity sets, includes one or more logical entities;

processing, by the device, the one or more logical entity sets to generate one or more metadata tags for the one or more logical entity sets;

transmitting the one or more metadata tags concurrently with the generation of the one or more logical entity sets, wherein the multimedia content file is associated with a live stream and the one or more metadata tags are included into the live stream; and

storing, by the device, a metadata file including the one or more metadata tags in connection with the multimedia content file;

wherein generating a metadata tag for a logical entity set, of the one or more logical entity sets, is based on a position of the logical entity set within the one or more logical entity sets.

16 . The method of claim 15 , wherein the multimedia content file does not include metadata labeling logical entities or logical entity sets.

17 . The method of claim 15 , wherein a single logical entity is divided into a plurality of logical entities based on a length of the single logical entity.

18 . The method of claim 15 , further comprising: receiving a request for the multimedia content file; and transmitting a message including the multimedia content file and the metadata file.

19 . The method of claim 15 , wherein the machine learning model includes at least one of a zero-shot video transformer model or a modality transformer model.