IP Library Patent Application 18581328
Patent Application
App. No. 18/581,328

MULTI-MODAL METADATA EXTRACTION SYSTEM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/581,328
Abstract

A multimodal metadata extraction system may be provided with a scene detector having a video content input and an output representing scene boundaries. The metadata extractor may be responsive to the content of a scene to extract metadata corresponding to several, plural, or multiple extraction modes. A metadata embedding may be used for each of the modes. An embedding aggregator responsive to the embedding operates to formulate an aggregated embedding for each scene thereby indexing the content of the scene. The scene detector may include a frame analyzer for identifying consecutive frames having similar characteristics. A boundary detector may be provided to identify boundaries of consecutive frames having sufficiently similar characteristics that they likely belong to the same shot. An embedding system may be provided to formulate a composite distance matrix capturing the distance between shot embeddings. A temporal clustering system may be connected to the composite distance matrix. An output of the temporal clustering system identifies the scene boundaries of the content. An embedding database may be connected to the embedding aggregator for storing the aggregated embedding for use as a search index for scenes identified in the content.

Claims (14)

1 . A multimodal metadata extraction system comprising:

a scene detector having a video content input and an output representing scene boundaries;

a metadata extractor responsive to content of a scene as identified by said scene boundaries to extract metadata corresponding to several extraction modes;

a metadata embedding for each extraction mode; and

an embedding aggregator response to said metadata embedding is to formulate aggregated embedding for each scene indexing said content.

2 . The multimodal metadata extraction system according to claim 1 wherein said output representing identified scenes is a set of video clips in each scene.

3 . The multimodal metadata extraction system according to claim 1 wherein said output representing identified scenes is an index to said video content corresponding to said identified scenes.

4 . The multimodal metadata extraction system according to claim 1 wherein said scene detector further comprising:

a frame analyzer for identifying consecutive frames having similar characteristics,

a boundary detector identifying boundaries of consecutive frames having such similar characteristics responsive to said frame analyzer;

an embedding system formulating a composite distance matrix capturing distance between shot embedding; and

a temporal clustering system connected to said composite distance matrix and having an output identifying scene boundaries of said content.

5 . The multimodal metadata extraction system according to claim 1 further comprising an embedding database connected to said embedding aggregator for storing said aggregated embedding for use as a search index for scenes of said content.

6 . The multimodal metadata extraction system according to claim 1 wherein said extraction modes include at least one of Audio (speech recognition, music recognition); Image recognition (feature recognition with temporal understanding); Text (caption, scene summarization, text recognition; and scene interpretation (sentiment, profanity, acting level).

Assignments (2)
SECURITY INTEREST Recorded Nov 19, 2024
From: ANOKI INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 069324/0514 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2024
From: GHOSE, SUSMITA; CHAUBEY, ASHUTOSH; SINHA ROY, SARTAKI; BALDUA, ASHISH
To: ANOKI INC.
Reel/Frame 068413/0925 →