IP Library Granted Patent US 12694430
Granted Patent B2
US 12694430 · App. 18/581,334 · Granted Jul 28, 2026

System for seamlessly stitching fully contextual ads to content for immersive advertising

Inventors: Susmita Ghose (Mountain View, CA); Ashutosh Chaubey (Chhattisgarh, IN); Sartaki Sinha Roy (Uttarpara, IN); Anoubhav Agarwaal (Karnataka, IN); Aayush Agrawal (Madhya, IN)
Assignee: ANOKI, INC.
G06Q30/0276G06Q30/0251G06Q30/0277H04N21/23418H04N21/812
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694430
App. No.
18/581,334
Granted
Jul 28, 2026
Kind
B2
Abstract

A system for contextual modification of content based on multimodal extraction of metadata from the content, wherein the metadata is extracted by processing one or more scenes in the content to extract metadata corresponding to multiple extraction modes, and an embedding model for each extraction mode wherein an aggregated embedding model responsive to the extracted metadata for each mode formulates an aggregated embedding. A process controller may include an embedding extractor responsive to a control input. The control input may specify one or more features appearing in the content defining a content modification opportunity. The embedding extractor may include an embedding model coordinated with the embedding model for one or more of the embedding modes to generate an opportunity embedding in the form of a vector. A vector comparison processor determines the distance between the opportunity embedding and the aggregated embedding, wherein the embeddings are in the form of vectors. The process controller is responsive to the vector comparison processor to generate edit control instructions indicating a modification of the content upon detection of the content modification opportunity. A content editor is responsive to the edit control instructions to modify the content. The content editor uses generative AI techniques to modify the content by replacing an element appearing within the content during the content modification opportunity with an element correlated with the element appearing in said content.

Claims (23)

1 . A system for contextual modification of content based on multimodal extraction of metadata from said content comprising:

a metadata extractor processing one or more scenes in said content to extract metadata corresponding to multiple extraction modes, and a first embedding model for each extraction mode wherein an aggregated embedding model responsive to said extracted metadata for each mode formulates an aggregated embedding;

a process controller including an embedding extractor responsive to a control input wherein said control input specifies one or more features defining a content modification opportunity and wherein said embedding extractor includes a second embedding model coordinated with said first embedding model for one or more said multiple extraction modes to generate an opportunity embedding in the form of a vector;

a vector comparison processor for determining the distance between said opportunity embedding and said aggregated embedding to determine a content modification opportunity when said distance is within a threshold;

wherein said processor controller is responsive to said vector comparison processor to generate edit control instructions indicating a modification of said content upon detection of said content modification opportunity; and

a content editor responsive to said edit control instructions for a content modification opportunity to modify said content and having a modified content output wherein said content editor uses generative AI techniques to modify said content by replacing an element appearing within said content during said content modification opportunity with an element correlated with said element appearing in said content.

2 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 1 wherein said element appearing during said content modification opportunity represents an object that appears in one or more frames of said content and wherein said element correlated with an element of said content represents an alternative object and said generative AI techniques exercise at least an area of said one or more frames corresponding to said element and inserting said element correlated with said element appearing in said content in said one or more frames.

3 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 2 wherein said generative AI techniques operate to fill any gaps between said element appearing within said content during said content modification opportunity and said element correlated with said element appearing in said content.

4 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 3 further comprising a creative library, wherein said creative library stores one or more templates creatives and said edit control instructions specify said generative AI techniques to said template creative retrieved from said creative library application for use by said content editor.

5 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 4 wherein said edit control instructions cause said content editor to apply said generative AI techniques based on triggering embeddings.

6 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 5 wherein said edit control instructions include triggering embeddings.

7 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 6 wherein said triggering embeddings are based in part on said aggregated embeddings.

8 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 7 wherein said triggering embeddings are based in part on metadata concerning viewers, and further comprising a viewer database for viewer metadata.

9 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 8 wherein said edit control instructions are dynamically generated based on said viewer database and scene content embeddings in said modified content.

10 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 9 further comprising a modification selection server responsive to said opportunity to select a modification to apply to said content.

11 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 10 wherein said modification selection server is a competitive bid processor.

12 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 1 further comprising a modification selection server responsive to said opportunity to select a modification to apply to said content wherein said modification selection server is a competitive bid processor wherein said element appearing during said content modification opportunity represents an object that appears in one or more frames of said content and wherein said element correlated with an element of said content represents an alternative object and said generative AI techniques insert one or more frames of modified content into said content at said content modification opportunity.

13 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 12 wherein said generative AI techniques operate to fill any gaps between said element appearing within said content during said content modification opportunity and said element correlated with said element appearing in said content.

14 . The system for contextual modification of media content based on multimodal extraction of metadata from said content according to claim 13 wherein said edit control instructions cause said content editor to apply said generative AI techniques based on triggering embeddings.

15 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 14 wherein said edit control instructions include triggering embeddings.

16 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 15 wherein said triggering embeddings are based in part on said aggregated embeddings.

17 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 16 wherein said triggering embeddings are based in part on metadata concerning viewers, and further comprising a viewer database for viewer metadata.

18 . The system for contextual modification of content based on multimodal extraction of metadata from said content according to claim 17 wherein said edit control instructions are dynamically generated based on said viewer database and scene content embeddings in said modified content.