IP Library › Granted Patent US 12,726,668
Granted Patent B2
US 12,726,668 · App. 19/009,721 · Granted Sep 1, 2026

Personalized video stream with diffusible features from other video streams

Inventors: Aaron Keith Baughman (Cary, NC); Chandankumar Johakhim Patel (Fairborn, OH); Eduardo Morales (Key Biscayne, FL); Rahul Agarwal (Jersey City, NJ)
Assignee: International Business Machines Corporation
H04N21/25891G06V10/44H04N21/251G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,726,668
App. No.
19/009,721
Granted
Sep 1, 2026
Kind
B2
Abstract

A method, according to one embodiment, includes receiving a plurality of video streams, performing feature identification to identify objects within scenes of the plurality of video streams, and correlating the identified objects with event outcomes within the video streams. The method further includes selecting a first of the video streams with a relatively highest level of correlation to the event outcomes, diffusing features from a remainder of the video streams into the first video stream to generate a personalized video stream, and causing the personalized video stream to be transmitted to a user device used by a target user. A computer program product, according to one embodiment, includes one or more computer readable storage media, and program instructions stored on the one or more storage media to perform the foregoing method.

Claims (49)

1 . A method comprising:

receiving a plurality of video streams;

performing feature identification to identify objects within scenes of the plurality of video streams;

correlating the identified objects with event outcomes within the video streams;

selecting a first of the video streams with a relatively highest level of correlation to the event outcomes;

diffusing features from a remainder of the video streams into the first video stream to generate a personalized video stream; and

causing the personalized video stream to be transmitted to a user device used by a target user.

2 . The method of claim 1 , wherein the correlating is based on directed acyclic graph structures.

3 . The method of claim 1 , wherein the correlating is based on adjacency matrices across events associated with the event outcomes.

4 . The method of claim 3 , further comprising:

creating a Directed Acyclic Graph (DAG) based on the identified objects and the event outcomes identified within the video streams, wherein the adjacency matrices are based on the DAG.

5 . The method of claim 1 , wherein the diffusing is based on a causal strength associated with the features.

6 . The method of claim 1 , wherein the diffusing is based on prompts to alter the features that are diffused from the remainder of the video streams into the first video stream.

7 . The method of claim 6 , further comprising:

extracting, from a user profile of the target user, information about preferences of the target user; and

causing the information to be analyzed by a trained artificial intelligence (AI) model, wherein an output of the trained AI model includes the prompts.

8 . The method of claim 1 , wherein the performed feature identification is feature pyramidal identification, wherein the feature pyramidal identification is based on a narrowing of features windows to identify the identified objects within frames that the scenes of the plurality of video streams are made up of.

9 . A computer program product comprising:

one or more computer readable storage media; and

program instructions stored on the one or more storage media to perform operations comprising:

receiving a plurality of video streams;

performing feature identification to identify objects within scenes of the plurality of video streams;

correlating the identified objects with event outcomes within the video streams;

selecting a first of the video streams with a relatively highest level of correlation to the event outcomes;

diffusing features from a remainder of the video streams into the first video stream to generate a personalized video stream; and

causing the personalized video stream to be transmitted to a user device used by a target user.

10 . The computer program product of claim 9 , wherein the correlating is based on directed acyclic graph structures.

11 . The computer program product of claim 9 , wherein the correlating is based on adjacency matrices across events associated with the event outcomes.

12 . The computer program product of claim 11 , wherein the operations further comprise:

creating a Directed Acyclic Graph (DAG) based on the identified objects and the event outcomes identified within the video streams, wherein the adjacency matrices are based on the DAG.

13 . The computer program product of claim 9 , wherein the diffusing is based on a causal strength associated with the features.

14 . The computer program product of claim 9 , wherein the diffusing is based on prompts to alter the features that are diffused from the remainder of the video streams into the first video stream.

15 . The computer program product of claim 14 , wherein the operations further comprise:

extracting, from a user profile of the target user, information about preferences of the target user; and

causing the information to be analyzed by a trained artificial intelligence (AI) model, wherein an output of the trained AI model includes the prompts.

16 . The computer program product of claim 9 , wherein the performed feature identification is feature pyramidal identification, wherein the feature pyramidal identification is based on a narrowing of features windows to identify the identified objects within frames that the scenes of the plurality of video streams are made up of.

17 . A computer system comprising:

a processor set;

one or more computer readable storage media; and

program instructions stored on the one or more storage media to cause the processor set to perform operations comprising:

receiving a plurality of video streams;

performing feature identification to identify objects within scenes of the plurality of video streams;

correlating the identified objects with event outcomes within the video streams;

selecting a first of the video streams with a relatively highest level of correlation to the event outcomes;

diffusing features from a remainder of the video streams into the first video stream to generate a personalized video stream; and

causing the personalized video stream to be transmitted to a user device used by a target user.

18 . The computer system of claim 17 , wherein the correlating is based on directed acyclic graph structures.

19 . The computer system of claim 17 , wherein the correlating is based on adjacency matrices across events associated with the event outcomes.

20 . The computer system of claim 17 , wherein the diffusing is based on a causal strength associated with the features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2025
From: BAUGHMAN, AARON KEITH; PATEL, CHANDANKUMAR JOHAKHIM; MORALES, EDUARDO; AGARWAL, RAHUL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069787/0493 →
Continuity (1)
Related Publication 20260197513A1 · Jul 9, 2026
References Cited (33)
US 8913835B2 · Kumar et al. · 2014 [cited by applicant]
US 8949235B2 · Peleg et al. · 2015 [cited by applicant]
US 10032113B2 · Codella et al. · 2018 [cited by applicant]
US 10074013B2 · Newman et al. · 2018 [cited by applicant]
US 10970554B2 · Oz et al. · 2021 [cited by applicant]
US 11896847B2 · Hibbard · 2024 [cited by applicant]
US 20020157116A1 · Jasinschi · 2002 [cited by applicant]
US 20150020119A1 · Kim · 2015 [cited by examiner]
US 20170068992A1 · Chen · 2017 [cited by examiner]
US 20210308487A1 · Hibbard · 2021 [cited by applicant]
US 20220318956A1 · Xu · 2022 [cited by applicant]
US 20230140125A1 · Glesinger et al. · 2023 [cited by applicant]
US 20230267736A1 · Green · 2023 [cited by examiner]
US 20240185498A1 · Francis · 2024 [cited by examiner]
US 20250078336A1 · Westcott · 2025 [cited by examiner]
US 20250104290A1 · Ding · 2025 [cited by examiner]
US 20250267315A1 · Maalej · 2025 [cited by examiner]
“Fine tuning vs Prompt Engineering: What's the difference?”, Datascientest, Mar. 19, 2024, 5 pages. [cited by applicant]
“Generative adversarial network”, Wikipedia, retrieved from web https://en.wikipedia.org/wiki/Generative_adversarial_network, dated Mar. 5, 2025, 31 pages. [cited by applicant]
“Knowledge distillation”, Wikipedia, retrieved from web https://en.wikipedia.org/wiki/Knowledge_distillation, dated Mar. 5, 2025, 6 pages. [cited by applicant]
“Transformer (deep learning architecture)”, Wikipedia, retrieved from web https://en.wikipedia.org/wiki/Transformer (deep_learning_architecture), dated Mar. 5, 2025, 31 pages. [cited by applicant]
“Variational autoencoder”, Wikipedia, retrieved from web https://en.wikipedia.org/wiki/Variational_autoencoder, dated Mar. 5, 2025, 9 pages. [cited by applicant]
Bird et al., “Typology of Risks of Generative Text-to-Image Models”, AIES '23, Aug. 8-10, 2023, 15 pages. [cited by applicant]
Cai et al., “DesignAID: Using Generative AI and Semantic Diversity for Design Inspiration”, CI '23, Nov. 6-9, 2023, pp. 1-11. [cited by applicant]
Croitoru et al., “Diffusion Models in Vision: A Survey”, arXiv:2209.04747, Jan. 16, 2025, 25 pages. [cited by applicant]
Issa Ali. “Transformer, GPT-3,GPT-J, T5 and BERT”, Medium, Jan. 26, 2023, 19 pages. [cited by applicant]
Kahatapitiya et al., “Object-Centric Diffusion for Efficient Video Editing”, arXiv:2401.05735v3, Aug. 30, 2024, 31 pages. [cited by applicant]
Kanungo et al., “An efficient k-means clustering algorithm: analysis and implementation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, Jul. 2002, pp. 881-892. [cited by applicant]
Karras et al., “Elucidating the Design Space of Diffusion-Based Generative Models”, NeurIPS Proceedings, 2022, 13 pages. [cited by applicant]
Universität Heidelberg “A Unified Architecture for Instance and Semantic Segmentation” http://presentations.cocodataset.org/COCO17-Stuff-FAIR.pdf, 2017, 48 pages. [cited by applicant]
Wu et al., “Promptus: Can Prompts Streaming Replace Video Streaming with Stable Diffusion”, arXiv:2405.20032v1, May 30, 2024, 15 pages. [cited by applicant]
Xu et al., “Versatile Diffusion: Text, Images and Variations All in One Diffusion Model”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 7754-7765. [cited by applicant]
Zhang et al., “Text-to-image Diffusion Models in Generative AI: A Survey”, arXiv:2303.07909v3, Nov. 8, 2024, 40 pages. [cited by applicant]