IP Library › Granted Patent US 12,626,727
Granted Patent B2
US 12,626,727 · App. 18/911,777 · Granted May 12, 2026

Syncing commentary with videos

Inventors: Aaron Keith Baughman (Cary, NC); Leonid Karlinsky (Acton, MA); Gozde Akay (Fredericton, CA); Eduardo Morales (Key Biscayne, FL)
Assignee: International Business Machines Corporation
G11B27/10G06F40/169G06F40/30G06V10/7715G06V10/82G06V20/41G06V20/46G11B27/036H04N21/8133
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,727
App. No.
18/911,777
Granted
May 12, 2026
Kind
B2
Abstract

Systems, methods, and computer program products for automatically syncing commentary with videos are described herein. A method comprises reading a sequence of frames of a video; generating frame documents based on the sequence of frames; reading commentary associated with the video; providing the commentary as input to a language model; reading embeddings generated by the language model based on the commentary; generating a commentary document in accordance with the embeddings; determining a semantic distance between the commentary document and each of the frame documents; selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document; identifying a consecutive subsequence of the sequence of frames associated with the subset; providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and reading a first frame generated by the diffusion model.

Claims (87)

1 . A computer-implemented method comprising:

reading a sequence of frames of a video;

generating a set of frame documents based on the sequence of frames, wherein each frame document of the set of frame documents corresponds to at least one of the frames of the sequence of frames;

reading commentary associated with the video;

providing the commentary as input to a language model;

reading embeddings generated by the language model based on the commentary, wherein the embeddings characterize the commentary;

generating a commentary document in accordance with the embeddings;

determining a semantic distance between the commentary document and each of the frame documents;

selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document;

identifying a consecutive subsequence of the sequence of frames associated with the subset;

providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and

reading a first frame generated by the diffusion model.

2 . The computer-implemented method of claim 1 , wherein reading the sequence of frames comprises individually receiving each frame during a stream of the video.

3 . The computer-implemented method of claim 1 , wherein a first frame document corresponding to a first frame is a distribution over topics associated with the first frame.

4 . The computer-implemented method of claim 1 , wherein providing each frame of the sequence of frames as input to the machine learning model comprises providing the at least two consecutive frames as input to the machine learning model.

5 . The computer-implemented method of claim 1 , wherein generating the set of frame documents comprises:

providing each frame of the sequence of frames as input to a machine learning model;

reading a feature map generated by the machine learning model based on at least one frame of the sequence of frames, wherein the feature map characterizes objects depicted in the at least one frame and semantic relationships among the objects; and

generating the frame document in accordance with the feature map.

6 . The computer-implemented method of claim 1 , the computer-implemented method further comprising:

providing the first frame and an end frame as input to the diffusion model, wherein the at least two consecutive frames comprise the end frame;

reading a second frame generated by the diffusion model;

determining whether the first frame and the second frame are equivalent; and

determining whether a number of frames greater than or equal to an insertion threshold have been generated.

7 . The computer-implemented method of claim 1 , the computer-implemented method further comprising:

inserting the first frame into the sequence of frames such that the video is modified.

8 . The computer-implemented method of claim 7 , the computer-implemented method further comprising:

synchronizing an audio representation of the commentary with the video in accordance with the consecutive subsequence; and

transmitting the video for presentation via a client computing platform.

9 . The computer-implemented method of claim 1 , wherein the machine learning model is a Feature Pyramid Network.

10 . The computer-implemented method of claim 1 , the computer-implemented method further comprising:

providing a prompt and a characterization of the consecutive subsequence as input to a generative machine learning model, wherein the prompt indicates a duration of a shortened commentary to be generated.

11 . The computer-implemented method of claim 1 , the computer-implemented method further comprising determining whether a length of the commentary is greater than a length of the consecutive subsequence, wherein the at least two consecutive frames of the sequence of frames and the embeddings are provided as input to the diffusion model responsive to determining the length of the commentary is greater than the length of the consecutive subsequence.

12 . A computer program product for syncing commentary with videos, the computer program product comprising:

one or more non-transitory computer-readable storage media;

program instructions stored on the one or more non-transitory computer-readable storage media to perform operations comprising:

reading a sequence of frames of a video;

generating a set of frame documents generated by the machine learning model based on the sequence of frames, wherein each frame document of the set of frame documents corresponds to at least one of the frames of the sequence of frames;

reading commentary associated with the video;

providing the commentary as input to a language model;

reading embeddings generated by the language model based on the commentary, wherein the embeddings characterize the commentary,

generating a commentary document in accordance with the embeddings;

determining a semantic distance between the commentary document and each of the frame documents

selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document;

identifying a consecutive subsequence of the sequence of frames associated with the subset;

providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and

reading a first frame generated by the diffusion model.

13 . The computer program product of claim 12 , wherein generating the set of frame documents comprises:

providing each frame of the sequence of frames as input to a machine learning model;

reading a feature map generated by the machine learning model based on at least one frame of the sequence of frames, wherein the feature map characterizes objects depicted in the at least one frame and semantic relationships among the objects; and

generating the frame document in accordance with the feature map.

14 . The computer program product of claim 12 , wherein the operations further comprise:

providing the first frame and an end frame as input to the diffusion model, wherein the at least two consecutive frames comprise the end frame;

reading a second frame generated by the diffusion model;

determining whether the first frame and the second frame are equivalent; and

determining whether a number of frames greater than or equal to an insertion threshold have been generated.

15 . The computer program product of claim 12 , wherein the operations further comprise:

inserting the first frame into the sequence of frames such that the video is modified.

16 . The computer program product of claim 12 , wherein the operations further comprise:

synchronizing an audio representation of the commentary with the video in accordance with the consecutive subsequence; and

transmitting the video for presentation via a client computing platform.

17 . A computer system for syncing commentary with videos, the computer system comprising:

a processor set;

one or more computer-readable storage media;

program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations comprising:

reading a sequence of frames of a video;

generating a set of frame documents based on the sequence of frames, wherein each frame document of the set of frame documents corresponds to at least one of the frames of the sequence of frames;

reading commentary associated with the video;

providing the commentary as input to a language model;

reading embeddings generated by the language model based on the commentary, wherein the embeddings characterize the commentary;

generating a commentary document in accordance with the embeddings;

determining a semantic distance between the commentary document and each of the frame documents;

selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document;

identifying a consecutive subsequence of the sequence of frames associated with the subset;

providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and

reading a first frame generated by the diffusion model.

18 . The computer system of claim 17 , wherein generating the set of frame documents comprises:

providing each frame of the sequence of frames as input to a machine learning model;

reading a feature map generated by the machine learning model based on at least one frame of the sequence of frames, wherein the feature map characterizes objects depicted in the at least one frame and semantic relationships among the objects; and

generating the frame document in accordance with the feature map.

19 . The computer system of claim 17 , wherein the operations further comprise:

providing the first frame and an end frame as input to the diffusion model, wherein the at least two consecutive frames comprise the end frame;

reading a second frame generated by the diffusion model;

determining whether the first frame and the second frame are equivalent; and

determining whether a number of frames greater than or equal to an insertion threshold have been generated.

20 . The computer program product of claim 17 , wherein the operations further comprise:

inserting the first frame into the sequence of frames such that the video is modified.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2024
From: BAUGHMAN, AARON KEITH; KARLINSKY, LEONID; AKAY, GOZDE; MORALES, EDUARDO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 068943/0452 →
Continuity (1)
Related Publication 20260105936A1 · Apr 16, 2026
References Cited (23)
US 10678855B2 · Vaughn et al. · 2020 [cited by applicant]
US 10825480B2 · Marco et al. · 2020 [cited by applicant]
US 11012662B1 · Baughman et al. · 2021 [cited by applicant]
US 11455576B2 · Dalli et al. · 2022 [cited by applicant]
US 11521655B2 · Baughman et al. · 2022 [cited by applicant]
US 20220172050A1 · Dalli et al. · 2022 [cited by applicant]
US 20230018621A1 · Lin · 2023 [cited by examiner]
US 20230214422A1 · Kwatra et al. · 2023 [cited by applicant]
US 20250111803A1 · Mace · 2025 [cited by examiner]
CN 109769132A · 2019 [cited by applicant]
Chan et al., “Automatic Linguistic Resolution: Framework and Applications,” IEEE Xplore, 1996, pp. 625-630. [cited by applicant]
Chan et al., “Symbolic Connetionism in Natural language Disambiguation,” IEEE Transactions on Neural Networks, Sep. 1998, pp. 739-755, vol. 9, No. 5. [cited by applicant]
Hassid et al., “More than words: In-the-wild visually-driven prosody for text-to-speech.” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10587-10597 (2022). [cited by applicant]
Li et al., “Scallop: A Language for Neurosymbolic Programming,” Proc. ACM Program. Lang. 7, PLDI, Article 166, 25 pages, (2023). [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” arXiv 1612.03144v2 (2017): 10 pages. [cited by applicant]
Lu et al., “High-Quality Automatic Voice Over with Alignment: Supervision through Self-Supervised Discrete Speech Units.” arXiv preprint arXiv:2306.17005 (2023). [cited by applicant]
Lu et al., “Visualtts: TTS With Accurate Lip-Speech Synchronization for Automatic Voice Over,” 2022 International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 8032-8036 (2022). [cited by applicant]
Yurochkin et al., “Hierarchical optimal transport for document representation.” Advances in neural information processing systems 32 (2019). [cited by applicant]
Jack Bantoc, “The Masters app, website feature AI commentary for tournament coverage”, https://edition.cnn.com/2023/04/07/golf/the-masters-ai-commentary-spt-intl/index.html, Apr. 7, 2023, 9 pages. [cited by applicant]
Moss et al., “Sounding liquids: Automatic sound synthesis from fluid simulation”, ACM Transactions on Graphics (TOG), Jul. 2, 2010, 13 pages. [cited by applicant]
No Author, “Text generation strategies”, https://web.archive.org/web/20240222011056/https://huggingface.co/docs/transformers/v4.27.2/en/generation_strategies, Feb. 22, 2024, 9 pages. [cited by applicant]
Pranav Dixit, “Game, set and AI: Wimbledon 2023 will see AI commentary for the first time in tennis with help of IBM”, https://www.businesstoday.in/technology/news/story/game-set-and-ai-wimbledon-2023-will-see-ai-commen… [cited by applicant]
Saha et al., “TinyNS: Platform-aware Neurosymbolic Auto Tiny Machine Learning”, ACM Transactions on Embedded Computing Systems, May 11, 2024, 48 pages. [cited by applicant]