Syncing commentary with videos
Systems, methods, and computer program products for automatically syncing commentary with videos are described herein. A method comprises reading a sequence of frames of a video; generating frame documents based on the sequence of frames; reading commentary associated with the video; providing the commentary as input to a language model; reading embeddings generated by the language model based on the commentary; generating a commentary document in accordance with the embeddings; determining a semantic distance between the commentary document and each of the frame documents; selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document; identifying a consecutive subsequence of the sequence of frames associated with the subset; providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and reading a first frame generated by the diffusion model.
1 . A computer-implemented method comprising:
reading a sequence of frames of a video;
generating a set of frame documents based on the sequence of frames, wherein each frame document of the set of frame documents corresponds to at least one of the frames of the sequence of frames;
reading commentary associated with the video;
providing the commentary as input to a language model;
reading embeddings generated by the language model based on the commentary, wherein the embeddings characterize the commentary;
generating a commentary document in accordance with the embeddings;
determining a semantic distance between the commentary document and each of the frame documents;
selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document;
identifying a consecutive subsequence of the sequence of frames associated with the subset;
providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and
reading a first frame generated by the diffusion model.
2 . The computer-implemented method of claim 1 , wherein reading the sequence of frames comprises individually receiving each frame during a stream of the video.
3 . The computer-implemented method of claim 1 , wherein a first frame document corresponding to a first frame is a distribution over topics associated with the first frame.
4 . The computer-implemented method of claim 1 , wherein providing each frame of the sequence of frames as input to the machine learning model comprises providing the at least two consecutive frames as input to the machine learning model.
5 . The computer-implemented method of claim 1 , wherein generating the set of frame documents comprises:
providing each frame of the sequence of frames as input to a machine learning model;
reading a feature map generated by the machine learning model based on at least one frame of the sequence of frames, wherein the feature map characterizes objects depicted in the at least one frame and semantic relationships among the objects; and
generating the frame document in accordance with the feature map.
6 . The computer-implemented method of claim 1 , the computer-implemented method further comprising:
providing the first frame and an end frame as input to the diffusion model, wherein the at least two consecutive frames comprise the end frame;
reading a second frame generated by the diffusion model;
determining whether the first frame and the second frame are equivalent; and
determining whether a number of frames greater than or equal to an insertion threshold have been generated.
7 . The computer-implemented method of claim 1 , the computer-implemented method further comprising:
inserting the first frame into the sequence of frames such that the video is modified.
8 . The computer-implemented method of claim 7 , the computer-implemented method further comprising:
synchronizing an audio representation of the commentary with the video in accordance with the consecutive subsequence; and
transmitting the video for presentation via a client computing platform.
9 . The computer-implemented method of claim 1 , wherein the machine learning model is a Feature Pyramid Network.
10 . The computer-implemented method of claim 1 , the computer-implemented method further comprising:
providing a prompt and a characterization of the consecutive subsequence as input to a generative machine learning model, wherein the prompt indicates a duration of a shortened commentary to be generated.
11 . The computer-implemented method of claim 1 , the computer-implemented method further comprising determining whether a length of the commentary is greater than a length of the consecutive subsequence, wherein the at least two consecutive frames of the sequence of frames and the embeddings are provided as input to the diffusion model responsive to determining the length of the commentary is greater than the length of the consecutive subsequence.
12 . A computer program product for syncing commentary with videos, the computer program product comprising:
one or more non-transitory computer-readable storage media;
program instructions stored on the one or more non-transitory computer-readable storage media to perform operations comprising:
reading a sequence of frames of a video;
generating a set of frame documents generated by the machine learning model based on the sequence of frames, wherein each frame document of the set of frame documents corresponds to at least one of the frames of the sequence of frames;
reading commentary associated with the video;
providing the commentary as input to a language model;
reading embeddings generated by the language model based on the commentary, wherein the embeddings characterize the commentary,
generating a commentary document in accordance with the embeddings;
determining a semantic distance between the commentary document and each of the frame documents
selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document;
identifying a consecutive subsequence of the sequence of frames associated with the subset;
providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and
reading a first frame generated by the diffusion model.
13 . The computer program product of claim 12 , wherein generating the set of frame documents comprises:
providing each frame of the sequence of frames as input to a machine learning model;
reading a feature map generated by the machine learning model based on at least one frame of the sequence of frames, wherein the feature map characterizes objects depicted in the at least one frame and semantic relationships among the objects; and
generating the frame document in accordance with the feature map.
14 . The computer program product of claim 12 , wherein the operations further comprise:
providing the first frame and an end frame as input to the diffusion model, wherein the at least two consecutive frames comprise the end frame;
reading a second frame generated by the diffusion model;
determining whether the first frame and the second frame are equivalent; and
determining whether a number of frames greater than or equal to an insertion threshold have been generated.
15 . The computer program product of claim 12 , wherein the operations further comprise:
inserting the first frame into the sequence of frames such that the video is modified.
16 . The computer program product of claim 12 , wherein the operations further comprise:
synchronizing an audio representation of the commentary with the video in accordance with the consecutive subsequence; and
transmitting the video for presentation via a client computing platform.
17 . A computer system for syncing commentary with videos, the computer system comprising:
a processor set;
one or more computer-readable storage media;
program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations comprising:
reading a sequence of frames of a video;
generating a set of frame documents based on the sequence of frames, wherein each frame document of the set of frame documents corresponds to at least one of the frames of the sequence of frames;
reading commentary associated with the video;
providing the commentary as input to a language model;
reading embeddings generated by the language model based on the commentary, wherein the embeddings characterize the commentary;
generating a commentary document in accordance with the embeddings;
determining a semantic distance between the commentary document and each of the frame documents;
selecting a subset of the set of frame documents having the lowest semantic distance to the commentary document;
identifying a consecutive subsequence of the sequence of frames associated with the subset;
providing at least two frames of the consecutive subsequence and the embeddings as input to a diffusion model; and
reading a first frame generated by the diffusion model.
18 . The computer system of claim 17 , wherein generating the set of frame documents comprises:
providing each frame of the sequence of frames as input to a machine learning model;
reading a feature map generated by the machine learning model based on at least one frame of the sequence of frames, wherein the feature map characterizes objects depicted in the at least one frame and semantic relationships among the objects; and
generating the frame document in accordance with the feature map.
19 . The computer system of claim 17 , wherein the operations further comprise:
providing the first frame and an end frame as input to the diffusion model, wherein the at least two consecutive frames comprise the end frame;
reading a second frame generated by the diffusion model;
determining whether the first frame and the second frame are equivalent; and
determining whether a number of frames greater than or equal to an insertion threshold have been generated.
20 . The computer program product of claim 17 , wherein the operations further comprise:
inserting the first frame into the sequence of frames such that the video is modified.