Generation of manifests for image presentation during video presentation
Techniques for generating manifests for image presentation during video presentation are described herein. In an example, a system generates an input to an artificial intelligence (AI) model on a set of frames of video content. The system receives, based on the input, an output of the AI model indicating a set of images corresponding to the set of frames that are to be presented during a scrubbing operation associated with a presentation of the video content. The output further indicates durations for individual images of the set of images. A first image of the set of images is to be presented for a first duration that is different from a second duration that a second image is to be presented. The system generates a manifest including the video content, an indication of the set of images, and the durations. The system stores the manifest in a storage location.
1 . A computer-implemented method, comprising:
receiving a set of frames of video content;
generating an input to an artificial intelligence model based at least in part on the set of frames;
receiving, based at least in part on the input, an output of the artificial intelligence model indicating a set of images corresponding to the set of frames that are to be presented during a scrubbing operation associated with a presentation of the video content, wherein the output further indicates durations for individual images of the set of images, and wherein a first image of the set of images is to be presented for a first duration that is different from a second duration that a second image of the set of images is to be presented;
generating a manifest including the video content, an indication of the set of images, and the durations;
receiving, via a user interface of a user device, a user input to initiate a beginning of the scrubbing operation by adjusting a timeline feature of the user interface during the presentation of the video content at the user device;
determining a timestamp of the video content associated with the beginning of the scrubbing operation;
determining, based at least in part on the manifest, that the timestamp corresponds to the first image; and
causing presentation of the first image in connection with the scrubbing operation and for the first duration relative to the timeline feature during the presentation of the video content.
2 . The computer-implemented method of claim 1 , wherein the manifest includes a first portion associated with the video content, a second portion associated with interactive content that is to be presented in associated with the video content during the presentation of the video content, and a third portion corresponding to targeted content that is to be presented in associated with the video content during the presentation of the video content.
3 . The computer-implemented method of claim 2 , wherein the set of images is a first set of images, wherein the first portion is associated with the first set of images, the second portion is associated with a second set of images that are to be presented during the scrubbing operation associated with the interactive content, and the third portion is associated with a third set of images that are to be presented during the scrubbing operation associated with the targeted content.
4 . The computer-implemented method of claim 1 , further comprising:
generating a mapping between a timeline of the video content and the set of images based at least in part on the durations;
determining a point within the timeline corresponding to the timestamp; and
determining that the point corresponds to the first image based at least in part on the mapping.
5 . One or more non-transitory computer-readable media comprising computer-executable instructions that, when executed by one or more computer systems, cause the one or more computer systems to perform operations comprising:
generating an input to an artificial intelligence model based at least in part on a set of frames of video content;
receiving, based at least in part on the input, an output of the artificial intelligence model indicating a set of images corresponding to the set of frames that are to be presented during a scrubbing operation associated with presentation of the video content, and wherein the output further indicates durations for individual images of the set of images, and wherein a first image of the set of images is to be presented for a first duration that is different from a second duration that a second image of the set of images is to be presented, and wherein the scrubbing operation involves adjusting a timeline feature of a user interface of a user device;
generating a manifest including the video content, an indication of the set of images, and the durations; and
storing the manifest in a storage location such that the set of images are configured to be presented in connection with the scrubbing operation and for the durations relative to the timeline feature during presentation of the video content based at least in part on the manifest.
6 . The one or more non-transitory computer-readable media of claim 5 , wherein the operations further comprise:
receiving, from the user device, information that represents a user selection to initiate a beginning of the scrubbing operation during the presentation of the video content at the user device;
determining a timestamp of the video content associated with the beginning of the scrubbing operation;
determining, based at least in part on the manifest, that the timestamp corresponds to the first image; and
causing presentation of the first image in connection with the scrubbing operation during the presentation of the video content.
7 . The one or more non-transitory computer-readable media of claim 5 , wherein the indication of the set of images indicates a location of an MP4 file that includes the set of images, and wherein the manifest further indicates a byte range of the MP4 file corresponding to first image.
8 . The one or more non-transitory computer-readable media of claim 7 , wherein the operations further comprise:
receiving, from the user device, information that represents a user selection to initiate a beginning of the scrubbing operation during the presentation of the video content at the user device;
determining a timestamp of the video content associated with the beginning of the scrubbing operation; and
retrieving, in response to determining the timestamp corresponds to the first image, the first image from the location of the MP4 file based at least in part on the byte range for the first image.
9 . The one or more non-transitory computer-readable media of claim 5 , wherein the operations further comprise:
generating a mapping between a timeline of the video content and the set of images based at least in part on the durations.
10 . The one or more non-transitory computer-readable media of claim 9 , wherein the operations further comprise:
receiving, from the user device, information that represents a user selection to initiate a beginning of the scrubbing operation during the presentation of the video content at the user device;
determining a timestamp of the video content associated with the beginning of the scrubbing operation;
determining a point within the timeline corresponding to the timestamp; and
determining that the point corresponds to the first image based at least in part on the mapping.
11 . The one or more non-transitory computer-readable media of claim 5 , wherein the operations further comprise:
determining an updated set of images that are to be presented during the scrubbing operation;
determining updated durations for presenting individual images of the updated set of images; and
generating an updated manifest including the video content, an updated indication of the updated set of images, and the updated durations.
12 . The one or more non-transitory computer-readable media of claim 5 , wherein the manifest includes a first portion associated with the video content, a second portion associated with interactive content that is to be presented in associated with the video content during the presentation of the video content, and a third portion corresponding to targeted content that is to be presented in associated with the video content during the presentation of the video content.
13 . A system, comprising:
one or more memories configured to store computer-executable instructions;
one or more processors configured to access the one or more memories and execute the computer-executable instructions to at least:
receive a request to initiate a beginning of a scrubbing operation by adjusting a timeline feature of a user interface during a presentation of video content at a user device;
determine a timestamp of the video content associated with the beginning of the scrubbing operation;
receive a manifest associated with the video content, wherein the manifest indicates a set of images corresponding to a set of frames of the video content that are to be presented during the scrubbing operation associated with the presentation of the video content, wherein the manifest further indicates durations for individual images of the set of images, and wherein a first image of the set of images is to be presented for a first duration that is different from a second duration that a second image of the set of images is to be presented;
determine, based at least in part on the manifest, that the timestamp corresponds to the first image of the set of images; and
cause presentation of the first image in connection with the scrubbing operation and for the first duration relative to the timeline feature during the presentation of the video content.
14 . The system of claim 13 , wherein the indication of the set of images indicates a first location of a first JPEG image corresponding to the first image.
15 . The system of claim 14 , wherein the indication of the set of images further indicates a second location of a second JPEG image corresponding to the second image, wherein the first location is different from the second location.
16 . The system of claim 13 , wherein the set of images are stored in tiles, and wherein the indication of the set of images indicates a location of the tiles.
17 . The system of claim 16 , wherein the one or more processors are configured to access the one or more memories and execute additional computer-executable instructions to at least:
generate a mapping between a timeline of the video content and the tiles based at least in part on the durations.
18 . The system of claim 17 , wherein the one or more processors are configured to access the one or more memories and execute additional computer-executable instructions to at least:
determine a point within the timeline corresponding to the timestamp; and
determine that the point corresponds to the first image based at least in part on the mapping.
19 . The system of claim 13 , wherein the indication of the set of images indicates a location of a BIF file that includes the set of images, and wherein the manifest further indicates a byte range of the BIF file corresponding to first image.
20 . The system of claim 19 , wherein the one or more processors are configured to access the one or more memories and execute additional computer-executable instructions to at least:
retrieve, in response to determining the timestamp corresponds to the first image, the first image from the location of the BIF file based at least in part on the byte range for the first image.