Facilitating video generation
Features described herein generally relate to content production. Particularly, the present disclosure relates to facilitating video generation. Using machine-learning models, a storyboard can be generated from an inspirational video, video attributes can be determined for the storyboard, editing scores and actions can be determined for candidate videos, candidate videos can be edited based on the editing scores and actions, and the edited candidate videos can be combined to generate a video.
1 . A method for facilitating video generation, the method comprising:
identifying a first video;
segmenting the first video into a plurality of video segments;
automatically generating a plurality of storyboard panels based on the plurality of video segments using a machine-learning model, each storyboard panel of the plurality of storyboard panels representing at least one video segment of the plurality of video segments;
automatically determining a plurality of video attributes by processing each video segment of the plurality of video segments, each video attribute of the plurality of video attributes corresponding to a storyboard panel of the plurality of storyboard panels;
identifying a plurality of second videos, each second video of the plurality of second videos corresponding to at least one video attribute of the plurality of video attributes;
determining a plurality of editing scores or a plurality of editing actions for the plurality of second videos by processing the plurality of second videos using another machine-learning model; and
editing the plurality of second videos based on the plurality of editing scores or the plurality of editing actions;
combining the second videos to generate a third video; and
outputting the third video.
2 . The method of claim 1 , wherein each storyboard panel of the plurality of storyboard panels includes a textual label describing at least one of content depicted in a video segment of the plurality of video segments and a video attribute of the plurality of video attributes.
3 . The method of claim 1 , wherein the plurality of video attributes is automatically determined.
4 . The method of claim 1 , wherein the plurality of editing scores or the plurality of editing actions includes
the plurality of editing scores.
5 . The method of claim 4 , wherein the other machine-learning model used to determine the plurality of editing scores or the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.
6 . The method of claim 1 , wherein the plurality of editing scores or the plurality of editing actions includes
the plurality of editing actions.
7 . The method of claim 6 , wherein the other machine-learning model used to determine the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.
8 . A system for facilitating video generation comprising:
one or more processors; and
one or more memories, the one or more memories storing instructions which, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
identifying a first video;
segmenting the first video into a plurality of video segments;
automatically generating a plurality of storyboard panels based on the plurality of video segments using a machine-learning model, each storyboard panel of the plurality of storyboard panels representing at least one video segment of the plurality of video segments;
automatically determining a plurality of video attributes by processing each video segment of the plurality of video segments, each video attribute of the plurality of video attributes corresponding to a storyboard panel of the plurality of storyboard panels;
identifying a plurality of second videos, each second video of the plurality of second videos corresponding to at least one video attribute of the plurality of video attributes;
determining a plurality of editing scores or a plurality of editing actions for the plurality of second videos by processing the plurality of second videos using another machine-learning model; and
editing the plurality of second videos based on the plurality of editing scores or the plurality of editing actions;
combining the second videos to generate a third video; and
outputting the third video.
9 . The system of claim 8 , wherein each storyboard panel of the plurality of storyboard panels includes a textual label describing at least one of content depicted in a video segment of the plurality of video segments and a video attribute of the plurality of video attributes.
10 . The system of claim 8 , wherein the plurality of video attributes is automatically determined.
11 . The system of claim 8 , wherein the plurality of editing scores or the plurality of editing actions includes
the plurality of editing scores.
12 . The system of claim 11 , wherein the other machine-learning model used to determine the plurality of editing scores or the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.
13 . The system of claim 8 , wherein the plurality of editing scores or the plurality of editing actions includes
the plurality of editing actions.
14 . The system of claim 13 , wherein the other machine-learning model used to determine the plurality of editing actions or the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.
15 . One or more non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:
identifying a first video;
segmenting the first video into a plurality of video segments;
automatically generating a plurality of storyboard panels based on the plurality of video segments using a machine-learning model, each storyboard panel of the plurality of storyboard panels representing at least one video segment of the plurality of video segments;
automatically determining a plurality of video attributes by processing each video segment of the plurality of video segments, each video attribute of the plurality of video attributes corresponding to a storyboard panel of the plurality of storyboard panels;
identifying a plurality of second videos, each second video of the plurality of second videos corresponding to at least one video attribute of the plurality of video attributes;
determining a plurality of editing scores or a plurality of editing actions for the plurality of second videos by processing the plurality of second videos using another machine-learning model; and
editing the plurality of second videos based on the plurality of editing scores or the plurality of editing actions;
combining the second videos to generate a third video; and
outputting the third video.
16 . The one or more non-transitory computer-readable media of claim 15 , wherein each storyboard panel of the plurality of storyboard panels includes a textual label describing at least one of content depicted in a video segment of the plurality of video segments and a video attribute of the plurality of video attributes.
17 . The one or more non-transitory computer-readable media of claim 15 , wherein the plurality of video attributes is automatically determined using the other machine-learning model.
18 . The one or more non-transitory computer-readable media of claim 15 , wherein the plurality of editing scores or the plurality of editing actions includes
the plurality of editing scores.
19 . The one or more non-transitory computer-readable media of claim 18 , wherein the plurality of editing scores or the plurality of editing actions includes
the plurality of editing actions.
20 . The one or more non-transitory computer-readable media of claim 19 , wherein the other machine-learning model used to determine the plurality of editing scores and used to determine the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.