Systems and methods for automating video editing
Provided are systems and methods for automatic video processing that employ machine learning models to process input video and understand user video content in a semantic and cultural context. This recognition enables the processing system to recognize interesting temporal events, and build narrative video sequences automatically, for example, by linking or interleaving temporal events or other content with film-based categorizations. In further embodiments, the implementation of the processing system is adapted to mobile computing platforms which can be distributed as an “app” within various app stores. In various example, the mobile apps turn everyday users into professional videographers. In further embodiments, music selection and dialog based editing can likewise be automated via machine learning models to create dynamic and professional quality video segments.
1 . A video processing system, comprising:
a video processing component, executed by at least one processor, configured to perform operations comprising:
accepting a first user sourced video input generated by the first user;
executing a first machine learning process to perform operations comprising:
analyzing the user sourced video input and decomposing the user sourced video input into video segments comprising the user sourced video input, the decomposing into the video segments being based, at least in part, on determining importance of sections and content within the user sourced video input, the video segments comprising trimmed clip segments having variable durations;
transforming the video segments into a semantic embedding space comprising feature vectors of the video segments in a multi-dimensional space;
classifying the transformed video segments into at least one of cinematic categories or spatial cinematic layout categories based on determining similarity between classified video segments having associated cinematic categories and spatial cinematic layout categories and the transformed video segments;
in response to selection of an automatic editing function, editing automatically at least a first or second video segment of the first user sourced video input, the edit execution including at least one of altering the duration of the first or second video segment and introducing at least one visual effect into the first or second video segment based, at least in part, on the cinematic categories or the spatial cinematic layout categories and a machine learning model trained or fine-tuned on approved edits, the edits including altering the duration and introducing at least one visual effect to one or more video segments; and
generating a rough-cut video output including a sequence of video including edited versions of the first or second video segment and at least some of the plurality of video segments from the first user sourced video.
2 . The system of claim 1 , wherein the operations further comprise:
automatically identifying using a second machine learning process a narrative goal based on analysis of the first user sourced video input; and
defining a new sequencing of the first user sourced video to convey the narrative goal based on a third machine learning algorithm trained on film-based categorizations of video segments, the film-based categorizations including at least cinematic style.
3 . The system of claim 1 , wherein the video processing component includes at least a first neural network configured to transform the first user sourced video input into a semantic embedding space.
4 . The system of claim 3 , wherein the first neural network comprises a convolutional neural network.
5 . The system of claim 3 , wherein the first neural network is configured to classify user video into visual concept categories.
6 . The system of claim 4 , wherein the video processing component further comprises a second neural network configured to determine a narrative goal associated with the first user sourced video input or the sequence of video to be displayed.
7 . The system of claim 6 , wherein the second neural network comprises a long term short term memory recurrent network.
8 . The system of claim 1 , further comprising a second neural network configured to classify visual beats within user sourced video.
9 . The system of claim 8 , wherein the operations further comprise automatically selecting at least one soundtrack for the first user sourced video input.
10 . The system of claim 1 , wherein the semantic embedding space comprises respective numerical representation of the respective video segments in a multiple dimensioned space, and the numerical values from the sematic embedding space are input into a neural network to output a matching film idiom from the neural network.
11 . A computer implemented method for automatic video processing, the method comprising:
accepting, by at least one processor, a first user sourced video input generated by the first user;
analyzing, by the at least one processor, the user sourced video input and decomposing the user sourced video input into video segments comprising the user sourced video input, the decomposing being based, at least in part, on determining importance of sections and content within the user sourced video input, the video segments comprising trimmed clip segments having variable durations;
transforming, by the at least one processor, the video segments into a semantic embedding space comprising feature vectors of the video segments in a multi-dimensional space;
classifying, by the at least one processor, the transformed video segments into at least one of contextual cinematic categories or spatial cinematic layout categories based on determining similarity between classified video segments having associated cinematic categories and spatial cinematic layout categories and the transformed video segments;
editing, by the at least one processor, automatically at least a first or second video segment of the first user sourced video input in response to selection of an automatic editing function, the editing including at least one of altering the duration of the first or second video segment and introducing at least one visual effect into the first or second video segment based at least in part on the cinematic categories or the spatial cinematic layout categories and a machine learning model trained on reviewed or approved edits, the edits including altering the duration and introducing at least one visual effect to one or more video segments; and
generating, by the at least one processor, a rough-cut video output including a sequence of video including edited versions of first or second video segment and at least some of the plurality of video segments from the first user sourced video.
12 . The method of claim 11 , wherein the method further comprises:
automatically identifying, by the at least one processor, a narrative goal using a second machine learning process based on analysis of the first user sourced video input; and
defining, by the at least one processor, a new sequence of the first user sourced video input to convey the narrative goal based on a third machine learning algorithm trained on film-based categorizations of video segments, the film-based categorizations including at least cinematic style.
13 . The method of claim 11 , wherein the method further comprises executing at least a first neural network configured to transform the first user sourced video input into a semantic embedding space.
14 . The method of claim 13 , wherein the first neural network comprises a convolutional neural network.
15 . The method of claim 14 , wherein the method further comprises determining, by a second neural network, a narrative goal associated with the first user sourced video or the sequence of video to be displayed.
16 . The method of claim 15 , wherein the second neural network comprises a long term short term memory recurrent network.
17 . The method of claim 13 , wherein the method further comprises classifying user video into visual concept categories with the first neural network.
18 . The method of claim 11 , wherein the method further comprises classifying, by a third neural network, visual beats within the first user sourced video.
19 . The method of claim 18 , wherein the method further comprises automatically selecting at least one soundtrack for the user sourced video.
20 . The method of claim 11 , wherein the semantic embedding space comprises respective numerical representations of the respective video segments in a multiple dimensioned space, and the method includes processing at least some of the numerical values from the sematic embedding space as input into a neural network to output a matching film idiom from the neural network.