TEXT-DRIVEN EDITOR FOR AUDIO AND VIDEO ASSEMBLY
The disclosed technology is a system and computer-implemented method for assembling and editing a video program from spoken words or soundbites. The disclosed technology imports source audio/video clips and any of multiple formats. Spoken audio is transcribed into searchable text. The text transcript is synchronized to the video track by timecode markers. Each spoken word corresponds to a timecode marker, which in turn corresponds to a video frame or frames. Using word processing operations and text editing functions, a user selects video segments by selecting corresponding transcribed text segments. By selecting text and arranging that text, a corresponding video program is assembled. The selected video segments are assembled on a timeline display in any chosen order by the user. The sequence of video segments may be re-ordered and edited, as desired, to produce a finished video program for export.
1 . A computer-implemented method comprising:
generating a transcript map that associates video frames of a digital video with words of a transcription of an audio track of the digital video;
detecting an indication of a keyword within the transcription of the audio track of the digital video;
identifying one or more instances of the keyword within the transcription of the audio track of the digital video; and
removing, from the digital video and utilizing the transcript map, one or more video frames corresponding to the one or more instances of the keyword within the transcription of the audio track of the digital video.
2 . The computer-implemented method as recited in claim 1 , further comprising generating the transcription of the audio track of the digital video by:
determining time codes for increments of the audio track based on metadata of the digital video;
generating a transcription of the audio track; and
assigning the time codes for the increments of the audio track to corresponding increments of the transcription of the audio track.
3 . The computer-implemented method as recited in claim 2 , wherein generating the transcript map comprises:
determining a start time code and an end time code for every word in the transcription of the audio track; and
generating the transcript map comprising the words of the transcription of the audio track correlated with corresponding start time codes and end time codes.
4 . The computer-implemented method as recited in claim 3 , wherein identifying the one or more instances of the keyword within the transcription of the audio track of the digital video comprises:
identifying the keyword within the transcript map;
identifying one or more pairs of start time codes and end time codes correlated with the keyword within the transcript map; and
generating a listing of the one or more pairs of start time codes and end time codes.
5 . The computer-implemented method as recited in claim 4 , wherein removing the one or more video frames corresponding to the one or more instances of the keyword within the transcription of the audio track of the digital video comprises, for each of the one or more pairs of start time codes and end time codes in the listing:
identifying a first video frame in the digital video corresponding to the start time code;
identifying a second video frame in the digital video corresponding to the end time code; and
removing, from the digital video, video frames in between the first video frame and the second video frame.
6 . The computer-implemented method as recited in claim 1 , wherein detecting the indication of the keyword within the transcription of the audio track of the digital video comprises at least one of:
detecting a user selection of a word within a display comprising the transcription of the audio track,
detecting a user input of the keyword in a text box associated with the display comprising the transcription of the audio track, or
utilizing a machine learning model to automatically detect the keyword based on user trends.
7 . The computer-implemented method as recited in claim 1 , further comprising:
detecting an indication of an additional keyword within the transcription of the audio track of the digital video;
determining one or more instances of the additional keyword within the transcription of the audio track of the digital video; and
removing, from the digital video and utilizing the transcript map, one or more video frames corresponding to the one or more instances of the additional keyword within the transcription of the audio track of the digital video.
8 . A system comprising:
at least one physical processor; and
physical memory comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform acts comprising:
generating a transcript map that associates video frames of a digital video with words of a transcription of an audio track of the digital video;
detecting an indication of a keyword within the transcription of the audio track of the digital video;
identifying one or more instances of the keyword within the transcription of the audio track of the digital video; and
removing, from the digital video and utilizing the transcript map, one or more video frames corresponding to the one or more instances of the keyword within the transcription of the audio track of the digital video.
9 . The system as recited in claim 8 , further comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform an act comprising generating the transcription of the audio track of the digital video by:
determining time codes for increments of the audio track based on metadata of the digital video;
generating a transcription of the audio track; and
assigning the time codes for the increments of the audio track to corresponding increments of the transcription of the audio track.
10 . The system as recited in claim 9 , further comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform the act comprising generating the transcript map by:
determining a start time code and an end time code for every word in the transcription of the audio track; and
generating the transcript map comprising the words of the transcription of the audio track correlated with corresponding start time codes and end time codes.
11 . The system as recited in claim 10 , further comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform the act comprising identifying the one or more instances of the keyword within the transcription of the audio track of the digital video by:
identifying the keyword within the transcript map;
identifying one or more pairs of start time codes and end time codes correlated with the keyword within the transcript map; and
generating a listing of the one or more pairs of start time codes and end time codes.
12 . The system as recited in claim 11 , further comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform the act comprising removing the one or more video frames corresponding to the one or more instances of the keyword within the transcription of the audio track of the digital video by, for each of the one or more pairs of start time codes and end time codes in the listing:
identifying a first video frame in the digital video corresponding to the start time code;
identifying a second video frame in the digital video corresponding to the end time code; and
removing, from the digital video, video frames in between the first video frame and the second video frame.
13 . The system as recited in claim 8 , further comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform the act comprising detecting the indication of the keyword within the transcription of the audio track of the digital video by at least one of:
detecting a user selection of a word within a display comprising the transcription of the audio track,
detecting a user input of the keyword in a text box associated with the display comprising the transcription of the audio track, or
utilizing a machine learning model to automatically detect the keyword based on user trends.
14 . The system as recited in claim 8 , further comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform acts comprising:
detecting an indication of an additional keyword within the transcription of the audio track of the digital video;
determining one or more instances of the additional keyword within the transcription of the audio track of the digital video; and
removing, from the digital video and utilizing the transcript map, one or more video frames corresponding to the one or more instances of the additional keyword within the transcription of the audio track of the digital video.
15 . A non-transitory computer-readable medium comprising one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to perform acts comprising:
generating a transcript map that associates video frames of a digital video with words of a transcription of an audio track of the digital video;
detecting an indication of a keyword within the transcription of the audio track of the digital video;
identifying one or more instances of the keyword within the transcription of the audio track of the digital video; and
removing, from the digital video and utilizing the transcript map, one or more video frames corresponding to the one or more instances of the keyword within the transcription of the audio track of the digital video.
16 . The non-transitory computer-readable medium as recited in claim 15 , further comprising computer-executable instructions that, when executed by the at least one processor of the computing device, cause the computing device to perform an act comprising generating the transcription of the audio track of the digital video by:
determining time codes for increments of the audio track based on metadata of the digital video;
generating a transcription of the audio track; and
assigning the time codes for the increments of the audio track to corresponding increments of the transcription of the audio track.
17 . The non-transitory computer-readable medium as recited in claim 16 , further comprising computer-executable instructions that, when executed by the at least one processor of the computing device, cause the computing device to perform the act comprising generating the transcript map by:
determining a start time code and an end time code for every word in the transcription of the audio track; and
generating the transcript map comprising the words of the transcription of the audio track correlated with corresponding start time codes and end time codes.
18 . The non-transitory computer-readable medium as recited in claim 17 , further comprising computer-executable instructions that, when executed by the at least one processor of the computing device, cause the computing device to perform the act comprising identifying the one or more instances of the keyword within the transcription of the audio track of the digital video by:
identifying the keyword within the transcript map;
identifying one or more pairs of start time codes and end time codes correlated with the keyword within the transcript map; and
generating a listing of the one or more pairs of start time codes and end time codes.
19 . The non-transitory computer-readable medium as recited in claim 18 , further comprising computer-executable instructions that, when executed by the at least one processor of the computing device, cause the computing device to perform the act comprising removing the one or more video frames corresponding to the one or more instances of the keyword within the transcription of the audio track of the digital video by, for each of the one or more pairs of start time codes and end time codes in the listing:
identifying a first video frame in the digital video corresponding to the start time code;
identifying a second video frame in the digital video corresponding to the end time code; and
removing, from the digital video, video frames in between the first video frame and the second video frame.
20 . The non-transitory computer-readable medium as recited in claim 15 , further comprising computer-executable instructions that, when executed by the at least one processor of the computing device, cause the computing device to perform the act comprising detecting the indication of the keyword within the transcription of the audio track of the digital video by at least one of:
detecting a user selection of a word within a display comprising the transcription of the audio track,
detecting a user input of the keyword in a text box associated with the display comprising the transcription of the audio track, or utilizing a machine learning model to automatically detect the keyword based on user trends.