Script based video effects for live video
In various examples, a video effect is displayed in a live video stream in response to determining a portion of an audio stream of the live video stream that corresponds to a text segment of a script associated with the video effect. For example, during presentation of the script, the audio stream is obtained to determine if a portion of the audio stream corresponds to the text segment.
1 . A method comprising:
obtaining a script and a video effect associated with a text segment of the script;
obtaining an audio stream corresponding to a live video stream, the live video stream capturing a live performance of the script by a user using a user device;
determining a portion of the audio stream corresponds to the text segment by at least matching a set of words spoken by the user in the portion of the audio stream to a portion of the text segment of the script;
responsive to determining the portion of the audio stream corresponds to the text segment, causing the video effect to be displayed in the live video stream; and
causing the live video stream including the video effect to be presented to at least one other user device during a real-time presentation.
2 . The method of claim 1 , wherein the method further comprises advancing a cursor location within a teleprompter displaying the script on the user device based on a location within the script.
3 . The method of claim 2 , wherein the method further comprises determining the location within the script by at least matching a sliding window including a first plurality of words obtained from the audio stream to a second plurality of words included in the script.
4 . The method of claim 3 , wherein the sliding window further comprises three words.
5 . The method of claim 3 , wherein the method further comprises causing a notification to be displayed in a presentation user interface indicating that the location with the script is undetermined based on a second portion of the audio stream and the script.
6 . The method of claim 1 , wherein the method further comprises causing a script authoring interface to be displayed by a user device enabling a user to provide the script and the video effect to apply to the text segment of the script.
7 . The method of claim 6 , wherein the method further comprises obtaining, from an input device associated with the user device a set of inputs to the script authoring interface, the set of inputs including at least one of: a set of words to be included in the script, a first selection of the text segment in the script, and a second selection of the video effect to be applied during the live video stream.
8 . A non-transitory computer-readable medium storing executable instructions embodied thereon, which, when executed by a processing device, cause the processing device to perform operations comprising:
obtaining a script index corresponding to a script and a video effect to be applied to a video stream in response to a text segment included in the script, the script index including words in the script and location information corresponding to the words in the script, where the video stream captures a user presenting the script;
obtaining an audio stream associated with the video stream;
determining a location within the script based on the script index and a portion of the audio stream by at least matching a first plurality of words within the script to a second plurality of words obtained from the portion of the audio stream; and
applying the video effect to the video stream as a result of the location corresponding to the text segment included in the script.
9 . The medium of claim 8 , wherein determining the location within the script further comprises advancing a cursor location within a teleprompter indicating the location with a presentation interface.
10 . The medium of claim 9 , wherein the cursor location further comprises the location information included in the script index corresponding to a word in the script associated with the location.
11 . The medium of claim 8 , wherein the script index further comprises a key-value store where keys of the key-value store correspond to the location information and values of the key-value store correspond to the words of the script.
12 . The medium of claim 8 , wherein determining the location within the script further comprises obtaining, from a model, a set of locations and corresponding probabilities, where inputs to the model include the script and a transcript generated based on the audio stream.
13 . The medium of claim 8 , wherein the computer-readable medium further stores executable instructions that cause the processing device to perform operations comprising:
determining a cadence associated with a user speaking in the audio stream; and
advancing a cursor location within a teleprompter indicating the location with a presentation interface based on the cadence.
14 . The medium of claim 8 , wherein the location further comprises at least one of a word in the script, a sentence in the script, and a paragraph in the script.
15 . A system comprising:
a memory component; and
a processing device coupled to the memory component, the processing device to perform operations comprising:
obtaining a script index including a set of words in a script and a set of locations corresponding to words of the set of words in the script;
obtaining a video stream and an audio stream, the video stream and the audio stream are captured by a user device and include a live performance of the script by the user;
determining a location of the set of locations included in the script index based on a portion of the audio stream by at least matching a first subset of words of the set of words in the script to a second set of words extracted from the portion of the audio stream; and
as a result of determining the location corresponds to a text segment, applying a video effect to the video stream.
16 . The system of claim 15 , wherein the script index further comprises a key-value store, wherein locations of the set of locations correspond to keys of the key-value store and words of the set of words correspond to values of the key-value store.
17 . The system of claim 16 , wherein determining the location further comprises modifying a cursor location to indicate a key of the key-value store associated with the location.
18 . The system of claim 17 , the processing device further performs operations comprising causing a teleprompter to be updated based on the cursor location.
19 . The system of claim 15 , wherein determining the location further comprises determining a first set of words obtained from the portion of the audio stream matches a second set of words of the set of words in the script.
20 . The medium of claim 8 , wherein the second plurality of words are converted from the portion of the audio stream.