Generative sound effects for video editing
A method, apparatus, non-transitory computer readable medium, and system for generating sound effects for a video includes obtaining a selection input indicating an element of a video. Embodiments then generate, using an audio generation model, a synthetic audio clip based on the selection input. Embodiments subsequently generate a multimedia file including the video and the synthetic audio clip.
1 . A method comprising:
obtaining a video and a selection input indicating a portion of the video;
generating, using an audio generation model, a synthetic audio clip corresponding to the portion of the video by performing a diffusion denoising process; and
generating a multimedia file including the video and the synthetic audio clip, wherein the multimedia file includes the synthetic audio clip at the portion of the video indicated by the selection input.
2 . The method of claim 1 , wherein obtaining the selection input comprises:
displaying a video timeline interface for the video, wherein the selection input is obtained via the video timeline interface.
3 . The method of claim 1 , further comprising:
obtaining a text input describing an audio element, wherein the synthetic audio clip is generated based on the text input.
4 . The method of claim 3 , wherein obtaining the text input comprises:
generating the text input based on the portion of the video.
5 . The method of claim 1 , further comprising:
identifying a set of frames including the portion of the video, wherein the synthetic audio clip is generated based on the set of frames.
6 . The method of claim 1 , further comprising:
obtaining a semantic adjustment input indicating a volume for the portion of the video; and
generating an additional synthetic audio clip based on the semantic adjustment input.
7 . The method of claim 1 , further comprising:
obtaining a sound dynamics input, wherein the synthetic audio clip is generated based on the sound dynamics input.
8 . The method of claim 1 , further comprising:
generating an additional synthetic audio clip, wherein the multimedia file includes the synthetic audio clip and the additional synthetic audio clip.
9 . A non-transitory computer readable medium storing code for sound processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
displaying a video timeline interface for a video;
obtaining, via the video timeline interface, a selection input indicating a portion of the video;
generating, using an audio generation model, a synthetic audio clip at to the portion of the video indicated by the selection input by performing a diffusion denoising process.
10 . The non-transitory computer readable medium of claim 9 , code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
determining a foreground subject in a frame of the video upon obtaining the selection input.
11 . The non-transitory computer readable medium of claim 9 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
obtaining a text input describing an audio element, wherein the synthetic audio clip is generated based on the text input.
12 . The non-transitory computer readable medium of claim 11 , wherein obtaining the text input comprises:
generating the text input based on the portion of the video.
13 . The non-transitory computer readable medium of claim 9 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
identifying a set of frames including the portion of the video, wherein the synthetic audio clip is generated based on the set of frames.
14 . The non-transitory computer readable medium of claim 9 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
obtaining a semantic adjustment input indicating a volume for the portion of the video; and
generating an additional synthetic audio clip based on the semantic adjustment input.
15 . The non-transitory computer readable medium of claim 9 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
generating a multimedia file including the video and the synthetic audio clip.
16 . A system for sound processing, comprising:
a memory component;
a processing device coupled to the memory component, the processing device configured to perform operations comprising:
obtaining a video and a selection input indicating a portion of the video;
generating, using an audio generation model, a synthetic audio clip corresponding to the portion of the video by performing a diffusion denoising process; and
generating a multimedia file including the video and the synthetic audio clip, wherein the multimedia file includes the synthetic audio clip at the portion of the video indicated by the selection input.
17 . The system of claim 16 , the system further comprising:
displaying a video timeline interface for the video, wherein the selection input is obtained via the video timeline interface.
18 . The system of claim 16 , the system further comprising:
obtaining a text input describing an audio element, wherein the synthetic audio clip is generated based on the text input.
19 . The system of claim 16 , the system further comprising:
identifying a set of frames including the portion of the video, wherein the synthetic audio clip is generated based on the set of frames.
20 . The system of claim 16 , the system further comprising:
obtaining a semantic adjustment input indicating a volume for the portion of the video; and
generating an additional synthetic audio clip based on the semantic adjustment input.