AUTOMATED VOICEOVER MIXING AND COMPONENTS THEREFOR
Voiceover mixing is provided by receiving a voiceover file and a music file. The voiceover file is audio processed to generate a processed voiceover file and a music file is audio processed to generate a processed music file. The processed voiceover file and the processed music file are weight summed to generate a weighted combination of the processed voiceover file and the processed music file. Single band compressing is performed on the weighted combination. A creative file that contains a compressed and weighted combination of the processed voiceover file and the processed music file is then generated.
1 . A computer-implemented method for voiceover mixing, comprising:
receiving a voiceover file and a music file;
audio processing a voiceover file to generate a processed voiceover file;
audio processing a music file to generate a processed music file;
weighted summing the processed voiceover file and the processed music file to generate a weighted combination of the processed voiceover file and the processed music file;
single band compressing the weighted combination; and
generating a creative file containing a compressed and weighted combination of the processed voiceover file and the processed music file.
2 . The computer-implemented method for voiceover mixing according to claim 1 , further comprising:
measuring the energy level of the voice file within a frequency range; and
filtering the frequency range if the energy level exceeds a predetermined threshold.
3 . The computer-implemented method for voiceover mixing according to claim 1 :
wherein audio processing the voiceover file includes normalizing, compressing and equalizing the voiceover file;
wherein audio processing the music file includes normalizing, compressing and equalizing the music file; and
wherein the voiceover file and the music file are normalized, compressed and equalized asynchronously.
4 . The computer-implemented method for voiceover mixing according to claim 1 , further comprising:
storing, in a voice activations store, a curve corresponding to when a voice is present in the voiceover file.
5 . The computer-implemented method for voiceover mixing according to claim 1 , further comprising:
setting an advertisement duration time;
setting a start time for the voiceover file;
trimming the music file according to the advertisement duration time; and
mixing the voiceover file and the music file according to the start time and the advertisement duration time.
6 . The computer-implemented method for voiceover mixing according to claim 1 , further comprising:
generating a script;
converting the script to voice content; and
saving the voice content in the voiceover file.
7 . The computer-implemented method for voiceover mixing according to claim 1 , further comprising:
mapping each track in a library of tracks to a point in an embedding space;
computing an acoustic embedding based on a query track within the embedding space;
obtaining a track from the library of tracks with acoustically similar content; and
saving the track from the library of tracks with acoustically similar content in the music file.
8 . A system for voiceover mixing, comprising:
a voice processor operable to:
receive a voiceover file, and
generate a processed voiceover file from the voiceover file;
a music processor operable to:
receive a music file, and
generate a processed music file from the music file; and
a mixing processor operable to:
weight sum the processed voiceover file and the processed music file to generate a weighted combination of the processed voiceover file and the processed music file,
single band compress the weighted combination, and
generate a creative file containing a compressed and weighted combination of the processed voiceover file and the processed music file.
9 . The system for voiceover mixing according to claim 8 , further comprising:
the voice processor further operable to:
measure the energy level of the voice file within a frequency range; and
filter the frequency range if the energy level exceeds a predetermined threshold.
10 . The system for voiceover mixing according to claim 8 ,
the voice processor further operable to normalize, compress and equalize the voiceover file; and
the music processor further operable to normalize, compress and equalize the music file,
wherein the voiceover file and the music file are normalized, compressed and equalized asynchronously.
11 . The system for voiceover mixing according to claim 8 , further comprising:
a voice activations store operable to store a curve corresponding to when a voice is present in the voiceover file.
12 . The system for voiceover mixing according to claim 8 , further comprising:
an advertisement store operable to store an advertisement duration time;
the voice processor further operable to set a start time for the voiceover file;
the music processor further operable to trim the music file according to the advertisement duration time; and
the mixing processor further operable to mix the voiceover file and the music file according to the start time and the advertisement duration time.
13 . The system for voiceover mixing according to claim 8 , further comprising:
a script processor:
operable to generate a script from at least one script section;
a text to voice processor operable to convert the script to voice content; and
a voiceover store configured to save the voice content in the voiceover file.
14 . The system for voiceover mixing according to claim 8 , further comprising:
a background music search processor operable to:
map each track in a library of tracks to a point in an embedding space;
compute an acoustic embedding based on a query track within the embedding space;
obtain a track from the library of tracks with acoustically similar content; and
save the track from the library of tracks with acoustically similar content in the music file.
15 . A non-transitory computer-readable medium having stored thereon one or more sequences of instructions for causing one or more processors to perform:
receiving a voiceover file and a music file;
audio processing a voiceover file to generate a processed voiceover file;
audio processing a music file to generate a processed music file;
weighted summing the processed voiceover file and the processed music file to generate a weighted combination of the processed voiceover file and the processed music file;
single band compressing the weighted combination; and
generating a creative file containing a compressed and weighted combination of the processed voiceover file and the processed music file.
16 . The computer-readable medium of claim 15 , further having stored thereon a sequence of instructions for causing the one or more processors to perform:
measuring the energy level of the voice file within a frequency range; and
filtering the frequency range if the energy level exceeds a predetermined threshold.
17 . The computer-readable medium of claim 15 :
wherein audio processing the voiceover file includes normalizing, compressing and equalizing the voiceover file; and
wherein audio processing the music file includes normalizing, compressing and equalizing the music file,
wherein the voiceover file and the music file are normalized, compressed and equalized asynchronously.
18 . The computer-readable medium of claim 15 , further having stored thereon a sequence of instructions for causing the one or more processors to perform:
storing, in a voice activations store, a curve corresponding to when a voice is present in the voiceover file.
19 . The computer-readable medium of claim 15 , further having stored thereon a sequence of instructions for causing the one or more processors to perform:
setting an advertisement duration time;
setting a start time for the voiceover file;
trimming the music file according to the advertisement duration time; and
mixing the voiceover file and the music file according to the start time and the advertisement duration time.
20 . The computer-readable medium of claim 15 , further having stored thereon a sequence of instructions for causing the one or more processors to perform:
generating a script;
converting the script to voice content; and
saving the voice content in the voiceover file.
21 . The computer-readable medium of claim 15 , further having stored thereon a sequence of instructions for causing the one or more processors to perform:
mapping each track in a library of tracks to a point in an embedding space;
computing an acoustic embedding based on a query track within the embedding space;
obtaining a track from the library of tracks with acoustically similar content; and
saving the track from the library of tracks with acoustically similar content in the music file.