System and method for dynamic mixing of audio
A system for dynamically mixing audio content, the system comprising: a receiving unit configured to receive input audio; an analysis unit configured to analyse the input audio to determine one or more masking patterns; an attenuation unit configured to attenuate one or more channels of the input audio in accordance with the one or more masking patterns, and an output unit configured to output attenuated audio.
1 . A system for dynamically mixing audio content, the system comprising:
one or more processors;
one or more computer-readable media having stored thereon instructions that when executed with the one or more processors cause the system to, at least:
receive input audio;
analyze, with the one or more processors, the input audio to determine one or more masking patterns, wherein the input audio comprises a plurality of audio channels and individual masking patterns quantify an extent to which an audio channel of the plurality of audio channels causes psychoacoustic masking with respect to one or more other audio channels of the plurality of audio channels;
attenuate one or more lower priority audio channels of the input audio in accordance with the one or more masking patterns such that, after attenuation, individual masking patterns corresponding to individual lower priority audio channels indicate reduced psychoacoustic masking with respect to one or more higher priority audio channels; and
output attenuated audio.
2 . A system according to claim 1 , wherein:
the input audio comprises at least a first audio channel and a second audio channel;
the analyzing of the input audio comprises determining one or more masking patterns of the first audio channel with respect to the second audio channel, and
the attenuating of the one or more lower priority audio channels comprises attenuating at least part of the second channel based on the one or more masking patterns of the first channel.
3 . A system according to claim 2 , wherein the system is further caused to determine one or more critical frequency bands masked by the first audio channel, and attenuate the second audio channel across the one or more critical frequency bands.
4 . A system according to claim 3 , wherein the critical frequency bands are determined based on a set of BARK filters and/or a set of mel scaled filters.
5 . A system according to claim 1 , wherein the system is further caused to identify spectral regions in which the one or more masking patterns substantially completely mask an audio channel of the input audio, and remove the audio from that audio channel in the identified spectral region.
6 . A system according to claim 1 , wherein the system is further caused to:
receive a mixing modifier comprising one or more selected from the list consisting of:
i. a user input; and
ii. a predefined configuration; and
attenuate the one or more audio channels of the input audio in accordance with the mixing modifier.
7 . A system according to claim 6 , wherein the system is further caused to determine a protected frequency band based on the mixing modifier, and leave the protected band unattenuated on one or more of the plurality of audio channels.
8 . A system according to claim 1 , wherein the system is further caused to determine one or more boost bands, and boost amplitude of one or more audio channels over the one or more boost bands.
9 . A system according to claim 1 , wherein the system is further caused to determine priority levels of one or more of the plurality of audio channels in the input audio; and determine one or more masking patterns of a highest priority channel.
10 . A system according to claim 9 , wherein the system is further caused to attenuate a lowest priority channel.
11 . A system according to claim 10 , wherein the system is further caused to receive video game data, and derive the priority levels from the received video game data.
12 . A system according to claim 1 , wherein the system is further caused to receive an audio stream from a currently running video game, and the output processed audio in real-time to the video game.
13 . A system in accordance with claim 1 , further comprising determining a priority for individual ones of the plurality of audio channels.
14 . A system in accordance with claim 13 , wherein the priority of an audio channel is determined based on metadata accompanying the input audio.
15 . A system in accordance with claim 13 , wherein the priority of an audio channel is determined based on a machine learning classification of the audio channel.
16 . A system in accordance with claim 13 , wherein the priority of an audio channel is determined based on a priority schedule with respect to a timeline of the input audio.
17 . A method for dynamically mixing audio content, the method comprising the steps of:
receiving input audio;
analyzing, with one or more processors of a computerized device, the input audio to determine one or more masking patterns, wherein the input audio comprises a plurality of audio channels and individual masking patterns quantify an extent to which an audio channel of the plurality of audio channels causes psychoacoustic masking with respect to one or more other audio channels of the plurality of audio channels;
attenuating one or more lower priority audio channels of the input audio in accordance with the one or more masking patterns such that, after attenuation, individual masking patterns corresponding to individual lower priority audio channels indicate reduced psychoacoustic masking with respect to one or more higher priority audio channels; and,
outputting the attenuated audio comprising the one or more attenuated channels.
18 . A method according to claim 17 , wherein the input audio comprises at least a first audio channel and a second audio channel;
the step of analyzing the input audio comprises determining one or more masking patterns of the first audio channel with respect to the second audio channel, and
the step of attenuating the one or more lower priority audio channels of the input audio comprises attenuating at least part of the second channel based on the one or more masking patterns of the first channel.
19 . A method according to claim 18 , further comprising the steps of:
receiving video game data from a video game system; and
determining priority levels of audio channels according to the received video game data, and wherein the step of attenuating one or more audio channels of the input audio comprises attenuating a lowest priority channel based on one or more masking patterns of a highest priority channel.
20 . A non-transitory, computer readable storage medium containing a computer program comprising computer executable instructions that when executed by a computer system, cause the computer system to perform a method for dynamically mixing audio content, the method comprising the steps of:
receiving input audio;
analyzing the input audio to determine one or more masking patterns, wherein the input audio comprises a plurality of audio channels and individual masking patterns quantify an extent to which an audio channel of the plurality of audio channels causes psychoacoustic masking with respect to one or more other audio channels of the plurality of audio channels;
attenuating one or more lower priority audio channels of the input audio in accordance with the one or more masking patterns such that, after attenuation, individual masking patterns corresponding to individual lower priority audio channels indicate reduced psychoacoustic masking with respect to one or more higher priority audio channels; and,
outputting the attenuated audio comprising the one or more attenuated channels.