Method and system for encoding loudness metadata of audio components
A method that includes receiving an audio component associated with an audio scene, the audio component including an audio signal, determining a loudness level of the audio component based on the audio signal, receiving a target loudness level for the audio component, producing a bitstream with the audio component by encoding the audio signal and including metadata that has the loudness level and the target loudness level, and transmitting the bitstream to an electronic device.
1 . A method performed by a programmed processor of an encoder side, the method comprising:
receiving a first audio component associated with an audio scene, the first audio component comprising a first audio signal;
receiving a second audio component associated with the audio scene, the second audio component comprising a second audio signal;
determining a first source loudness of the first audio component based on the first audio signal;
determining a second source loudness of the second audio component based on the second audio signal;
receiving a first target loudness for the first audio component and a second target loudness for the second audio component;
determining an audio scene loudness for the audio scene using at least the first source loudness of the first audio component and the first target loudness;
producing a bitstream with the first audio component and the second audio component by encoding the first audio signal and the second audio signal and including metadata that has the first source loudness, the first target loudness, the second source loudness, the second target loudness, and the audio scene loudness; and
transmitting the bitstream to an electronic device.
2 . The method of claim 1 , wherein the first audio signal is a portion of an entire audio signal that makes up the first audio component, wherein the first source loudness is an average loudness across the portion of the entire audio signal that is received.
3 . The method of claim 2 , wherein the portion is a first portion, wherein the method further comprises:
receiving a second portion of the entire audio signal that is received after the first portion;
determining a third source loudness based on the first portion and the second portion; and
transmitting the third source loudness as metadata in the bitstream that includes an encoded second portion of the entire audio signal.
4 . The method of claim 3 , wherein the third source loudness converges closer or is equal to an overall loudness of the entire audio signal than the first source loudness.
5 . The method of claim 1 , wherein the first audio component and the second audio component are part of a live audio program that is received in real-time as the live audio program is recorded.
6 . The method of claim 1 , wherein determining the audio scene loudness comprises:
determining a scalar gain based on a difference between the first target loudness and the first source loudness; and
producing a gain-adjusted audio signal by applying the scalar gain to at least the first audio signal, wherein the audio scene loudness is determined using the gain-adjusted audio signal.
7 . The method of claim 6 , wherein the scalar gain is a first scalar gain, wherein the method further comprising:
producing a first gain-adjusted audio signal by applying the first scalar gain to the first audio signal, wherein the first scalar gain is based on a first difference between the first target loudness and the first source loudness;
producing a second gain-adjusted audio signal by applying a second scalar gain to the second audio signal, wherein the second scalar gain is based on a second difference between the second target loudness and the second source loudness;
determining an audio scene loudness level for the audio scene based on the first gain-adjusted audio signal and the second gain-adjusted audio signal; and
adding the audio scene loudness level to the metadata.
8 . An audio encoder device comprising:
at least one processor; and
memory having stored therein instructions which when executed by the at least one processor causes the audio encoder device to:
receive a first audio component associated with an audio scene, the first audio component comprising a first audio signal;
receive a second audio component associated with the audio scene, the second audio component comprising a second audio signal;
determine a first source loudness of the first audio component based on the first audio signal;
determine a second source loudness of the second audio component based on the second audio signal;
receive a first target loudness for the first audio component and a second target loudness for the second audio component;
determine an audio scene loudness for the audio scene using at least the first source loudness of the first audio component and the first target loudness for the first audio component; and
encode the first audio component and the second audio component and metadata that comprises the first source loudness, the second source loudness, the first target loudness, the second target loudness, and the audio scene loudness into a bitstream for an electronic device.
9 . The audio encoder device of claim 8 , wherein determining the first source loudness comprises retrieving the first source loudness from the memory, wherein the first source loudness is an overall loudness that spans a duration of the first audio signal.
10 . The audio encoder device of claim 8 , wherein determining the first source loudness of the first audio component comprises applying the first audio signal to a loudness model.
11 . The audio encoder device of claim 8 , wherein encoding the metadata comprises converting each of the first source loudness, the second source loudness, the first target loudness, and the second target loudness into respective 8-bit integers and storing each of the respective 8-bit integers into the bitstream.
12 . The audio encoder device of claim 8 , wherein the bitstream comprises a first encoded audio signal of the first audio signal and a second encoded audio signal of the first audio signal with the metadata, wherein a first signal level of the first encoded audio signal is the same as a first corresponding signal level of the first audio signal of the received first audio component and a second signal level of the second encoded audio signal is the same as a second corresponding signal level of the second audio signal of the received second audio component.
13 . A non-transitory machine-readable medium having instructions stored therein which when executed by at least one processor of a first electronic device causes the first electronic device to:
determine a first source loudness of a first audio component and a second source loudness of a second audio component associated with an audio scene, the first audio component comprising a first audio signal and the second audio component comprising a second audio signal;
receiving a first target loudness for the first audio component and a second target loudness for the second audio component;
determine an audio scene loudness for the audio scene using at least the first source loudness of the first audio component and the first target loudness;
produce a bitstream with the first audio component and the second audio component by encoding the first audio signal and the second audio signal and including metadata that has the first source loudness, the second source loudness, the first target loudness, the second target loudness, and the audio scene loudness; and
transmit the bitstream to a second electronic device.
14 . The non-transitory machine-readable medium of claim 13 , wherein the first audio signal is a portion of an entire audio signal that makes up the first audio component, wherein the first source loudness is an average loudness across the portion of the entire audio signal.
15 . The non-transitory machine-readable medium of claim 14 , wherein the portion is a first portion, wherein the non-transitory machine-readable medium comprises further instructions to:
receive a second portion of the entire audio signal that is subsequent the first portion;
determine a third source loudness based on the first portion and the second portion; and
transmit the third source loudness as metadata in the bitstream that includes an encoded second portion of the entire audio signal.
16 . The non-transitory machine-readable medium of claim 15 , wherein the third source loudness converges closer or is equal to an overall loudness of the entire audio signal than the first source loudness.
17 . The non-transitory machine-readable medium of claim 13 , wherein the instructions to produce the bitstream with the first audio component and the second audio component by encoding the first audio signal and the second audio signal and including metadata comprises instructions to convert each of the first source loudness, the second source loudness. the first target loudness, the second target loudness, and the audio scene loudness into respective 8-bit integers and storing each 8-bit integer into the bitstream.
18 . The non-transitory machine-readable medium of claim 13 , wherein the instructions to determine the audio scene loudness comprises instructions to:
determine a scalar gain based on a difference between the first target loudness and the first source loudness; and
produce a gain-adjusted audio signal by applying the scalar gain to at least the first audio signal, wherein the audio scene loudness is determined using the gain-adjusted audio signal.
19 . The non-transitory machine-readable medium of claim 18 , wherein the scalar gain is a first scalar gain, wherein the non-transitory machine-readable medium comprises further instructions to:
produce a first gain-adjusted audio signal by applying the first scalar gain to the first audio signal, wherein the first scalar gain is based on a first difference between the first target loudness and the first source loudness;
produce a second gain-adjusted audio signal by applying a second scalar gain to the second audio signal, wherein the second scalar gain is based on a second difference between the second target loudness and the second source loudness;
determine an audio scene loudness level for the audio scene based on the first gain-adjusted audio signal and the second gain-adjusted audio signal; and
add the audio scene loudness level to the metadata.