Method, apparatus, device and storage medium for processing playing loudness of media data
The disclosure relates to the technical field of computer processing, and discloses a method, an apparatus, a device and a storage medium for processing playing loudness of media data. The method according to the disclosure comprises obtaining media data to be played and an audio feature of the media data to be played; obtaining loudness requirement information, wherein the loudness requirement information comprises a current playing environment and/or an attribute of a playing device; determining loudness processing information for the media data to be played based on the audio feature and the loudness requirement information; and processing the media data to be played based on the loudness processing information to obtain target media data for playing.
1 . A method for processing playing loudness of media data, comprising:
obtaining media data and an audio feature of the media data;
obtaining loudness requirement information, wherein the loudness requirement information comprises a current playing environment and/or an attribute of a playing device;
determining loudness processing information for the media data based on the audio feature and the loudness requirement information; and
processing the media data based on the loudness processing information to obtain first media data for playing,
wherein determining the loudness processing information for the media data based on the audio feature and the loudness requirement information comprises:
obtaining a loudness compensation gain, wherein the loudness compensation gain is obtained based on a loudness offset,
wherein obtaining the loudness compensation gain comprises:
determining the loudness offset based on the audio feature and a first loudness for the media data, and
wherein, in a case where duration of the media data is longer than preset duration, determining the loudness offset based on the audio feature and the first loudness comprises:
obtaining a slope in a dynamic range control parameter and a starting point of a dynamic range in the audio feature;
conducting loudness estimation based on the first loudness, the slope and the starting point to determine estimated loudness; and
determining the loudness offset based on a difference between the first loudness and the estimated loudness.
2 . The method according to claim 1 , wherein determining the loudness processing information for the media data based on the audio feature and the loudness requirement information comprises:
determining the first loudness for the media data based on the loudness requirement information;
determining a loudness gain based on a difference between source loudness in the audio feature and the first loudness;
determining the dynamic range control parameter for the media data based on the audio feature and the loudness requirement information,
wherein the loudness processing information comprises the loudness gain, the dynamic range control parameter and the loudness compensation gain.
3 . The method according to claim 2 , wherein the attribute of the playing device comprises a playing capability, and determining the first loudness for the media data based on the loudness requirement information comprises:
determining initial first loudness based on the playing capability; and
adjusting the initial first loudness based on the audio feature to obtain the first loudness.
4 . The method according to claim 3 , wherein determining the initial first loudness based on the playing capability comprises:
obtaining a first correspondence between the playing capability of the playing device and initial loudness; and
inquiring the first correspondence based on the playing capability of the playing device in the loudness requirement information to obtain the initial first loudness.
5 . The method according to claim 3 , wherein adjusting the initial first loudness based on the audio feature to obtain the first loudness comprises:
determining a first loudness increment based on the audio feature, wherein the audio feature comprises a type of audio content or a speech ratio; and
adjusting the initial first loudness based on the first loudness increment to obtain the first loudness.
6 . The method according to claim 2 , wherein determining the dynamic range control parameter for the media data based on the audio feature and the loudness requirement information comprises:
obtaining a second correspondence between a loudness requirement and a characteristic of dynamic range control;
inquiring the second correspondence based on the loudness requirement information to determine a first characteristic of the dynamic range control; and
determining, in a case where the first characteristic is a compression demand, the slope of a curve of the dynamic range control based on the audio feature and the first loudness, wherein the dynamic range control parameter comprises the slope.
7 . The method according to claim 6 , wherein the audio feature comprises a loudness peak value, and determining the slope of the curve of the dynamic range control based on the audio feature and the first loudness comprises:
conducting peak normalization processing on the source loudness based on the loudness peak value and a preset loudness peak value to obtain processed loudness;
determining an initial slope of the curve of the dynamic range control based on a ratio of the processed loudness to the first loudness;
obtaining, in a case where the loudness requirement information comprises the current playing environment or a user preference, a preset slope corresponding to the first characteristic; and
determining the slope of the curve of the dynamic range control based on a maximum value of the preset ratio and the initial ratio.
8 . The method according to claim 6 , wherein in a case where the first characteristic is the compression demand, determining the dynamic range control parameter for the media data based on the audio feature and the loudness requirement information further comprises:
determining, based on a dynamic range in the audio feature, the starting point of the dynamic range; and
determining the starting point of the dynamic range as a threshold of a static characteristic in the curve of the dynamic range control, and determining a knee width of the curve as a preset value greater than zero.
9 . The method according to claim 6 , wherein determining the dynamic range control parameter for the media data based on the audio feature and the loudness requirement information further comprises:
executing, in the case where duration of the media data is longer than preset duration, a step of obtaining the second correspondence between the loudness requirement and the characteristic of the dynamic range control;
obtaining, in a case where the duration is shorter than the preset duration, maximum short-time loudness in the audio feature, and determining a loudness difference between the maximum short-time loudness and the first loudness; and
determining, in a case where the loudness difference is greater than a preset loudness difference, the slope of the curve of the dynamic range control based on the loudness difference and the preset loudness difference.
10 . The method according to claim 9 , wherein determining the slope of the curve of the dynamic range control based on the loudness difference and the preset loudness difference comprises:
determining the slope of the curve of the dynamic range control based on a ratio of the loudness difference to the preset loudness difference.
11 . The method according to claim 9 , wherein determining the dynamic range control parameter for the media data based on the audio feature and the loudness requirement information further comprises:
determining, in a case where the loudness difference is smaller than the preset loudness difference, that the first characteristic of the dynamic range control is no compression demand.
12 . The method according to claim 2 , wherein obtaining the loudness compensation gain further comprises:
determining the loudness compensation gain based on the loudness offset.
13 . The method according to claim 12 , wherein determining the loudness compensation gain based on the loudness offset comprises:
obtaining the loudness compensation gain based on a sum of the loudness offset and a first compensation value.
14 . The method according to claim 2 , wherein in a case where duration of the media data is shorter than preset duration, determining the loudness offset based on the audio feature and the first loudness comprises:
obtaining maximum short-time loudness in the audio feature; and
determining the loudness offset based on a difference between the maximum short-time loudness and the first loudness.
15 . The method according to claim 14 , wherein determining the loudness compensation gain based on the loudness offset comprises:
obtaining, in a case where the loudness offset is greater than a preset loudness difference, the loudness compensation gain based on a difference between the loudness offset and a second compensation value; and
determining, in a case where the loudness offset is smaller than the preset loudness difference, the loudness compensation gain as zero.
16 . The method according to claim 2 , wherein processing the media data based on the loudness processing information to obtain the first media data for playing comprises:
conducting loudness processing on the media data based on the loudness gain to obtain second media data;
conducting dynamic range processing on the second media data based on the dynamic range control parameter to obtain third media data;
conducting loudness compensation on the third media data based on the loudness compensation gain to obtain fourth media data; and
conducting peak limiting on the fourth media data to obtain the first media data.
17 . A computer device, comprising:
a memory and a processor, wherein the memory is in communication connection with the processor, the memory stores a computer instruction, and the processor executes the computer instruction to:
obtain media data and an audio feature of the media data;
obtain loudness requirement information, wherein the loudness requirement information comprises a current playing environment and/or an attribute of a playing device;
determine loudness processing information for the media data based on the audio feature and the loudness requirement information; and
process the media data based on the loudness processing information to obtain first media data for playing,
wherein determining the loudness processing information for the media data based on the audio feature and the loudness requirement information comprises:
obtaining a loudness compensation gain, wherein the loudness compensation gain is obtained based on a loudness offset,
wherein obtaining the loudness compensation gain comprises:
determining the loudness offset based on the audio feature and a first loudness for the media data, and
wherein in a case where duration of the media data is longer than preset duration, determining the loudness offset based on the audio feature and the first loudness comprises:
obtaining a slope in a dynamic range control parameter and a starting point of a dynamic range in the audio feature;
conducting loudness estimation based on the first loudness, the slope and the starting point to determine estimated loudness; and
determining the loudness offset based on a difference between the first loudness and the estimated loudness.
18 . The computer device according to claim 17 , wherein determining the loudness processing information for the media data based on the audio feature and the loudness requirement information comprises:
determining the first loudness for the media data based on the loudness requirement information;
determining a loudness gain based on a difference between source loudness in the audio feature and the first loudness;
determining the dynamic range control parameter for the media data based on the audio feature and the loudness requirement information,
wherein the loudness processing information comprises the loudness gain, the dynamic range control parameter and the loudness compensation gain.
19 . A non-transitory computer-readable storage medium, storing a computer instruction, wherein the computer instruction is configured to cause a computer to:
obtain media data and an audio feature of the media data;
obtain loudness requirement information, wherein the loudness requirement information comprises a current playing environment and/or an attribute of a playing device;
determine loudness processing information for the media data based on the audio feature and the loudness requirement information; and
process the media data based on the loudness processing information to obtain first media data for playing,
wherein determining the loudness processing information for the media data based on the audio feature and the loudness requirement information comprises:
obtaining a loudness compensation gain, wherein the loudness compensation gain is obtained based on a loudness offset,
wherein obtaining the loudness compensation gain comprises:
determining the loudness offset based on the audio feature and a first loudness for the media data, and
wherein in a case where duration of the media data is longer than preset duration, determining the loudness offset based on the audio feature and the first loudness comprises:
obtaining a slope in a dynamic range control parameter and a starting point of a dynamic range in the audio feature;
conducting loudness estimation based on the first loudness, the slope and the starting point to determine estimated loudness; and
determining the loudness offset based on a difference between the first loudness and the estimated loudness.