Three-dimensional audio signal processing method and apparatus
Embodiments of this application disclose a three-dimensional audio signal processing method and apparatus, to implement bit allocation of a signal. The method includes: performing spatial coding on a to-be-coded three-dimensional audio signal, to obtain a transmission channel signal and transmission channel attribute information, where the transmission channel signal includes at least one virtual speaker signal group and at least one residual signal group; and determining a bit allocation ratio of the virtual speaker signal group and a bit allocation ratio of the residual signal group based on the transmission channel attribute information.
1 . A three-dimensional audio signal processing method, comprising:
performing spatial coding on a to-be-coded three-dimensional audio signal to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal comprises a plurality of virtual speaker signals grouped into at least one virtual speaker signal group and a plurality of residual signals grouped into at least one residual signal group;
determining a bit allocation ratio of the virtual speaker signal group and a bit allocation ratio of the residual signal group based on the transmission channel attribute information; and
encoding the transmission channel signal, the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and a bit allocation ratio for each transmission channel of the transmission channel signal-into a bitstream, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal.
2 . The method according to claim 1 , wherein the transmission channel attribute information comprises a virtual speaker coding efficiency, the method further comprising:
performing signal reconstruction on the to-be-coded three-dimensional audio signal using a virtual speaker to obtain a reconstructed three-dimensional audio signal;
obtaining an energy representation value of the reconstructed three-dimensional audio signal and an energy representation value of the to-be-coded three-dimensional audio signal; and
obtaining the virtual speaker coding efficiency based on the energy representation value of the reconstructed three-dimensional audio signal and the energy representation value of the to-be-coded three-dimensional audio signal.
3 . The method according to claim 1 , wherein the transmission channel attribute information comprises an energy ratio of the virtual speaker signal group, the method further comprising:
obtaining an energy representation value of the virtual speaker signal group based on an energy representation value of each virtual speaker signal in the virtual speaker signal group;
obtaining an energy representation value of the residual signal group based on an energy representation value of each residual signal in the residual signal group; and
obtaining the energy ratio of the virtual speaker signal group based on the energy representation value of the virtual speaker signal group and the energy representation value of the residual signal group.
4 . The method according to claim 1 , wherein the transmission channel attribute information comprises a virtual speaker code identifier that indicates whether bit allocation of the virtual speaker signal group is dominant, the method further comprising:
performing spatial coding on the to-be-coded three-dimensional audio signal to obtain a quantity of anisotropic sound sources of the transmission channel signal and virtual speaker coding efficiency; and
obtaining the virtual speaker code identifier based on the quantity of anisotropic sound sources of the transmission channel signal and the virtual speaker coding efficiency.
5 . The method according to claim 4 , further comprising:
when the quantity of anisotropic sound sources of the transmission channel signal is less than or equal to a preset threshold of the quantity of anisotropic sound sources and the virtual speaker coding efficiency is greater than or equal to a preset first virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is dominant; or
when the quantity of anisotropic sound sources of the transmission channel signal is greater than a preset threshold of the quantity of anisotropic sound sources or the virtual speaker coding efficiency is less than a preset first virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is not dominant.
6 . The method according to claim 5 , wherein dominance comprises sub-dominance or pre-dominance, the method further comprising:
when the virtual speaker coding efficiency is greater than or equal to the preset first virtual speaker coding efficiency threshold and the virtual speaker coding efficiency is less than or equal to a preset second virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is sub-dominant; or
when the virtual speaker coding efficiency is greater than or equal to the preset first virtual speaker coding efficiency threshold and the virtual speaker coding efficiency is greater than a preset second virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is pre-dominant, wherein
the preset second virtual speaker coding efficiency threshold is greater than the preset first virtual speaker coding efficiency threshold.
7 . The method according to claim 1 , wherein the transmission channel attribute information comprises an energy ratio of the virtual speaker signal group and/or a virtual speaker code identifier, the method further comprising:
determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group according to a preset first signal group bit allocation algorithm when the energy ratio of the virtual speaker signal group is greater than or equal to a preset first energy ratio threshold and/or the virtual speaker code identifier is pre-dominant; or
determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group according to a preset second signal group bit allocation algorithm when the energy ratio of the virtual speaker signal group is greater than or equal to a preset second energy ratio threshold and less than a preset first energy ratio threshold and/or the virtual speaker code identifier is sub-dominant, wherein the preset second energy ratio threshold is less than the preset first energy ratio threshold; or
determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group according to a preset third signal group bit allocation algorithm when the energy ratio of the virtual speaker signal group is less than a preset first energy ratio threshold or the virtual speaker code identifier is not dominant.
8 . The method according to claim 7 , further comprising:
when directionalNrgRatio≥TH1, and/or S≤TH0 and η>TH2 are met, the plurality of virtual speaker signals are grouped into one virtual speaker signal group, and the plurality of residual signals are grouped into one residual signal group, calculating the bit allocation ratio of the virtual speaker signal group in the following manner:
Ratio1_1=FAC1*directionalNrgRatio+(1−FAC1) * maxdirectionalNrgRatio, wherein
directionalNrgRatio represents the energy ratio of the virtual speaker signal group, S is a quantity of anisotropic sound sources, n represents a virtual speaker coding efficiency, maxdirectionalNrgRatio is a preset maximum bit allocation ratio of the virtual speaker signal group, FAC1 is a preset first adjustment factor, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, * represents a multiplication operation, TH1 is the preset first energy ratio threshold, TH0 is a threshold of the quantity of anisotropic sound sources, and TH2 is a second virtual speaker coding efficiency threshold; and
calculating the bit allocation ratio of the residual signal group in the following manner:
Ratio2=1−Ratio1_1, wherein
Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.
9 . The method according to claim 8 , wherein after the bit allocation ratio of the virtual speaker signal group is obtained, the method further comprises:
updating the bit allocation ratio of the virtual speaker signal group in the following manner:
Ratio1_2=min (Ratio1_1, maxdirectionalNrgRatio+FAC2*Ratio1_1), wherein
Ratio1_2 represents an updated bit allocation ratio of the virtual speaker signal group, FAC2 is a preset second adjustment factor, maxdirectionalNrgRatio is the preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio that is of the virtual speaker signal group and that exists before updating, * represents a multiplication operation, and min is a minimization operation.
10 . The method according to claim 7 , further comprising:
when TH3≤directionalNrgRatio<TH1 is met, and/or S≤TH0 and TH4≤η≤TH2 are met, the plurality of virtual speaker signals are grouped into one virtual speaker signal group, and the plurality of residual signals are grouped into one residual signal group, calculating Ratio1_1 in the following manner:
Ratio1_1=FAC3*directionalNrgRatio+ (1−FAC3) * maxdirectionalNrgRatio, wherein
maxdirectionalNrgRatio is a preset bit allocation ratio of the virtual speaker signal group, FAC3 is a preset third adjustment factor, directionalNrgRatio represents the energy ratio of the virtual speaker signal group, S is a quantity of anisotropic sound sources, n represents a virtual speaker coding efficiency, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, * represents a multiplication operation, TH0 is a threshold of the quantity of anisotropic sound sources, TH1 is the preset first energy ratio threshold, TH2 is a second virtual speaker coding efficiency threshold, TH3 is the preset second energy ratio threshold, and TH4 is a first virtual speaker coding efficiency threshold; and
calculating the bit allocation ratio of the residual signal group in the following manner:
Ratio2=1−Ratio1_1, wherein
Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.
11 . A three-dimensional audio signal processing method, comprising:
receiving a bitstream, wherein the bitstream comprises a bit allocation ratio of a virtual speaker signal group of a transmission channel signal, a bit allocation ratio of a residual signal group of the transmission channel signal, and a bit allocation ratio for each transmission channel of the transmission channel signal, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal;
decoding the bitstream to obtain the bit allocation ratio for each transmission channel, the bit allocation ratio of a virtual speaker signal group and the bit allocation ratio of a residual signal group; and
decoding a virtual speaker signal and a residual signal in the bitstream based on the bit allocation ratio for each transmission channel, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain a three-dimensional audio signal through decoding.
12 . The method according to claim 11 , further comprising:
determining a quantity of available bits;
determining a bit quantity of the virtual speaker signal group based on the quantity of available bits and the bit allocation ratio of the virtual speaker signal group;
decoding the virtual speaker signal in the bitstream based on the bit quantity of the virtual speaker signal group and a bit allocation ratio of a corresponding transmission channel;
determining a bit quantity of the residual signal group based on the quantity of available bits and the bit allocation ratio of the residual signal group; and
decoding the residual signal in the bitstream based on the bit quantity of the residual signal group and a bit allocation ratio of a corresponding transmission channel.
13 . A three-dimensional audio signal processing apparatus, comprising:
a memory; and
at least one processor coupled to the memory that stores instructions that, when executed by the at least one processor, cause the apparatus to:
perform spatial coding on a to-be-coded three-dimensional audio signal to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal comprises a plurality of virtual speaker signals grouped into at least one virtual speaker signal group and a plurality of residual signals grouped into at least one residual signal group;
determine a bit allocation ratio of the virtual speaker signal group and a bit allocation ratio of the residual signal group based on the transmission channel attribute information; and
encode the transmission channel signal, the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and a bit allocation ratio for each transmission channel of the transmission channel signal into a bitstream, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal.
14 . The three-dimensional audio signal processing apparatus according to claim 13 , wherein the three-dimensional audio signal processing apparatus further comprises the memory.
15 . The three-dimensional audio signal processing apparatus according to claim 13 , wherein the apparatus is further to:
perform signal reconstruction on the to-be-coded three-dimensional audio signal by using a virtual speaker, to obtain a reconstructed three-dimensional audio signal;
obtain an energy representation value of the reconstructed three-dimensional audio signal and an energy representation value of the to-be-coded three-dimensional audio signal; and
obtain the virtual speaker coding efficiency based on the energy representation value of the reconstructed three-dimensional audio signal and the energy representation value of the to-be-coded three-dimensional audio signal.
16 . A three-dimensional audio signal processing apparatus, comprising:
a memory; and
at least one processor coupled to the memory that stores instructions that, when executed by the at least one processor, cause the apparatus to:
receive a bitstream, wherein the bitstream comprises a bit allocation ratio of a virtual speaker signal group of a transmission channel signal, a bit allocation ratio of a residual signal group of the transmission channel signal, and a bit allocation ratio for each transmission channel of the transmission channel signal, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal;
decode the bitstream, to obtain the bit allocation ratio for each transmission channel, the bit allocation ratio of a virtual speaker signal group and the bit allocation ratio of a residual signal group; and
decode a virtual speaker signal and a residual signal in the bitstream based on the bit allocation ratio for each transmission channel, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, to obtain a three-dimensional audio signal through decoding.
17 . The three-dimensional audio signal processing apparatus according to claim 16 , wherein the three-dimensional audio signal processing apparatus further comprises the memory.
18 . The three-dimensional audio signal processing apparatus according to claim 16 , wherein the apparatus is further to:
determine a quantity of available bits;
determine a bit quantity of the virtual speaker signal group based on the quantity of available bits and the bit allocation ratio of the virtual speaker signal group, and decoding the virtual speaker signal in the bitstream based on the bit quantity of the virtual speaker signal group and a bit allocation ratio of a corresponding transmission channel; and
determine a bit quantity of the residual signal group based on the quantity of available bits and the bit allocation ratio of the residual signal group, and decoding the residual signal in the bitstream based on the bit quantity of the residual signal group and a bit allocation ratio of a corresponding transmission channel.
19 . A non-transitory computer-readable storage medium, comprising instructions, wherein when the instructions run on a computer, the computer is enabled to perform the method according to claim 16 .
20 . A non-transitory computer-readable storage medium, comprising a bitstream generated in the method according to claim 1 .