IP Library Granted Patent US 12,633,293
Granted Patent B2
US 12,633,293 · App. 18/532,085 · Granted May 19, 2026

Three-dimensional audio signal processing method and apparatus

Inventors: Shuai Liu (Beijing, CN); Yuan Gao (Beijing, CN); Bingyin Xia (Beijing, CN); Bin Wang (Shenzhen, CN); Zhe Wang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L19/002G10L19/008H04S7/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,293
App. No.
18/532,085
Granted
May 19, 2026
Kind
B2
Abstract

Embodiments of this application disclose a three-dimensional audio signal processing method and apparatus, to implement bit allocation of a signal. The method includes: performing spatial coding on a to-be-coded three-dimensional audio signal, to obtain a transmission channel signal and transmission channel attribute information, where the transmission channel signal includes at least one virtual speaker signal group and at least one residual signal group; and determining a bit allocation ratio of the virtual speaker signal group and a bit allocation ratio of the residual signal group based on the transmission channel attribute information.

Claims (78)

1 . A three-dimensional audio signal processing method, comprising:

performing spatial coding on a to-be-coded three-dimensional audio signal to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal comprises a plurality of virtual speaker signals grouped into at least one virtual speaker signal group and a plurality of residual signals grouped into at least one residual signal group;

determining a bit allocation ratio of the virtual speaker signal group and a bit allocation ratio of the residual signal group based on the transmission channel attribute information; and

encoding the transmission channel signal, the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and a bit allocation ratio for each transmission channel of the transmission channel signal-into a bitstream, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal.

2 . The method according to claim 1 , wherein the transmission channel attribute information comprises a virtual speaker coding efficiency, the method further comprising:

performing signal reconstruction on the to-be-coded three-dimensional audio signal using a virtual speaker to obtain a reconstructed three-dimensional audio signal;

obtaining an energy representation value of the reconstructed three-dimensional audio signal and an energy representation value of the to-be-coded three-dimensional audio signal; and

obtaining the virtual speaker coding efficiency based on the energy representation value of the reconstructed three-dimensional audio signal and the energy representation value of the to-be-coded three-dimensional audio signal.

3 . The method according to claim 1 , wherein the transmission channel attribute information comprises an energy ratio of the virtual speaker signal group, the method further comprising:

obtaining an energy representation value of the virtual speaker signal group based on an energy representation value of each virtual speaker signal in the virtual speaker signal group;

obtaining an energy representation value of the residual signal group based on an energy representation value of each residual signal in the residual signal group; and

obtaining the energy ratio of the virtual speaker signal group based on the energy representation value of the virtual speaker signal group and the energy representation value of the residual signal group.

4 . The method according to claim 1 , wherein the transmission channel attribute information comprises a virtual speaker code identifier that indicates whether bit allocation of the virtual speaker signal group is dominant, the method further comprising:

performing spatial coding on the to-be-coded three-dimensional audio signal to obtain a quantity of anisotropic sound sources of the transmission channel signal and virtual speaker coding efficiency; and

obtaining the virtual speaker code identifier based on the quantity of anisotropic sound sources of the transmission channel signal and the virtual speaker coding efficiency.

5 . The method according to claim 4 , further comprising:

when the quantity of anisotropic sound sources of the transmission channel signal is less than or equal to a preset threshold of the quantity of anisotropic sound sources and the virtual speaker coding efficiency is greater than or equal to a preset first virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is dominant; or

when the quantity of anisotropic sound sources of the transmission channel signal is greater than a preset threshold of the quantity of anisotropic sound sources or the virtual speaker coding efficiency is less than a preset first virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is not dominant.

6 . The method according to claim 5 , wherein dominance comprises sub-dominance or pre-dominance, the method further comprising:

when the virtual speaker coding efficiency is greater than or equal to the preset first virtual speaker coding efficiency threshold and the virtual speaker coding efficiency is less than or equal to a preset second virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is sub-dominant; or

when the virtual speaker coding efficiency is greater than or equal to the preset first virtual speaker coding efficiency threshold and the virtual speaker coding efficiency is greater than a preset second virtual speaker coding efficiency threshold, determining that the virtual speaker code identifier is pre-dominant, wherein

the preset second virtual speaker coding efficiency threshold is greater than the preset first virtual speaker coding efficiency threshold.

7 . The method according to claim 1 , wherein the transmission channel attribute information comprises an energy ratio of the virtual speaker signal group and/or a virtual speaker code identifier, the method further comprising:

determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group according to a preset first signal group bit allocation algorithm when the energy ratio of the virtual speaker signal group is greater than or equal to a preset first energy ratio threshold and/or the virtual speaker code identifier is pre-dominant; or

determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group according to a preset second signal group bit allocation algorithm when the energy ratio of the virtual speaker signal group is greater than or equal to a preset second energy ratio threshold and less than a preset first energy ratio threshold and/or the virtual speaker code identifier is sub-dominant, wherein the preset second energy ratio threshold is less than the preset first energy ratio threshold; or

determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group according to a preset third signal group bit allocation algorithm when the energy ratio of the virtual speaker signal group is less than a preset first energy ratio threshold or the virtual speaker code identifier is not dominant.

8 . The method according to claim 7 , further comprising:

when directionalNrgRatio≥TH1, and/or S≤TH0 and η>TH2 are met, the plurality of virtual speaker signals are grouped into one virtual speaker signal group, and the plurality of residual signals are grouped into one residual signal group, calculating the bit allocation ratio of the virtual speaker signal group in the following manner:

Ratio1_1=FAC1*directionalNrgRatio+(1−FAC1) * maxdirectionalNrgRatio, wherein

directionalNrgRatio represents the energy ratio of the virtual speaker signal group, S is a quantity of anisotropic sound sources, n represents a virtual speaker coding efficiency, maxdirectionalNrgRatio is a preset maximum bit allocation ratio of the virtual speaker signal group, FAC1 is a preset first adjustment factor, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, * represents a multiplication operation, TH1 is the preset first energy ratio threshold, TH0 is a threshold of the quantity of anisotropic sound sources, and TH2 is a second virtual speaker coding efficiency threshold; and

calculating the bit allocation ratio of the residual signal group in the following manner:

Ratio2=1−Ratio1_1, wherein

Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.

9 . The method according to claim 8 , wherein after the bit allocation ratio of the virtual speaker signal group is obtained, the method further comprises:

updating the bit allocation ratio of the virtual speaker signal group in the following manner:

Ratio1_2=min (Ratio1_1, maxdirectionalNrgRatio+FAC2*Ratio1_1), wherein

Ratio1_2 represents an updated bit allocation ratio of the virtual speaker signal group, FAC2 is a preset second adjustment factor, maxdirectionalNrgRatio is the preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio that is of the virtual speaker signal group and that exists before updating, * represents a multiplication operation, and min is a minimization operation.

10 . The method according to claim 7 , further comprising:

when TH3≤directionalNrgRatio<TH1 is met, and/or S≤TH0 and TH4≤η≤TH2 are met, the plurality of virtual speaker signals are grouped into one virtual speaker signal group, and the plurality of residual signals are grouped into one residual signal group, calculating Ratio1_1 in the following manner:

Ratio1_1=FAC3*directionalNrgRatio+ (1−FAC3) * maxdirectionalNrgRatio, wherein

maxdirectionalNrgRatio is a preset bit allocation ratio of the virtual speaker signal group, FAC3 is a preset third adjustment factor, directionalNrgRatio represents the energy ratio of the virtual speaker signal group, S is a quantity of anisotropic sound sources, n represents a virtual speaker coding efficiency, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, * represents a multiplication operation, TH0 is a threshold of the quantity of anisotropic sound sources, TH1 is the preset first energy ratio threshold, TH2 is a second virtual speaker coding efficiency threshold, TH3 is the preset second energy ratio threshold, and TH4 is a first virtual speaker coding efficiency threshold; and

calculating the bit allocation ratio of the residual signal group in the following manner:

Ratio2=1−Ratio1_1, wherein

Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.

11 . A three-dimensional audio signal processing method, comprising:

receiving a bitstream, wherein the bitstream comprises a bit allocation ratio of a virtual speaker signal group of a transmission channel signal, a bit allocation ratio of a residual signal group of the transmission channel signal, and a bit allocation ratio for each transmission channel of the transmission channel signal, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal;

decoding the bitstream to obtain the bit allocation ratio for each transmission channel, the bit allocation ratio of a virtual speaker signal group and the bit allocation ratio of a residual signal group; and

decoding a virtual speaker signal and a residual signal in the bitstream based on the bit allocation ratio for each transmission channel, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain a three-dimensional audio signal through decoding.

12 . The method according to claim 11 , further comprising:

determining a quantity of available bits;

determining a bit quantity of the virtual speaker signal group based on the quantity of available bits and the bit allocation ratio of the virtual speaker signal group;

decoding the virtual speaker signal in the bitstream based on the bit quantity of the virtual speaker signal group and a bit allocation ratio of a corresponding transmission channel;

determining a bit quantity of the residual signal group based on the quantity of available bits and the bit allocation ratio of the residual signal group; and

decoding the residual signal in the bitstream based on the bit quantity of the residual signal group and a bit allocation ratio of a corresponding transmission channel.

13 . A three-dimensional audio signal processing apparatus, comprising:

a memory; and

at least one processor coupled to the memory that stores instructions that, when executed by the at least one processor, cause the apparatus to:

perform spatial coding on a to-be-coded three-dimensional audio signal to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal comprises a plurality of virtual speaker signals grouped into at least one virtual speaker signal group and a plurality of residual signals grouped into at least one residual signal group;

determine a bit allocation ratio of the virtual speaker signal group and a bit allocation ratio of the residual signal group based on the transmission channel attribute information; and

encode the transmission channel signal, the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and a bit allocation ratio for each transmission channel of the transmission channel signal into a bitstream, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal.

14 . The three-dimensional audio signal processing apparatus according to claim 13 , wherein the three-dimensional audio signal processing apparatus further comprises the memory.

15 . The three-dimensional audio signal processing apparatus according to claim 13 , wherein the apparatus is further to:

perform signal reconstruction on the to-be-coded three-dimensional audio signal by using a virtual speaker, to obtain a reconstructed three-dimensional audio signal;

obtain an energy representation value of the reconstructed three-dimensional audio signal and an energy representation value of the to-be-coded three-dimensional audio signal; and

obtain the virtual speaker coding efficiency based on the energy representation value of the reconstructed three-dimensional audio signal and the energy representation value of the to-be-coded three-dimensional audio signal.

16 . A three-dimensional audio signal processing apparatus, comprising:

a memory; and

at least one processor coupled to the memory that stores instructions that, when executed by the at least one processor, cause the apparatus to:

receive a bitstream, wherein the bitstream comprises a bit allocation ratio of a virtual speaker signal group of a transmission channel signal, a bit allocation ratio of a residual signal group of the transmission channel signal, and a bit allocation ratio for each transmission channel of the transmission channel signal, wherein each transmission channel is used to transmit a virtual speaker signal or a residual signal;

decode the bitstream, to obtain the bit allocation ratio for each transmission channel, the bit allocation ratio of a virtual speaker signal group and the bit allocation ratio of a residual signal group; and

decode a virtual speaker signal and a residual signal in the bitstream based on the bit allocation ratio for each transmission channel, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, to obtain a three-dimensional audio signal through decoding.

17 . The three-dimensional audio signal processing apparatus according to claim 16 , wherein the three-dimensional audio signal processing apparatus further comprises the memory.

18 . The three-dimensional audio signal processing apparatus according to claim 16 , wherein the apparatus is further to:

determine a quantity of available bits;

determine a bit quantity of the virtual speaker signal group based on the quantity of available bits and the bit allocation ratio of the virtual speaker signal group, and decoding the virtual speaker signal in the bitstream based on the bit quantity of the virtual speaker signal group and a bit allocation ratio of a corresponding transmission channel; and

determine a bit quantity of the residual signal group based on the quantity of available bits and the bit allocation ratio of the residual signal group, and decoding the residual signal in the bitstream based on the bit quantity of the residual signal group and a bit allocation ratio of a corresponding transmission channel.

19 . A non-transitory computer-readable storage medium, comprising instructions, wherein when the instructions run on a computer, the computer is enabled to perform the method according to claim 16 .

20 . A non-transitory computer-readable storage medium, comprising a bitstream generated in the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: LIU, SHUAI; GAO, YUAN; XIA, BINGYIN; WANG, BIN; WANG, ZHE
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 073483/0480 →
Priority Claims (2)
CN 202110657283.7 · Jun 11, 2021 · national
CN 202110700570.1 · Jun 23, 2021 · national
Continuity (2)
Continuation PCTCN2022096546 · Jun 1, 2022
Related Publication 20240112684A1 · Apr 4, 2024
References Cited (10)
US 10075802B1 · Kim et al. · 2018 [cited by applicant]
US 11057731B2 · Tsingos · 2021 [cited by examiner]
US 12142285B2 · Olivieri · 2024 [cited by examiner]
US 20070127733A1 · Henn · 2007 [cited by examiner]
US 20200296535A1 · Tsingos · 2020 [cited by examiner]
US 20200402522A1 · Olivieri · 2020 [cited by examiner]
US 20210219084A1 · Laitinen · 2021 [cited by examiner]
CN 107493542A · 2017 [cited by applicant]
CN 114582357A · 2022 [cited by applicant]
3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Codec for Immersive Voice and Audio Services; General overview (Release 18). 3GPP TS 26.250 V1.0.0 (Sep. 2023). total 12 pag… [cited by applicant]