IP Library › Granted Patent US 12,555,586
Granted Patent B2
US 12,555,586 · App. 18/538,708 · Granted Feb 17, 2026

Method and apparatus for encoding three-dimensional audio signal, encoder, and system

Inventors: Yuan Gao (Beijing, CN); Shuai Liu (Beijing, CN); Bingyin Xia (Beijing, CN); Bin Wang (Shenzhen, CN); Zhe Wang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L19/008G10L25/21H04S7/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,586
App. No.
18/538,708
Granted
Feb 17, 2026
Kind
B2
Abstract

A method for encoding a three-dimensional audio signal is provided. The method includes: An encoder obtains a current frame of a three-dimensional audio signal; obtains coding efficiency of an initial virtual speaker for the current frame based on the current frame of the three-dimensional audio signal; and when the coding efficiency of the initial virtual speaker for the current frame meets a preset condition, determines an updated virtual speaker for the current frame from a set of candidate virtual speakers; encodes the current frame based on the updated virtual speaker for the current frame, to obtain a first bitstream; or when the coding efficiency of the initial virtual speaker for the current frame does not meet the preset condition, encodes the current frame based on the initial virtual speaker for the current frame, to obtain a second bitstream.

Claims (74)

1 . A method for encoding a three-dimensional audio signal, comprising:

obtaining a current frame of a three-dimensional audio signal;

obtaining coding efficiency of an initial virtual speaker for the current frame based on the current frame of the three-dimensional audio signal, wherein the initial virtual speaker for the current frame belongs to a set of candidate virtual speakers;

when the coding efficiency of the initial virtual speaker for the current frame meets a preset condition, determining an updated virtual speaker for the current frame from the set of candidate virtual speakers, and encoding the current frame based on the updated virtual speaker for the current frame, to obtain a first bitstream; and

when the coding efficiency of the initial virtual speaker for the current frame does not meet the preset condition, encoding the current frame based on the initial virtual speaker for the current frame, to obtain a second bitstream;

wherein the preset condition comprises the coding efficiency of the initial virtual speaker for the current frame is less than a first threshold.

2 . The method according to claim 1 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

obtaining a reconstructed current frame of a reconstructed three-dimensional audio signal based on the initial virtual speaker for the current frame; and

determining the coding efficiency of the initial virtual speaker for the current frame based on energy of the reconstructed current frame and energy of the current frame.

3 . The method according to claim 2 , wherein the energy of the reconstructed current frame is determined based on a coefficient of the reconstructed current frame, and the energy of the current frame is determined based on a coefficient of the current frame.

4 . The method according to claim 1 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

obtaining a reconstructed current frame of a reconstructed three-dimensional audio signal based on the initial virtual speaker for the current frame;

obtaining a residual signal of the current frame based on the current frame of the three-dimensional audio signal and the reconstructed current frame of the reconstructed three-dimensional audio signal;

obtaining an energy sum of energy of a virtual speaker signal of the current frame and energy of the residual signal; and

determining the coding efficiency of the initial virtual speaker for the current frame based on a ratio of the energy of the virtual speaker signal of the current frame to the energy sum.

5 . The method according to claim 2 , wherein obtaining the reconstructed current frame of the reconstructed three-dimensional audio signal comprises:

determining a virtual speaker signal of the current frame based on the initial virtual speaker for the current frame; and

determining the reconstructed current frame based on the virtual speaker signal of the current frame.

6 . The method according to claim 1 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

determining a quantity of sound sources based on the current frame of the three-dimensional audio signal; and

determining the coding efficiency of the initial virtual speaker for the current frame based on a quantity of initial virtual speakers for the current frame and the quantity of sound sources.

7 . The method according to claim 1 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

determining a quantity of sound sources based on the current frame of the three-dimensional audio signal;

determining a virtual speaker signal of the current frame based on the initial virtual speaker for the current frame; and

determining the coding efficiency of the initial virtual speaker for the current frame based on a quantity of virtual speaker signals of the current frame and the quantity of sound sources of the three-dimensional audio signal.

8 . The method according to claim 1 , wherein determining the updated virtual speaker for the current frame from the set of candidate virtual speakers comprises:

when the coding efficiency of the initial virtual speaker for the current frame is less than a second threshold, using a preset virtual speaker in the set of candidate virtual speakers as the updated virtual speaker for the current frame, wherein the second threshold is less than the first threshold; or

when the coding efficiency of the initial virtual speaker for the current frame is less than the first threshold and greater than the second threshold, using a virtual speaker for a previous frame as the updated virtual speaker for the current frame, wherein the virtual speaker for the previous frame is a virtual speaker used for encoding the previous frame of the three-dimensional audio signal.

9 . A system, comprising:

an encoder comprising at least one processor and a memory coupled to the at least one processor to store instructions, which when executed by the at least one processor, cause the encoder to perform the method according to claim 1 ; and

a decoder comprising at least one processor and a memory coupled to the at least one processor to store instructions, which when executed by the at least one processor, cause the decoder to decode a bitstream generated by the encoder.

10 . An encoder, comprising:

at least one processor; and

a memory configured to store a computer program, which when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

obtaining a current frame of a three-dimensional audio signal;

obtaining coding efficiency of an initial virtual speaker for the current frame based on the current frame of the three-dimensional audio signal, wherein the initial virtual speaker for the current frame belongs to a set of candidate virtual speakers;

when the coding efficiency of the initial virtual speaker for the current frame meets a preset condition, determining an updated virtual speaker for the current frame from the set of candidate virtual speakers, and encoding the current frame based on the updated virtual speaker for the current frame, to obtain a first bitstream; and

when the coding efficiency of the initial virtual speaker for the current frame does not meet the preset condition, encoding the current frame based on the initial virtual speaker for the current frame, to obtain a second bitstream;

wherein the preset condition comprises the coding efficiency of the initial virtual speaker for the current frame is less than a first threshold.

11 . The encoder according to claim 10 ,

wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

obtaining a reconstructed current frame of a reconstructed three-dimensional audio signal based on the initial virtual speaker for the current frame; and

determining the coding efficiency of the initial virtual speaker for the current frame based on energy of the reconstructed current frame and energy of the current frame.

12 . The encoder according to claim 11 , wherein the energy of the reconstructed current frame is determined based on a coefficient of the reconstructed current frame, and the energy of the current frame is determined based on a coefficient of the current frame.

13 . The encoder according to claim 10 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

obtaining a reconstructed current frame of a reconstructed three-dimensional audio signal based on the initial virtual speaker for the current frame;

obtaining a residual signal of the current frame based on the current frame of the three-dimensional audio signal and the reconstructed current frame of the reconstructed three-dimensional audio signal;

obtaining an energy sum of energy of a virtual speaker signal of the current frame and energy of the residual signal; and

determining the coding efficiency of the initial virtual speaker for the current frame based on a ratio of the energy of the virtual speaker signal of the current frame to the energy sum.

14 . The encoder according to claim 11 , wherein obtaining the reconstructed current frame of the reconstructed three-dimensional audio signal comprises:

determining a virtual speaker signal of the current frame based on the initial virtual speaker for the current frame; and

determining the reconstructed current frame based on the virtual speaker signal of the current frame.

15 . The encoder according to claim 10 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

determining a quantity of sound sources based on the current frame of the three-dimensional audio signal; and

determining the coding efficiency of the initial virtual speaker for the current frame based on a quantity of initial virtual speakers for the current frame and the quantity of sound sources.

16 . The encoder according to claim 10 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

determining a quantity of sound sources based on the current frame of the three-dimensional audio signal;

determining a virtual speaker signal of the current frame based on the initial virtual speaker for the current frame; and

determining the coding efficiency of the initial virtual speaker for the current frame based on a quantity of virtual speaker signals of the current frame and the quantity of sound sources of the three-dimensional audio signal.

17 . A non-transitory computer-readable storage medium comprising computer software instructions, which when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising:

obtaining a current frame of a three-dimensional audio signal;

obtaining coding efficiency of an initial virtual speaker for the current frame based on the current frame of the three-dimensional audio signal, wherein the initial virtual speaker for the current frame belongs to a set of candidate virtual speakers;

when the coding efficiency of the initial virtual speaker for the current frame meets a preset condition, determining an updated virtual speaker for the current frame from the set of candidate virtual speakers, and encoding the current frame based on the updated virtual speaker for the current frame, to obtain a first bitstream; and

when the coding efficiency of the initial virtual speaker for the current frame does not meet the preset condition, encoding the current frame based on the initial virtual speaker for the current frame, to obtain a second bitstream;

wherein the preset condition comprises the coding efficiency of the initial virtual speaker for the current frame is less than a first threshold.

18 . The non-transitory computer-readable storage medium according to claim 17 , wherein obtaining the coding efficiency of the initial virtual speaker for the current frame comprises:

obtaining a reconstructed current frame of a reconstructed three-dimensional audio signal based on the initial virtual speaker for the current frame; and

determining the coding efficiency of the initial virtual speaker for the current frame based on energy of the reconstructed current frame and energy of the current frame.

19 . A non-transitory computer-readable storage medium comprising a bitstream obtained by using a method for encoding a three-dimensional audio signal, the method comprising:

obtaining a current frame of a three-dimensional audio signal;

obtaining coding efficiency of an initial virtual speaker for the current frame based on the current frame of the three-dimensional audio signal, wherein the initial virtual speaker for the current frame belongs to a set of candidate virtual speakers;

when the coding efficiency of the initial virtual speaker for the current frame meets a preset condition, determining an updated virtual speaker for the current frame from the set of candidate virtual speakers, and encoding the current frame based on the updated virtual speaker for the current frame, to obtain a first bitstream; and

when the coding efficiency of the initial virtual speaker for the current frame does not meet the preset condition, encoding the current frame based on the initial virtual speaker for the current frame, to obtain a second bitstream;

wherein the preset condition comprises the coding efficiency of the initial virtual speaker for the current frame is less than a first threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2025
From: GAO, YUAN; LIU, SHUAI; XIA, BINGYIN; WANG, BIN; WANG, ZHE
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 072794/0988 →
Priority Claims (1)
CN 202110680341.8 · Jun 18, 2021 · national
Continuity (2)
Continuation PCTCN2022096476 · May 31, 2022
Related Publication 20240119950A1 · Apr 11, 2024
References Cited (51)
US 9002716B2 · Spille · 2015 [cited by examiner]
US 9154877B2 · Kim · 2015 [cited by examiner]
US 9530421B2 · Jot · 2016 [cited by examiner]
US 9980074B2 · Sen · 2018 [cited by examiner]
US 10672405B2 · Hines · 2020 [cited by examiner]
US 11395083B2 · Schevciw · 2022 [cited by examiner]
US 11587572B2 · Wang · 2023 [cited by examiner]
US 11783843B2 · Fuchs · 2023 [cited by examiner]
US 20030138108A1 · Gentle · 2003 [cited by examiner]
US 20040088169A1 · Smith · 2004 [cited by examiner]
US 20080140426A1 · Kim · 2008 [cited by examiner]
US 20090177479A1 · Yoon · 2009 [cited by examiner]
US 20090210238A1 · Kim · 2009 [cited by examiner]
US 20100191537A1 · Breebaart · 2010 [cited by examiner]
US 20100241439A1 · Mouhssine · 2010 [cited by examiner]
US 20100305952A1 · Mouhssine · 2010 [cited by examiner]
US 20110173011A1 · Geiger · 2011 [cited by examiner]
US 20110202355A1 · Grill · 2011 [cited by examiner]
US 20110246207A1 · Choi · 2011 [cited by examiner]
US 20110264456A1 · Koppens · 2011 [cited by examiner]
US 20110305344A1 · Sole · 2011 [cited by examiner]
US 20140025386A1 · Xiang · 2014 [cited by examiner]
US 20140310010A1 · Seo · 2014 [cited by examiner]
US 20140350944A1 · Jot · 2014 [cited by examiner]
US 20140355796A1 · Xiang · 2014 [cited by examiner]
US 20140358564A1 · Sen et al. · 2014 [cited by applicant]
US 20150131824A1 · Nguyen · 2015 [cited by examiner]
US 20150142453A1 · Oomen · 2015 [cited by examiner]
US 20150149187A1 · Kastner · 2015 [cited by examiner]
US 20150154965A1 · Wuebbolt · 2015 [cited by examiner]
US 20150154971A1 · Boehm · 2015 [cited by examiner]
US 20150170657A1 · Thompson · 2015 [cited by examiner]
US 20150194161A1 · Najaf-Zadeh · 2015 [cited by examiner]
US 20150199973A1 · Borsum · 2015 [cited by examiner]
US 20150213803A1 · Peters · 2015 [cited by examiner]
US 20150332683A1 · Kim · 2015 [cited by examiner]
US 20150340043A1 · Koppens · 2015 [cited by examiner]
US 20160035356A1 · Morrell · 2016 [cited by examiner]
US 20160163321A1 · Arnott · 2016 [cited by examiner]
US 20180308500A1 · Krueger et al. · 2018 [cited by applicant]
US 20190147892A1 · Setiawan · 2019 [cited by examiner]
US 20190239015A1 · Schevciw et al. · 2019 [cited by applicant]
US 20200176001A1 · Wang · 2020 [cited by examiner]
US 20220139409A1 · Fuchs · 2022 [cited by examiner]
US 20230133252A1 · Gao · 2023 [cited by examiner]
US 20230298601A1 · Gao · 2023 [cited by examiner]
US 20240119950A1 · Gao · 2024 [cited by examiner]
CN 105940447B · 2020 [cited by applicant]
CN 114582357A · 2022 [cited by applicant]
WO 2020105423A1 · 2020 [cited by applicant]
3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Codec for Immersive Voice and Audio Services; General overview (Release 18). 3GPP TS 26.250 V1.0.0 (Sep. 2023). total 12 pag… [cited by applicant]