IP Library › Granted Patent US 12,731,593
Granted Patent B2
US 12,731,593 · App. 18/640,724 · Granted Sep 8, 2026

Method and apparatus for neural network-based audio encoding using subband decomposition

Inventors: Meng Wang (Shenzhen, CN); Shan Yang (Shenzhen, CN); Qingbo Huang (Shenzhen, CN); Yuyong Kang (Shenzhen, CN); Yupeng Shi (Shenzhen, CN); Wei Xiao (Shenzhen, CN); Shidong Shang (Shenzhen, CN); Dan Su (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L19/032G10L19/0204G10L21/038G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,593
App. No.
18/640,724
Granted
Sep 8, 2026
Kind
B2
Abstract

An audio processing method and apparatus, including decomposing an audio signal into a low-frequency subband signal and a high-frequency subband signal, obtaining a low-frequency feature of the low-frequency subband signal, obtaining a high-frequency feature of the high-frequency subband signal, feature dimensionality of the high-frequency feature being lower than feature dimensionality of the low-frequency feature, performing quantization encoding on the low-frequency feature to obtain a low-frequency bitstream of the audio signal, and performing quantization encoding on the high-frequency feature to obtain a high-frequency bitstream of the audio signal.

Claims (105)

1 . An audio processing method, performed by an electronic device, comprising:

decomposing an audio signal into a low-frequency subband signal and a high-frequency subband signal;

obtaining a low-frequency feature of the low-frequency subband signal,

wherein the obtaining the low-frequency feature comprises:

performing convolution on the low-frequency subband signal to obtain a convolution feature of the low-frequency subband signal, and

performing pooling on the convolution feature to obtain a pooling feature of the low-frequency subband signal;

obtaining a high-frequency feature of the high-frequency subband signal,

wherein a feature dimensionality of the high-frequency feature is lower than a feature dimensionality of the low-frequency feature;

performing quantization encoding on the low-frequency feature to obtain a low-frequency bitstream of the audio signal; and

performing quantization encoding on the high-frequency feature to obtain a high-frequency bitstream of the audio signal.

2 . The audio processing method according to claim 1 , wherein

the decomposing comprises:

obtaining a sampled signal of the audio signal, the sampled signal comprising a plurality of sample points obtained through sampling;

performing low-pass filtering on the sampled signal to obtain a low-pass filtered signal;

downsampling the low-pass filtered signal to obtain the low-frequency subband signal of the audio signal;

performing high-pass filtering on the sampled signal to obtain a high-pass filtered signal; and

downsampling the high-pass filtered signal to obtain the high-frequency subband signal of the audio signal.

3 . The audio processing method according to claim 1 , wherein the obtaining the low-frequency feature further comprises:

downsampling the pooling feature to obtain a downsampling feature of the low-frequency subband signal; and

performing convolution on the downsampling feature to obtain the low-frequency feature of the low-frequency subband signal.

4 . The audio processing method according to claim 3 , wherein

the downsampling the pooling feature is implemented through a plurality of concatenated encoding layers; and

wherein the downsampling the pooling feature comprises:

downsampling the pooling feature through a first encoding layer of the plurality of concatenated encoding layers;

outputting a downsampling result of the first encoding layer to a subsequent concatenated encoding layer, and continuing to perform downsampling and outputting for each subsequent concatenated encoding layer of remaining concatenated encoding layers of the plurality of concatenated encoding layers, until a last encoding layer of the plurality of concatenated encoding layers performs output of the downsampling result; and

setting the downsampling result outputted by the last encoding layer as the downsampling feature of the low-frequency subband signal.

5 . The audio processing method according to claim 1 , wherein the obtaining the high-frequency feature comprises:

a first neural network model to extract the high-frequency feature of the high-frequency subband signal; or

performing bandwidth extension on the high-frequency subband signal to obtain the high-frequency feature of the high-frequency subband signal.

6 . The audio processing method according to claim 5 , wherein the performing the bandwidth extension on the high-frequency subband signal comprises:

performing frequency domain transform based on a plurality of sample points comprised in the high-frequency subband signal to obtain transform coefficients respectively corresponding to the plurality of sample points;

dividing the transform coefficients into a plurality of subbands;

calculating a mean based on the transform coefficients comprised in each subband, setting the mean as an average energy corresponding to each subband, and setting the average energy as a subband spectral envelope corresponding to each subband; and

setting subband spectral envelopes respectively corresponding to the plurality of subbands as the high-frequency feature of the high-frequency subband signal.

7 . The audio processing method according to claim 6 , wherein the performing the frequency domain transform comprises:

obtaining a reference high-frequency subband signal of a reference audio signal, the reference audio signal being another audio signal adjacent to the audio signal; and

performing, based on a plurality of sample points comprised in the reference high-frequency subband signal and the plurality of sample points comprised in the high-frequency subband signal, discrete cosine transform on the plurality of sample points comprised in the high-frequency subband signal to obtain the transform coefficients respectively corresponding to the plurality of sample points comprised in the high-frequency subband signal.

8 . The audio processing method according to claim 6 , wherein the calculating the mean comprises:

determining a sum of squares of the transform coefficients corresponding to the sample points comprised in each subband; and

setting a ratio of the sum of squares of the transform coefficients to a quantity of the sample points comprised in each subband as the average energy corresponding to each subband.

9 . The audio processing method according to claim 1 , wherein the performing the quantization encoding on the low-frequency feature comprises:

quantizing the low-frequency feature into an index value of the low-frequency feature; and

performing entropy encoding on the index value of the low-frequency feature to obtain the low-frequency bitstream of the audio signal; and

wherein the performing the quantization encoding on the high-frequency feature comprises:

quantizing the high-frequency feature into an index value of the high-frequency feature; and

performing entropy encoding on the index value of the high-frequency feature to obtain the high-frequency bitstream of the audio signal.

10 . An audio processing apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

decomposition code configured to cause at least one of the at least one processor to decompose an audio signal into a low-frequency subband signal and a high-frequency subband signal;

feature extraction code configured to cause at least one of the at least one processor to:

perform convolution on the low-frequency subband signal to obtain a convolution feature of the low-frequency subband signal;

perform pooling on the convolution feature to obtain a pooling feature of the low-frequency subband signal; and

obtain a low-frequency feature of the low-frequency subband signal based on the pooling feature of the low-frequency subband signal;

high-frequency analysis code configured to cause at least one of the at least one processor to obtain a high-frequency feature of the high-frequency subband signal,

wherein a feature dimensionality of the high-frequency feature is lower than a feature dimensionality of the low-frequency feature; and

encoding code configured to cause at least one of the at least one processor to perform quantization encoding on the low-frequency feature to obtain a low-frequency bitstream of the audio signal, and perform quantization encoding on the high-frequency feature to obtain a high-frequency bitstream of the audio signal.

11 . The audio processing apparatus according to claim 10 , wherein the decomposition code is further configured to cause at least one of the at least one processor to:

obtain a sampled signal of the audio signal, the sampled signal comprising a plurality of sample points obtained through sampling;

perform low-pass filtering on the sampled signal to obtain a low-pass filtered signal;

downsample the low-pass filtered signal to obtain the low-frequency subband signal of the audio signal;

perform high-pass filtering on the sampled signal to obtain a high-pass filtered signal; and

downsample the high-pass filtered signal to obtain the high-frequency subband signal of the audio signal.

12 . The audio processing apparatus according to claim 10 , wherein the feature extraction code is further configured to cause at least one of the at least one processor to:

downsample the pooling feature to obtain a downsampling feature of the low-frequency subband signal; and

perform convolution on the downsampling feature to obtain the low-frequency feature of the low-frequency subband signal.

13 . The audio processing apparatus according to claim 12 , wherein the downsampling is implemented through a plurality of concatenated encoding layers;

wherein the pooling feature is downsampled through a first encoding layer of the plurality of concatenated encoding layers; and

wherein the feature extraction code is further configured to cause at least one of the at least one processor to:

output a downsampling result of the first encoding layer to a subsequent concatenated encoding layer, and continue to perform the downsample and the output for each subsequent concatenated encoding layer of remaining concatenated encoding layers of the plurality of concatenated encoding layers, until a last encoding layer of the plurality of concatenated encoding layers performs output of the downsampling result; and

set the downsampling result outputted by the last encoding layer as the downsampling feature of the low-frequency subband signal.

14 . The audio processing apparatus according to claim 10 , wherein the high-frequency analysis code is further configured to cause at least one of the at least one processor to:

call a first neural network model to extract the high-frequency feature of the high-frequency subband signal; or

perform bandwidth extension on the high-frequency subband signal to obtain the high-frequency feature of the high-frequency subband signal.

15 . The audio processing apparatus according to claim 14 , wherein the high-frequency analysis code is further configured to cause at least one of the at least one processor to:

perform frequency domain transform based on a plurality of sample points comprised in the high-frequency subband signal to obtain transform coefficients respectively corresponding to the plurality of sample points;

divide the transform coefficients into a plurality of subbands;

calculate a mean based on the transform coefficients comprised in each subband, set the mean as an average energy corresponding to each subband, and set the average energy as a subband spectral envelope corresponding to each subband; and

set subband spectral envelopes respectively corresponding to the plurality of subbands as the high-frequency feature of the high-frequency subband signal.

16 . The audio processing apparatus according to claim 15 , wherein the high-frequency analysis code is further configured to cause at least one of the at least one processor to:

obtain a reference high-frequency subband signal of a reference audio signal, the reference audio signal being another audio signal adjacent to the audio signal; and

perform, based on a plurality of sample points comprised in the reference high-frequency subband signal and the plurality of sample points comprised in the high-frequency subband signal, discrete cosine transform on the plurality of sample points comprised in the high-frequency subband signal to obtain the transform coefficients respectively corresponding to the plurality of sample points comprised in the high-frequency subband signal.

17 . The audio processing apparatus according to claim 15 , wherein the high-frequency analysis code is further configured to cause at least one of the at least one processor to:

determine a sum of squares of the transform coefficients corresponding to the sample points comprised in each subband; and

set a ratio of the sum of squares of the transform coefficients to a quantity of the sample points comprised in each subband as the average energy corresponding to each subband.

18 . The audio processing apparatus according to claim 10 , wherein the encoding code is further configured to cause at least one of the at least one processor to:

quantize the low-frequency feature into an index value of the low-frequency feature;

perform entropy encoding on the index value of the low-frequency feature to obtain the low-frequency bitstream of the audio signal;

quantize the high-frequency feature into an index value of the high-frequency feature; and

perform entropy encoding on the index value of the high-frequency feature to obtain the high-frequency bitstream of the audio signal.

19 . A non-transitory computer-readable storage medium, storing computer code which, when executed by at least one processor, causes the at least one processor to at least:

decompose an audio signal into a low-frequency subband signal and a high-frequency subband signal;

perform convolution on the low-frequency subband signal to obtain a convolution feature of the low-frequency subband signal;

perform pooling on the convolution feature to obtain a pooling feature of the low-frequency subband signal;

obtain a low-frequency feature of the low-frequency subband signal based on the pooling feature of the low-frequency subband signal;

obtain a high-frequency feature of the high-frequency subband signal,

wherein a feature dimensionality of the high-frequency feature is lower than a feature dimensionality of the low-frequency feature;

perform quantization encoding on the low-frequency feature to obtain a low-frequency bitstream of the audio signal; and

perform quantization encoding on the high-frequency feature to obtain a high-frequency bitstream of the audio signal.

20 . The non-transitory computer-readable storage medium according to claim 19 , wherein the computer code, when executed by the at least one processor, causes the at least one processor to at least:

obtaining a sampled signal of the audio signal, the sampled signal comprising a plurality of sample points obtained through sampling;

performing low-pass filtering on the sampled signal to obtain a low-pass filtered signal;

downsampling the low-pass filtered signal to obtain the low-frequency subband signal of the audio signal;

performing high-pass filtering on the sampled signal to obtain a high-pass filtered signal; and

downsampling the high-pass filtered signal to obtain the high-frequency subband signal of the audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2024
From: WANG, MENG; YANG, SHAN; HUANG, QINGBO; KANG, YUYONG; SHI, YUPENG; XIAO, WEI; SHANG, SHIDONG; SU, DAN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 067173/0154 →
Priority Claims (1)
CN 202210681365.X · Jun 15, 2022 · national
Continuity (2)
Continuation PCTCN2023088638 · Apr 17, 2023
Related Publication 20240265929A1 · Aug 8, 2024
References Cited (26)
US 5765127A · Nishiguchi et al. · 1998 [cited by applicant]
US 9443534B2 · Gao · 2016 [cited by examiner]
US 9818419B2 · Atti · 2017 [cited by examiner]
US 11985179B1 · Tacer · 2024 [cited by examiner]
US 20100063812A1 · Gao · 2010 [cited by examiner]
US 20140074489A1 · Chong · 2014 [cited by examiner]
US 20160255452A1 · Nowak · 2016 [cited by examiner]
US 20170188175A1 · Oh · 2017 [cited by examiner]
US 20180277097A1 · Li · 2018 [cited by examiner]
US 20190258917A1 · Chai · 2019 [cited by examiner]
US 20210005209A1 · Beack · 2021 [cited by examiner]
US 20210166705A1 · Chang · 2021 [cited by examiner]
US 20230016637A1 · Schmidt · 2023 [cited by examiner]
CN 102158692A · 2011 [cited by applicant]
CN 108417219A · 2018 [cited by applicant]
CN 110556123A · 2019 [cited by applicant]
CN 113470667A · 2021 [cited by applicant]
CN 113903345A · 2022 [cited by applicant]
CN 115116456A · 2022 [cited by applicant]
WO 2016023322A1 · 2016 [cited by applicant]
Li, Kehuang, and Chin-Hui Lee. “A deep neural network approach to speech bandwidth expansion.” 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015. [cited by examiner]
J. Su, Y. Wang, A. Finkelstein and Z. Jin, “Bandwidth Extension is All You Need,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 2021, pp. 696-70… [cited by examiner]
Schmidt, Konstantin, and Bernd Edler. “Blind bandwidth extension based on convolutional and recurrent deep neural networks.” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, … [cited by examiner]
Office Action issued Jun. 15, 2024 in Chinese Application No. 202210681365.X. [cited by applicant]
International Search Report for PCT/CN2023/088638 dated Aug. 3, 2023 (PCT/ISA/210). [cited by applicant]
Written Opinion for PCT/CN2023/088638 dated Aug. 3, 2023 (PCT/ISA/237). [cited by applicant]