IP Library › Granted Patent US 12,738,287
Granted Patent B2
US 12,738,287 · App. 18/643,717 · Granted Sep 15, 2026

Audio decoding method, electronic device, and computer-readable storage medium based on label information vector

Inventors: Yupeng Shi (Shenzhen, CN); Wei Xiao (Shenzhen, CN); Meng Wang (Shenzhen, CN); Yuyong Kang (Shenzhen, CN); Qingbo Huang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L19/06G10L19/0204G10L21/02G10L19/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,287
App. No.
18/643,717
Granted
Sep 15, 2026
Kind
B2
Abstract

Embodiments of this application provide an audio coding method and apparatus, an audio decoding method and apparatus, an electronic device, and a storage medium, applied to an on board scene. The audio decoding method includes obtaining a bitstream of an audio signal; performing label extraction processing on a predicted value of a feature vector of the audio signal associated with the bitstream to obtain a label information vector, a dimension of the label information vector being the same as a dimension of the predicted value of the feature vector; performing signal reconstruction based on the predicted value of the feature vector and the label information vector; and identifying a predicted value of the audio signal obtained by the signal reconstruction as a decoding result of the bitstream.

Claims (111)

1 . An audio decoding method, executed by an electronic device, and comprising:

obtaining a bitstream of an audio signal;

performing label extraction processing on a predicted value of a feature vector of the audio signal associated with the bitstream to obtain a label information vector, comprising:

performing convolution processing on the predicted value of the feature vector to obtain a first tensor having a same dimension as the predicted value of the feature vector;

performing feature extraction processing on the first tensor to obtain a second tensor having a same dimension as the first tensor;

performing full-connection processing on the second tensor to obtain a third tensor having a same dimension as the second tensor; and

activating the third tensor to obtain the label information vector, a dimension of the label information vector being the same as a dimension of the predicted value of the feature vector;

performing signal reconstruction based on the predicted value of the feature vector and the label information vector; and

identifying a predicted value of the audio signal obtained by the signal reconstruction as a decoding result of the bitstream.

2 . The method according to claim 1 , wherein the performing signal reconstruction based on the predicted value of the feature vector and the label information vector comprises:

splicing the predicted value of the feature vector and the label information vector to obtain a spliced vector; and

compressing the spliced vector to obtain the predicted value of the audio signal.

3 . The method according to claim 1 , further comprising:

decoding the bitstream to obtain an index value of the feature vector of the audio signal; and

querying a quantization table based on the index value to obtain the predicted value of the feature vector of the audio signal.

4 . The method according to claim 2 , wherein the compressing the spliced vector to obtain the predicted value of the audio signal comprises:

performing first convolution processing on the spliced vector to obtain a convolution feature of the audio signal;

upsampling the convolution feature to obtain an upsampled feature of the audio signal;

performing pooling processing on the upsampled feature to obtain a pooled feature of the audio signal; and

performing second convolution processing on the pooled feature to obtain the predicted value of the audio signal.

5 . The method according to claim 4 , wherein

the upsampling process uses a plurality of cascaded decoding layers, and sampling factors of different decoding layers are different; and

the upsampling the convolution feature to obtain an upsampled feature of the audio signal comprises:

upsampling the convolution feature by using the first decoding layer among the plurality of cascaded decoding layers;

outputting an upsampling result of the first decoding layer to a subsequent cascaded decoding layer, and repeating upsampling processing and outputting upsampling result by using the subsequent cascaded decoding layer until an output reaches the last decoding layer; and

identifying an upsampling result output by the last decoding layer as the upsampled feature of the audio signal.

6 . The method according to claim 3 , wherein

the bitstream comprises a low-frequency bitstream and a high-frequency bitstream, the low-frequency bitstream being obtained by coding a low-frequency sub-band signal obtained by decomposing the audio signal, and the high-frequency bitstream being obtained by coding a high-frequency sub-band signal obtained by decomposing the audio signal; and

the decoding the bitstream to obtain a predicted value of a feature vector of the audio signal comprises:

decoding the low-frequency bitstream to obtain a predicted value of a feature vector of the low-frequency sub-band signal; and

decoding the high-frequency bitstream to obtain a predicted value of a feature vector of the high-frequency sub-band signal.

7 . The method according to claim 6 , wherein the performing label extraction processing on the predicted value of the feature vector to obtain a label information vector comprises:

performing label extraction processing on the predicted value of the feature vector of the low-frequency sub-band signal to obtain a first label information vector, a dimension of the first label information vector being the same as a dimension of the predicted value of the feature vector of the low-frequency sub-band signal; and

performing label extraction processing on the predicted value of the feature vector of the high-frequency sub-band signal to obtain a second label information vector, a dimension of the second label information vector being the same as a dimension of the predicted value of the feature vector of the high-frequency sub-band signal.

8 . The method according to claim 6 , wherein the performing label extraction processing on the predicted value of the feature vector of the low-frequency sub-band signal to obtain a first label information vector comprises:

invoking a first enhancement network to perform the following processing:

performing convolution processing on the predicted value of the feature vector of the low-frequency sub-band signal to obtain a fourth tensor having a same dimension as the predicted value of the feature vector of the low-frequency sub-band signal;

performing feature extraction processing on the fourth tensor to obtain a fifth tensor having a same dimension as the fourth tensor;

performing full-connection processing on the fifth tensor to obtain a sixth tensor having a same dimension as the fifth tensor; and

activating the sixth tensor to obtain the first label information vector.

9 . The method according to claim 6 , wherein the performing label extraction processing on the predicted value of the feature vector of the high-frequency sub-band signal to obtain a second label information vector comprises:

invoking a second enhancement network to perform the following processing:

performing convolution processing on the predicted value of the feature vector of the high-frequency sub-band signal to obtain a seventh tensor having a same dimension as the predicted value of the feature vector of the high-frequency sub-band signal;

performing feature extraction processing on the seventh tensor to obtain an eighth tensor having a same dimension as the seventh tensor;

performing full-connection processing on the eighth tensor to obtain a ninth tensor having a same dimension as the eighth tensor; and

activating the ninth tensor to obtain the second label information vector.

10 . The method according to claim 7 , wherein

the predicted value of the feature vector comprises: the predicted value of the feature vector of the low-frequency sub-band signal and the predicted value of the feature vector of the high-frequency sub-band signal; and

the performing signal reconstruction based on the predicted value of the feature vector and the label information vector comprises:

splicing the predicted value of the feature vector of the low-frequency sub-band signal and the first label information vector to obtain a first spliced vector;

invoking, based on the first spliced vector, a first synthesis network for signal reconstruction to obtain a predicted value of the low-frequency sub-band signal;

splicing the predicted value of the feature vector of the high-frequency sub-band signal and the second label information vector to obtain a second spliced vector;

invoking, based on the second spliced vector, a second synthesis network for signal reconstruction to obtain a predicted value of the high-frequency sub-band signal; and

synthesizing the predicted value of the low-frequency sub-band signal and the predicted value of the high-frequency sub-band signal to obtain the predicted value of the audio signal.

11 . The method according to claim 10 , wherein the invoking, based on the first spliced vector, a first synthesis network for signal reconstruction to obtain a predicted value of the low-frequency sub-band signal comprises:

invoking the first synthesis network to perform the following processing:

performing first convolution processing on the first spliced vector to obtain a convolution feature of the low-frequency sub-band signal;

upsampling the convolution feature to obtain an upsampled feature of the low-frequency sub-band signal;

performing pooling processing on the upsampled feature to obtain a pooled feature of the low-frequency sub-band signal; and

performing second convolution processing on the pooled feature to obtain the predicted value of the low-frequency sub-band signal,

the upsampling process being implemented by using a plurality of cascaded decoding layers, and sampling factors of different decoding layers being different.

12 . The method according to claim 10 , wherein the invoking, based on the second spliced vector, a second synthesis network for signal reconstruction to obtain a predicted value of the high-frequency sub-band signal comprises:

invoking the second synthesis network to perform the following processing:

performing first convolution processing on the second spliced vector to obtain a convolution feature of the high-frequency sub-band signal;

upsampling the convolution feature to obtain an upsampled feature of the high-frequency sub-band signal;

performing pooling processing on the upsampled feature to obtain a pooled feature of the high-frequency sub-band signal; and

performing second convolution processing on the pooled feature to obtain the predicted value of the high-frequency sub-band signal,

the upsampling process being implemented by using the plurality of cascaded decoding layers, and sampling factors of different decoding layers being different.

13 . The method according to claim 3 , wherein

the bitstream comprises N sub-bitstreams, the N sub-bitstreams corresponding to different frequency bands and being obtained by coding N sub-band signals obtained by decomposing the audio signal, and N being an integer greater than 2; and

the decoding the bitstream to obtain a predicted value of a feature vector of the audio signal comprises:

decoding the N sub-bitstreams respectively to obtain predicted values of feature vectors corresponding to the N sub-band signals, respectively.

14 . The method according to claim 13 , wherein the performing label extraction processing on the predicted value of the feature vector to obtain a label information vector comprises:

performing label extraction processing on the predicted values of the feature vectors corresponding to the N sub-band signals respectively to obtain N label information vectors, a dimension of each label information vector being the same as a dimension of the predicted value of the feature vector corresponding to the sub-band signal.

15 . The method according to claim 14 , wherein the performing label extraction processing on the predicted values of the feature vectors corresponding to the N sub-band signals respectively to obtain N label information vectors comprises:

invoking, based on a predicted value of a feature vector of an i th sub-band signal, an i th enhancement network for label extraction processing to obtain an i th label information vector,

a value range of i satisfying that i is greater than or equal to 1 and is smaller than or equal to N, and a dimension of the i th label information vector being the same as a dimension of the predicted value of the feature vector of the i th sub-band signal.

16 . The method according to claim 15 , wherein the invoking, based on a predicted value of a feature vector of an i th sub-band signal, an i th enhancement network for label extraction processing to obtain an i th label information vector comprises:

invoking the i th enhancement network to perform the following processing:

performing convolution processing on the predicted value of the feature vector of the i th sub-band signal to obtain a tenth tensor having a same dimension as the predicted value of the feature vector of the i th sub-band signal;

performing feature extraction processing on the tenth tensor to obtain an eleventh tensor having a same dimension as the tenth tensor;

performing full-connection processing on the eleventh tensor to obtain a twelfth tensor having a same dimension as the eleventh tensor; and

activating the twelfth tensor to obtain the i th label information vector.

17 . The method according to claim 14 , wherein the performing signal reconstruction based on the predicted value of the feature vector and the label information vector comprises:

splicing the predicted values of the feature vectors corresponding to the N sub-band signals respectively and the N label information vectors one-to-one to obtain N spliced vectors;

invoking, based on a j th spliced vector, a j th synthesis network for signal reconstruction to obtain a predicted value of a j th sub-band signal, a value range of j satisfying that j is greater than or equal to 1 and is smaller than or equal to N;

performing first convolution processing on the j th spliced vector to obtain a convolution feature of the j th sub-band signal;

upsampling the convolution feature to obtain an upsampled feature of the j th sub-band signal;

performing pooling processing on the upsampled feature to obtain a pooled feature of the j th sub-band signal; and

performing second convolution processing on the pooled feature to obtain the predicted value of the j th sub-band signal, the upsampling process being implemented by using the plurality of cascaded decoding layers, and sampling factors of different decoding layers being different; and

synthesizing predicted values corresponding to the N sub-band signals respectively to obtain the predicted value of the audio signal.

18 . An electronic device, comprising:

a memory, configured to store computer-executable instructions; and

a processor, configured, when executing the computer-executable instructions stored in the memory, to implement:

obtaining a bitstream of an audio signal;

performing label extraction processing on a predicted value of a feature vector of the audio signal associated with the bitstream to obtain a label information vector, comprising:

performing convolution processing on the predicted value of the feature vector to obtain a first tensor having a same dimension as the predicted value of the feature vector;

performing feature extraction processing on the first tensor to obtain a second tensor having a same dimension as the first tensor;

performing full-connection processing on the second tensor to obtain a third tensor having a same dimension as the second tensor; and

activating the third tensor to obtain the label information vector, a dimension of the label information vector being the same as a dimension of the predicted value of the feature vector;

performing signal reconstruction based on the predicted value of the feature vector and the label information vector; and

identifying a predicted value of the audio signal obtained by the signal reconstruction as a decoding result of the bitstream.

19 . A non-transitory computer-readable storage medium, having computer-executable instructions stored thereon, the computer-executable instructions, when being executed by a processor, causing the processor to implement:

obtaining a bitstream of an audio signal;

performing label extraction processing on a predicted value of a feature vector of the audio signal associated with the bitstream to obtain a label information vector, comprising:

performing convolution processing on the predicted value of the feature vector to obtain a first tensor having a same dimension as the predicted value of the feature vector;

performing feature extraction processing on the first tensor to obtain a second tensor having a same dimension as the first tensor;

performing full-connection processing on the second tensor to obtain a third tensor having a same dimension as the second tensor; and

activating the third tensor to obtain the label information vector, a dimension of the label information vector being the same as a dimension of the predicted value of the feature vector;

performing signal reconstruction based on the predicted value of the feature vector and the label information vector; and

identifying a predicted value of the audio signal obtained by the signal reconstruction as a decoding result of the bitstream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2024
From: SHI, YUPENG; XIAO, WEI; WANG, MENG; KANG, YUYONG; HUANG, QINGBO
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 067213/0075 →
Priority Claims (1)
CN 202210676984.X · Jun 15, 2022 · national
Continuity (2)
Continuation PCTCN2023092246 · May 5, 2023
Related Publication 20240274144A1 · Aug 15, 2024
References Cited (37)
US 11929085B2 · Biswas · 2024 [cited by examiner]
US 20050261897A1 · Jelinek · 2005 [cited by examiner]
US 20150179182A1 · Vinton · 2015 [cited by applicant]
US 20210166701A1 · Lim et al. · 2021 [cited by applicant]
US 20210327445A1 · Biswas · 2021 [cited by examiner]
US 20210343301A1 · Gao · 2021 [cited by examiner]
US 20210343303A1 · Gao · 2021 [cited by examiner]
US 20220108681A1 · Chang · 2022 [cited by examiner]
US 20220148613A1 · Xiao et al. · 2022 [cited by applicant]
US 20220180881A1 · Xiao et al. · 2022 [cited by applicant]
US 20220223161A1 · Fuchs · 2022 [cited by examiner]
US 20230099343A1 · Wang et al. · 2023 [cited by applicant]
US 20230154474A1 · Feng · 2023 [cited by examiner]
US 20230298593A1 · Ramos · 2023 [cited by examiner]
US 20240021210A1 · Biswas · 2024 [cited by examiner]
US 20240055006A1 · Biswas · 2024 [cited by examiner]
US 20250022477A1 · Omran · 2025 [cited by examiner]
AU 2020271965A1 · 2021 [cited by applicant]
CN 101202043A · 2008 [cited by applicant]
CN 101572586A · 2009 [cited by applicant]
CN 108986835A · 2018 [cited by applicant]
CN 113140225A · 2021 [cited by applicant]
CN 113470667A · 2021 [cited by applicant]
CN 113488063A · 2021 [cited by applicant]
CN 113763973A · 2021 [cited by applicant]
CN 113990347A · 2022 [cited by applicant]
CN 114550732A · 2022 [cited by applicant]
CN 115116451A · 2022 [cited by applicant]
WO 2020208137A1 · 2020 [cited by applicant]
Kolbãlk, Morten, Zheng-Hua Tan, and Jesper Jensen. “Speech intelligibility potential of general and specialized deep neural network based speech enhancement systems.” IEEE/ACM Transactions on Audio, Speech, and Language… [cited by examiner]
Omran, Ahmed, et al. “Disentangling speech from surroundings in a neural audio codec.” arXiv preprint ArXiv:2203.15578, Mar. 2022, pp. 1-5. (Year: 2022). [cited by examiner]
Park, Se Rim, and Jinwon Lee. “A fully convolutional neural network for speech enhancement.” arXiv preprint arXiv:1609.07132, Sep. 2016, pp. 1-6. (Year: 2016). [cited by examiner]
Qi, Jun, et al. “Exploring deep hybrid tensor-to-vector network architectures for regression based speech enhancement.” arXiv preprint arXiv:2007.13024, Aug. 2020, pp. 1-5. (Year: 2020). [cited by examiner]
Xian, Yang, et al. “Convolutional fusion network for monaural speech enhancement.” Neural Networks 143, Nov. 2021, pp. 97-107. (Year: 2021). [cited by examiner]
Tan, Ke, et al. “Compressing deep neural networks for efficient speech enhancement.” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 8358-8362. (Year: … [cited by examiner]
The European Patent Office (EPO) The Extended European Search Report for Application No. 23822825.8, Jan. 21, 2025 8 Pages. [cited by applicant]
The World Intellectual Property Organization (WIPO) International Search Report for PCT/CN2023/092246 Aug. 7, 2023 7 Pages (including translation). [cited by applicant]