IP Library › Granted Patent US 12,711,973
Granted Patent B2
US 12,711,973 · App. 18/689,052 · Granted Aug 18, 2026

Systems and methods for multi-band audio coding

Inventors: Zisis Iason Skordilis (San Diego, CA); Vivek Rajendran (San Diego, CA); Duminda Dewasurendra (San Diego, CA); Guillaume Konrad Sautiere (Amsterdam, NL)
Assignee: QUALCOMM Incorporated
G10L19/08G10L19/087G10L19/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,973
App. No.
18/689,052
Granted
Aug 18, 2026
Kind
B2
Abstract

Systems and techniques are described for audio coding. An audio system receives feature(s) corresponding an audio signal, for example from an encoder and/or a speech synthesis engine. The audio system generates an excitation signal, such as a harmonic signal and/or a noise signal, based on the feature(s). The audio system uses a filterbank to generate band-specific signals from the excitation signal. The band-specific signals correspond to frequency bands. The audio system inputs the feature(s) into a machine learning (ML) filter estimator to generate parameter(s) associated with linear filter(s). The audio system inputs the feature(s) into a voicing estimator to generate gain value(s). The audio system generates an output audio signal based on modification of the band-specific signals, application of the linear filter(s) according to the parameter(s), and amplification using the gain amplifier(s) according to the gain value(s).

Claims (82)

1 . An apparatus for audio coding, the apparatus comprising:

a memory; and

one or more processors coupled to the memory, the one or more processors configured to:

receive one or more features corresponding an audio signal;

generate an excitation signal based on the one or more features;

use a filterbank to generate a plurality of band-specific signals from the excitation signal, wherein the plurality of band-specific signals correspond to a plurality of frequency bands;

use a machine learning (ML) filter estimator to generate one or more parameters associated with one or more linear filters in response to input of the one or more features to the ML filter estimator;

use a voicing estimator to generate one or more gain values associated with one or more gain amplifiers in response to input of the one or more features to the voicing estimator;

combine the plurality of band-specific signals into a combined signal;

use a second filterbank to generate a second plurality of band-specific signals from the combined signal, wherein the second plurality of band-specific signals correspond to a second plurality of frequency bands;

modify the second plurality of band-specific signals based on application of at least one of the one or more gain amplifiers to each of the second plurality of band-specific signals according to the one or more gain values; and

combine the second plurality of band-specific signals to generate an output audio signal.

2 . The apparatus of claim 1 , wherein the audio signal is a speech signal, and wherein the output audio signal is a reconstructed speech signal that is a reconstructed variant of the speech signal.

3 . The apparatus of claim 1 , wherein, to receive the one or more features, the one or more processors are configured to receive the one or more features from at least one of:

an encoder configured to generate the one or more features at least in part by encoding the audio signal; or

a speech synthesizer configured to generate the one or more features at least in part based on a text input, wherein the audio signal is an audio representation of a voice reading the text input.

4 . The apparatus of claim 1 , wherein the excitation signal is one of:

a harmonic excitation signal corresponding to a harmonic component of the audio signal; or

a noise excitation signal corresponding to a noise component of the audio signal.

5 . The apparatus of claim 1 , wherein the ML filter estimator includes one of:

one or more trained ML models; or

one or more trained neural networks.

6 . The apparatus of claim 1 , wherein the voicing estimator includes one of:

one or more trained ML models; or

one or more trained neural networks.

7 . The apparatus of claim 1 , wherein, to generate the output audio signal, the one or more processors are configured to combine the plurality of band-specific signals using a synthesis filterbank.

8 . The apparatus of claim 1 , wherein, to generate the output audio signal, the one or more processors are configured to modify the plurality of band-specific signals by applying at least one of the one or more linear filters to each of the plurality of band-specific signals according to the one or more parameters.

9 . The apparatus of claim 8 , wherein the combined signal comprises a filtered signal.

10 . The apparatus of claim 1 , wherein, to generate the output audio signal, the one or more processors are configured to modify the plurality of band-specific signals by applying at least one of the one or more gain amplifiers to each of the plurality of band-specific signals according to the one or more gain values.

11 . The apparatus of claim 10 , wherein the combined signal comprises an amplified signal.

12 . The apparatus of claim 1 , wherein the one or more processors are configured to:

modify the output audio signal using a first additional linear filter.

13 . The apparatus of claim 1 , wherein the one or more processors are configured to:

modify the excitation signal using a second additional linear filter before using the filterbank to generate the plurality of band-specific signals from the excitation signal.

14 . The apparatus of claim 1 , wherein the one or more features include one or more log-mel-frequency spectrum features.

15 . The apparatus of claim 1 , wherein the one or more parameters associated with one or more linear filters include at least one of:

an impulse response associated with the one or more linear filters;

a frequency response associated with the one or more linear filters; or

a rational transfer function coefficient associated with the one or more linear filters.

16 . A method for audio coding, the method comprising:

receiving one or more features corresponding an audio signal;

generating an excitation signal based on the one or more features;

using a filterbank to generate a plurality of band-specific signals from the excitation signal, wherein the plurality of band-specific signals correspond to a plurality of frequency bands;

using a machine learning (ML) filter estimator to generate one or more parameters associated with one or more linear filters in response to input of the one or more features to the ML filter estimator;

using a voicing estimator to generate one or more gain values associated with one or more gain amplifiers in response to input of the one or more features to the voicing estimator;

combining the plurality of band-specific signals into a combined signal;

using a second filterbank to generate a second plurality of band-specific signals from the combined signal, wherein the second plurality of band-specific signals correspond to a second plurality of frequency bands;

modifying the second plurality of band-specific signals by applying at least one of the one or more gain amplifiers to each of the second plurality of band-specific signals according to the one or more gain values; and

combining the second plurality of band-specific signals to generate an output audio signal.

17 . The method of claim 16 , wherein the audio signal is a speech signal, and wherein the output audio signal is a reconstructed speech signal that is a reconstructed variant of the speech signal.

18 . The method of claim 16 , wherein receiving the one or more features includes receiving the one or more features from at least one of:

an encoder that generates the one or more features at least in part by encoding the audio signal; or

a speech synthesizer that generates the one or more features at least in part based on a text input, wherein the audio signal is an audio representation of a voice reading the text input.

19 . The method of claim 16 , wherein the excitation signal is one of:

a harmonic excitation signal corresponding to a harmonic component of the audio signal; or

a noise excitation signal corresponding to a noise component of the audio signal.

20 . The method of claim 16 , wherein the ML filter estimator includes one of:

one or more trained ML models; or

one or more trained neural networks.

21 . The method of claim 16 , wherein the voicing estimator includes one of:

one or more trained ML models; or

one or more trained neural networks.

22 . The method of claim 16 , wherein generating the output audio signal includes combining the plurality of band-specific signals using a synthesis filterbank.

23 . The method of claim 16 , wherein generating the output audio signal includes modifying the plurality of band-specific signals by applying at least one of the one or more linear filters to each of the plurality of band-specific signals according to the one or more parameters.

24 . The method of claim 23 , wherein the combined signal comprises a filtered signal.

25 . The method of claim 16 , wherein generating the output audio signal includes modifying the plurality of band-specific signals by applying at least one of the one or more gain amplifiers to each of the plurality of band-specific signals according to the one or more gain values.

26 . The method of claim 25 , Wherein the combined signal comprises an amplified signal.

27 . The method of claim 16 , further comprising:

modifying the output audio signal using a first additional linear filter.

28 . The method of claim 16 , further comprising:

modifying the excitation signal using an additional linear filter before using the filterbank to generate the plurality of band-specific signals from the excitation signal.

29 . The method of claim 16 , wherein the one or more features include one or more log-mel-frequency spectrum features.

30 . A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to:

receive one or more features corresponding an audio signal;

generate an excitation signal based on the one or more features;

use a filterbank to generate a plurality of band-specific signals from the excitation signal, wherein the plurality of band-specific signals correspond to a plurality of frequency bands;

use a machine learning (ML) filter estimator to generate one or more parameters associated with one or more linear filters in response to input of the one or more features to the ML filter estimator;

use a voicing estimator to generate one or more gain values associated with one or more gain amplifiers in response to input of the one or more features to the voicing estimator;

combine the plurality of band-specific signals into a combined signal;

use a second filterbank to generate a second plurality of band-specific signals from the combined signal, wherein the second plurality of band-specific signals correspond to a second plurality of frequency bands;

modify the second plurality of band-specific signals based on application of at least one of the one or more gain amplifiers to each of the second plurality of band-specific signals according to the one or more gain values; and

combine the second plurality of band-specific signals to generate an output audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2024
From: SKORDILIS, ZISIS IASON; RAJENDRAN, VIVEK; DEWASURENDRA, DUMINDA; SAUTIERE, GUILLAUME KONRAD
To: QUALCOMM INCORPORATED
Reel/Frame 066756/0529 →
Priority Claims (1)
GR 20210100699 · Oct 14, 2021 · national
Continuity (1)
Related Publication 20240371384A1 · Nov 7, 2024
References Cited (30)
US 5455888A · Iyengar · 1995 [cited by examiner]
US 5682407A · Funaki · 1997 [cited by examiner]
US 5956674A · Smyth · 1999 [cited by examiner]
US 6691082B1 · Aguilar · 2004 [cited by examiner]
US 10741192B2 · Rajendran · 2020 [cited by examiner]
US 12347445B2 · Liang · 2025 [cited by examiner]
US 20080040104A1 · Ide · 2008 [cited by examiner]
US 20100057476A1 · Sudo · 2010 [cited by examiner]
US 20110295598A1 · Yang · 2011 [cited by examiner]
US 20120029926A1 · Krishnan et al. · 2012 [cited by applicant]
US 20130024191A1 · Krutsch · 2013 [cited by examiner]
US 20130144614A1 · Myllyla · 2013 [cited by examiner]
US 20150228288A1 · Subasingha · 2015 [cited by examiner]
US 20160019902A1 · Lamblin et al. · 2016 [cited by applicant]
US 20160133273A1 · Kaniewska · 2016 [cited by examiner]
US 20180308505A1 · Chebiyyam · 2018 [cited by examiner]
US 20200251119A1 · Yang · 2020 [cited by examiner]
US 20210074308A1 · Skordilis et al. · 2021 [cited by applicant]
US 20230016637A1 · Schmidt · 2023 [cited by examiner]
US 20230035504A1 · Lin · 2023 [cited by examiner]
US 20250240590A1 · Marquardt · 2025 [cited by examiner]
EP 1887566A1 · 2008 [cited by applicant]
TW 201205557A · 2012 [cited by applicant]
Valin, Jean-Marc, and Jan Skoglund. “A real-time wideband neural vocoder at 1.6 kb/s using LPCNet.” arXiv preprint arXiv: 1903.12087 (2019). [cited by examiner]
Valin, Jean-Marc, and Jan Skoglund. “LPCNet: Improving neural speech synthesis through linear prediction.” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019. [cited by examiner]
Schmidt, Konstantin. Blind Bandwidth Extension of Speech. Diss. Dissertation, Erlangen, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), 2022, 2023. [cited by examiner]
E. Song, K. Byun and H.-G. Kang, “ExcitNet Vocoder: A Neural Excitation Model for Parametric Speech Synthesis Systems,” 2019 27th European Signal Processing Conference (EUSIPCO), A Coruna, Spain, 2019, pp. 1-5, doi: 10.… [cited by examiner]
International Search Report and Written Opinion—PCT/US2022/077868—ISA/EPO—Dec. 6, 2022. [cited by applicant]
Mccree A.V., et al., “A Mixed Excitation LPC Vocoder Model for Low Bit Rate Speech Coding”, IEEE Transactions on Speech and Audio Processing, IEEE Service Center, New York, NY, US, vol. 3, No. 4, Jul. 1, 1995 (Jul. 1, 1… [cited by applicant]
Taiwan Search Report—TW111138882—TIPO—Apr. 30, 2026. [cited by applicant]