IP Library Granted Patent US 12,688,860
Granted Patent B2
US 12,688,860 · App. 18/689,054 · Granted Jul 21, 2026

Audio coding using combination of machine learning based time-varying filter and linear predictive coding filter

Inventors: Duminda Dewasurendra (San Diego, CA); Guillaume Konrad Sautiere (Amsterdam, NL); Zisis Iason Skordilis (San Diego, CA); Vivek Rajendran (San Diego, CA)
Assignee: QUALCOMM Incorporated
G10L19/08G10L19/12G10L25/24G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,860
App. No.
18/689,054
Granted
Jul 21, 2026
Kind
B2
Abstract

Systems and techniques are described for coding audio signals. For example, a voice decoder can generate, using a neural network, an excitation signal for at least one sample of an audio signal based on one or more inputs to the neural network, the excitation signal being configured to excite a linear predictive coding (LPC) filter. The voice decoder can further generate, using the LPC filter based on the excitation signal, at least one sample of a reconstructed audio signal. For example, the neural network can generate coefficients for one or more linear time-varying filters (e.g., a linear time-varying harmonic filter and a linear time-varying noise filter). The voice decoder can use the one or more linear time-varying filters including the generated coefficients to generate the excitation signal.

Claims (68)

1 . An apparatus for reconstructing one or more audio signals, comprising:

at least one memory configured to store audio data; and

at least one processor coupled to the at least one memory, the at least one processor configured to:

generate, using a neural network, coefficients for one or more linear time-varying filters;

generate, using the one or more linear time-varying filters and based on the coefficients, an excitation signal for at least one sample of an audio signal based on one or more inputs to the neural network, the excitation signal being configured to excite a linear predictive coding (LPC) filter; and

generate, using the LPC filter based on the excitation signal, at least one sample of a reconstructed audio signal.

2 . The apparatus of claim 1 , wherein the one or more inputs to the neural network include features associated with the audio signal.

3 . The apparatus of claim 2 , wherein the features include log-Mel-frequency spectrum features.

4 . The apparatus of claim 1 , wherein the LPC filter is a time-varying LPC filter.

5 . The apparatus of claim 1 , wherein the at least one processor is configured to:

use filter coefficients of the LPC filter to generate the at least one sample of the reconstructed audio signal.

6 . The apparatus of claim 5 , wherein the filter coefficients of the LPC filter are generated based on an autocorrelation of an input audio signal in a voice encoder.

7 . The apparatus of claim 5 , wherein the at least one processor is configured to:

derive the filter coefficients of the LPC filter based on features received from a voice encoder.

8 . The apparatus of claim 7 , wherein the features include Mel spectrum features.

9 . The apparatus of claim 1 , wherein the at least one processor is configured to:

input a pulse train signal based on pitch features to a harmonic filter generated using the neural network;

generate a harmonic filter output;

input a random noise signal to a noise filter generated using the neural network; and

generate a noise filter output; and

wherein, to generate the excitation signal, the at least one processor is configured to combine the harmonic filter output with the noise filter output.

10 . The apparatus of claim 1 , wherein, to generate the excitation signal for the at least one sample of the audio signal using the neural network, the at least one processor is configured to:

generate, using the neural network, coefficients for one or more linear time-varying filters; and

generate, using the one or more linear time-varying filters including the generated coefficients, the excitation signal.

11 . The apparatus of claim 10 , wherein the one or more linear time-varying filters include a linear time-varying harmonic filter and a linear time-varying noise filter.

12 . The apparatus of claim 1 , wherein, to generate the excitation signal for the at least one sample of the audio signal using the neural network, the at least one processor is configured to:

generate, using the neural network, an additional excitation signal for a linear time-invariant filter; and

generate, using the linear time-invariant filter based on the additional excitation signal, the excitation signal.

13 . A method of reconstructing one or more audio signals, the method comprising:

generating, using a neural network, coefficients for one or more linear time-varying filters;

generating, using the one or more linear time-varying filters and based on the coefficients, an excitation signal for at least one sample of an audio signal based on one or more inputs to the neural network, the excitation signal being configured to excite a linear predictive coding (LPC) filter; and

generating, using the LPC filter based on the excitation signal, at least one sample of a reconstructed audio signal.

14 . The method of claim 13 , wherein the one or more inputs to the neural network include features associated with the audio signal.

15 . The method of claim 14 , wherein the features include log-Mel-frequency spectrum features.

16 . The method of claim 13 , wherein the LPC filter is a time-varying LPC filter.

17 . The method of claim 13 , further comprising:

using filter coefficients of the LPC filter to generate the at least one sample of the reconstructed audio signal.

18 . The method of claim 17 , wherein the filter coefficients of the LPC filter are generated based on an autocorrelation of an input audio signal in a voice encoder.

19 . The method of claim 17 , further comprising:

deriving the filter coefficients of the LPC filter based on features received from a voice encoder.

20 . The method of claim 19 , wherein the features include Mel spectrum features.

21 . The method of claim 13 , further comprising:

inputting a pulse train signal based on pitch features to a harmonic filter generated using the neural network;

generating a harmonic filter output;

inputting a random noise signal to a noise filter generated using the neural network;

generating a noise filter output; and

generating the excitation signal at least in part by combining the harmonic filter output with the noise filter output.

22 . The method of claim 13 , wherein generating the excitation signal for the at least one sample of the audio signal using the neural network includes:

generating, using the neural network, coefficients for one or more linear time-varying filters; and

generating, using the one or more linear time-varying filters including the generated coefficients, the excitation signal.

23 . The method of claim 22 , wherein the one or more linear time-varying filters include a linear time-varying harmonic filter and a linear time-varying noise filter.

24 . The method of claim 13 , wherein generating the excitation signal for the at least one sample of the audio signal using the neural network includes:

generating, using the neural network, an additional excitation signal for a linear time-invariant filter; and

generating, using the linear time-invariant filter based on the additional excitation signal, the excitation signal.

25 . An apparatus for reconstructing one or more audio signals, comprising:

at least one memory configured to store audio data; and

at least one processor coupled to the at least one memory, the at least one processor configured to:

generate, using a linear predictive coding (LPC) filter based on an excitation signal, a predicted signal for at least one sample of an audio signal, the predicted signal being configured to excite a linear time-varying filter;

generate, using a neural network, coefficients for the linear time-varying filter; and

generate, using the linear time-varying filter based on the coefficients, at least one sample of a reconstructed audio signal.

26 . The apparatus of claim 25 , wherein one or more inputs to the neural network include features associated with the audio signal.

27 . The apparatus of claim 25 , wherein the LPC filter is a time-varying LPC filter.

28 . A method of reconstructing one or more audio signals, comprising:

generating, using a linear predictive coding (LPC) filter based on an excitation signal, a predicted signal for at least one sample of an audio signal, the predicted signal being configured to excite a linear time-varying filter;

generating, using a neural network, coefficients for the linear time-varying filter; and

generating, using the linear time-varying filter based on the coefficients, at least one sample of a reconstructed audio signal.

29 . The method of claim 28 , wherein one or more inputs to the neural network include features associated with the audio signal.

30 . The method of claim 28 , wherein the LPC filter is a time-varying LPC filter.