IP Library › Granted Patent US 12,354,615
Granted Patent B2
US 12,354,615 · App. 18/381,866 · Granted Jul 8, 2025

Audio decoder, method and computer program using a zero-input-response to obtain a smooth transition

Inventors: Emmanuel Ravelli (Erlangen, DE); Guillaume Fuchs (Bubenreuth, DE); Sascha Disch (Fuerth, DE); Markus Multrus (Nuremberg, DE); Grzegorz Pietrzyk (Nuremberg, DE); Benjamin Schubert (Nuremberg, DE)
Assignee: Fraunhofer-Gesellschaft zur Foerderung der angewandten Forschung E.V.
G10L19/20G10L19/02G10L19/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,615
App. No.
18/381,866
Granted
Jul 8, 2025
Kind
B2
Abstract

An audio decoder for providing a decoded audio information on the basis of an encoded audio information includes a linear-prediction-domain decoder configured to provide a first decoded audio information on the basis of an audio frame encoded in a linear prediction domain, a frequency domain decoder configured to provide a second decoded audio information on the basis of an audio frame encoded in a frequency domain, and a transition processor. The transition processor is configured to obtain a zero-input-response of a linear predictive filtering, wherein an initial state of the linear predictive filtering is defined depending on the first decoded audio information and the second decoded audio information, and modify the second decoded audio information depending on the zero-input-response, to obtain a smooth transition between the first and the modified second decoded audio information.

Claims (218)

1. An audio decoder for providing a decoded audio information on the basis of an encoded audio information, the audio decoder comprising:

a linear-prediction-domain decoder configured to provide a first decoded audio information on the basis of an audio frame encoded in a linear prediction domain;

a frequency domain decoder configured to provide a second decoded audio information on the basis of an audio frame encoded in a frequency domain,

wherein the frequency-domain decoder is configured to perform an inverse lapped transform and

a transition processor,

wherein the transition processor is configured to obtain a zero-input-response of a linear predictive filtering, wherein an initial state of the linear predictive filtering is defined in dependence on the first decoded audio information, and

wherein the transition processor is configured to obtain a windowed and time-mirrored version of the first decoded audio information, and

wherein the transition processor is configured to modify the second decoded audio information, which is provided on the basis of an audio frame encoded in the frequency domain following an audio frame encoded in the linear prediction domain, in dependence on the zero-input-response and in dependence on the windowed and time-mirrored version of the first decoded audio information.

2. The audio decoder according to claim 1 ,

wherein the transition processor is configured to obtain a first zero-input-response of a linear predictive filter in response to a first initial state of the linear predictive filter defined by the first decoded audio information, and

wherein the transition processor is configured to obtain a second zero-input-response of the linear predictive filter in response to a second initial state of the linear predictive filter defined by a modified version of the first decoded audio information, which is provided with an artificial aliasing, and which comprises a contribution of a portion of the second decoded audio information, or

wherein the transition processor is configured to obtain a combined zero-input-response of the linear predictive filter in response to an initial state of the linear predictive filter defined by a combination of the first decoded audio information and of a modified version of the first decoded audio information, which is provided with an artificial aliasing, and which comprises a contribution of a portion of the second decoded audio information;

wherein the transition processor is configured to modify the second decoded audio information, which is provided on the basis of an audio frame encoded in the frequency domain following an audio frame encoded in the linear prediction domain, in dependence on the first zero-input-response and the second zero-input-response, or in dependence on the combined zero-input-response, to obtain a smooth transition between the first decoded audio information and the modified second decoded audio information.

3. The audio decoder according to claim 1 , wherein the frequency-domain decoder is configured to perform an inverse lapped transform, such that the second decoded audio information comprises an aliasing in a time portion which is temporally overlapping with a time portion for which the linear-prediction-domain decoder provides a first decoded audio information, and such that the second decoded audio information is aliasing-free for a time portion following the time portion for which the linear-prediction-domain decoder provides a first decoded audio information.

4. The audio decoder according to claim 1 , wherein the portion of the second decoded audio information, which is used to obtain the modified version of the first decoded audio information, comprises an aliasing.

5. The audio decoder according to claim 4 , wherein the artificial aliasing, which is used to obtain the modified version of the first decoded audio information, at least partially compensates an aliasing which is comprised in the portion of the second decoded audio information, which is used to obtain the modified version of the first decoded audio information.

6. The audio decoder according to claim 1 , wherein the transition processor is configured to obtain the first zero-input-response, or a first component of the combined zero-input-response, according to

s

Z

1

(

n

)

=

-

∑

m

=

1

M

a

m

⁢

s

Z

1

(

n

-

m

)

,

n

=

0

,

...

,

N

-

1

or according to

s

Z

1

(

n

)

=

+

∑

m

=

1

M

a

m

⁢

s

Z

1

(

n

-

m

)

,

n

=

0

,

...

,

N

-

1

with

s Z 1 ( n )= S C ( n ), n=−L, . . . ,− 1

M≤L

wherein n designates a time index,

wherein s Z 1 (n) for n=0, . . . , N−1 designates the first zero input response for time index n, or a first component of the combined zero-input-response for time index n;

wherein s Z 1 (n) for n=−L, . . . , −1 designates the first initial state for time index n, or a first component of the initial state for time index n;

wherein m designates a running variable,

wherein M designates a filter length of the linear predictive filter;

wherein a m designates filter coefficients of the linear predictive filter;

wherein S c (n) designates a previously decoded value of the first decoded audio information for time index n;

wherein N designates a processing length.

7. The audio decoder according to claim 1 , wherein the transition processor is configured to apply a first windowing to the first decoded audio information, to obtain a windowed version of the first decoded audio information, and to apply a second windowing to a time-mirrored version of the first decoded audio information, to obtain a windowed version of the time-mirrored version of the first decoded audio information, and

wherein the transition processor is configured to combine the windowed version of the first decoded audio information and the windowed version of the time-mirrored version of the first decoded audio information, in order to obtain the modified version of the first decoded audio information.

8. The audio decoder according to claim 1 , wherein the transition processor is configured to obtain the modified version of the first decoded audio information according to

( n )= S C ( n ) w (− n− 1) w (− n− 1)+ S C (− n−L− 1) w ( n+L ) w (− n− 1), n=−L, . . . ,− 1

wherein n designates a time index,

wherein w(−n−1) designates a value of a window function for time index (−n−1);

wherein w(n+L) designates a value of a window function for time index (n+L);

wherein S c (n) designates a previously decoded value of the first decoded audio information for time index (n);

wherein S c (−n−L−1) designates a previously decoded value of the first decoded audio information for time index (−n−L−1);

wherein S M (n) designates a decoded value of the second decoded audio information for time index n; and

wherein L describes a length of a window.

9. The audio decoder according to claim 1 , wherein the transition processor is configured to obtain the second zero-input-response, or a second component of the combined zero-input-response according to

s

Z

2

(

n

)

=

-

∑

m

=

1

M

a

m

⁢

s

Z

2

(

n

-

m

)

,

n

=

0

,

...

,

N

-

1

or according to

s

Z

2

(

n

)

=

+

∑

m

=

1

M

a

m

⁢

s

Z

2

(

n

-

m

)

,

n

=

0

,

...

,

N

-

1

with

s Z 2 ( n )= ( n ), n=−L, . . . ,− 1

M≤L

wherein n designates a time index,

wherein s Z 2 (n) for n=0, . . . , N−1 designates the second zero input response for time index n, or a second component of the combined zero-input-response for time index n;

wherein s Z 2 (n) for n=−L, . . . , −1 designates the second initial state for time index n, or a second component of the initial state for time index n;

wherein m designates a running variable,

wherein M designates a filter length of the linear predictive filter;

wherein a m designates filter coefficients of the linear predictive filter;

wherein (n) designates values of the modified version of the first decoded audio information for time index n;

wherein N designates a processing length.

10. The audio decoder according to claim 1 , wherein the transition processor is configured to linearly combine the second decoded audio information with the first zero-input-response and the second zero-input-response, or with the combined zero-input-response, for a time portion for which no first decoded audio information is provided by the linear-prediction-domain decoder, in order to obtain the modified second decoded audio information.

11. The audio decoder according to claim 1 , wherein the transition processor is configured to obtain the modified second decoded audio information according to

( n )= S M ( n )− s Z 2 ( n )+ s Z 1 ( n ), for n= 0, . . . , N− 1,

or according to

( n )= S M ( n )− v ( n ) s Z 2 ( n )+ v ( n ) s Z 1 ( n ), for n= 0, . . . , N− 1,

wherein

wherein n designates a time index;

wherein S M (n) designates values of the second decoded audio information for time index n;

wherein s Z 1 (n) for n=0, . . . , N−1 designates the first zero input response for time index n, or a first component of the combined zero-input-response for time index n; and

wherein s Z 2 (n) for n=0, . . . , N−1 designates the second zero input response for time index n, or a second component of the combined zero-input-response for time index n;

wherein v(n) designates values of a window function;

wherein N designates a processing length.

12. The audio decoder according to claim 1 , wherein the transition processor is configured to leave the first decoded audio information unchanged by the second decoded audio information when providing a decoded audio information for an audio frame encoded in a linear-prediction domain, such that the decoded audio information provided for an audio frame encoded in the linear-prediction-domain is provided independent from decoded audio information provided for a subsequent audio frame encoded in the frequency domain.

13. The audio decoder according to claim 1 , wherein the audio decoder is configured to provide a fully decoded audio information for an audio frame encoded in the linear-prediction domain, which is followed by an audio frame encoded in the frequency domain, before decoding the audio frame encoded in the frequency domain.

14. The audio decoder according to claim 1 , wherein the transition processor is configured to window the first zero-input-response and the second zero-input-response, or the combined zero-input-response, before modifying the second decoded audio information in dependence on the windowed first zero-input-response and the windowed second zero-input-response, or in dependence on the windowed combined zero-input-response.

15. The audio decoder according to claim 14 , wherein the transition processor is configured to window the first zero-input-response and the second zero-input-response, or the combined zero-input-response, using a linear window.

16. A method for providing a decoded audio information on the basis of an encoded audio information, the method comprising:

providing a first decoded audio information on the basis of an audio frame encoded in a linear prediction domain;

providing a second decoded audio information on the basis of an audio frame encoded in a frequency domain using an inverse lapped transform; and

obtaining a zero-input-response of a linear predictive filtering, wherein an initial state of the linear predictive filtering is defined in dependence on the first decoded audio information, and

obtaining a windowed and time-mirrored version of the first decoded audio information; and

modifying the second decoded audio information, which is provided on the basis of an audio frame encoded in the frequency domain following an audio frame encoded in the linear prediction domain, in dependence on the zero-input-response and in dependence on the windowed and time-mirrored version of the first decoded audio information.

17. A non-transitory digital storage medium having a computer program stored thereon to perform the method for providing a decoded audio information on the basis of an encoded audio information, the method comprising:

providing a first decoded audio information on the basis of an audio frame encoded in a linear prediction domain;

providing a second decoded audio information on the basis of an audio frame encoded in a frequency domain using an inverse lapped transform; and

obtaining a zero-input-response of a linear predictive filtering, wherein an initial state of the linear predictive filtering is defined in dependence on the first decoded audio information, and

obtaining a windowed and time-mirrored version of the first decoded audio information; and

modifying the second decoded audio information, which is provided on the basis of an audio frame encoded in the frequency domain following an audio frame encoded in the linear prediction domain, in dependence on the zero-input-response and in dependence on the windowed and time-mirrored version of the first decoded audio information;

when said computer program is run by a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2023
From: RAVELLI, EMMANUEL; FUCHS, GUILLAUME; DISCH, SASCHA; MULTRUS, MARKUS; PIETRZYK, GRZEGORZ; SCHUBERT, BENJAMIN
To: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 065281/0598 →
Priority Claims (1)
EP 14178830 · Jul 28, 2014 · regional
Continuity (5)
Continuation 17479151 · Sep 20, 2021
Continuation 16427488 · May 31, 2019
Continuation 15416052 · Jan 26, 2017
Continuation PCTEP2015066953 · Jul 23, 2015
Related Publication 20240046941A1 · Feb 8, 2024
References Cited (139)
US 5040217A · Brandenburg et al. · 1991 [cited by applicant]
US 5657422A · Janiszewski et al. · 1997 [cited by applicant]
US 6108621A · Nishiguchi et al. · 2000 [cited by applicant]
US 6134518A · Cohen · 2000 [cited by applicant]
US 6963842B2 · Goodwin · 2005 [cited by applicant]
US 7406410B2 · Kikuiri · 2008 [cited by applicant]
US 7454330B1 · Nishiguchi et al. · 2008 [cited by applicant]
US 7873064B1 · Li et al. · 2011 [cited by applicant]
US 7873510B2 · Kurniawati et al. · 2011 [cited by applicant]
US 8484038B2 · Bessette · 2013 [cited by examiner]
US 8515767B2 · Reznik · 2013 [cited by applicant]
US 8560329B2 · Qi et al. · 2013 [cited by applicant]
US 8630862B2 · Geiger · 2014 [cited by examiner]
US 8700388B2 · Edler et al. · 2014 [cited by applicant]
US 8725503B2 · Bessette · 2014 [cited by applicant]
US 8744843B2 · Geiger et al. · 2014 [cited by applicant]
US 8744863B2 · Neuendorf et al. · 2014 [cited by applicant]
US 9280982B1 · Kushner · 2016 [cited by applicant]
US 9489962B2 · Chong et al. · 2016 [cited by applicant]
US 9583110B2 · Fuchs et al. · 2017 [cited by applicant]
US 9583114B2 · Lombard et al. · 2017 [cited by applicant]
US 10157621B2 · Atti · 2018 [cited by examiner]
US 10325611B2 · Ravelli et al. · 2019 [cited by applicant]
US 10839814B2 · Atti · 2020 [cited by applicant]
US 11084954B2 · Wutti et al. · 2021 [cited by applicant]
US 20030004711A1 · Koishida · 2003 [cited by applicant]
US 20030009325A1 · Kirchherr et al. · 2003 [cited by applicant]
US 20050154584A1 · Jelinek et al. · 2005 [cited by applicant]
US 20060149537A1 · Shiramizo · 2006 [cited by applicant]
US 20060173605A1 · Pfaeffle · 2006 [cited by applicant]
US 20060271373A1 · Khalil et al. · 2006 [cited by applicant]
US 20070206645A1 · Sundqvist et al. · 2007 [cited by applicant]
US 20070239462A1 · Makinen et al. · 2007 [cited by applicant]
US 20080027717A1 · Rajendran et al. · 2008 [cited by applicant]
US 20080049795A1 · Lakaniemi · 2008 [cited by applicant]
US 20080147414A1 · Son et al. · 2008 [cited by applicant]
US 20080172223A1 · Oh et al. · 2008 [cited by applicant]
US 20080294429A1 · Su et al. · 2008 [cited by applicant]
US 20090187409A1 · Krishman et al. · 2009 [cited by applicant]
US 20090234644A1 · Reznik et al. · 2009 [cited by applicant]
US 20090299757A1 · Guo et al. · 2009 [cited by applicant]
US 20090326930A1 · Kawashima et al. · 2009 [cited by applicant]
US 20100049511A1 · Ma et al. · 2010 [cited by applicant]
US 20110119054A1 · Lee et al. · 2011 [cited by applicant]
US 20110125505A1 · Vaillancourt et al. · 2011 [cited by applicant]
US 20110153333A1 · Bessette · 2011 [cited by applicant]
US 20110173008A1 · Lecomte et al. · 2011 [cited by applicant]
US 20110173010A1 · Lecomte et al. · 2011 [cited by applicant]
US 20110196673A1 · Sharma et al. · 2011 [cited by applicant]
US 20110200198A1 · Grill et al. · 2011 [cited by applicant]
US 20110202353A1 · Neuendorf et al. · 2011 [cited by applicant]
US 20120022880A1 · Bessette · 2012 [cited by applicant]
US 20120101813A1 · Vaillancourt et al. · 2012 [cited by applicant]
US 20120209604A1 · Sehlstedt · 2012 [cited by applicant]
US 20120245947A1 · Neuendorf · 2012 [cited by examiner]
US 20120265541A1 · Geiger et al. · 2012 [cited by applicant]
US 20120271644A1 · Bessette et al. · 2012 [cited by applicant]
US 20130030798A1 · Mittal et al. · 2013 [cited by applicant]
US 20130144632A1 · Sung · 2013 [cited by applicant]
US 20130289981A1 · Ragot et al. · 2013 [cited by applicant]
US 20130332177A1 · Helmrich et al. · 2013 [cited by applicant]
US 20140188465A1 · Choo · 2014 [cited by applicant]
US 20160293173A1 · Faure et al. · 2016 [cited by applicant]
US 20170133026A1 · Ravelli et al. · 2017 [cited by applicant]
US 20170270935A1 · Atti · 2017 [cited by applicant]
US 20200160874A1 · Ravelli · 2020 [cited by applicant]
US 20220076685A1 · Ravelli et al. · 2022 [cited by applicant]
AU 2013200680 · 2013 [cited by applicant]
CN 1187665 · 1998 [cited by applicant]
CN 1672192 · 2005 [cited by applicant]
CN 1705979 · 2005 [cited by applicant]
CN 1849648 · 2006 [cited by applicant]
CN 101025918 · 2007 [cited by applicant]
CN 101197134 · 2008 [cited by applicant]
CN 101256771 · 2008 [cited by applicant]
CN 101364854 · 2009 [cited by applicant]
CN 101523486 · 2009 [cited by applicant]
CN 102089758 · 2011 [cited by applicant]
CN 102105930 · 2011 [cited by applicant]
CN 102150205 · 2011 [cited by applicant]
CN 102737642 · 2012 [cited by applicant]
CN 103282959 · 2013 [cited by applicant]
CN 103703512 · 2014 [cited by applicant]
CN 101836251 · 2020 [cited by applicant]
EP 0747884 · 1996 [cited by applicant]
EP 0966102 · 1999 [cited by applicant]
FR 2830970 · 2003 [cited by applicant]
JP 2010517083 · 2010 [cited by applicant]
JP 2017504677 · 2017 [cited by applicant]
RU 2483365C2 · 2012 [cited by applicant]
RU 2483366C2 · 2013 [cited by applicant]
WO 9847313 · 1998 [cited by applicant]
WO 0063885 · 2000 [cited by applicant]
WO 2004010416 · 2004 [cited by applicant]
WO 2005027095 · 2005 [cited by applicant]
WO 2009059333A1 · 2009 [cited by applicant]
WO 2010101190 · 2010 [cited by applicant]
WO 2011042464A1 · 2011 [cited by applicant]
WO 2011048094A1 · 2011 [cited by applicant]
WO 2013168414 · 2016 [cited by applicant]
ISO/IEC FDIS 23003-3:2011(E), “Information Technology—MPEG Audio Techology—MPEG Audio Technologies—Part 3: Unified Speech and Audio Coding”, ISO/IEC JTC 1/SC 29/WG 11, Sep. 20, 2011. [cited by applicant]
Parallel Korean Patent Application No. 10-2017-7004348 Office Action dated Sep. 17, 2018. [cited by applicant]
Jeremie Lecomte et al.: “Efficient Cross-Fade Windows for Transitions between LPC-based and non-LPC based audio coding”: 126th AES Convention; May 2009; paper 7712. [cited by applicant]
Office Action in parallel Russian Application No. 2017106091 dated Feb. 12, 2018. [cited by applicant]
Fraunhofer IIS: Tdoc S4-130345, Qualification Deliverables for the Fraunhofer IIS Candidate for EVS (including Technical Description and Report on Compliance to Design Constraints), TSG SA4#72bis meeting, Mar. 11-15, 20… [cited by applicant]
3GPP TS 26.290 V2.0.0 (Sep. 2004). [cited by applicant]
3GPP TS 26.290 V10.0.0 (Mar. 2011). [cited by applicant]
3GPP TS 26.290 V6.1.0 (Dec. 2004). [cited by applicant]
3GPP TS 26.403 V6.0.0 (Sep. 2004). [cited by applicant]
3GPP TS 26.442 V14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.443 14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.445 V12.0.0 (Sep. 2014). [cited by applicant]
3GPP TS 26.445 V14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.445 V14.2.0 (Dec. 2017). [cited by applicant]
3GPP TS 26.445 V16.2.0 (Dec. 2021). [cited by applicant]
3GPP TS 26.445 V17.0.0 (Apr. 2022). [cited by applicant]
3GPP TS 26.447 V14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.447 V14.2.0 (Jun. 2020). [cited by applicant]
3GPP TS 26.447 V16.0.0 (Mar. 2019). [cited by applicant]
3GPP TS 26.952 v17.0.0 (Apr. 2022). [cited by applicant]
Fuchs et al, MDCT-Based Coder for Highly Adaptive Speech and Audio Coding, 17th European Signal Processing Conference (EUSIPCO 2009), Glasgow, Scotland, Aug. 24-28, 2009. [cited by applicant]
Schnell et al., Proposed Core Experiment on AAC-ELD, Apr. 18, 2007. [cited by applicant]
Schnell et al., Low Delay Filter banks for Enhanced Low Delay Audio Coding, 2007 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics Oct. 21, 2007. [cited by applicant]
Henrique S. Malvar, Signal Processing with Lapped Transforms, Computer Science Engineering, 1992 , Chapter. [cited by applicant]
Marina Bosi and Richard E. Goldberg, Introduction to Digital Audio Coding and Standards, Springer 2003. [cited by applicant]
Convolution theorem—Wikipedia. [cited by applicant]
G. Clark, S. Parker and S. Mitra, “A unified approach to time- and frequency-domain realization of FIR adaptive digital filters,” in IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 31, No. 5, pp. 107… [cited by applicant]
ITU-T G.722 (Jul. 2003). [cited by applicant]
ITU-T G.722.2 (Jan. 2002) of the Telecommunication Standardization Sector of the International Telecommunication Union (“G.722.2”), Annex A. [cited by applicant]
Kondoz, Digital Speech: Coding for Low bit Rate Communication Systems (John Wiley & Sons 2004). [cited by applicant]
Lecomte et al., “An Improved Low Complexity AMR-WB+ Encoder using Neural Networks for Mode Selection” (AES 123rd Convention, New York, NY, USA, Oct. 5-8, 2007m Convention Paper 7294 section 2.1.2. [cited by applicant]
ITU-T G.723.1. [cited by applicant]
Ravishankar, C., Hughes Network Systems, Germantown, MD. Speech coding. United States, https://doi.org/10.2172/325392. [cited by applicant]
Huan Hou and Weibei Dou, Real-time Audio Error Concealment Method Based on Sinusoidal Model, International Conference on audio Language and Image Processing, IEEE, Jul. 2008, Shanghai, P.R. China, DOI: 10.1109/ICALIP.20… [cited by applicant]
ITU-T G.718 (Jun. 2008), Series G: Transmission Systems and Media, Digital Systems and Networks, Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit/s Digital… [cited by applicant]
J.D. Warren, et al., Analysis of the spectral envelope of sounds by the human brain, NEUROIMAGE. Feb. 15, 2005;24(4):1052-7, https://pubmed.ncbi.nlm.nih.gov/15670682/#:˜:text=Spectral%20envelope%20is%20the%20shape,of%20… [cited by applicant]
Virette, D., Low Delay Transform for High Quality Low Delay Audio Coding, Signal and Image Processing, (Université de Rennes 1, 2012), 40-41. [cited by applicant]
Ostergaard, J., et al., Real-time perceptual moving-horizon multiple-description audio coding, IEEE Transactions on Signal Processing, 4286 (2011). [cited by applicant]
Oh, H., et al., A Fast Quantization Loop Algorithm for MP3/AAC Encoders, AES 29th International Conference (2006). [cited by applicant]