IP Library › Granted Patent US 12,738,283
Granted Patent B2
US 12,738,283 · App. 18/405,369 · Granted Sep 15, 2026

Processor for generating a prediction spectrum based on long-term prediction and/or harmonic post-filtering

Inventors: Goran Markovic (Erlangen, DE); Bernd Edler (Erlangen, DE); Stefan Bayer (Erlangen, DE); Jan Frederik Kiene (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V.
G10L19/02G10L19/09G10L25/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,283
App. No.
18/405,369
Granted
Sep 15, 2026
Kind
B2
Abstract

A processor for processing an (encoded) audio signal, the processor comprising: an LTP buffer configured to receive samples derived from a frame of the encoded audio signal; an interval splitter configured to divide a time interval associated with a subsequent frame of the encoded audio signal into sub-intervals depending on the encoded pitch parameter; calculation means configured to derive sub-interval parameters from the encoded pitch parameter dependent on a position of the sub-intervals within the time interval associated with the subsequent frame of the encoded audio signal; a predictor configured for generating a prediction signal from the LTP buffer dependent on the sub-interval parameters; and a frequency domain transformer configured for generating a prediction spectrum (X P ) based on the prediction signal.

Claims (91)

1 . A processor for processing an encoded audio signal, the encoded audio signal comprising at least an encoded pitch parameter, the processor comprising:

an LTP buffer configured to receive samples derived from a frame of the encoded audio signal;

an interval splitter configured to divide a time interval associated with a subsequent frame of the encoded audio signal into sub-intervals depending on the encoded pitch parameter;

a calculation unit configured to derive sub-interval parameters from the encoded pitch parameter dependent on a position of the sub-intervals within the time interval associated with the subsequent frame of the encoded audio signal;

a predictor configured for generating a prediction signal from the LTP buffer dependent on the sub-interval parameters; and

a frequency domain transformer configured for generating a prediction spectrum based on the prediction signal.

2 . The processor according to claim 1 , wherein there are more sub-intervals than temporarily distinct encoded pitch parameters; and/or

wherein there are more distinct sub-interval parameters than temporarily distinct encoded pitch parameters; and/or

wherein there are more than one temporarily distinct encoded pitch parameters in the frame.

3 . The processor according to claim 1 , further comprising a combiner configured to combine at least a portion of a derivation of the prediction spectrum with an error spectrum to generate a combined spectrum; and/or

wherein a derivation of the prediction spectrum is derived from the prediction spectrum by perceptually flattening the prediction spectrum.

4 . The processor according to claim 3 , further comprising a unit for putting all samples from a block of aliased time domain audio signal being not different from the audio signal into the LTP buffer; or

further comprising a unit for putting samples from the block of aliased time domain audio signal not different from a time domain audio signal into the LTP buffer, wherein the samples are used for producing the subsequent frame of audio signal; or

further comprising a unit for putting samples from the block of aliased time domain audio signal not different from a current frame into the LTP buffer, wherein the samples are used for producing the subsequent frame of time domain audio signal, wherein a selection of a portion of the current frame or of the samples selected from the block of aliased time domain audio signal is adapted by the unit for putting samples.

5 . The processor according to claim 1 , wherein the processor further comprises an inverse frequency domain transformer; and/or

wherein the processor further comprises an inverse frequency domain transformer configured for generating a block of aliased time domain audio signal from a derivation of an error spectrum, where the prediction spectrum is obtained from the frame of the encoded audio signal and/or where an error spectrum is obtained from the subsequent frame of the encoded audio signal subsequent to the frame and the derivation of the error spectrum is derived from the error spectrum; or

wherein the processor further comprises the inverse frequency domain transformer configured for generating the block of aliased time domain audio signal from the derivation of the error spectrum, where the prediction spectrum is obtained from the frame of the encoded audio signal and/or where the error spectrum is obtained from the subsequent frame of the encoded audio signal subsequent to the frame and the derivation of the error spectrum is derived from the error spectrum; and a unit for generating a frame of time domain audio signal using at least two blocks of the aliased time domain audio signal, where at least some portions of the aliased time domain audio signal are different from the time domain audio signal and the received samples, respectively.

6 . The processor according to claim 4 , further comprising an entity configured for zero filling based on a signal received from a band-wise parametric decoder and a combined spectrum to obtain the derivation of the error spectrum where the combined spectrum is obtained based on at least a portion of a derivation of the prediction spectrum and the error spectrum; and an entity configured for spectral shaping a spectral envelope of a signal modified by an entity configured for temporal shaping and taking into account a coded information for the spectral shaping to obtain the derivation of the error spectrum and the entity configured for temporal shaping the signal taking into account the coded information for temporal shaping to obtain the derivation of the error spectrum.

7 . The processor according to claim 1 , further comprising a combiner configured to combine at least a portion of the prediction spectrum X P with an error spectrum XD to generate a combined spectrum X DT ; and/or

further comprising a combiner configured to combine at least a portion of the prediction spectrum X P or at least a portion of a derivation of a prediction spectrum X PS with an error spectrum X D , wherein the portion is determined based on the encoded pitch parameter; and/or

further comprising a combiner configured to combine at least a portion of the prediction spectrum X P or at least a portion of a derivation of the prediction spectrum X PS with an error spectrum X D , wherein if the LTP buffer is active, then first └(n LTP +0.5)iF0┘ coefficients of the prediction spectrum or the derivation of the prediction spectrum, except a zeroth coefficient, are added to the error spectrum to produce a combined spectrum X DT ; and/or wherein the zeroth and the coefficients above └(n LTP +0.5)iF0┘ are copied from the error spectrum to the combined spectrum), wherein “└ ┘” indicates a use of a floor function;

where n LTP is a parameter from the encoded audio signal and/or where n LTP is a number of predictable harmonics; and

where iF0 is derived from the encoded pitch parameter.

8 . The processor according to claim 1 , wherein in each sub-interval the predicted signal is constructed using the LAP buffer and/or using a decoded audio signal out of the LTP buffer and a filter whose parameters are derived from the encoded pitch parameter and the sub-interval position within the time interval associated with the subsequent frame of the encoded audio signal.

9 . The processor according to claim 1 , wherein the calculation unit are configured to derive sub-interval parameters from the encoded pitch parameter, wherein the sub-interval parameters comprise at least a sub-interval pitch parameter, as follows:

obtaining a sub-interval pitch lag associated with a center of the sub-interval from a pitch contour, wherein the pitch contour comprises multiple values, comprising:

setting the sub-interval pitch lag to the pitch contour value at the position of the sub-interval center,

determining a sub-interval end,

comparing the sub-interval pitch lag to the sub-interval end producing a comparison result, and/or

adapting the sub-interval pitch lag for the pitch contour value at position derived from the sub-interval pitch lag depending on the comparison result

and

further comprising the calculation unit configured to derive a pitch contour from the encoded pitch parameter; where the pitch contour is obtained from the encoded pitch parameters using an interpolation; or

further comprising the calculation unit configured to derive a pitch contour from the encoded pitch parameter; where the pitch contour is obtained from the encoded pitch parameters using an interpolation.

10 . The processor according to claim 1 , further comprising a unit for smoothing the prediction signal across and/or at borders of at least two sub-intervals of a plurality of sub-intervals and/or

further comprising a unit for smoothing the prediction signal across and/or at borders of at least two sub-intervals of the plurality of sub-intervals, wherein at least the at least two sub-intervals are overlapping.

11 . The processor according to claim 1 , further comprising a unit for modifying the predicted spectrum, or a derivative of the predicted spectrum, dependent on a parameter derived from the encoded pitch parameter in order to generate a modified predicted spectrum; and/or

further comprising a unit for modifying the predicted spectrum, or a derivative of the predicted spectrum, wherein the unit for modifying are configured to adapt magnitudes of MDCT coefficients at least n Fsafeguard away from harmonics in X P or in X PS by setting to zero or multiplying with a positive factor smaller than 1 magnitudes of the MDCT coefficients; or further comprising a unit for modifying the predicted spectrum, or a derivative of the predicted spectrum, wherein the unit for modifying are configured to reduce magnitudes of the predicted spectrum, or magnitudes of the derivative of the predicted spectrum, between harmonics.

12 . The processor according to claim 1 , further comprising a unit for deriving a modified pitch parameter from the encoded pitch parameter dependent on a content of the LTP buffer; or

wherein the predicted spectrum is generated dependent on a modified pitch parameter.

13 . A processing unit comprising a processor claim 1 , and the processor comprising:

a splitter configured for splitting a time interval associated with a frame of the audio signal into a plurality of sub-intervals, each comprising a respective length, the respective length of the plurality of sub-intervals being dependent on a pitch lag value;

a harmonic post-filter configured for filtering the plurality of sub-intervals, wherein the harmonic post-filter is based on a transfer function comprising a numerator and a denominator, where the numerator comprises a harmonicity value, and wherein the denominator comprises a sub-interval pitch lag value and the harmonicity value and/or a gain value;

wherein the associated harmonicity value and/or the sub-interval pitch lag value and/or the gain value is different in at least two different sub-intervals of the plurality of sub-intervals;

wherein the sub-interval pitch lag value, the harmonicity value and/or the gain value are obtained based on the audio signal in each sub-interval of the plurality of sub-intervals.

14 . A processor for processing an audio signal, the processor comprising:

a splitter configured for splitting a time interval associated with a frame of the audio signal into a plurality of sub-intervals, each comprising a respective length, the respective length of the plurality of sub-intervals being dependent on a pitch lag value;

a harmonic post-filter configured for filtering the plurality of sub-intervals, wherein the harmonic post-filter is based on a transfer function comprising a numerator and a denominator, where the numerator comprises a harmonicity value, and wherein the denominator comprises a sub-interval pitch lag value and the harmonicity value and/or a gain value;

wherein the associated harmonicity value and/or the sub-interval pitch lag value and/or the gain value is different in at least two different sub-intervals of the plurality of sub-intervals;

wherein the sub-interval pitch lag value, the harmonicity value and/or the gain value are obtained based on the audio signal in each sub-interval of the plurality of sub-intervals.

15 . The processor according to claim 14 , wherein at least two subintervals or the plurality of sub-intervals are overlapping.

16 . The processor according to claim 14 , wherein the harmonicity value is proportional to a desired intensity of the harmonic post-filter and/or independent of amplitude changes in the audio signal; and/or

wherein the gain value is dependent on the amplitude changes in the audio signal).

17 . The processor according to according to claim 14 , wherein the harmonic post-filter changes from a sub-interval to a subsequent sub-interval; and/or

wherein the harmonicity value and/or the gain value and/or the sub-interval pitch lag value in the subsequent sub-interval are derived using an output of the harmonic post-filter in the sub-interval.

18 . The processor according to claim 14 , wherein the harmonic post-filter is different in at least two different sub-intervals of the plurality of sub-intervals; or

wherein the harmonic post-filter is different in at least two different sub-intervals of the plurality of sub-intervals or wherein the associated harmonicity value and/or the sub-interval pitch lag value and/or the gain value is different in at least two different sub-intervals of the plurality of sub-intervals, the in at least two different sub-intervals of the plurality of sub-intervals belonging to a same frame.

19 . The processor according to claim 14 , further comprising a unit for smoothing an output of the harmonic post-filter in the plurality of sub-intervals across and/or at sub-interval borders.

20 . The processor according to claim 14 , wherein there are at least two sub-intervals within the frame.

21 . The processor according to claim 14 , wherein the respective length is dependent on an average pitch; and/or

wherein an average pitch is obtained from an encoded pitch parameter; and/or

wherein the encoded pitch parameter comprises higher time resolution than a codec framing and/or wherein the encoded pitch parameter comprises lower time resolution then a pitch contour.

22 . The processor according to claim 14 , further comprising a domain converter configured for converting on a frame basis a first domain representation of the audio signal into a second domain representation of the audio signal; or

further comprising a domain converter configured for converting on a frame basis a frequency domain representation of the audio signal into a time domain representation of the audio signal.

23 . The processor claim 14 , wherein the harmonicity value and/or the gain value and/or the sub-interval pitch lag value for each sub-interval is derived dependent on a position of the respective sub-interval within the time interval associated with the frame.

24 . A decoder for decoding an encoded audio signal which comprises a processor claim 1 and/or a processor comprising:

a splitter configured for splitting a time interval associated with a frame of the audio signal into a plurality of sub-intervals, each comprising a respective length, the respective length of the plurality of sub-intervals being dependent on a pitch lag value;

a harmonic post-filter configured for filtering the plurality of sub-intervals, wherein the harmonic post-filter is based on a transfer function comprising a numerator and a denominator, where the numerator comprises a harmonicity value, and wherein the denominator comprises a sub-interval pitch lag value and the harmonicity value and/or a gain value;

wherein the associated harmonicity value and/or the sub-interval pitch lag value and/or the gain value is different in at least two different sub-intervals of the plurality of sub-intervals; wherein the sub-interval pitch lag value, the harmonicity value and/or the gain value are obtained based on the audio signal in each sub-interval of the plurality of sub-intervals.

25 . The decoder claim 24 , further comprising a frequency domain decoder or a decoder based on an inverse MDCT.

26 . An encoder for encoding an audio signal, comprising the processor claim 1 .

27 . A method for processing an encoded audio signal, the encoded audio signal comprising at least an encoded pitch parameter, the method comprising following steps:

receiving samples derived from a frame of the encoded audio signal using an LTP buffer;

dividing a time interval associated with a subsequent frame of the encoded audio signal subsequent to the frame into sub-intervals depending on the encoded pitch parameter;

deriving sub-interval parameters from the encoded pitch parameter dependent on a position of the sub-intervals within the time interval associated with the subsequent frame of the encoded audio signal;

generating a prediction signal from the LTP buffer dependent on the sub-interval parameters; and

generating a prediction spectrum based on the prediction signal.

28 . A method for processing an audio signal, the method comprising following steps:

splitting a time interval associated with a frame of the audio signal into a plurality of sub-intervals, each comprising a respective length, the respective lengths of at least two of the plurality of sub-intervals being dependent on a pitch lag value; and

filtering the plurality of sub-intervals using a harmonic post-filter, wherein the harmonic post-filter is based on a transfer function comprising a numerator and a denominator, where the numerator comprises a harmonicity value, and wherein the denominator comprises a sub-interval pitch lag value and the harmonicity value and/or a gain value,

wherein the associated harmonicity value and/or the sub-interval pitch lag value and/or the gain value is different in at least two different sub-intervals of the plurality of sub-intervals, and

wherein the sub-interval pitch lag value, the harmonicity value and/or the gain value are obtained based on the audio signal in each sub-interval of the plurality of sub-intervals.

29 . A non-transitory digital storage medium having a computer program stored thereon to perform a method for processing an encoded audio signal, the encoded audio signal comprising at least an encoded pitch parameter and the method comprising following steps:

receiving samples derived from a frame of the encoded audio signal using an LTP buffer;

dividing a time interval associated with a subsequent frame of the encoded audio signal subsequent to the frame into sub-intervals depending on the encoded pitch parameter;

deriving sub-interval parameters from the encoded pitch parameter dependent on a position of the sub-intervals within the time interval associated with the subsequent frame of the encoded audio signal;

generating a prediction signal from the LTP buffer dependent on the sub-interval parameters; and

generating a prediction spectrum based on the prediction signal, when said computer program is run by a computer.

30 . A non-transitory digital storage medium having a computer program stored thereon to perform a method for processing an audio signal, the method comprising following steps:

splitting a time interval associated with a frame of the audio signal into a plurality of sub-intervals, each comprising a respective length, the respective lengths of at least two of the plurality of sub-intervals being dependent on a pitch lag value; and

filtering the plurality of sub-intervals using a harmonic post-filter, wherein the harmonic post-filter is based on a transfer function comprising a numerator and a denominator, where the numerator comprises a harmonicity value, and wherein the denominator comprises a sub-interval pitch lag value and the harmonicity value and/or a gain value,

wherein the associated harmonicity value and/or the sub-interval pitch lag value and/or the gain value is different in at least two different sub-intervals of the plurality of sub-intervals, and wherein the sub-interval pitch lag value, the harmonicity value and/or the gain value are obtained based on the audio signal in each sub-interval of the plurality of sub-intervals, when said computer program is run by a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2024
From: MARKOVIC, GORAN; EDLER, BERND; BAYER, STEFAN; KIENE, JAN FREDERIK
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 067885/0831 →
Priority Claims (1)
EP 21185662 · Jul 14, 2021 · regional
Continuity (2)
Continuation PCTEP2022069751 · Jul 14, 2022
Related Publication 20240177720A1 · May 30, 2024
References Cited (107)
US 5886276A · Levine et al. · 1999 [cited by applicant]
US 6064954A · Cohen · 2000 [cited by examiner]
US 8060363B2 · Anssi et al. · 2011 [cited by applicant]
US 11043226B2 · Ravelli et al. · 2021 [cited by applicant]
US 20050228648A1 · Heikkinen · 2005 [cited by examiner]
US 20060047522A1 · Ojanpera · 2006 [cited by applicant]
US 20060089832A1 · Ojanpera · 2006 [cited by applicant]
US 20060239473A1 · Kjorling et al. · 2006 [cited by applicant]
US 20070239440A1 · Garudadri et al. · 2007 [cited by applicant]
US 20080027711A1 · Rajendran · 2008 [cited by examiner]
US 20080126081A1 · Geiser et al. · 2008 [cited by applicant]
US 20090070118A1 · Den Brinker et al. · 2009 [cited by applicant]
US 20090234644A1 · Reznik et al. · 2009 [cited by applicant]
US 20100250260A1 · Laaksonen et al. · 2010 [cited by applicant]
US 20100262420A1 · Herre et al. · 2010 [cited by applicant]
US 20110270616A1 · Garudadri et al. · 2011 [cited by applicant]
US 20110320212A1 · Tsujino et al. · 2011 [cited by applicant]
US 20120146831A1 · Eksler · 2012 [cited by applicant]
US 20130151262A1 · Lohwasser et al. · 2013 [cited by applicant]
US 20130346073A1 · Laaksonen et al. · 2013 [cited by applicant]
US 20140244274A1 · Yamanashi et al. · 2014 [cited by applicant]
US 20140310009A1 · Lee · 2014 [cited by examiner]
US 20150106108A1 · Baeckstroem et al. · 2015 [cited by applicant]
US 20150228289A1 · Oh et al. · 2015 [cited by applicant]
US 20150262588A1 · Tsutsumi · 2015 [cited by examiner]
US 20160027450A1 · Gao · 2016 [cited by applicant]
US 20160133265A1 · Disch et al. · 2016 [cited by applicant]
US 20160155451A1 · Baeckstroem et al. · 2016 [cited by applicant]
US 20160284357A1 · Kawashima et al. · 2016 [cited by applicant]
US 20160293175A1 · Atti et al. · 2016 [cited by applicant]
US 20170069328A1 · Kawashima et al. · 2017 [cited by applicant]
US 20170256267A1 · Disch et al. · 2017 [cited by applicant]
US 20180182400A1 · Choo et al. · 2018 [cited by applicant]
US 20180190303A1 · Ghido et al. · 2018 [cited by applicant]
US 20190158833A1 · Sung et al. · 2019 [cited by applicant]
US 20190272835A1 · Adami et al. · 2019 [cited by applicant]
US 20190295561A1 · Disch et al. · 2019 [cited by applicant]
US 20200294518A1 · Ravelli et al. · 2020 [cited by applicant]
US 20210125624A1 · Ravelli · 2021 [cited by examiner]
EP 2980794A1 · 2016 [cited by applicant]
EP 3671741A1 · 2020 [cited by applicant]
JP H09261064A · 1997 [cited by applicant]
JP 2007004050A · 2007 [cited by applicant]
JP 2008513848A · 2008 [cited by applicant]
JP 2010530079A · 2010 [cited by applicant]
JP 2017523473A · 2017 [cited by applicant]
RU 2409874C9 · 2011 [cited by applicant]
RU 2428748C2 · 2011 [cited by applicant]
RU 2482554C1 · 2013 [cited by applicant]
RU 2662921C2 · 2018 [cited by applicant]
RU 2667376C2 · 2018 [cited by applicant]
WO 2008151755A1 · 2008 [cited by applicant]
WO 2010003556A1 · 2010 [cited by applicant]
WO 2008072701A1 · 2010 [cited by applicant]
WO 2013107602A1 · 2013 [cited by applicant]
WO 2014056705A1 · 2014 [cited by applicant]
WO 2014118171A1 · 2014 [cited by applicant]
WO 2014118175A1 · 2014 [cited by applicant]
WO 2014128197A1 · 2014 [cited by applicant]
WO 2015010949A1 · 2015 [cited by applicant]
WO 2015010952A1 · 2015 [cited by applicant]
WO 2015010953A1 · 2015 [cited by applicant]
WO 2015010954A1 · 2015 [cited by applicant]
WO 2016016121A1 · 2016 [cited by applicant]
WO 2016142357A1 · 2016 [cited by applicant]
WO 2019091573A1 · 2019 [cited by applicant]
WO 2019091904A1 · 2019 [cited by applicant]
WO 2019092220A1 · 2019 [cited by applicant]
WO 2021104623A1 · 2021 [cited by applicant]
Examiner, “Office Action for Japanese Application No. 2024-501939 ”, Mar. 18, 2025, JPO, Japan. [cited by applicant]
Examiner, “Office Action for Japanese Application No. 2024-501938 ”, Mar. 18, 2025, JPO, Japan. [cited by applicant]
Volkov P.A . . . “Office Action for RU Application No. 2024103467”, Apr. 12, 2024, Rospatent, Russia. [cited by applicant]
Kenichi Makino et al., “Hybrid audio coding for speech and audio below medium bit rate”, Consumer Electronics, 2000, ICCE, 2000 Digest of Technical Papers, International Conference on 2000, pp. 264-265. [cited by applicant]
Juha Ojanperä et al., “Long term predictor for transform domain perceptual audio coding”, An Audio Engineering Society Convention 107th, Sep. 24-27, 1999, New York. [cited by applicant]
Sean A. Ramprashad, “A multimode transform predictive coder (MTPC) for speech and audio”, Speech Coding Proceedings, 1999 IEEE Workshop, 1999, pp. 10-12. [cited by applicant]
Lars Villemoes et al., “Speech coding with transform domain prediction”, 2017 IEEE Workshop on Applications of signal Processing to Audio and Acoustics(WASPAA), 2017, pp. 324-328. [cited by applicant]
Ronald Howell Frazier, “An adaptive filtering approach toward speech enhancement,” Citeseer, 1975. [cited by applicant]
D. Malah et al., “A generalized comb filtering technique for speech enhancement,” Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP' 82., 1982, vol. 7, pp. 160-163. [cited by applicant]
Jeongook Song et al., “Harmonic Enhancement in Low Bitrate Audio Coding Using an Efficient Long-Term Predictor,” EURASIP Journal on Advances in Signal Processing, Feb. 8, 1010, vol. 2010, Article ID 939542, 9 pages. [cited by applicant]
Universal Mobile Telecommunications System (UMTS), LTE, EVS Codec General Overview, 3GPP TS 26.441 version 12.0.0 Release 12, Oct. 2014. [cited by applicant]
Ning Guo et al., “Frequency Domain Long-Term Prediction for Low Delay General Audio Coding”, IEEE Signal Processing Letters, pp. 1185-1189, vol. 28, 2021. [cited by applicant]
Tejaswi Nanjundaswamy et al., “Cascaded Long Term Prediction for Enhanced Compression of Polyphonic Audio Signals,” IEEE/ACM Transactions On Audio, Speech, And Language Processing, Mar. 2014, pp. 697-710, vol. 22, No. 3. [cited by applicant]
“Low Complexity Communication Codec”, Bluetooth SIG Proprietary, v 1.0, Sep. 15, 2020, Hearing Aid Working Group. [cited by applicant]
“Digital Enhanced Cordless Telecommunications (DECT), Low Complexity Communication Codec plus (LC3plus)”, ETSI TS 103 634 V1.3.1, Oct. 2021. [cited by applicant]
Juin-Hwey Chen et al., “Adaptive postfiltering for quality enhancement of coded speech”, IEEE Transactions on Speech and Audio Processing, Jan. 1, 1995, pp. 59-71, vol. 3, No. 1, DOI: 10. 1109/89.365380, URL: http://iee… [cited by applicant]
Jean-Marc Valin et al., “A High-Quality Speech and Audio Codec With Less Than 10 ms Delay”, arxiv.org, Cornell University Library, Feb. 17, 2016, DOI: 10.1109/TASL.2009.2023186, XP080684285, 201 OLIN Library Cornell Uni… [cited by applicant]
Ajit B. Rao et al., “Pitch adaptive windows for improved excitation coding in low-rate CELP coders”, IEEE Transactions on Speech and Audio Processing, Nov. 1, 2003, pp. 648-659, vol. 11, No. 6, DOI: 10.1109/TSA.2003.815… [cited by applicant]
3GPP TS 26.290 V16.0.0, “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Audio codec processing functions; Extended Adaptive Multi-Rate—Wideband (AMR-WB+) codec; Transcodin… [cited by applicant]
Jurgen Herre et al., “Extending the MPEG-4 AAC Codec by Perceptual Noise Substitution”, Audio Engineering Society Convention 104, May 16-19, 1998, pp. 1-14, Amsterdam. [cited by applicant]
Frederik Nagel et al.,“A continuous modulated single sideband bandwidth extension”, 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 2010, pp. 357-360. [cited by applicant]
Christian Neukam et al., “A MDCT based harmonic spectral bandwidth extension method”, 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, 2013, pp. 566-570. [cited by applicant]
Sascha Disch et al.,“Intelligent Gap Filling in Perceptual Transform Coding of Audio”, Convention Paper 9661, the Journal of the Audio Engineering Society 141 st Convention,, Sep. 29, 2016, Los Angeles, CA, US. [cited by applicant]
Sascha Disch et al., “Improved Psychoacoustic Model for Efficient Perceptual Audio Codecs”, Convention Paper 10029, the Journal of the Audio Engineering Society 145 th Convention, Oct. 18, 2018, New York, NY, US. [cited by applicant]
Christian R. Helmrich et al., “Spectral envelope reconstruction via IGF for audio transform coding”, 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 389-393. [cited by applicant]
Marie Oger et al., “Model-based deadzone optimization for stack-run audio coding with uniform scalar quantization”, 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, 2008, pp. 4761-4764. [cited by applicant]
Oliver Niemeyer et al., “Detection and Extraction of Transients for Audio Coding”, Audio Engineering Society Convention Paper 6811, the Journal of the AES 120 th Convention, May 20, 2006, XP040507705, New York, NY, US. [cited by applicant]
Florin Ghido et al., “Coding Of Fine Granular Audio Signals Using High Resolution Envelope Processing (HREP)”, 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 701-705. [cited by applicant]
Alexander Adami et al., “Transient-to-noise ratio restoration of coded applause-like signals”, 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), Oct. 15, 2017, pp. 349-353, New Pal… [cited by applicant]
Richard Füg et al., “Harmonic-percussive-residual sound separation using the structure tensor on spectrograms”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 445-449. [cited by applicant]
L. Daudet, B. Torresani et al., “An hybrid audio coder for very low bit rate with two-level psychoacoustic modeling”, Multimedia Signal Processing, 1999 IEEE 3rd Workshop on Copenhagen, Denmark, Sep. 13-15, 1999, ISBN 9… [cited by applicant]
Volkov P.A., “Office Action for RU Application No. 2024103466”, Aug. 23, 2024, Rospatent, Russia. [cited by applicant]
Volkov P.A . . . “Office Action for RU Application No. 2024103244”, Mar. 6, 2024, Rospatent, Russia. [cited by applicant]
Schmidt Konstantin et al., “Low complexity tonality control in the Intelligent Gap Filling tool”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Mar. 20, 2016 (Mar. 20, 201… [cited by applicant]
Luc Krembel, “Extended European Search Report for EP Application No. 25197110.7”, Feb. 9, 2026, EPO, Germany. [cited by applicant]
Ghido, Florin, et al., “Coding of Fine Granular Audio Signals Using High Resolution Envelope Processing (HREP)” Proceedings of The 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), … [cited by applicant]
Niemeyer, Oliver, et al., “Detection and Extraction of Transients for Audio Coding,” AES 120th Convention, 2006, Paris, France. [cited by applicant]
Oliver Niemeyer et al., “Detection and extraction of transients for audio coding”, Audio Engineering Society, Convention Paper 6811, pp. 1-8, May 20-23, 2006. [cited by applicant]