IP Library Granted Patent US 12,488,804
Granted Patent B2
US 12,488,804 · App. 18/540,819 · Granted Dec 2, 2025

Frequency-domain audio coding supporting transform length switching

Inventors: Sascha Dick (Nuremberg, DE); Christian Helmrich (Erlangen, DE); Andreas Hoelzer (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Foerderung der angewandten Forschung e.V.
G10L19/022G10L19/03G10L19/008G10L19/028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,804
App. No.
18/540,819
Granted
Dec 2, 2025
Kind
B2
Abstract

A frequency-domain audio codec is provided with the ability to additionally support a certain transform length in a backward-compatible manner, by the following: the frequency-domain coefficients of a respective frame are transmitted in an interleaved manner irrespective of the signalization signaling for the frames as to which transform length actually applies, and additionally the frequency-domain coefficient extraction and the scale factor extraction operate independent from the signalization. By this measure, old-fashioned frequency-domain audio coders/decoders, insensitive for the signalization, would be able to nevertheless operate without faults and with reproducing a reasonable quality. Concurrently, frequency-domain audio coders/decoders able to support the additional transform length would offer even better quality despite the backward compatibility. As far as coding efficiency penalties due to the coding of the frequency domain coefficients in a manner transparent for older decoders are concerned, same are of comparatively minor nature due to the interleaving.

Claims (50)

1 . An audio decoder, comprising an electronic circuit, programmed computer or microprocessor configured to:

extract a sequence of frequency-domain coefficients relating to a frame of an audio signal from a data stream;

extract scale factors for the frame from the data stream;

subject the frequency-domain coefficients, scaled according to the scale factors, to inverse transformation to obtain a time-domain portion of the audio signal corresponding to the frame;

subject the time-domain portion to an overlap-add process to obtain the audio signal,

wherein the audio decoder is responsive to a signalization within the data stream for the frame of the audio signal so as to, depending on the signalization,

form one transform out of the sequence of frequency-domain coefficients by maintaining a sequential order of the frequency-domain coefficients in the sequence of frequency-domain coefficients and subject the one transform, scaled according to the scale factors, to an inverse transformation of a first transform length, or

form more than one transform by de-interleaving the frequency-domain coefficients from the sequence of frequency-domain coefficients and subject, scaled according to the scale factors, each of the more than one transforms to an inverse transformation of a second transform length, shorter than the first transform length,

wherein the frequency-domain coefficients are grouped into a number of scale factor bands which is independent from the signalization, and a scale factor is extracted from the data stream for each scale factor band.

2 . The audio decoder of claim 1 , configured to use context-based entropy decoding to extract the sequence of frequency-domain coefficients from the data stream, with assigning, for each frequency-domain coefficient, a context to the respective frequency-domain coefficient in a manner independent from the signalization.

3 . The audio decoder of claim 1 , configured to subject the frequency-domain coefficients to scaling according to the scale factors at a spectral resolution independent from the signalization.

4 . The audio decoder of claim 1 , configured to subject the sequence of frequency-domain coefficients to noise filling, at a spectral resolution which is independent from the signalization.

5 . The audio decoder of claim 1 , configured to:

in the formation of the one transform, apply inverse temporal noise shaping filtering on the sequence of frequency-domain coefficients using the sequential order, and

in the formation of the more than one transforms, apply inverse temporal noise shaping filtering on the sequence of frequency-domain coefficients by de-interleaving the frequency-domain coefficients of the more than one transforms, concatenating of the more than one transforms spectrally transform-wise to yield a concatenation of the more than one transforms and applying the inverse temporal noise shaping filtering on the concatenation of the more than one transforms.

6 . The audio decoder of claim 1 , configured to support joint-stereo coding with or without inter-channel stereo prediction and to use the sequence of frequency-domain coefficients as a sum or difference spectrum or prediction residual of the inter-channel stereo prediction.

7 . The audio decoder of claim 1 , wherein the number of the more than one transforms equals 2, and the first transform length is twice the second transform length.

8 . The audio decoder of claim 1 , wherein the inverse transformation is an inverse modified discrete cosine transform (MDCT).

9 . An audio encoder, comprising an electronic circuit, programmed computer or microprocessor configured to:

subject time-domain portions of an audio signal to transformation to obtain, for each time-domain portion, frequency-domain coefficients;

inversely scale the frequency-domain coefficients according to scale factors,

wherein the audio encoder configured to switch, for a predetermined frame, between

performing one transform of a first transform length, and

performing more than one transform of a second transform length, shorter than the first transform length;

signal the switching for the predetermined frame by a signalization within the data stream for the predetermined frame,

wherein the audio encoder is configured to:

arrange, for the predetermined frame, the frequency-domain coefficients, inversely scaled according to scale factors, at a sequential order so as to obtain a sequence of frequency-domain coefficients by, depending on the signalization,

sequentially arranging the frequency-domain coefficients of the one transform in case of one transform performed for the respective frame, and

by interleaving the frequency-domain coefficients of the more than one transform of the respective frame in case of more than one transform performed for the respective frame, so that spectrally corresponding frequency-domain coefficients of the more than two transforms immediately follow each other; and

wherein the frequency-domain coefficients are grouped into a number of scale factor bands which is independent from the signalization, and a scale factor is inserted into the data stream for each scale factor band.

10 . A method for audio decoding, comprising:

extracting a sequence of frequency-domain coefficients relating to a frame of an audio signal from a data stream;

extracting scale factors for the frame from the data stream;

subjecting the frequency-domain coefficients, scaled according to scale factors, to inverse transformation to obtain a time-domain portion of the audio signal; and

subjecting the time-domain portion to an overlap-add process to obtain the audio signal,

wherein the subjection to inverse transformation is responsive to a signalization within the data stream for the frame so as to, depending on the signalization, comprise:

forming one transform out of the sequence of frequency-domain coefficients by maintaining a sequential order of the frequency-domain coefficients in the sequence of frequency-domain coefficients and subjecting the one transform, scaled according to the scale factors, to an inverse transformation of a first transform length, or

forming more than one transform by de-interleaving the frequency-domain coefficients from the sequence of frequency-domain coefficients and subjecting, scaled according to the scale factors, each of the more than one transforms to an inverse transformation of a second transform length, shorter than the first transform length,

wherein the frequency-domain coefficients are grouped into a number of scale factor bands which is independent from the signalization, and a scale factor is extracted from the data stream for each scale factor band.

11 . A method for audio encoding, comprising:

subjecting time-domain portions of an audio signal to transformation to obtain, for each time-domain portion, frequency-domain coefficients;

inversely scaling the frequency-domain coefficients according to scale factors;

switching for a predetermined frame between:

performing one transform of a first transform length, and

performing more than one transform of a second transform length, shorter than the first transform length;

signaling the switching by a signalization for the predetermined frame within the data stream;

arranging, for the predetermined frame, the frequency-domain coefficients, inversely scaled according to scale factors, at a sequential order so as to obtain a sequence of frequency-domain coefficients by, depending on the signalization,

sequentially arranging the frequency-domain coefficients of the one transform in case of one transform performed for the respective frame, and

by interleaving the frequency-domain coefficients of the more than one transform of the respective frame in case of more than one transform performed for the respective frame, so that spectrally corresponding frequency-domain coefficients of the more than two transforms immediately follow each other; and

wherein the frequency-domain coefficients are grouped into a number of scale factor bands which is independent from the signalization, and a scale factor is inserted into the data stream for each scale factor band.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: DICK, SASCHA; HELMRICH, CHRISTIAN; HOELZER, ANDREAS
To: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 067735/0407 →
Priority Claims (2)
EP 13177373 · Jul 22, 2013 · regional
EP 13189334 · Oct 18, 2013 · regional
Continuity (5)
Continuation 17227178 · Apr 9, 2021
Continuation 16284534 · Feb 25, 2019
Continuation 15004563 · Jan 22, 2016
Continuation PCTEP2014065169 · Jul 15, 2014
Related Publication 20240127836A1 · Apr 18, 2024
References Cited (102)
US 5394473A · Davidson · 1995 [cited by applicant]
US 5848391A · Bosi et al. · 1998 [cited by applicant]
US 6131084A · Hardwick · 2000 [cited by applicant]
US 6353807B1 · Tsutsui et al. · 2002 [cited by applicant]
US 6424936B1 · Shen et al. · 2002 [cited by applicant]
US 6950794B1 · Subramaniam et al. · 2005 [cited by applicant]
US 6978236B1 · Liljeryd et al. · 2005 [cited by applicant]
US 7283968B2 · Youn · 2007 [cited by applicant]
US 7860709B2 · Maekinen · 2010 [cited by applicant]
US 7953595B2 · Xie et al. · 2011 [cited by applicant]
US 8428957B2 · Garudadri et al. · 2013 [cited by applicant]
US 10242268B2 · Harris et al. · 2019 [cited by applicant]
US 10984809B2 · Dick et al. · 2021 [cited by applicant]
US 20020138259A1 · Kawahara · 2002 [cited by applicant]
US 20040131204A1 · Vinton · 2004 [cited by applicant]
US 20050185850A1 · Vinton et al. · 2005 [cited by applicant]
US 20050267744A1 · Nettre et al. · 2005 [cited by applicant]
US 20060074642A1 · You · 2006 [cited by applicant]
US 20060122825A1 · Oh et al. · 2006 [cited by applicant]
US 20080059202A1 · You · 2008 [cited by applicant]
US 20080065373A1 · Oshikiri et al. · 2008 [cited by applicant]
US 20080140428A1 · Choo et al. · 2008 [cited by applicant]
US 20080253440A1 · Srinivasan et al. · 2008 [cited by applicant]
US 20090012797A1 · Boehm et al. · 2009 [cited by applicant]
US 20090319278A1 · Yoon et al. · 2009 [cited by applicant]
US 20100017213A1 · Edler et al. · 2010 [cited by applicant]
US 20100076754A1 · Kovesi et al. · 2010 [cited by applicant]
US 20100094637A1 · Vinton · 2010 [cited by applicant]
US 20100114583A1 · Lee et al. · 2010 [cited by applicant]
US 20100286991A1 · Hedelin et al. · 2010 [cited by applicant]
US 20110046966A1 · Dalimba · 2011 [cited by applicant]
US 20110238425A1 · Neuendorf et al. · 2011 [cited by applicant]
US 20110238426A1 · Fuchs et al. · 2011 [cited by applicant]
US 20110257982A1 · Smithers · 2011 [cited by applicant]
US 20120271644A1 · Bessette et al. · 2012 [cited by applicant]
US 20130030819A1 · Purnhagen et al. · 2013 [cited by applicant]
US 20130182862A1 · Disch · 2013 [cited by applicant]
US 20130253938A1 · You · 2013 [cited by applicant]
US 20140257824A1 · Taleb et al. · 2014 [cited by applicant]
US 20140310011A1 · Biswas et al. · 2014 [cited by applicant]
US 20160050420A1 · Helmrich et al. · 2016 [cited by applicant]
CN 1247415A · 2000 [cited by applicant]
CN 2482427Y · 2002 [cited by applicant]
CN 1625768A · 2005 [cited by applicant]
CN 1677493A · 2005 [cited by applicant]
CN 1735925A · 2006 [cited by applicant]
CN 1926609A · 2007 [cited by applicant]
CN 101061533A · 2007 [cited by applicant]
CN 101494054A · 2009 [cited by applicant]
CN 101939781A · 2011 [cited by applicant]
CN 102113051A · 2011 [cited by applicant]
CN 102177426A · 2011 [cited by applicant]
CN 102483923A · 2012 [cited by applicant]
EP 2304719A1 · 2011 [cited by applicant]
JP 10293600A · 1998 [cited by applicant]
JP 2003510644A · 2003 [cited by applicant]
JP 2005338637A · 2005 [cited by applicant]
JP 2008129250A · 2008 [cited by applicant]
JP 2009500682A · 2009 [cited by applicant]
JP 2009500683A · 2009 [cited by applicant]
JP 2009500684A · 2009 [cited by applicant]
JP 2010500631A · 2010 [cited by applicant]
JP 4731775B2 · 2011 [cited by applicant]
JP 2013508765A · 2013 [cited by applicant]
RU 2455709C2 · 2012 [cited by applicant]
RU 2483365C2 · 2013 [cited by applicant]
WO 0036753A1 · 2000 [cited by applicant]
WO 0122403A1 · 2001 [cited by applicant]
WO 2004079923A2 · 2004 [cited by applicant]
WO 2004082288A1 · 2004 [cited by applicant]
WO 2005034080A2 · 2005 [cited by applicant]
WO 2007008000A2 · 2007 [cited by applicant]
WO 2007008001A2 · 2007 [cited by applicant]
WO 2007008002A2 · 2007 [cited by applicant]
WO 2010003556A1 · 2010 [cited by applicant]
WO 2010040522A2 · 2010 [cited by applicant]
WO 2011147950A1 · 2011 [cited by applicant]
WO 2012161675A1 · 2012 [cited by applicant]
WO 2013079524A2 · 2013 [cited by applicant]
Herre, Jürgen, and Sascha Disch. “Perceptual audio coding.” Academic press library in Signal processing. vol. 4. Elsevier, 2014. 757-800. (Year: 2014). [cited by examiner]
Johnston, James D., et al. “MPEG audio coding.” Wavelet, subband and block transforms in communications and multimedia. Boston, MA: Springer US, 1999. 207-253. (Year: 1999). [cited by examiner]
Bosi, M, et al., “Final Text of ISO/IEC 13818-7 AAC”, 39. MPEG Meeting; Apr. 7, 1997-Apr. 11, 1997; Bristol; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. N1650, Apr. 11, 1997 (Apr. 11, 1997), XP030010416… [cited by applicant]
Bosi, Marina, et al. “ISO/IEC MPEG-2 advanced audio coding.” Journal of the Audio engineering society 45.10 (1997): 789-814. ( Year: 1997). [cited by applicant]
Brandenburg, Karlheinz, and Gerhard Stoll. “ISO/MPEG-1 audio: a generic standard for coding of high-quality digital audio.” Journal of the Audio Engineering Society 42.10 (1994): 780-792. (Year: 1994). [cited by applicant]
Noll, Peter. “MPEG digital audio coding.” IEEE signal processing magazine 14.5 (1997): 59-81. (Year: 1997). [cited by applicant]
“ATSC Standard: Digital Audio Compression (AC-3, E-AC-3)”, Advanced Television Systems Committee. Doc.A/52:2012, Dec. 17, 2012, pp. 1-270. [cited by applicant]
“Information technology—Generic coding of moving pictures and associated audio information—Part 7: Advanced Audio Coding (AAC)”, ISO/IEC 13818-7:2004(E), Third edition, Oct. 15, 2004, 206 pp. [cited by applicant]
Bosi, M. , et al., “Final Text of ISO/IEG 13818-7 AAC”, 39. MPEG Meeting Apr. 7, 1997-Apr. 11, 1997; Bristol, Motion Picture Expert Group or ISO/IEG JTG1/SG29/WG11, No. N1650, pp. 1-106. [cited by applicant]
Bosi, Marina , et al., “ISO/IEC MPEG-2 Advanced Audio Coding”, J. Audio Eng. Soc., vol. 45, No. 10—Oct. 1997, pp. 789-814. [cited by applicant]
Davidson, G.A. , et al., “Digital Audio Coding: Dolby AC-3, Digital Signal Processing Handbook”, CRC Press LLC-IEEE Press, 1999. [cited by applicant]
Davidson, Grant A, “Digital Audio Coding: Dolby AC-3”, In: The Digital Signal Processing Handbook, CRC Press LLC, IEEE Press, XP055140739, ISBN: 978-0-84-938572-8, p. 41. 3, line 12, paragraph 41.1—p. 41.4, line 6; figu… [cited by applicant]
Herre, Jurgen , et al., “Enhancing the Performance of Perceptual Audio Coders by Using Temporal Noise Shaping (TNS)”, Audio Engineering Society Convention 101. Audio Engineering Society 1996. [cited by applicant]
Hui, Dai , “Digital Video Technology”, Bejing, Dec. 2012, 30 pages, 1-30. [cited by applicant]
ISO/IEC 14496-3:2009 , “Information Technology—Coding of Audio-Visual Objects—Part 3: Audio”, International Organization for Standardization, Geneva, Switzerland, Aug. 2009. [cited by applicant]
ISO/IEC 23003-3 , “Information Technology—MPEG audio technologies—Part 3: Unified Speech and Audio Coding”, International Organization for Standardization, Geneva, Jan. 2012, 286 pages. [cited by applicant]
ITU-T , “G.719: Low-complexity, full-band audio coding for high-quality, conversational applications”, Recommendation ITU-T G.719, Telecommunication Standardization Sector of ITU,, Jun. 2008, 58 pages. [cited by applicant]
Johnston, J.D. , et al., “Sum-Difference Stereo Transform Coding”, in Proc. IEEE ICASSP-92, vol. 2, pp. II-569-II-572. [cited by applicant]
Johnston, James D, et al., “MPEG Audio Coding”, Wavelet, subband and block transforms in communications and multimedia. Springer, Boston, MA 2002. 207-253, 2002, pp. 207-253. [cited by applicant]
Neuendorf, Max , et al., “MPEG Unified Speech and Audio Coding—The ISO/MPEG Standard for High-Efficiency Audio Coding of all Content Types”, Audio Engineering Society Convention Paper 8654, Presented at the 132nd Conven… [cited by applicant]
Ravelli, Emmanuel , et al., “Union of MDCT Bases for Audio Coding”, IEEE Transactions on Audio, Speech and Language Processing, IEEE Service Center, vol. 16, No. 8, XP011236278, ISSN: 1558-7916, DOI: 10.1109/TASL.2008.2… [cited by applicant]
Sperschneider, Ralph , “Text of ISO/IEC13818-7:2004 (MPEG-2 AAC 3rd edition)”, ISO/IEC JTC1/SC29/WG11 N6428, Munich, Germany,, pp. 1-198. [cited by applicant]
Valin, JM , et al., “Defintion of the Opus Audio Codec”, IETF, pp. 1-326. [cited by applicant]