IP Library › Granted Patent US 12,322,399
Granted Patent B2
US 12,322,399 · App. 18/502,973 · Granted Jun 3, 2025

MDCT-based complex prediction stereo coding

Inventors: Heiko Purnhagen (Sundbyberg, SE); Pontus Carlsson (Bromma, SE); Lars Villemoes (Järfälla, SE)
Assignee: DOLBY INTERNATIONAL AB
G10L19/008G10L19/0212G10L19/06G10L19/167G10L19/18H04S3/008G10L25/12H04S2400/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,399
App. No.
18/502,973
Granted
Jun 3, 2025
Kind
B2
Abstract

The invention provides methods and devices for stereo encoding and decoding using complex prediction in the frequency domain. In one embodiment, a decoding method, for obtaining an output stereo signal from an input stereo signal encoded by complex prediction coding and comprising first frequency-domain representations of two input channels, comprises the upmixing steps of: (i) computing a second frequency-domain representation of a first input channel; and (ii) computing an output channel on the basis of the first and second frequency-domain representations of the first input channel, the first frequency-domain representation of the second input channel and a complex prediction coefficient. The upmixing can be suspended responsive to control data.

Claims (42)

1. A decoder system for providing a stereo signal by complex prediction stereo coding, the decoder system comprising:

an upmix stage adapted to generate the stereo signal based on first frequency-domain representations of a downmix signal and a residual signal, each of the first frequency-domain representations comprising first spectral components representing spectral content of the corresponding signal expressed in a first subspace of a multidimensional space, wherein the upmix stage:

computes a second frequency-domain representation of the downmix signal based on the first frequency-domain representation thereof, the second frequency-domain representation comprising second spectral components representing spectral content of the signal expressed in a second subspace of the multidimensional space that includes a portion of the multidimensional space not included in the first subspace, wherein the second spectral components of the downmix signal are determined by applying a Finite Impulse Response (FIR) filter to the first spectral components of the downmix signal;

computes a first frequency-domain representation of a side signal(S), the first frequency-domain representation of the side signal(S) comprising first spectral components representing spectral content of the side signal expressed in the first subspace of the multidimensional space, on the basis of the first and second frequency-domain representations of the downmix signal, the first frequency-domain representation of the residual signal, and a complex prediction coefficient (a) encoded in a bit stream signal received by the decoder system; wherein each spectral component represents a range of frequencies, and wherein each of the first spectral components of the side signal is determined from spectral components of the downmix signal and the residual signal representing the same range of frequencies as the first spectral component of the side signal; and

computes the stereo signal on the basis of the first frequency-domain representation of the downmix signal and the side signal,

wherein the upmix stage is adapted to apply independent bandwidth limits for the downmix signal and the residual signal.

2. The decoder system of claim 1 , wherein an impulse response of the FIR filter is determined depending on a window function applied to determine the first frequency domain representation of the downmix signal.

3. The decoder system of claim 1 , wherein the bandwidth limits to be applied are signaled by two data fields, indicating for each of the signals a highest frequency band to be decoded.

4. The decoder system of claim 3 , adapted to receive an MPEG bit stream in which each of said data fields is encoded as a value of a max_sfb parameter of the MPEG bit stream.

5. The decoder system of claim 1 , further comprising:

a dequantization stage arranged upstream of the upmix stage, for providing said first frequency-domain representations of the downmix signal and residual signal based on a bit stream signal.

6. The decoder system of claim 1 , wherein:

the first spectral components have real values expressed in the first subspace;

the second spectral components have imaginary values expressed in the second subspace.

7. The decoder system of claim 1 , wherein the first spectral components are obtainable by one of the following:

a discrete cosine transform, DCT, or

a modified discrete cosine transform, MDCT.

8. The decoder system of claim 1 , further comprising at least one temporal noise shaping, TNS, module arranged upstream of the upmix stage; and

at least one further TNS module arranged downstream of the upmix stage; and

a selector arrangement for selectively activating either:

(a) said TNS module(s) upstream of the upmix stage, or

(b) said further TNS module(s) downstream of the upmix stage.

9. The decoder system of claim 6 , wherein:

the downmix signal is partitioned into successive time frames, each associated with a value of the complex prediction coefficient; and

computing a second frequency-domain representation of the downmix signal is deactivated responsive to the absolute value of the imaginary part of the complex prediction coefficient being smaller than a predetermined tolerance for a time frame.

10. The decoder system of claim 1 , said stereo signal being represented in the time domain and the decoder system further comprising:

a switching assembly arranged between said dequantization stage and said upmix stage, operable to function as either:

(a) a pass-through stage, or

(b) a sum-and-difference stage,

thereby enabling switching between directly and jointly coded stereo input signals;

an inverse transform stage adapted to compute a time-domain representation of the stereo signal; and

a selector arrangement arranged upstream of the inverse transform stage, adapted to selectively connect this to either:

(a) a point downstream of the upmix stage, whereby the stereo signal obtained by complex prediction is supplied to the inverse transform stage; or

(b) a point downstream of the switching assembly and upstream of the upmix stage, whereby a stereo signal obtained by direct stereo coding is supplied to the inverse transform stage.

11. A decoding method for upmixing an input stereo signal by complex prediction stereo coding into an output stereo signal, wherein:

said input stereo signal comprises first frequency-domain representations of a downmix signal and a residual signal and a complex prediction coefficient; and

each of said first frequency-domain representations comprises first spectral components representing spectral content of the corresponding signal expressed in a first subspace of a multidimensional space,

the method being performed by an upmix stage, and comprising:

computing a second frequency-domain representation of the downmix signal based on the first frequency-domain representation thereof, the second frequency-domain representation comprising second spectral components representing spectral content of the signal expressed in a second subspace of the multidimensional space that includes a portion of the multidimensional space not included in the first subspace, wherein computing a second frequency-domain representation of the downmix signal includes determining the second spectral components of the downmix signal by applying a Finite Impulse Response (FIR) filter to the first spectral components of the downmix signal; and

computing a first frequency-domain representation of a side signal(S), the first frequency-domain representation of the side signal(S) comprising first spectral components representing spectral content of the side signal expressed in the first subspace of the multidimensional space, on the basis of the first and second frequency-domain representations of the downmix signal, the first frequency-domain representation of the residual signal, and the complex prediction coefficient; wherein each spectral component represents a range of frequencies, and wherein each of the first spectral components of the side signal is determined from spectral components of the downmix signal and the residual signal representing the same range of frequencies as the first spectral component of the side signal, and

wherein independent bandwidth limits are applied for the downmix signal and the residual signal.

12. A computer-program product comprising a non-transitory computer-readable medium storing instructions which when executed by a general-purpose computer perform the method set forth in claim 11 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2024
From: PURNHAGEN, HEIKO; CARLSSON, PONTUS; VILLEMOES, LARS
To: DOLBY INTERNATIONAL AB
Reel/Frame 066038/0166 →
Continuity (10)
Continuation 17560295 · Dec 23, 2021
Continuation 16931377 · Jul 16, 2020
Continuation 16593235 · Oct 4, 2019
Continuation 16427735 · May 31, 2019
Continuation 16222721 · Dec 17, 2018
Continuation 15849622 · Dec 20, 2017
Continuation 14793297 · Jul 7, 2015
Division 13638898
Provisional Application 61322458 · Apr 9, 2010
Related Publication 20240144940A1 · May 2, 2024
References Cited (115)
US 6980933B2 · Cheng · 2005 [cited by applicant]
US 7693709B2 · Thumpudi · 2010 [cited by applicant]
US 8346547B1 · Tang · 2013 [cited by examiner]
US 8655670B2 · Purnhagen · 2014 [cited by applicant]
US 8948404B2 · Kim · 2015 [cited by applicant]
US 8977541B2 · Toguri · 2015 [cited by applicant]
US 9082395B2 · Heiko · 2015 [cited by applicant]
US 9159326B2 · Purnhagen · 2015 [cited by examiner]
US 10475459B2 · Purnhagen · 2019 [cited by examiner]
US 10734002B2 · Purnhagen · 2020 [cited by examiner]
US 20030091194A1 · Teichmann · 2003 [cited by examiner]
US 20050078941A1 · Takeda · 2005 [cited by examiner]
US 20050165587A1 · Cheng et al. · 2005 [cited by applicant]
US 20050197831A1 · Edler et al. · 2005 [cited by applicant]
US 20060047523A1 · Ojanpera · 2006 [cited by examiner]
US 20060253276A1 · Kang · 2006 [cited by applicant]
US 20070016415A1 · Thumpudi · 2007 [cited by applicant]
US 20070174062A1 · Mehrotra · 2007 [cited by applicant]
US 20070203697A1 · Pang · 2007 [cited by applicant]
US 20080046253A1 · Vinton · 2008 [cited by applicant]
US 20080126104A1 · Seefeldt · 2008 [cited by applicant]
US 20080136686A1 · Feiten · 2008 [cited by examiner]
US 20080255828A1 · Chesnutt · 2008 [cited by applicant]
US 20090006103A1 · Koishida · 2009 [cited by examiner]
US 20090048852A1 · Burns et al. · 2009 [cited by applicant]
US 20090063140A1 · Villemoes · 2009 [cited by applicant]
US 20090083040A1 · Myburg et al. · 2009 [cited by applicant]
US 20090198356A1 · Goodwin · 2009 [cited by applicant]
US 20090210234A1 · Sung · 2009 [cited by applicant]
US 20090216542A1 · Pang · 2009 [cited by applicant]
US 20090319263A1 · Gupta · 2009 [cited by examiner]
US 20100010807A1 · Oh · 2010 [cited by applicant]
US 20100014679A1 · Kim · 2010 [cited by applicant]
US 20100023335A1 · Szczerba · 2010 [cited by applicant]
US 20100046759A1 · Pang · 2010 [cited by applicant]
US 20100057447A1 · Ehara · 2010 [cited by applicant]
US 20100070284A1 · Oh · 2010 [cited by applicant]
US 20100086030A1 · Chen · 2010 [cited by examiner]
US 20100262427A1 · Chivukula et al. · 2010 [cited by applicant]
US 20110178795A1 · Bayer · 2011 [cited by applicant]
US 20110202355A1 · Grill · 2011 [cited by examiner]
US 20110257981A1 · Beack · 2011 [cited by applicant]
US 20120002818A1 · Heiko et al. · 2012 [cited by applicant]
US 20120245947A1 · Neuendorf · 2012 [cited by examiner]
US 20130012411A1 · Soito · 2013 [cited by applicant]
US 20130030819A1 · Purnhagen · 2013 [cited by applicant]
US 20150269950A1 · Schug · 2015 [cited by applicant]
US 20160125888A1 · Purnhagen · 2016 [cited by applicant]
US 20160133273A1 · Kaniewska · 2016 [cited by applicant]
US 20180108366A1 · Villemoes · 2018 [cited by examiner]
AU 2008314030B2 · 2009 [cited by applicant]
CN 1666572 · 2005 [cited by applicant]
CN 1677490 · 2005 [cited by applicant]
CN 1922656 · 2007 [cited by applicant]
CN 101031959 · 2007 [cited by applicant]
CN 101067931 · 2007 [cited by applicant]
CN 101160619 · 2008 [cited by applicant]
CN 101202043 · 2008 [cited by applicant]
CN 101253557 · 2008 [cited by applicant]
CN 101529501B · 2009 [cited by applicant]
CN 101583994 · 2009 [cited by applicant]
EP 1107232 · 2001 [cited by applicant]
EP 1803325B1 · 2008 [cited by applicant]
EP 2137725 · 2009 [cited by applicant]
EP 2144230 · 2010 [cited by applicant]
EP 2144231 · 2010 [cited by applicant]
JP H04506141 · 1992 [cited by applicant]
JP 2005521921 · 2005 [cited by applicant]
JP 2007520748A · 2007 [cited by applicant]
JP 2007525716 · 2007 [cited by applicant]
JP 2008517337 · 2008 [cited by applicant]
JP 2008516290A · 2008 [cited by applicant]
JP 2009536360A · 2009 [cited by applicant]
JP 2010507115A · 2010 [cited by applicant]
KR 20080027129 · 2008 [cited by applicant]
KR 101698439B1 · 2017 [cited by applicant]
RU 2174714 · 2001 [cited by applicant]
RU 2345506C2 · 2009 [cited by applicant]
RU 2346339C2 · 2009 [cited by applicant]
RU 2374703C2 · 2009 [cited by applicant]
RU 2380766 · 2010 [cited by applicant]
WO 2006091139A1 · 2006 [cited by applicant]
WO 2006091150A1 · 2006 [cited by applicant]
WO 20070128523 · 2007 [cited by applicant]
WO 2007140809A1 · 2007 [cited by applicant]
WO 2008046531A1 · 2008 [cited by applicant]
WO 20080063035A1 · 2008 [cited by applicant]
WO 2009038512A1 · 2009 [cited by applicant]
WO 2009049895A1 · 2009 [cited by applicant]
WO 2009049896A1 · 2009 [cited by applicant]
WO 2009084920A1 · 2009 [cited by applicant]
WO 2009141775 · 2009 [cited by applicant]
WO 2010019265W · 2010 [cited by applicant]
Avendano, C. et al “A Frequency-Domain Approach to Multichannel Upmix” JAES vol. 52, Issue 7/8, pp. 740-749, Jul. 2004, published on Jul. 15, 2004. [cited by applicant]
Baumgarte, Frank “Enhanced Dynamic Range Control Metadata”, MPEG Meeting, Apr. 17, 2013, Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11. [cited by applicant]
Breebaart, J. et al “Spatial Audio Processing—Ch. 6 MPEG Surround” in “Spatial Audio Processing”, Jan. 1, 2007, John Wiley & Sons, pp. 93-115. [cited by applicant]
Carlsson, P. et al. “Technical Description of CE on Improved Stereo Coding in USAC” MPEG meeting Jul. 22, 2010, Geneva, Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11. [cited by applicant]
Chen, S. et al. “Estimating Spatial Cues for Audio Coding in MDCT Domain” ICME 2009, pp. 53-56. [cited by applicant]
Chen, S. et al. “Spatial Parameters for Audio Coding: MDCT Domain Analysis and Synthesis” Springer, published in Jul. 2009. [cited by applicant]
Chen, S.et al., “Analysis and Synthesis of Spatial Parameters Using MDCT” Multimedia and Ubiquitous Engineering published in Jun. 2009, pp. 18-21. [cited by applicant]
Derrien, O. et al. “A New Model-Based Algorithm for Optimizing the MPEG-AAC in MS-Stereo” IEEE Signal Processing Society, published in Nov. 2008, vol. 16, Issue 8, pp. 1373-1382. [cited by applicant]
Goodwin, M. et al “A Frequency-Domain Framework for Spatial Audio Coding Based on Universal Spatial Cues” AES Convention Paper 6751, presented at the 120th Convention, May 20-23, 2006, Paris, France, pp. 1-12. [cited by applicant]
Heiko, P. et al. “MPEG 2009/M16921” Technical Description of Propposed Unified Stereo Coding in USAC, Oct. 30, 2009. [cited by applicant]
Herre, J. et al. “MPEG Surround—The ISO/MPEG Standard for Efficient and Compatible Multi-Channel Audio Coding” AES Convention Paper Presented at the 122nd Convention May 5-8, 2007, Vienna, Austria. [cited by applicant]
Johnston, J.D. et al “Sum-Difference Stereo Transform Coding” IEEE International Conference on Acoustics, Speech, and Signal Processing, Mar. 23-26, 1992, pp. 569-572, vol. 2. [cited by applicant]
Kruger, H. et al “A New Approach for Low-Delay Joint-Stereo Coding” ITG Facthtagung Sprachkommunikation Oct. 8-10, 2008, pp. 1-4. [cited by applicant]
Kuech, F. et al “Description of the Fraunhofer IIS Submission for the DRC CFP” MPEG Meeting Oct. 28-Nov. 1, 2013, Geneva, Motion Expert Group or ISO/IEC JTC1/SC29/WG11. [cited by applicant]
Neuendorf, M. et al “MPEG Unified Speech and Audio Coding—The ISO/MPEG Standard for High-Efficiency Audio Coding of All Content Types” Audio Engineering Society Convention 132, Apr. 26, 2012. [cited by applicant]
Neuendorf, M. et al “The ISO/MPEG Unified Speech and Audio Coding Standard—Consistent High Quality for All Content Types and at all Bit Rates” Journal of the Audio Engineering Society, vol. 61, Issue 12, pp. 956-977, De… [cited by applicant]
Ofir, H. et al. “Audio Packet Loss Concealment in a Combined MDCT-MDST Domain” IEEE Signal Processing Letter, vol. XX, No. Y, 2007. [cited by applicant]
Pang, Hee-Suk, “Clipping Prevention Scheme for MPEG Surround” ETRI Journal, Electronics and Telecommunications Research Institute, vol. 30, No. 4, Aug. 2008, pp. 606-608. [cited by applicant]
Purnhagen, Heiko “Low Complexity Parametric Stereo Coding in MPEG-4” Proc. of the 7th Int. Conference on Digital Audio Effects, Naples, Italy, Oct. 5-8, 2004, pp. 163-168. [cited by applicant]
Robinson, C. et al “Dynamic Range Control via Metadata 5028 (J-1) AN Audio Engineering Society Preprint, Dynamic Range Control via Metadata” AES, presented at the 107th Convention, Sep. 24-27, 1999. pp. 1-14. [cited by applicant]
Schuijers, E. et al. “Low Complexity Parametric Stereo Coding” AES presented at the 116th Convention, May 8-11, 2004, Berlin, Germany. [cited by applicant]
ISO/IEC FDIS 23003-3:2011(E), Information technology—MPEG audio technologies—Part 3: Unified speech and audio coding. ISO/IEC JTC 1/SC 29/WG 11. Sep. 20, 2011. [cited by applicant]