IP Library › Granted Patent US 12,243,541
Granted Patent B2
US 12,243,541 · App. 17/884,910 · Granted Mar 4, 2025

Apparatus and method for encoding or decoding a multichannel signal using a side gain and a residual gain

Inventors: Jan Buethe (Erlangen, DE); Guillaume Fuchs (Bubenreuth, DE); Wolfgang Jaegers (Erlangen, DE); Franz Reutelhuber (Erlangen, DE); Juergen Herre (Erlangen, DE); Eleni Fotopoulou (Nuremberg, DE); Markus Multrus (Nuremberg, DE); Srikanth Korse (Nuremberg, DE)
Assignee: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
G10L19/008H04S1/007H04S3/008H04S7/30H04S2400/01H04S2420/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,541
App. No.
17/884,910
Granted
Mar 4, 2025
Kind
B2
Abstract

An apparatus for encoding a multi-channel signal having at least two channels, has: a downmixer for calculating a downmix signal from the multi-channel signal; a parameter calculator for calculating a side gain from a first channel of the at least two channels and a second channel of the at least two channels and for calculating a residual gain from the first channel and the second channel; and an output interface for generating an output signal, the output signal having information on the downmix signal, and on the side gain and the residual gain.

Claims (253)

1. An apparatus for decoding an encoded multi-channel audio signal comprising a single-channel downmix audio signal, a side gain and a residual gain, the apparatus comprising:

an input interface configured for receiving the encoded multi-channel audio signal and for acquiring the single-channel downmix audio signal, the side gain and the residual gain from the encoded multi-channel audio signal;

a residual signal synthesizer configured for synthesizing a single-channel residual signal using the residual gain, wherein the residual signal synthesizer comprises a weighter configured for weighting a raw residual signal by the residual gain to obtain the single-channel residual signal; and

an upmixer configured for upmixing the single-channel downmix audio signal using the side gain and the single-channel residual signal to acquire a reconstructed first audio channel and a reconstructed second audio channel.

2. The apparatus of claim 1 , wherein the upmixer is configured

to perform a first weighting operation of the single-channel downmix audio signal gel using the side gain to acquire a first weighted downmix signal, and

to perform a second weighting operation using the side gain and the single-channel downmix audio signal to acquire a second weighted downmix signal, wherein the first weighting operation is different from the second weighting operation, so that the first weighted downmix signal is different from the second weighted downmix signal.

3. The apparatus of claim 2 ,

wherein the upmixer is configured to calculate the reconstructed first channel using a first combination of the first weighted downmix signal and the single-channel residual signal and to calculate the reconstructed second channel using a second combination of the second weighted downmix signal and the single-channel residual signal, wherein the first combination and the second combination are different from each other.

4. The apparatus of claim 3 ,

wherein one of the first and the second combinations is an adding operation and the other of the first and the second combinations is a subtracting operation.

5. The apparatus of claim 2 ,

wherein the upmixer is configured to perform the first weighting operation comprising a weighting factor derived from a sum of the side gain and a first predetermined number, and

wherein the upmixer is configured to perform the second weighting operation comprising a weighting factor derived from a difference between a second predetermined number and the side gain, wherein the first predetermined number and the second predetermined number are equal to each other or are different from each other.

6. The apparatus of claim 1 ,

wherein the residual signal synthesizer is configured to weight the single-channel residual signal of a preceding frame using the residual gain for a current frame to acquire the residual signal for the current frame, or

to weight a decorrelated signal derived from the current frame or from one or more preceding frames using the residual gain for the current frame to acquire the single-channel residual signal for the current frame.

7. The apparatus of claim 1 ,

wherein the residual signal synthesizer is configured to calculate the single-channel residual signal so that an energy of the single-channel residual signal is equal to a signal energy indicated by the residual gain.

8. The apparatus of claim 1 ,

wherein the residual signal synthesizer is configured to calculate the single channel residual signal so that values of the single channel residual signal are in a range of ±20% of values as obtainable by the following equation:

res

t

,

k

=

r

~

t

,

b

⁢

g

norm

⁢

ρ

~

t

,

k

2

,

wherein res t,k is the single channel residual signal for a frame t and a frequency bin k, wherein {tilde over (r)} t,b is the residual gain for the frame t and a sub-band h comprising the frequency bin k, and wherein {tilde over (ρ)} t,k is a raw signal for the single channel residual signal for the frame t and the frequency bin k, and wherein g norm is an energy-adjusting factor that can be present or not.

9. The apparatus of claim 8 ,

wherein the g norm as the energy-adjusting factor comprising values in the range of ±20% of values determined by the following equation:

E

M

~

,

t

,

b

E

ρ

~

,

t

,

b

,

wherein E {tilde over (M)},t,b is an energy of the single channel downmix audio signal for a frame t and a sub-band b, and wherein E {tilde over (ρ)},t,b is an energy of the single-channel residual signal for a sub-band b and the frame t, or

wherein a raw signal for the single-channel residual signal is as obtainable based on the following equation:

{tilde over (ρ)} t,k ={tilde over (M)} t−d b ,k ,

wherein {tilde over (ρ)} t,k is a raw signal for the single-channel residual signal t and a frequency bin k,

wherein {tilde over (M)} t−d b ,k is the single-channel downmix audio signal for the frame t−t b and the frequency bin k, wherein d b is a frame delay greater than 0, or

wherein the upmixer is configured to calculate the reconstructed first channel and the reconstructed second channel so that the reconstructed first channel and the reconstructed second channel comprise values that are in the range of ±20% with respect to values as obtainable by the following equations:

L

~

t

,

k

=

(

M

~

t

,

k

(

1

+

g

~

t

,

b

)

+

r

~

t

,

b

⁢

g

norm

⁢

ρ

~

t

,

k

)

2

⁢

R

~

t

,

k

=

(

M

~

t

,

k

(

1

+

g

~

t

,

b

)

+

r

~

t

,

b

⁢

g

norm

⁢

ρ

~

t

,

k

)

2

wherein {tilde over (M)} t,k is the single-channel downmix audio signal for a frame t and a frequency bin k, wherein {tilde over (L)} t,k is the reconstructed first channel for the frame t and the frequency bin k, wherein {tilde over (R)} t,k is the reconstructed second channel for the frame t and the frequency bin k, wherein {tilde over (g)} t,b is the side gain for the frame t and a sub-band b, wherein {tilde over (r)} t,b is the residual gain for the frame t and the suband b, wherein g norm is an energy adjusting factor that can be there or not, and wherein {tilde over (ρ)} t,k is a raw signal for the single channel residual signal for the frame t and the frequency bin k.

10. The apparatus of claim 1 ,

wherein the input interface is configured to acquire, from the encoded multichannel signal, inter-channel phase difference values, and

wherein the residual signal synthesizer or the upmixer is configured to apply the inter-channel phase difference values when calculating the single-channel residual signal or the reconstructed first channel and the reconstructed second channel.

11. The apparatus of claim 10 ,

wherein the upmixer is configured to calculate a phase rotation parameter from an inter-channel phase difference value and

to apply the phase rotation parameter when calculating the reconstructed first channel in a first manner and to apply the inter-channel phase difference value and/or the phase rotation parameter when calculating the reconstructed second channel in a second manner, wherein the first manner is different from the second manner.

12. The apparatus of claim 11 ,

wherein the upmixer is configured to calculate the phase rotation parameter so that the phase rotation parameter is within ±20% of values as obtainable by the following equation:

=

a

⁢

tan

⁢

2

⁢

(

sin

⁡

(

IPD

t

,

b

)

,

cos

⁡

(

IPD

t

,

b

)

+

A

⁢

1

+

g

~

t

,

b

1

-

g

~

t

,

b

)

,

wherein atan 2 is the atan 2 function, wherein β is the phase rotation parameter, wherein IPD is the inter-channel phase difference value, wherein t is a frame index, b is a sub-band index, and g t,b is the side gain for a frame with the frame index t and a sub-band with the sub-band index b, and wherein A is a value between 0.1 and 100 or between −0.1 and −100.

13. The apparatus of claim 1 ,

wherein the input interface is configured to extract code words, wherein a code word jointly comprises a quantized side gain and a quantized residual gain, and wherein the input interface is configured for dequantizing the joint code word using a predefined codebook to acquire the side gain and the residual gain.

14. The apparatus of claim 13 ,

wherein a code book used by the input interface comprises 16 groups of quantization points, each group of quantization points having 8 quantization points, and wherein a code word of the code book is an 8-bit code word, the 8-bit code word comprising:

a single sign bit;

a group of 4 bits identifying a group of quantization points among the 16 groups of quantization points; and

a group of 3 bits identifying a quantization point within the identified group of quantization points among the 16 groups of quantization points.

15. The apparatus of claim 1 ,

wherein the upmixer is configured to calculate the reconstructed first channel and the reconstructed second channel in a spectral domain,

wherein the apparatus further comprises a spectrum-time converter for converting the reconstructed first channel and the reconstructed second channel into a time domain.

16. The apparatus of claim 15 , wherein the spectrum-time converter is configured

to convert, for each one of the reconstructed first channel and the reconstructed second channel, subsequent frames into a time sequence of frames

to weight each time frame using a synthesis window; and

to overlap and add subsequent windowed time frames to acquire a time block of the reconstructed first channel and a time block of the reconstructed second channel.

17. A method of decoding an encoded multi-channel audio signal comprising a single-channel downmix audio signal, a side gain and a residual gain, the method comprising:

receiving the encoded multi-channel audio signal and acquiring the single-channel downmix audio signal, the side gain and the residual gain from the encoded multi-channel audio signal;

synthesizing a single-channel residual signal using the residual gain, wherein the synthesizing comprises weighting a raw residual signal by the residual gain to obtain the single-channel residual signal; and

upmixing the single-channel downmix audio signal using the side gain and the single-channel residual signal to acquire a reconstructed first audio channel and a reconstructed second audio channel.

18. A non-transitory digital storage medium having stored thereon a computer program for performing a method of decoding an encoded multi-channel audio signal comprising a single-channel downmix audio signal, a side gain and a residual gain, the method comprising:

receiving the encoded multi-channel audio signal and acquiring the single-channel downmix audio signal, the side gain and the residual gain from the encoded multi-channel audio signal;

synthesizing a single-channel residual signal using the residual gain, wherein the synthesizing comprises weighting a raw residual signal by the residual gain to obtain the single-channel residual signal; and

upmixing the single-channel downmix audio signal using the side gain and the single-channel residual signal to acquire a reconstructed first audio channel and a reconstructed second audio channel,

when said computer program is run by a computer.

19. An apparatus for encoding a multi-channel audio signal comprising at least two audio channels, comprising:

a downmixer configured for calculating a downmix signal from the multi-channel audio signal;

a parameter calculator configured for calculating a side gain from a first audio channel of the at least two audio channels and a second audio channel of the at least two audio channels and configured for calculating a residual gain from the first audio channel and the second audio channel,

wherein the parameter calculator is configured to calculate the side gain such that an energy of a residual signal being equal to a difference between a side signal of the first audio channel and the second audio channel and the downmix signal multiplied by the side gain is minimal, and to calculate the residual gain so that an energy of the downmix signal if having applied the residual gain is equal to the energy of the residual signal; and

an output interface configured for generating an output signal, the output signal comprising information on the downmix signal, and on the side gain and the residual gain.

20. The apparatus of claim 19 ,

wherein the parameter calculator is configured to calculate the side gain and the residual gain so that the residual gain depends on the side gain, and

wherein the output interface is configured to quantize the side gain and to then quantize the residual gain, wherein a quantization step for the residual gain depends on the value of the side gain.

21. The apparatus of claim 19 ,

wherein the parameter calculator is configured to calculate the side gain and the residual gain so that the residual gain depends on the side gain, and

wherein the output interface is configured to perform a joint quantization using groups of quantization points, each group of quantization points being defined by a fixed amplitude-related ratio between the first audio channel and the second audio channel.

22. The apparatus of claim 21 , wherein the output interface is configured:

to calculate an inter-channel level difference (ILD) between the first audio channel and the second audio channel,

to identify a group of quantization points matching with the inter-channel level difference (ILD) to obtain an identified group,

to only search within the identified group to obtain an identification of a point; and

to combine a sign bit, an identification of the identified group and the identification of the point within the identified group to acquire a code word representing the quantized side gain and the quantized residual gain.

23. The apparatus of claim 21 ,

wherein a code book used by the output interface comprises a code table with a multitude of entries, each entry being identified by a binary code word, each binary code word comprising a sign bit, a first group of bits identifying the group of quantization points, and a second group of bits identifying a quantization point within the group of quantization points.

24. The apparatus of claim 21 ,

wherein a code book used by the output interface comprises 16 groups of quantization points, 8 quantization points per group, and wherein a code word of the code book is an 8-bit code word, the 8-bit code word comprising:

a single sign bit;

a group of 4 bits identifying a group of quantization points among the 16 groups of quantization points; and

a group of 3 bits identifying a quantization point within the identified group of quantization points among the 16 groups of quantization points.

25. A method of encoding a multi-channel audio signal comprising at least two audio channels, comprising:

calculating a downmix signal from the multi-channel audio signal;

calculating a side gain from a first audio channel of the at least two audio channels and a second audio channel of the at least two audio channels and calculating a residual gain from the first audio channel and the second audio channel,

wherein the side gain is calculated such that an energy of a residual signal being equal to a difference between a side signal of the first audio channel and the second audio channel and the downmix signal multiplied by the side gain is minimal, and wherein the residual gain is calculated so that an energy of the downmix signal if having applied the residual gain is equal to an energy of the residual signal; and

generating an output signal, the output signal comprising information on the downmix signal, and on the side gain and the residual gain.

26. A non-transitory digital storage medium having stored thereon a computer program for performing, when said computer program is run by a computer, a method of encoding a multi-channel audio signal comprising at least two audio channels, comprising:

calculating a downmix signal from the multi-channel audio signal;

calculating a side gain from a first audio channel of the at least two audio channels and a second audio channel of the at least two audio channels and calculating a residual gain from the first audio channel and the second audio channel,

wherein the side gain is calculated such that an energy of a residual signal being equal to a difference between a side signal of the first audio channel and the second audio channel and the downmix signal multiplied by the side gain is minimal, and wherein the residual gain is calculated so that an energy of the downmix signal if having applied the residual gain is equal to an energy of the residual signal; and

generating an output signal, the output signal comprising information on the downmix signal, and on the side gain and the residual gain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2022
From: BUETHE, JAN; FUCHS, GUILLAUME; JAEGERS, WOLFGANG; REUTELHUBER, FRANZ; HERRE, JUERGEN; FOTOPOULOU, ELENI; MULTRUS, MARKUS; KORSE, SRIKANTH
To: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 060770/0663 →
Priority Claims (1)
EP 16197816 · Nov 8, 2016 · regional
Continuity (3)
Division 16405422 · May 7, 2019
Continuation PCTEP2017077822 · Oct 30, 2017
Related Publication 20220392464A1 · Dec 8, 2022
References Cited (97)
US 6161089A · Hardwick · 2000 [cited by applicant]
US 6226616B1 · You et al. · 2001 [cited by applicant]
US 6611212B1 · Craven et al. · 2003 [cited by applicant]
US 7573912B2 · Lindblom · 2009 [cited by applicant]
US 7613306B2 · Miyasaka et al. · 2009 [cited by applicant]
US 8036904B2 · Myburg et al. · 2011 [cited by applicant]
US 8811621B2 · Schuijers · 2014 [cited by applicant]
US 8817991B2 · Jaillet et al. · 2014 [cited by applicant]
US 8964994B2 · Jaillet et al. · 2015 [cited by applicant]
US 9191045B2 · Purnhagen et al. · 2015 [cited by applicant]
US 9269361B2 · Ragot et al. · 2016 [cited by applicant]
US 9398294B2 · Robillard et al. · 2016 [cited by applicant]
US 9672837B2 · Purnhagen et al. · 2017 [cited by applicant]
US 20050105598A1 · Kaewell · 2005 [cited by applicant]
US 20050149322A1 · Bruhn · 2005 [cited by examiner]
US 20050254446A1 · Breebaart · 2005 [cited by applicant]
US 20060165184A1 · Purnhagen et al. · 2006 [cited by applicant]
US 20060190247A1 · Lindblom · 2006 [cited by applicant]
US 20070016416A1 · Roden et al. · 2007 [cited by applicant]
US 20070063877A1 · Shmunk et al. · 2007 [cited by applicant]
US 20070140499A1 · Davis · 2007 [cited by applicant]
US 20070162278A1 · Miyasaka et al. · 2007 [cited by applicant]
US 20070258607A1 · Purnhagen et al. · 2007 [cited by applicant]
US 20080195397A1 · Myburg et al. · 2008 [cited by applicant]
US 20080253576A1 · Choo et al. · 2008 [cited by applicant]
US 20090043591A1 · Breebaart et al. · 2009 [cited by applicant]
US 20090222272A1 · Seefeldt et al. · 2009 [cited by applicant]
US 20090325524A1 · Oh et al. · 2009 [cited by applicant]
US 20100246832A1 · Villemoes et al. · 2010 [cited by applicant]
US 20110096932A1 · Schuijers · 2011 [cited by applicant]
US 20110255588A1 · Shim et al. · 2011 [cited by applicant]
US 20110255714A1 · Neusinger et al. · 2011 [cited by applicant]
US 20110264456A1 · Koppens · 2011 [cited by examiner]
US 20120002818A1 · Heiko et al. · 2012 [cited by applicant]
US 20120093321A1 · Shim et al. · 2012 [cited by applicant]
US 20120224702A1 · Den Brinker et al. · 2012 [cited by applicant]
US 20130030819A1 · Purnhagen et al. · 2013 [cited by applicant]
US 20130121411A1 · Robillard et al. · 2013 [cited by applicant]
US 20130262130A1 · Ragot et al. · 2013 [cited by applicant]
US 20140016787A1 · Neuendorf · 2014 [cited by examiner]
US 20140211947A1 · Wu et al. · 2014 [cited by applicant]
US 20140222441A1 · Kuntz et al. · 2014 [cited by applicant]
US 20160104491A1 · Lee et al. · 2016 [cited by applicant]
US 20160133262A1 · Fueg et al. · 2016 [cited by applicant]
US 20160217800A1 · Purnhagen et al. · 2016 [cited by applicant]
AU 2013206557A1 · 2013 [cited by applicant]
CA 2808226A1 · 2005 [cited by applicant]
CN 101552007A · 2009 [cited by applicant]
CN 102446507A · 2012 [cited by applicant]
EP 2790419A1 · 2014 [cited by applicant]
EP 3044788A1 · 2016 [cited by applicant]
EP 3067889A1 · 2016 [cited by applicant]
JP 2008020931A · 2008 [cited by applicant]
JP 2008530616A · 2008 [cited by applicant]
JP 2011522472A · 2011 [cited by applicant]
JP 2012512438A · 2012 [cited by applicant]
JP 2013511062A · 2013 [cited by applicant]
JP 2013528824A · 2013 [cited by applicant]
JP 2013546013A · 2013 [cited by applicant]
JP 2014535182A · 2014 [cited by applicant]
JP 2016531327A · 2016 [cited by applicant]
KR 20070001162A · 2007 [cited by applicant]
KR 20160033776A · 2016 [cited by applicant]
RU 2402160C2 · 2009 [cited by applicant]
RU 2369982C2 · 2009 [cited by applicant]
RU 2425040C1 · 2011 [cited by applicant]
RU 2443075C2 · 2012 [cited by applicant]
RU 2554844C2 · 2014 [cited by applicant]
RU 2577195C2 · 2014 [cited by applicant]
RU 2560790C2 · 2015 [cited by applicant]
WO WO2010097748A1 · 2010 [cited by applicant]
WO WO2011024616A1 · 2011 [cited by applicant]
WO WO2017125558A1 · 2017 [cited by applicant]
WO WO2017125559A1 · 2017 [cited by applicant]
WO WO2017125562A1 · 2017 [cited by applicant]
WO WO2017125563A1 · 2017 [cited by applicant]
TIPO, Office Action, Oct. 8, 2018, re Taiwanese Patent Application No. 106138472. [cited by applicant]
Herre, Jürgen. “From joint stereo to spatial audio coding—recent progress and standardization.” Sixth International Conference on Digital Audio Effects (DAFX04), Naples, Italy. 2004. [cited by applicant]
ISA/EP, International Search Report and Written Opinion, Dec. 20, 2018, re PCT International Patent Application No. PCT/EP2017/077822. [cited by applicant]
Herre, J. and M. Dietz, “MPEG-4 high-efficiency AAC coding [Standards in a Nutshell],” in IEEE Signal Processing Magazine, vol. 25, No. 3, pp. 137-142, May 2008. URL: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnum… [cited by applicant]
WIPO/IB, International Preliminary Report on Patentability, Nov. 15, 2018, re PCT International Patent Application No. PCT/EP2017/077824. [cited by applicant]
WIPO/IB, International Search Report and Written Opinion, Dec. 21, 2017, re PCT International Patent Application No. PCT/EP2017/077824. [cited by applicant]
Meltzer, Stefan, and Gerald Moser. “Mpeg-4 he-aac v2-audio coding for today's digital media world.” EBU technical Review 305 (2006): 37-38. [cited by applicant]
Kurniawati, E., et al. “A stereo to mono dowmixing scheme for MPEG-4 parametric stereo encoder.” 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings. vol. 5. IEEE, 2006. [cited by applicant]
RUPTO, Official Action with English Translation, Dec. 27, 2019 re Russian Patent Application No. 2019117764/08(033957). [cited by applicant]
RUPTO, Official Action with English Translation, Dec. 26, 2019 re Russian Patent Application No. 2019117749/08(033935). [cited by applicant]
Samsudin, E. K., et al. “A Stereo to Mono Dowmixing Scheme for MPEG-4 Parametric Stereo Encoder.” 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings. vol. 5. IEEE, 2006. [cited by applicant]
CIPO, Examination Report, Apr. 7, 2020, re Canadian Patent Application No. 3045948. [cited by applicant]
CIPO, Examination Report, Mar. 9, 2020 re Canadian Patent Application No. 3042580. [cited by applicant]
Fatus, Bertrand. “Parametric Coding for Spatial Audio.” Master's Thesis, KTH, Stockholm, Sweden. 2015. [cited by applicant]
Herre, Jürgen, et al. “MPEG surround-the ISO/MPEG standard for efficient and compatible multichannel audio coding.” Journal of the Audio Engineering Society 56.11 (2008): 932-955. [cited by applicant]
Kim, JungHoe et al. “Enhanced stereo coding with phase parameters for MPEG unified speech and audio coding.” Audio Engineering Society Convention 127. Audio Engineering Society, 2009. [cited by applicant]
Zhang, Shuhua et al., “Maximal Coherence Rotation for stereo coding.” 2010 IEEE International Conference on Multimedia and Expo. IEEE, pp. 1097-1101, 2010. [cited by applicant]
ISO/IEC DIS 23008-3: 2015(E). Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio. ISO/IEC JTC 1/SC 29/WG 11. Feb. 20, 2015. [cited by applicant]
ISA/EP, International Search Report and Written Opinion, Dec. 20, 2017, re PCT International Patent Application No. PCT/EP2017/077822. [cited by applicant]
Breebaart, Jeroen, et al. “High-quality parametric spatial audio coding at low bitrates.” Audio Engineering Society Convention 116. Audio Engineering Society, 2004. [cited by applicant]
ISO/IEC FDIS 23003-3: 2011(E). Information technology—MPEG audio technologies—Part 3: Unified speech and audio coding. ISO/IEC JTC 1/SC 29/WG 11. Sep. 20, 2011. [cited by applicant]
Cited By (1)
US 12,711,967