IP Library › Granted Patent US 12,266,372
Granted Patent B2
US 12,266,372 · App. 17/550,953 · Granted Apr 1, 2025

Parameter encoding and decoding

Inventors: Alexandre Bouthéon (Erlangen, DE); Guillaume Fuchs (Erlangen, DE); Markus Multrus (Erlangen, DE); Fabian Küch (Erlangen, DE); Oliver Thiergart (Erlangen, DE); Stefan Bayer (Erlangen, DE); Sascha Disch (Erlangen, DE); Jürgen Herre (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V.
G10L19/008G10L19/08H04S3/02H04S2400/01H04S2400/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,372
App. No.
17/550,953
Granted
Apr 1, 2025
Kind
B2
Abstract

There are disclosed several examples of encoding and decoding technique. In particular, an audio synthesizer for generating a synthesis signal from a downmix signal, includes: an input interface for receiving the downmix signal, the downmix signal having a number of downmix channels and side information, the side information including channel level and correlation information of an original signal, the original signal having a number of original channels; and a synthesis processor for generating, according to at least one mixing rule, the synthesis signal using: channel level and correlation information of the original signal; and covariance information associated with the downmix signal.

Claims (54)

1. An audio synthesizer for generating a synthesis signal from a downmix signal comprising a number of downmix channels, the synthesis signal comprising a number of synthesis channels, the downmix signal being a downmixed version of an original signal comprising a number of original channels, the audio synthesizer comprising:

a first path comprising:

a first mixing matrix block configured for synthesizing a first component of the synthesis signal according to a first mixing matrix calculated from:

a covariance matrix of the synthesis signal; and

a covariance matrix of the downmix signal,

a second path for synthesizing a second component of the synthesis signal, wherein the second component is a residual component, the second path comprising:

a prototype signal block configured for upmixing the downmix signal from the number of downmix channels to the number of synthesis channels;

a decorrelator configured for decorrelating the upmixed prototype signal;

a second mixing matrix block configured for synthesizing the second component of the synthesis signal according to a second mixing matrix from the decorrelated version of the downmix signal, the second mixing matrix being a residual mixing matrix,

wherein the audio synthesizer is configured to calculate the second mixing matrix from:

the residual covariance matrix provided by the first mixing matrix block; and

an estimate of the covariance matrix of the decorrelated prototype signals acquired from the covariance matrix of the downmix signal,

wherein the audio synthesizer further comprises an adder block for summing the first component of the synthesis signal with the second component of the synthesis signal.

2. The audio synthesizer of claim 1 , wherein the residual covariance matrix is acquired by subtracting, from the covariance matrix of the synthesis signal, a matrix acquired by applying the first mixing matrix to the covariance matrix of the downmix signal.

3. The audio synthesizer of claim 1 , configured to define the second mixing matrix from:

a second matrix which is acquired by decomposing the residual covariance matrix of the synthesis signal;

a first matrix which is the inverse, or the regularized inverse, of a diagonal matrix acquired from the estimate of the covariance matrix of the decorrelated prototype signals.

4. The audio synthesizer of claim 3 , wherein the diagonal matrix is acquired by applying the square root function to the main diagonal elements of the covariance matrix of the decorrelated prototype signals.

5. The audio synthesizer of claim 3 , wherein the second matrix is acquired by singular value decomposition, SVD, applied to the residual covariance matrix of the synthesis signal.

6. The audio synthesizer of claim 3 , configured to define the second mixing matrix by multiplication of the second matrix with the inverse, or the regularized inverse, of the diagonal matrix acquired from the estimate of the covariance matrix of the decorrelated prototype signals and a third matrix.

7. The audio synthesizer of claim 6 , configured to acquire the third matrix by SVP applied to a matrix acquired from a normalized version of the covariance matrix of the decorrelated prototype signals, where the normalization is to the main diagonal the residual covariance matrix, and the diagonal matrix and the second matrix.

8. The audio synthesizer of claim 1 , configured to define the first mixing matrix from a second matrix and the inverse, or regularized inverse, of a second matrix,

wherein the second matrix is acquired by decomposing the covariance matrix of the downmix signal, and

the second matrix is acquired by decomposing the reconstructed target covariance matrix of the downmix signal.

9. The audio synthesizer of claim 1 , configured to estimate the covariance matrix of the decorrelated prototype signals from the diagonal entries of the matrix acquired from applying, to the covariance matrix of the downmix signal, the prototype rule used at the prototype block for upmixing the downmix signal from the number of downmix channels to the number of synthesis channels.

10. The audio synthesizer of claim 1 , wherein the audio synthesizer is agnostic of the decoder.

11. The audio synthesizer of claim 1 , wherein bands are aggregated with each other into groups of aggregated bands, wherein information on the groups of aggregated bands is provided in the side information of the bitstream, wherein the channel level and correlation information of the original signal is provided per each group of bands, so as to calculate the same at least one mixing matrix for different bands of the same aggregated group of bands.

12. A method for generating a synthesis signal from a downmix signal comprising a number of downmix channels, the synthesis signal comprising a number of synthesis channels, the downmix signal being a downmixed version of an original signal comprising a number of original channels, the method comprising the following phases:

a first phase comprising:

synthesizing a first component of the synthesis signal according to a first mixing matrix calculated from:

a covariance matrix of the synthesis signal; and

a covariance matrix of the downmix signal,

a second phase for synthesizing a second component of the synthesis signal, wherein the second component is a residual component, the second phase comprising:

a prototype signal step upmixing the downmix signal from the number of downmix channels to the number of synthesis channels;

a decorrelator step decorrelating the upmixed prototype signal;

a second mixing matrix step synthesizing the second component of the synthesis signal according to a second mixing matrix from the decorrelated version of the downmix signal, the second mixing matrix being a residual mixing matrix,

wherein the method calculates the second mixing matrix from:

the residual covariance matrix provided by the first mixing matrix step; and

an estimate of the covariance matrix of the decorrelated prototype signals acquired from the covariance matrix of the downmix signal,

wherein the method further comprises an adder step summing the first component of the synthesis signal with the second component of the synthesis signal, thereby acquiring the synthesis signal.

13. A non-transitory digital storage medium having a computer program stored thereon to perform the method for generating a synthesis signal from a downmix signal comprising a number of downmix channels, the synthesis signal comprising a number of synthesis channels, the downmix signal being a downmixed version of an original signal comprising a number of original channels, the method comprising the following phases:

a first phase comprising:

synthesizing a first component of the synthesis signal according to a first mixing matrix calculated from:

a covariance matrix of the synthesis signal; and

a covariance matrix of the downmix signal,

a second phase for synthesizing a second component of the synthesis signal, wherein the second component is a residual component, the second phase comprising:

a prototype signal step upmixing the downmix signal from the number of downmix channels to the number of synthesis channels;

a decorrelator step decorrelating the upmixed prototype signal;

a second mixing matrix step synthesizing the second component of the synthesis signal according to a second mixing matrix from the decorrelated version of the downmix signal, the second mixing matrix being a residual mixing matrix,

wherein the method calculates the second mixing matrix from:

the residual covariance matrix provided by the first mixing matrix step; and

an estimate of the covariance matrix of the decorrelated prototype signals acquired from the covariance matrix of the downmix signal,

wherein the method further comprises an adder step summing the first component of the synthesis signal with the second component of the synthesis signal, thereby acquiring the synthesis signal,

when said computer program is run by a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2022
From: BOUTHÉON, ALEXANDRE; FUCHS, GUILLAUME; MULTRUS, MARKUS; KÜCH, FABIAN; THIERGART, OLIVER; BAYER, STEFAN; DISCH, SASCHA; HERRE, JÜRGEN
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 059158/0563 →
Priority Claims (1)
EP 19180385 · Jun 14, 2019 · regional
Continuity (2)
Continuation PCTEP2020066456 · Jun 15, 2020
Related Publication 20220122621A1 · Apr 21, 2022
References Cited (54)
US 8126152B2 · Taleb · 2012 [cited by examiner]
US 8155971B2 · Hellmuth et al. · 2012 [cited by applicant]
US 8804971B1 · Williams · 2014 [cited by examiner]
US 9165558B2 · Dressler · 2015 [cited by examiner]
US 9734833B2 · Disch et al. · 2017 [cited by applicant]
US 10089990B2 · Disch et al. · 2018 [cited by applicant]
US 20070019813A1 · Hilpert · 2007 [cited by examiner]
US 20070160218A1 · Jakka · 2007 [cited by examiner]
US 20070203697A1 · Pang · 2007 [cited by examiner]
US 20070223708A1 · Villemoes · 2007 [cited by examiner]
US 20080071549A1 · Chong · 2008 [cited by examiner]
US 20090110203A1 · Taleb · 2009 [cited by examiner]
US 20090171676A1 · Oh · 2009 [cited by examiner]
US 20120230497A1 · Dressler · 2012 [cited by examiner]
US 20140233762A1 · Vilkamo · 2014 [cited by examiner]
US 20140321652A1 · Schuijers · 2014 [cited by examiner]
US 20150221314A1 · Disch · 2015 [cited by examiner]
US 20150279377A1 · Disch · 2015 [cited by examiner]
US 20160247507A1 · Disch · 2016 [cited by examiner]
US 20160261967A1 · Villemoes · 2016 [cited by examiner]
US 20160275958A1 · Dick · 2016 [cited by examiner]
US 20170084285A1 · Engdegard · 2017 [cited by examiner]
US 20210377685A1 · Laitinen · 2021 [cited by examiner]
CN 101411214A · 2009 [cited by applicant]
EP 3022949A1 · 2016 [cited by examiner]
EP 3022949B1 · 2017 [cited by applicant]
RU 2409912C9 · 2011 [cited by applicant]
RU 2646375C2 · 2018 [cited by applicant]
TW I395204B · 2013 [cited by applicant]
TW 201423729A · 2014 [cited by applicant]
TW 201521469A · 2015 [cited by applicant]
TW I569260B · 2017 [cited by applicant]
WO WO2007111568A2 · 2007 [cited by examiner]
WO WO2014053548A1 · 2014 [cited by examiner]
WO 2015011015A1 · 2015 [cited by applicant]
WO 2021240053A1 · 2021 [cited by applicant]
Bertrand Fatus. Parametric Coding for Spatial Audio. Master's Thesis, KTH, Stockholm, Sweden. Dec. 2015 (70 pages). [cited by applicant]
ISO/IEC DIS 23008-3. Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio. ISO/IEC JTC 1/SC 29/WG 11. Aug. 5, 2014—435 pages)—Jul. 25, 2014. [cited by applicant]
ISO/IEC FDIS 23003-1:2006(E). Information technology—MPEG audio technologies Part 1: MPEG Surround. ISO/IEC JTC 1/SC 29/WG 11. Jul. 21, 2006 (288 pages). [cited by applicant]
ISO/IEC FDIS 23003-2:2010(E). Information technology—MPEG audio technologies—Part 2: Spatial Audio Object Coding (SAOC). ISO/IEC JTC 1/SC 29/WG 11. Mar. 10, 2010 (142 pages). [cited by applicant]
Wikipedia , “Combinatorial Number System”, https:/en.wikipedia.org/wiki/Combinatorial number system, Jan. 31, 2022, 6 pp. [cited by applicant]
Faller , et al., “Binaural Cue Coding—Part II: Schemes and Applications”, IEEE Transactions on Speech and Audio Processing, vol. 11, No. 6, Nov. 2003, pp. 520-531. [cited by applicant]
Hellmuth, O. , et al., “MPEG Spatial Audio Object Coding—The ISO/MPEG Standard for Efficient Coding of Interactive Audio Scenes”, O. Hellmuth et al; “MPEG Spatial Audio Object Coding—The ISO/MPEG Standard for Efficient … [cited by applicant]
Herre, Jurgen , et al., “MPEG Surround—The ISO/MPEG Standard for Efficient and Compatible Multichannel Audio Coding”, J. Herre et al; “MPEG Surround—The ISO/MPEG Standard for Efficient and Compatible Multichannel Audio … [cited by applicant]
Huffman, David A, “A Method for the Construction of Minimum-Redundancy Codes”, D. A. Huffman; “A Method for the Construction of Minimum-Redundancy Codes”; Proceedings of the IRE; vol. 40, No. 9, Sep. 1952, 1098-1101. [cited by applicant]
ITU-R , “[Part 1 of 4] Multichannel sound technology in home and broadcasting applications”, Report ITU-R BS.2159-4; Multichannel sound technology in home and broadcasting applications, BS Series; Broadcasting service (… [cited by applicant]
ITU-R , “[Part 2 of 4] Multichannel sound technology in home and broadcasting applications”, Report ITU-R BS.2159-4; Multichannel sound technology in home and broadcasting applications; BS Series; Broadcasting service (… [cited by applicant]
ITU-R , “[Part 3 of 4] Multichannel sound technology in home and broadcasting applications”, Report ITU-R BS.2159-4; Multichannel sound technology in home and broadcasting applications; BS Series; Broadcasting service (… [cited by applicant]
ITU-R , “[Part 4 of 4] Multichannel sound technology in home and broadcasting applications”, Report ITU-R BS.2159-4; Multichannel sound technology in home and broadcasting applications; BS Series; Broadcasting service (… [cited by applicant]
Karapetyan, A. , et al., “Active Multichannel Audio Downmix”, A. Karapetyan et al; “Active Multichannel Audio Downmix,”; in 145th Audio Engineering Society; New York, 2018, 10 pp. [cited by applicant]
Mikko-Ville, L. , et al., “Converting 5.1. Audio Recordings to B-Format for Directional Audio Coding Reproduction”, L. Mikko-Ville et al; “Converting 5.1. Audio Recordings to B-Format for Directional Audio Coding Reprod… [cited by applicant]
Neuendorf, M. , et al., “The ISO/MPEG Unified Speech and Audio Coding Standard-Consistent High Quality for All Content Types and at all Bit R”, M. Neuendorf et al; “The ISO/MPEG Unified Speech and Audio Coding Standard-… [cited by applicant]
Pulkki, V , “Spatial Sound Reproduction with Directional Audio Coding”, Journal of the AES. vol. 55, No. 6. New York, NY, USA., Jun. 2007, pp. 503-516. [cited by applicant]
Vilkamo, Juha , et al., “Optimized covariance domain framework for time-frequency processing of spatial audio”, J. Vilkamo et al; “Optimized Covariance Domain Framework for Time-Frequency Processing of Spatial Audio”; J… [cited by applicant]