IP Library › Granted Patent US 12,418,768
Granted Patent B2
US 12,418,768 · App. 18/363,978 · Granted Sep 16, 2025

Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to DirAC based spatial audio coding using diffuse compensation

Inventors: Guillaume Fuchs (Erlangen, DE); Oliver Thiergart (Erlangen, DE); Srikanth Korse (Erlangen, DE); Stefan Döhla (Erlangen, DE); Markus Multrus (Erlangen, DE); Fabian Küch (Erlangen, DE); Alexandre Bouthéon (Erlangen, DE); Andrea Eichenseer (Erlangen, DE); Stefan Bayer (Erlangen, DE)
Assignee: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
H04S7/307G10L19/02G10L19/0212H04N19/119H04N19/176H04N19/593H04S2400/01H04S2420/11H04S2420/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,418,768
App. No.
18/363,978
Granted
Sep 16, 2025
Kind
B2
Abstract

An apparatus generating a sound field description includes an input signal analyzer acquiring diffuseness data from the input signal; a sound component generator for generating, from the input signal, one or more sound field components of a first group of sound field components comprising for each sound field component a direct component and a diffuse component, and for generating, from the input signal, a second group of sound field components comprising only a direct component. The sound component generator is configured to perform an energy compensation when generating the first group of sound field components, the energy compensation depending on the diffuseness data and at least one of a number of sound field components in the second group, a number of diffuse components in the first group, a maximum order of sound field components of the first group and a maximum order of sound field components of the second group.

Claims (46)

1. An apparatus for generating a sound field description from an input signal comprising one or more channels, the apparatus comprising:

an input signal analyzer configured for acquiring diffuseness data from the input signal;

a sound component generator configured for generating, from the input signal, one or more sound field components of a first group of sound field components, the first group of sound field components comprising, for each sound field component of the first group of sound field components, a direct component and a diffuse component, and for generating, from the input signal, a second group of sound field components, a sound field component of the second group of sound field components comprising only a direct component,

wherein the sound component generator comprises an energy compensator configured for performing an energy compensation of the sound field components of the first group when generating the first group of sound field components, the energy compensator comprising a compensation gain calculator for

calculating a compensation gain using the diffuseness data, the maximum order of the sound field components of the first group and the maximum order of the sound field components of the second group, wherein the compensation gain calculator is configured to decrease the compensation gain with an increasing maximum order of the sound field components of the first group and to increase the compensation gain with an increasing maximum order of the sound field components of the second group, or

calculating a compensation gain using the diffuseness data, the number of the diffuse components in the first group, and the maximum order of the sound field components of the second group, wherein the compensation gain calculator is configured to decrease the compensation gain with an increasing number of the diffuse components in the first group and to decrease the compensation gain with an increasing maximum order of the sound field components of the first group.

2. The apparatus of claim 1 , wherein the sound component generator comprises a mid-order components generator comprising:

a reference signal provider configured for providing a reference signal for a sound field component of the first group of sound field components;

a decorrelator configured for generating a decorrelated signal from the reference signal,

wherein the direct component of the sound field component of the first group is derived from the reference signal, wherein the diffuse component of the sound field component of the first group is derived from the decorrelated signal, and

a mixer configured for mixing the direct component and the diffuse component using at least one of a direction of arrival data provided by the input signal analyzer and the diffuseness data.

3. The apparatus of claim 1 ,

wherein the input signal comprises only a single mono channel signal, and wherein the sound field components of the first group of sound field components are sound field components of a first order or a higher order, or wherein the input signal comprises two or more channels, and wherein the sound field components of the first group of sound field components are sound field components of a second order or a higher order.

4. The apparatus of claim 1 , wherein the input signal comprises a mono signal or at least two channels, and wherein the sound component generator comprises a low order components generator configured for generating the low-order sound field components by copying, or taking the input signal, or performing a weighted combination of the channels of the input signal.

5. The apparatus of claim 4 , wherein the input signal comprises the mono signal, and wherein the low order components generator is configured to generate a zero order Ambisonics signal by taking or copying the mono signal, or

wherein the input signal comprises at least two channels, and wherein the low-order components generator is configured to generate a zero order Ambisonics signal by adding the two channels and to generate a first order Ambisonics signal based on a difference of the two channels, or

wherein the input signal comprises a first order Ambisonics signal with three or four channels, and wherein the low order components generator is configured to generate a first order Ambisonics signal by taking or copying the three or four channels of the input signal, or

wherein the input signal comprises an A-format signal comprising four channels, and wherein the low-order components generator is configured to calculate a first order Ambisonics signal by performing a weighted linear combination of the four channels.

6. The apparatus of claim 1 , wherein the sound component generator comprises a high-order components generator configured for generating the sound field components of the second group, the sound field components of the second group comprising an order being higher than a truncation order used for generating the sound field components of the first group of sound field components.

7. The apparatus of claim 1 , wherein the compensation gain calculator is configured

to increase the compensation gain with an increasing number of the sound field components in the second group, or

to increase the compensation gain with an increasing diffuseness data.

8. The apparatus of claim 1 , wherein the compensation gain calculator is configured for calculating the compensation gain additionally using a first energy- or amplitude-related measure for an omnidirectional component derived from the input signal and using a second energy- or amplitude-related measure for a directional component derived from the input signal, the diffuseness data, and direction data acquired from the input signal.

9. The apparatus of claim 1 , wherein the compensation gain calculator is configured to calculate a first gain factor depending on the diffuseness data and at least one of the number of the sound field components in the second group, the number of the diffuse components in the first group, the maximum order of the sound field components of the first group, and the maximum order of the sound field components of the second group, to calculate a second gain factor depending on a first amplitude or energy-related measure for an omnidirectional component derived from the input signal, a second energy- or amplitude-related measure for a directional component derived from the input signal, the direction data and the diffuseness data, and to calculate the compensation gain using the first gain factor and the second gain factor.

10. The apparatus of claim 1 , wherein the compensation gain calculator is configured to perform a gain manipulation using a limitation with a fixed maximum threshold or a fixed minimum threshold or using a compression function for compressing low or high gain values towards medium gain values to acquire the compensation gain.

11. The apparatus of claim 1 , wherein the sound component generator comprises a compensation gain applicator configured for applying the compensation gain to at least one sound field component of the first group.

12. The apparatus of claim 11 , wherein the compensation gain applicator is configured to apply the compensation gain to each sound field component of the first group, or to only one or more sound field components of the first group with a diffuse portion, or to diffuse portions of the sound field components of the first group.

13. The apparatus of claim 1 , wherein the input signal analyzer is configured to extract the diffuseness data from metadata associated with the input signal or to extract the diffuseness data from the input signal by a signal analysis of the input signal comprising two or more channels or components.

14. The apparatus of claim 1 , wherein the input signal only comprises one or two sound field components, the one or two sound field components being up to an input order, wherein the sound component generator comprises a sound field components combiner configured for combining the sound field components of the first group and the sound field components of the second group to acquire a sound field description comprising sound field components being up to an output order, the output order being higher than the input order.

15. The apparatus of claim 1 , further comprising:

an analysis filter bank configured for generating the one or more sound field components of the first group and the second group for a plurality of different time-frequency tiles, wherein the input signal analyzer is configured to acquire a diffuseness data item for each time-frequency tile, and wherein the sound component generator is configured to perform the energy compensation separately for each time-frequency tile.

16. The apparatus of claim 1 , further comprising:

a high-order decoder configured for using the one or more sound field components of the first group and the one or more sound field components of the second group to generate a spectral domain or time domain representation of the sound field description generated from the input signal.

17. The apparatus of claim 1 , wherein the first group of sound field components and the second group of sound field components are orthogonal to each other, or wherein the sound field components are at least one of coefficients of orthogonal basis functions, coefficients of spatial basis functions, coefficients of spherical or circular harmonics, and Ambisonics coefficients.

18. A method for generating a sound field description from an input signal comprising one or more channels, comprising:

acquiring diffuseness data from the input signal;

generating, from the input signal, one or more sound field components of a first group of sound field components, the first group of sound field components comprising, for each sound field component of the first group of sound field components, a direct component and a diffuse component, and generating, from the input signal, a second group of sound field components, a sound field component of the second group of sound field components comprising only a direct component,

wherein the generating comprises performing an energy compensation of the sound field components of the first group when generating the first group of sound field components, the performing the energy compensation comprising

calculating a compensation gain using the diffuseness data, the maximum order of the sound field components of the first group and the maximum order of the sound field components of the second group, wherein the performing the energy compensation is configured to decrease the compensation gain with an increasing maximum order of the sound field components of the first group and to increase the compensation gain with an increasing maximum order of the sound field components of the second group, or

calculating a compensation gain using the diffuseness data, the number of the diffuse components in the first group, and the maximum order of the sound field components of the second group, wherein the performing the energy compensation is configured to decrease the compensation gain with an increasing number of the diffuse components in the first group and to decrease the compensation gain with an increasing maximum order of the sound field components of the first group.

19. A non-transitory digital storage medium having a computer program stored thereon to perform, when said computer program is run by a computer, the method for generating a sound field description from an input signal comprising one or more channels, comprising:

acquiring diffuseness data from the input signal;

generating, from the input signal, one or more sound field components of a first group of sound field components, the first group of sound field components comprising, for each sound field component of the first group of sound field components, a direct component and a diffuse component, and generating, from the input signal, a second group of sound field components, a sound field component of the second group of sound field components comprising only a direct component,

wherein the generating comprises performing an energy compensation of the sound field components of the first group, when generating the first group of sound field components, the performing the energy compensation comprising

calculating a compensation gain using the diffuseness data, the maximum order of the sound field components of the first group and the maximum order of the sound field components of the second group, wherein the performing the energy compensation is configured to decrease the compensation gain with an increasing maximum order of the sound field components of the first group and to increase the compensation gain with an increasing maximum order of the sound field components of the second group, or

calculating a compensation gain using the diffuseness data, the number of the diffuse components in the first group, and the maximum order of the sound field components of the second group, wherein the performing the energy compensation is configured to decrease the compensation gain with an increasing number of the diffuse components in the first group and to decrease the compensation gain with an increasing maximum order of the sound field components of the first group.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2023
From: FUCHS, GUILLAUME; THIERGART, OLIVER; KORSE, SRIKANTH; DÖHLA, STEFAN; MULTRUS, MARKUS; KÜCH, FABIAN; BOUTHÉON, ALEXANDRE; EICHENSEER, ANDREA; BAYER, STEFAN
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 064466/0722 →
Priority Claims (1)
EP 18211064 · Dec 7, 2018 · regional
Continuity (3)
Division 17332312 · May 27, 2021
Continuation PCTEP2019084053 · Dec 6, 2019
Related Publication 20230379652A1 · Nov 23, 2023
References Cited (127)
US 7031474B1 · Yuen et al. · 2006 [cited by applicant]
US 7974713B2 · Disch et al. · 2011 [cited by applicant]
US 8346565B2 · Uhle et al. · 2013 [cited by applicant]
US 8452019B1 · Fomin et al. · 2013 [cited by applicant]
US 8811621B2 · Schuijers · 2014 [cited by applicant]
US 8891797B2 · Thiergart et al. · 2014 [cited by applicant]
US 9691406B2 · Jax et al. · 2017 [cited by applicant]
US 9743215B2 · Uhle et al. · 2017 [cited by applicant]
US 9838821B2 · Kelloniemi · 2017 [cited by examiner]
US 9930464B2 · Kordon et al. · 2018 [cited by applicant]
US 9973874B2 · Stein et al. · 2018 [cited by applicant]
US 10136239B1 · Stefanakis et al. · 2018 [cited by applicant]
US 11272305B2 · Habets · 2022 [cited by examiner]
US 20080298597A1 · Turku et al. · 2008 [cited by applicant]
US 20090161880A1 · Hooley et al. · 2009 [cited by applicant]
US 20110106545A1 · Disch et al. · 2011 [cited by applicant]
US 20120114126A1 · Thiergart et al. · 2012 [cited by applicant]
US 20120155653A1 · Jax et al. · 2012 [cited by applicant]
US 20120177204A1 · Hellmuth et al. · 2012 [cited by applicant]
US 20130259243A1 · Herre et al. · 2013 [cited by applicant]
US 20140016802A1 · Sen · 2014 [cited by applicant]
US 20150049583A1 · Neal · 2015 [cited by examiner]
US 20150154965A1 · Wuebbolt et al. · 2015 [cited by applicant]
US 20150189436A1 · Kelloniemi · 2015 [cited by applicant]
US 20150221313A1 · Purnhagen et al. · 2015 [cited by applicant]
US 20150248889A1 · Dickins et al. · 2015 [cited by applicant]
US 20150271621A1 · Sen et al. · 2015 [cited by applicant]
US 20160057556A1 · Boehm · 2016 [cited by applicant]
US 20160064005A1 · Peters · 2016 [cited by applicant]
US 20160227337A1 · Goodwin · 2016 [cited by applicant]
US 20170032799A1 · Peters et al. · 2017 [cited by applicant]
US 20170078818A1 · Habets et al. · 2017 [cited by applicant]
US 20170078819A1 · Habets et al. · 2017 [cited by applicant]
US 20170154633A1 · Krueger · 2017 [cited by examiner]
US 20170180902A1 · Kordon et al. · 2017 [cited by applicant]
US 20170366914A1 · Stein · 2017 [cited by examiner]
US 20180156598A1 · Cable et al. · 2018 [cited by applicant]
US 20180192226A1 · Woelfl · 2018 [cited by examiner]
US 20180218740A1 · Kleijn · 2018 [cited by examiner]
US 20180234785A1 · Kordon et al. · 2018 [cited by applicant]
US 20180315432A1 · Boehm et al. · 2018 [cited by applicant]
US 20180333103A1 · Bardan et al. · 2018 [cited by applicant]
US 20190335291A1 · Baque et al. · 2019 [cited by applicant]
US 20200058311A1 · Goodwin · 2020 [cited by applicant]
US 20200154229A1 · Habets · 2020 [cited by examiner]
US 20200211521A1 · Voss · 2020 [cited by applicant]
US 20200221230A1 · Fuchs · 2020 [cited by examiner]
US 20210051435A1 · Yen et al. · 2021 [cited by applicant]
US 20210136489A1 · Janse · 2021 [cited by applicant]
US 20210295855A1 · Vasilache et al. · 2021 [cited by applicant]
US 20210319799A1 · Pihlajakuja et al. · 2021 [cited by applicant]
CN 1254153C · 2006 [cited by applicant]
CN 102422348A · 2012 [cited by applicant]
CN 102547549A · 2012 [cited by applicant]
CN 103583054A · 2014 [cited by applicant]
CN 104429102A · 2015 [cited by applicant]
CN 104471641A · 2015 [cited by applicant]
CN 105340008A · 2016 [cited by applicant]
CN 105917407A · 2016 [cited by applicant]
CN 106664485A · 2017 [cited by applicant]
CN 106664501A · 2017 [cited by applicant]
DE 102008004674A1 · 2009 [cited by applicant]
EP 1527655B1 · 2006 [cited by applicant]
EP 2510709A1 · 2012 [cited by applicant]
EP 2942982A1 · 2015 [cited by applicant]
KR 1020150134336A · 2015 [cited by applicant]
KR 1020170007749A · 2017 [cited by applicant]
RU 2388068C2 · 2010 [cited by applicant]
RU 2497204C2 · 2013 [cited by applicant]
RU 2558612C2 · 2015 [cited by applicant]
RU 2637990C1 · 2017 [cited by applicant]
RU 2663345C2 · 2018 [cited by applicant]
TW 1332192A · 2007 [cited by applicant]
TW 1352971A · 2008 [cited by applicant]
TW 1313857B · 2009 [cited by applicant]
TW 201810249A · 2018 [cited by applicant]
TW M564300U · 2018 [cited by applicant]
WO 2009077152A1 · 2009 [cited by applicant]
WO 2011069205A1 · 2011 [cited by applicant]
WO 2013141768A1 · 2013 [cited by applicant]
WO 2014013070A1 · 2014 [cited by applicant]
WO 2014194116A1 · 2014 [cited by applicant]
WO 2015116666A1 · 2015 [cited by applicant]
WO 2015175933A1 · 2015 [cited by applicant]
WO 2015175998A1 · 2015 [cited by applicant]
WO 2017157803A1 · 2017 [cited by applicant]
Chinese language Notice of Allowance dated Nov. 10, 2023, issued in application No. CN 201980091649.X. [cited by applicant]
English language translation of Notice of Allowance dated Nov. 10, 2023 (pp. 1-3 of attachment). [cited by applicant]
Rui, J.; “Conversion between Stereo and Multichannel Surround;” Xidian University; Dec. 2014; pp. 1-67. [cited by applicant]
English language translation of “Conversion between Stereo and Multichannel Surround” (p. 6 of publication). [cited by applicant]
Notice of Allowance dated Sep. 6, 2023, issued in U.S. Appl. No. 17/332,312. [cited by applicant]
Notice of Allowance dated Sep. 7, 2023, issued in U.S. Appl. No. 17/332,340. [cited by applicant]
International Search Report and Written Opinion dated Jan. 14, 2020, issued in application No. PCT/EP2019/084053. [cited by applicant]
Pulkki, V., et al.; “Directional audio coding—perception-based reproduction of spatial sound;” International Workshop on the Principles and Application on Spatial Hearing; Nov. 2009; pp. 1-4. [cited by applicant]
Pulkki, V.; “Spatial Sound Reproduction with Directional Audio Coding;” Journal of the Audio Engineering Society, Audio Engineering Society; vol. 55; No. 6; Jun. 2007; pp. 503-516. [cited by applicant]
Laitinen, M.V., et al.; “Converting 5.1 audio recordings to B-format for direction-al audio coding reproduction;” 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2011; pp. 61-64. [cited by applicant]
Furness, R.K.; “Ambisonics—An overview;” AES 8th International Conference; Apr. 1990; pp. 181-189. [cited by applicant]
Nachbar, C., et al.; “AMBIX—A Suggested Ambisonics Format;” Proceedings of the Ambisonics Symposium; 2011; pp. 1-12. [cited by applicant]
Russian language office action dated Oct. 20, 2021, issued in application No. RU 2021118694. [cited by applicant]
English language translation of office action dated Oct. 20, 2021, issued in application No. RU 2021118694. [cited by applicant]
Russian language office action dated Feb. 25, 2022, issued in application No. RU 2021118698. [cited by applicant]
English language translation of office action dated Feb. 25, 2022, issued in application No. RU 2021118698 (pp. 1-3 of attachment). [cited by applicant]
Russian language search report dated Feb. 25, 2022, issued in application No. RU 2021118698. [cited by applicant]
English language translation of search report dated Feb. 25, 2022, issued in application No. RU 2021118698 (pp. 1-2 of attachment). [cited by applicant]
Russian language office action dated Mar. 25, 2022, issued in application No. RU 2021118691. [cited by applicant]
English language translation of Russian language office action dated Mar. 25, 2022, issued in application No. RU 2021118691. [cited by applicant]
Chinese language office action dated Apr. 21, 2023, issued in application No. CN 201980091649.X. [cited by applicant]
Zhou, L.S., et al.; “2.5D Near-Field Sound Source Synthesis Using Higher Order Ambisonics;” China Academic Journal; vol. 45; No. 3; Mar. 2017; pp. 520-526. [cited by applicant]
English language abstract of “2.5D Near-Field Sound Source Synthesis Using Higher Order Ambisonics;” (p. 520 of publication). [cited by applicant]
Kallinger, M, et al.; “Enhanced Direction estimation using microphone arrays for directional audio coding;” IEEE; 2008; pp. 45-48. [cited by applicant]
Chinese language office action dated Apr. 21, 2023, issued in application No. CN 201980091648.5. [cited by applicant]
Zhu, T., et al.; “Evaluation of Ambisonics Reproduction System with Object Metric Based on Interaural Time Difference;” Journal of Nanjing University; vol. 53; No. 6; Nov. 2017; pp. 1153-1160. [cited by applicant]
Chinese language office action dated Apr. 27, 2023, issued in application No. CN 201980091619.9. [cited by applicant]
English language translation of office action dated Apr. 27, 2023 (pp. 8-17 of attachment). [cited by applicant]
Wang, X., et al.; “Reverberation suppression method based on diffuse information in a room sound field;” J. Tsinghua University; Jun. 2013; pp. 1-4. [cited by applicant]
English language translation of abstract of “Reverberation suppression method based on diffuse information in a room sound field” (p. 1 of attachment). [cited by applicant]
Wang, Y., et al.; “Low Frequency Sound Field Reproduction within a Cylindrical Cavity Using Higher Order Ambisonics;” Journal of Northwestern Polytechnical University; Aug. 2018; pp. 1-7. [cited by applicant]
English language translation of abstract of “Low Frequency Sound Field Reproduction within a Cylindrical Cavity Using Higher Order Ambisonics” (p. 7 of attachment). [cited by applicant]
Non-Final Office Action dated May 22, 2023, issued in application No. U.S. Appl. No. 17/332,340. [cited by applicant]
Notice of Allowance dated Nov. 1, 2023, issued in application No. U.S. Appl. No. 17/332,358. [cited by applicant]
Korean language Notice of Allowance dated Jul. 27, 2023, issued in application No. KR 10-2021-7020826. [cited by applicant]
Non-Final Office Action dated Apr. 23, 2024, issued in application No. U.S. Appl. No. 18/447,486. [cited by applicant]
Non-Final Office Action dated Apr. 15, 2024, issued in application No. U.S. Appl. No. 18/482,478. [cited by applicant]
Office Action dated Dec. 26, 2024, issued in application No. MY PI2021003012. [cited by applicant]
Final Office Action dated Sep. 16, 2024, issued in U.S. Appl. No. 18/447,486. [cited by applicant]
Notice of Allowance dated Mar. 18, 2025, issued in U.S. Appl. No. 18/482,478. [cited by applicant]
Korean language Notice of Allowance dated Mar. 6, 2025, issued in application No. KR 10-2023-7024795. [cited by applicant]