IP Library › Granted Patent US 12,744,046
Granted Patent B2
US 12,744,046 · App. 18/823,006 · Granted Sep 22, 2026

Apparatus and method for encoding or decoding directional audio coding parameters using different time/frequency resolutions

Inventors: Guillaume Fuchs (Bubenreuth, DE); Jürgen Herre (Erlangen, DE); Fabian Küch (Erlangen, DE); Stefan Döhla (Erlangen, DE); Markus Multrus (Nuremberg, DE); Oliver Thiergart (Erlangen, DE); Oliver Wübbolt (Hannover, DE); Florin Ghido (Nuremberg, DE); Stefan Bayer (Nuremberg, DE); Wolfgang Jaegers (Nuremberg, DE)
Assignee: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
G10L19/008G10L19/0204G10L19/032G10L19/038G10L19/167G10L19/26H03M7/3082H03M7/6005H03M7/6011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,046
App. No.
18/823,006
Granted
Sep 22, 2026
Kind
B2
Abstract

An apparatus for encoding directional audio coding parameters including diffuseness parameters and direction parameters includes: a parameter calculator for calculating the diffuseness parameters with a first time or frequency resolution and for calculating the direction parameters with a second time or frequency resolution; and a quantizer and encoder processor for generating a quantized and encoded representation of the diffuseness parameters and the direction parameters.

Claims (47)

1 . An apparatus for encoding, comprising:

an interface for receiving directional audio coding parameters comprising diffuseness parameters and direction parameters; and

a parameter calculator for calculating the diffuseness parameters with a first time or frequency resolution and for calculating the direction parameters with a second time or frequency resolution.

2 . The apparatus of claim 1 , wherein the parameter calculator is configured to calculate the diffuseness parameters and the direction parameters so that the first time resolution is lower than the second time resolution, or the second frequency resolution Is greater than the first frequency resolution, or the first time resolution is lower than the second time resolution and the first frequency resolution is equal to the second frequency resolution, or

wherein the parameter calculator is configured to calculate the diffuseness parameters and the direction parameters for a set of frequency bands, wherein a band having a lower center frequency is narrower than a band having a higher center frequency, or

wherein the parameter calculator is configured to obtain initial diffuseness parameters having a third time or frequency resolution and to obtain initial direction parameters having a fourth time or frequency resolution, and wherein the parameter calculator is configured to group and average the initial diffuseness parameters so that the third time or frequency resolution is higher than the first time or frequency resolution, or wherein the parameter calculator is configured to group and average the initial direction parameters so that the fourth time or frequency resolution is higher than the second time or frequency resolution.

3 . The apparatus of claim 1 ,

wherein the parameter calculator is configured to calculate the initial direction parameters so that the initial direction parameters each comprise a Cartesian vector having a component for each of two or three directions, and wherein the parameter calculator is configured to perform the averaging for each individual component of the Cartesian vector separately, or wherein the components are normalized so that the sum of squared components of the Cartesian vector for a direction parameter is equal to unity.

4 . The apparatus of claim 3 , further comprising:

a time-frequency decomposer for decomposing an input signal having a plurality of input channels into a time-frequency representation for each input channel, or

wherein the time-frequency decomposer is configured for decomposing the input signal having a plurality of input channels into a time-frequency representation for each input channel having the third time or frequency resolution or the fourth time or frequency resolution.

5 . The apparatus of claim 1 ,

wherein the apparatus is configured to associate an indication of the first or the second time or frequency resolution into the quantized and encoded representation for transmission to a decoder or for storage, or

comprising a quantizer and encoder processor for generating a quantized and encoded representation of the diffuseness parameters and the direction parameters, wherein the quantizer and encoder processor comprises a parameter quantizer for quantizing the diffuseness parameters and the direction parameters and a parameter encoder for encoding quantized diffuseness parameters and quantized direction parameters.

6 . A method for encoding, the method comprising:

receiving directional audio coding parameters comprising diffuseness parameters and direction parameters; and

calculating the diffuseness parameters with a first time or frequency resolution and calculating the direction parameters with a second time or frequency resolution.

7 . A decoder for decoding, comprising:

a parameter processor for decoding encoded directional audio coding parameters comprising encoded diffuseness parameters and encoded direction parameters to obtain decoded diffuseness parameters with a first time or frequency resolution and decoded direction parameters with a second time or frequency resolution, the second time or frequency resolution being different from the first time or frequency resolution; and

a parameter resolution converter for converting the encoded or decoded diffuseness parameters or the encoded or decoded direction parameters into converted diffuseness parameters or converted direction parameters having a third time or frequency resolution.

8 . The decoder of claim 7 , further comprising an audio renderer operating in a spectral domain, the spectral domain comprising, for a frame, a first number of time slots and a second number of frequency bands, so that a frame comprises a number of time/frequency bins being equal to a multiplication result of the first number and the second number, wherein the first number and the second number define the third time or frequency resolution, or

further comprising an audio renderer operating in a spectral domain, the spectral domain comprising, for a frame, a first number of time slots and a second number of frequency bands, so that a frame comprises a number of time/frequency bins being equal to a multiplication result of the first number and the second number, wherein the first number and the second number define a fourth time-frequency resolution, wherein the fourth time or frequency resolution is higher than the third time or frequency resolution, or

wherein the first time or frequency resolution is lower than the second time or frequency resolution, and wherein the parameter resolution converter is configured to generate, from a decoded diffuseness parameter, a first multitude of converted diffuseness parameters and to generate, from a decoded direction parameter, a second multitude of converted direction parameters, wherein the second multitude is greater than the first multitude, or

wherein the encoded audio signal comprises a sequence of frames, wherein each frame is organized in frequency bands, wherein each frame comprises only one encoded diffuseness parameter per frequency band and at least two time-sequential direction parameters per frequency band, and wherein the parameter resolution converter is configured to associate the decoded diffuseness parameter to all time bins in the frequency band or to each time/frequency bin included in the frequency band in the frame, and to associate one direction parameter of the at least two time-sequential direction parameters of the frequency band to a first group of time bins included in the frequency band, and to associate a second direction parameter of the at least two direction parameters to a second group of the time bins included in the frequency band, wherein the second group of the time bins does not include any of the time bins in the first group of the time bins, or

wherein the encoded audio signal comprises an encoded audio transport signal, wherein the decoder comprises: an audio decoder for decoding the encoded transport audio signal to obtain a decoded audio signal, and a time/frequency converter for converting the decoded audio signal into a frequency representation having the third time or frequency resolution.

9 . The decoder of claim 8 , comprising:

an audio renderer for applying the converted diffuseness parameters and the converted direction parameters to the frequency representation of the decoded audio signal in the third time or frequency resolution to obtain a synthesis spectrum representation; and

a spectrum/time converter for converting the synthesis spectrum representation in the third or fourth time or frequency resolution to obtain a synthesized time domain spatial audio signal having a time resolution being higher than the resolution of the third time or frequency resolution.

10 . The decoder of claim 7 ,

wherein the parameter resolution converter is configured to copy a decoded direction parameter or to copy a decoded diffuseness parameter or to smooth or low pass filter a set of copied direction parameters or a set of copied diffuseness parameters.

11 . The decoder of claim 7 , wherein the second time or frequency resolution is different from the first time or frequency resolution, or

wherein the first time resolution is lower than the second time resolution, or the second frequency resolution is greater than the first frequency resolution, or the first time resolution is lower than the second time resolution and the first frequency resolution is equal to the second frequency resolution, or

wherein the parameter resolution converter is configured to copy the decoded diffuseness parameters and decoded direction parameters into a corresponding number of frequency adjacent converted parameters for a set of bands, wherein a band having a lower center frequency receives less copied parameters than a band having a higher center frequency, or

wherein the parameter processor is configured to decode an encoded diffuseness parameter for a frame of the encoded audio signal to obtain a quantized diffuseness parameter for the frame, and wherein the parameter processor is configured to determine a dequantization precision for the dequantization of at least one direction parameter for the frame using the quantized or dequantized diffuseness parameter, and wherein the parameter processor is configured to dequantize a quantized direction parameter using the dequantization precision.

12 . The decoder of claim 7 , wherein the parameter processor is configured to determine, from a dequantization precision, to be used by the parameter processor for dequantizing, a decoding alphabet for decoding an encoded direction parameter for a frame, and

wherein the parameter processor is configured to decode the encoded direction parameter using the determined decoding alphabet and to determine a dequantized direction parameter.

13 . The decoder of claim 7 , wherein the parameter processor is configured to determine, from a dequantization precision to be used by the parameter processor for dequantizing the direction parameter, an elevation alphabet for the processing of an encoded elevation parameter and to determine, from an elevation index obtained using the elevation alphabet, an azimuth alphabet, and

wherein the parameter processor is configured to dequantize an encoded azimuth parameter using the azimuth alphabet.

14 . A method of decoding, the method comprising:

decoding encoded directional audio coding parameters comprising encoded diffuseness parameters and encoded direction parameters to obtain decoded diffuseness parameters with a first time or frequency resolution and decoded direction parameters with a second time or frequency resolution, the second time or frequency resolution being different from the first time or frequency resolution; and

converting the encoded or decoded diffuseness parameters or the encoded or decoded direction parameters into converted diffuseness parameters or converted direction parameters having a third time or frequency resolution.

15 . A non-transitory storage medium having stored there on a computer program for performing, when running on a computer or a processor, a method of encoding, the method comprising:

receiving directional audio coding parameters comprising diffuseness parameters and direction parameters; and

calculating the diffuseness parameters with a first time or frequency resolution and calculating the direction parameters with a second time or frequency resolution.

16 . A non-transitory storage medium having stored there on a computer program for performing, when running on a computer or a processor, a method of decoding, the method comprising:

decoding encoded directional audio coding parameters comprising encoded diffuseness parameters and encoded direction parameters to obtain decoded diffuseness parameters with a first time or frequency resolution and decoded direction parameters with a second time or frequency resolution, the second time or frequency resolution being different from the first time or frequency resolution; and

converting the encoded or decoded diffuseness parameters or the encoded or decoded direction parameters into converted diffuseness parameters or converted direction parameters having a third time or frequency resolution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2024
From: FUCHS, GUILLAUME; HERRE, JÜRGEN; KÜCH, FABIAN; DÖHLA, STEFAN; MULTRUS, MARKUS; THIERGART, OLIVER; WÜBBOLT, OLIVER; GHIDO, FLORIN; BAYER, STEFAN; JAEGERS, WOLFGANG
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 068471/0397 →
Priority Claims (1)
EP 17202393 · Nov 17, 2017 · regional
Continuity (4)
Continuation 18456670 · Aug 28, 2023
Continuation 16871223 · May 11, 2020
Continuation PCTEP2018081620 · Nov 16, 2018
Related Publication 20240428806A1 · Dec 26, 2024
References Cited (123)
US 6446037B1 · Fielder · 2002 [cited by applicant]
US 6678647B1 · Edler · 2004 [cited by applicant]
US 8712059B2 · Del Galdo et al. · 2014 [cited by applicant]
US 8804970B2 · Grill et al. · 2014 [cited by applicant]
US 8891797B2 · Thiergart et al. · 2014 [cited by applicant]
US 8897455B2 · Visser et al. · 2014 [cited by applicant]
US 9154896B2 · Mahabub et al. · 2015 [cited by applicant]
US 9196257B2 · Schultz-Amling et al. · 2015 [cited by applicant]
US 9466305B2 · Sen et al. · 2016 [cited by applicant]
US 9530421B2 · Jot et al. · 2016 [cited by applicant]
US 9805729B2 · Shi et al. · 2017 [cited by applicant]
US 10483913B2 · Najafi et al. · 2019 [cited by applicant]
US 12112762B2 · Fuchs · 2024 [cited by applicant]
US 20050108002A1 · Nagai et al. · 2005 [cited by applicant]
US 20070043575A1 · Onuma et al. · 2007 [cited by applicant]
US 20080232601A1 · Pulkki · 2008 [cited by applicant]
US 20080232616A1 · Pulkki · 2008 [cited by applicant]
US 20090164226A1 · Boehm et al. · 2009 [cited by applicant]
US 20090296808A1 · Regunathan et al. · 2009 [cited by applicant]
US 20100166191A1 · Herre · 2010 [cited by applicant]
US 20100286990A1 · Biswas et al. · 2010 [cited by applicant]
US 20100286991A1 · Hedelin et al. · 2010 [cited by applicant]
US 20110106545A1 · Disch et al. · 2011 [cited by applicant]
US 20110112843A1 · Shimada et al. · 2011 [cited by applicant]
US 20120101610A1 · Ojala et al. · 2012 [cited by applicant]
US 20120114126A1 · Thiergart et al. · 2012 [cited by applicant]
US 20130003998A1 · Kirkeby et al. · 2013 [cited by applicant]
US 20130022206A1 · Thiergart et al. · 2013 [cited by applicant]
US 20140355766A1 · Morrell et al. · 2014 [cited by applicant]
US 20140358565A1 · Peters et al. · 2014 [cited by applicant]
US 20140358567A1 · Koppens et al. · 2014 [cited by applicant]
US 20150127354A1 · Peters et al. · 2015 [cited by applicant]
US 20150332682A1 · Kim et al. · 2015 [cited by applicant]
US 20200273473A1 · Fuchs et al. · 2020 [cited by applicant]
US 20220036906A1 · Vasilache · 2022 [cited by applicant]
CN 101925950A · 2010 [cited by applicant]
CN 102422348A · 2012 [cited by applicant]
CN 102760437A · 2012 [cited by applicant]
CN 103428497A · 2013 [cited by applicant]
CN 103733622A · 2014 [cited by applicant]
CN 104464742A · 2015 [cited by applicant]
CN 104520925A · 2015 [cited by applicant]
CN 106023999A · 2016 [cited by applicant]
CN 111656442A · 2020 [cited by applicant]
CN 113228168A · 2021 [cited by applicant]
EP 2154910A1 · 2010 [cited by applicant]
EP 2249334A1 · 2010 [cited by applicant]
EP 2346028A1 · 2011 [cited by applicant]
EP 2437257A1 · 2012 [cited by applicant]
GB 2554446A · 2018 [cited by applicant]
JP 2005148274A · 2005 [cited by applicant]
JP 2006003580A · 2006 [cited by applicant]
JP 2007034230A · 2007 [cited by applicant]
JP 2007532934A · 2007 [cited by applicant]
JP 2008517339A · 2008 [cited by applicant]
JP 2011530720A · 2011 [cited by applicant]
JP 2012526296A · 2012 [cited by applicant]
JP 2013514696A · 2013 [cited by applicant]
JP 2013524267A · 2013 [cited by applicant]
JP 2014506416A · 2014 [cited by applicant]
JP 2020526987A · 2020 [cited by applicant]
JP 7175979B2 · 2022 [cited by applicant]
JP 7372360B2 · 2023 [cited by applicant]
RU 2388068C2 · 2010 [cited by applicant]
RU 2011145865C2 · 2013 [cited by applicant]
RU 2586842C2 · 2016 [cited by applicant]
TW 201007702A · 2010 [cited by applicant]
TW 201142830A · 2011 [cited by applicant]
TW 201303851A · 2013 [cited by applicant]
TW 201637000A · 2016 [cited by applicant]
WO 2001097511A1 · 2001 [cited by applicant]
WO 2009086919A1 · 2009 [cited by applicant]
WO 2010005050A1 · 2010 [cited by applicant]
WO 2011120800A1 · 2011 [cited by applicant]
WO 2014135235A1 · 2014 [cited by applicant]
WO 2014192602A1 · 2014 [cited by applicant]
Indian Office Action, dated Mar. 13, 2021, parallel patent application No. 202017020129. [cited by applicant]
Li, G., et al.; “The perceptual lossless quantization of spatial parameter for 3D audio signals;” International conference on multimedia modeling, MMM 2017: MultiMedia Modeling; Jan. 2017, Proceedings, Part II; pp. 381-… [cited by applicant]
Russian Office Action, dated Dec. 30, 2021, parallel patent application No. 2020119761. [cited by applicant]
English Translation of Russian Office Action, dated Dec. 30, 2021, parallel patent application No. 2020119761. [cited by applicant]
Russian Office Action, dated Dec. 25, 2021, parallel patent application No. 2020119762. [cited by applicant]
English Translation of Russian Office Action, dated Dec. 25, 2021, parallel patent application No. 2020119762. [cited by applicant]
Search Report, accompanying Office Action dated Dec. 30, 2021, in parallel patent application No. 2020119761. [cited by applicant]
6 Search Report, accompanying Office Action dated Dec. 25, 2021, parallel patent application No. 2020119762. [cited by applicant]
Japanese language Decision of Rejection and Decision to reject the amendments dated Oct. 21, 2025, issued in application No. JP 2023-179870 (English language translation, pp. 1-5 of attachment). [cited by applicant]
Non-Final Office Action dated Aug. 10, 2021, issued in application No. U.S. Appl. No. 16/867,856. [cited by applicant]
Japanese language office action dated Jul. 12, 2021, issued in application No. JP 2020-526987. [cited by applicant]
English language translation of office action dated Jul. 12, 2021, issued in application No. JP 2020-526987. [cited by applicant]
Japanese language office action dated Jul. 12, 2021, issued in application No. JP 2020-526994. [cited by applicant]
English language translation of office action dated Jul. 12, 2021, issued in application No. JP 2020-526994. [cited by applicant]
Japanese language Pre-Appeal Examination Report dated Dec. 16, 2024, issued in application No. JP 2022-133236. [cited by applicant]
Chinese language office action dated Mar. 15, 2023, issued in application No. CN 201880086690.3. [cited by applicant]
English language translation of office action dated Mar. 15, 2023 (pp. 1-6 of attachment). [cited by applicant]
Xu, J., et al.; “DRA Multi-channel Digital Audio Signal Coding Algorithm;” Digital Signal Pocessing; Jun. 2013; pp. 54-57. [cited by applicant]
English language translation of abstract of “DRA Multi-channel Digital Audio Signal Coding Algorithm” (p. 1 of attachment). [cited by applicant]
Ahonen J et al: “Diffuseness estimation using temporal variation of intensity vectors”; Applications of Signal Processing to Audio and Acoustics, 2009. WASPAA '09. IEEE Workshop on, IEEE, Piscataway, NJ, USA, Oct. 18, 2… [cited by applicant]
Robert L. Bleidt et al: “Development of the MPEG-H TV Audio System for ATSC 3.0”; IEEE Transactions on Broadcasting., vol. 63, No. 1, Mar. 2, 2017 (Mar. 2, 2017), pp. 202-236, XP055545453. [cited by applicant]
V. Pulkki, M-V. Laitinen, J. Vilkamo, J. Ahonen, T. Lokki, and T. Pihlajamäki, “Directional audio coding—perception-based reproduction of spatial sound”, International Workshop on the Principles and Application on Spati… [cited by applicant]
V. Pulkki, “Virtual source positioning using vector base amplitude panning”, J. Audio Eng. Soc., 45(6):456-466, Jun. 1997. [cited by applicant]
T. Hirvonen, J. Ahonen, and V. Pulkki, “Perceptual compression methods for metadata in Directional Audio Coding applied to audiovisual teleconference”, AES 126th Convention, May 7-10, 2009, Munich, Germany. [cited by applicant]
Chinese language Official Letter and Search Report, Jun. 10, 2019 in Taiwan application No. 107141081. [cited by applicant]
English Translation of Official Letter and Search Report, Jun. 10, 2019 in Taiwan application No. 107141081. [cited by applicant]
Chinese language First Office Action, Jun. 16, 2020 from Taiwan application No. 107141081. [cited by applicant]
English Translation of First Office Action, Jun. 16, 2020 from Taiwan application No. 107141081. [cited by applicant]
Official Letter and Search Report, Jun. 10, 2019 from Taiwan application No. 107141079. [cited by applicant]
English Translation of Official Letter and Search Report, Jun. 10, 2019 from Taiwan application No. 107141079. [cited by applicant]
Written opinion, in PCT/2018/081620, dated Feb. 19, 2019. [cited by applicant]
International Search Report, in PCT/2018/081620, dated Feb. 19, 2019. [cited by applicant]
Korean language office action dated Apr. 1, 2022, issued in application No. KR 10-2020-7017247. [cited by applicant]
Gray, R.M., et al.; “Quantization;” IEEE transactions on information theory; vol. 44; No. 6; Oct. 1998; pp. 2325-2383. [cited by applicant]
Japanese language office action dated Aug. 14, 2023, issued in application No. JP 2022-133236. [cited by applicant]
English language translation of office action dated Aug. 14, 2023 (pp. 1-10 of attachment). [cited by applicant]
Korean language Notice of Allowance dated Aug. 4, 2023, issued in application No. KR 10-2020-7017247. [cited by applicant]
Chinese language office action dated Dec. 13, 2022, issued in application No. CN 201880086689.0. [cited by applicant]
English language translation of office action dated Dec. 13, 2022, issued in application No. CN 201880086689.0 (pp. 14-23 of attachment). [cited by applicant]
Pulkki, V.; Directional Audio Coding in Spatial Sound Reproduction and Stereo Upmixing; Conference: Audio Engineering Society Conference: 28th International Conference: The Future of Audio Technology-Surround and Beyond… [cited by applicant]
Pulkki, V., et al.; “Directional audio coding-perception-based reproduction of spatial sound;” International Workshop on the Principles and Application on Spatial Hearing; Nov. 2009; pp. 1-4. [cited by applicant]
Non-Final Office Action dated Jul. 27, 2023, issued in application No. U.S. Appl. No. 17/571,970. [cited by applicant]
Japanese language office action dated Jan. 14, 2025, issued in application No. JP 2023-179870 (English language translation included—pp. 1-8 of attachment). [cited by applicant]
Korean language Notice of Allowance dated Feb. 22, 2023, issued in application No. KR 10-2020-7017280. [cited by applicant]
Del Galdo, G., et al.; “Efficient methods for high quality merging of spatial audio streams in directional audio coding;” Audio Engineering Society Convention; May 2009; pp. 1-14. [cited by applicant]
Pulkki, V.; “Spatial sound reproduction with directional audio coding;” Journal of the Audio Engineering Society; Apr. 2007; pp. 503-516. [cited by applicant]
Chinese language office action dated May 12, 2026, issued in application No. CN 202311255126.9 (English language translation, pp. 1-10 of attachment). [cited by applicant]