IP Library Granted Patent US 12,058,501
Granted Patent B2
US 12,058,501 · App. 17/585,169 · Granted Aug 6, 2024

Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to DirAC based spatial audio coding

Inventors: Guillaume Fuchs (Bubenreuth, DE); Jürgen Herre (Erlangen, DE); Fabian Küch (Erlangen, DE); Stefan Döhla (Erlangen, DE); Markus Multrus (Nuremberg, DE); Oliver Thiergart (Erlangen, DE); Oliver Wübbolt (Hannover, DE); Florin Ghido (Nuremberg, DE); Stefan Bayer (Nuremberg, DE); Wolfgang Jaegers (Forchheim, DE)
Assignee: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
H04R5/04H04S7/30H04S7/40H04R2205/024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,058,501
App. No.
17/585,169
Granted
Aug 6, 2024
Kind
B2
Abstract

An apparatus for generating a description of a combined audio scene, includes: an input interface for receiving a first description of a first scene in a first format and a second description of a second scene in a second format, wherein the second format is different from the first format; a format converter for converting the first description into a common format and for converting the second description into the common format, when the second format is different from the common format; and a format combiner for combining the first description in the common format and the second description in the common format to obtain the combined audio scene.

Claims (56)

1. An apparatus for generating a description of an enhanced audio scene, comprising:

an input interface for receiving a first description of a first scene in a first format and a second description of a second scene in a second format, wherein the second format is different from the first format;

a format converter for converting the first description into a common format and for converting the second description into the common format, when the second format is different from the common format; and

a format combiner for combining the first description in the common format and the second description in the common format to acquire a description of a combined audio scene comprising combined audio scene metadata and a combined audio scene transport signal;

a metadata generator for generating combined metadata, the combined metadata comprising the combined audio scene metadata and object metadata of a separate object description for a specific audio object,

wherein the separate object description for the specific audio object comprises, in the object metadata of the separate object description for the specific audio object, a single direction throughout all frequency bands of the specific audio object, wherein the specific audio object is either static or moving slower than a velocity threshold, and

wherein the separate object description comprises an object waveform signal,

a transport encoder for encoding the transport signal and the object waveform signal to obtain an encoded transport signal; and

an output interface for outputting the description of the enhanced audio scene, wherein the enhanced audio scene comprises the encoded transport signal and the combined metadata.

2. The apparatus of claim 1 ,

wherein the first format is selected from a group of formats comprising a first order Ambisonics format, a high order Ambisonics format, a DirAC format, an audio object format and a multi-channel format, and

wherein the second format is selected from a group of formats comprising a first order Ambisonics format, a high order Ambisonics format, the common format, a DirAC format, an audio object format, and a multi-channel format.

3. The apparatus of claim 1 ,

wherein the format converter is configured to convert the first description into a first B-format signal representation and to convert the second description into a second B-format signal representation, and

wherein the format combiner is configured to combine the first B-format signal representation and the second B-format signal representation by individually combining the individual components of the first B-format signal representation and the second B-format signal representation.

4. The apparatus of claim 1 ,

wherein the format converter is configured to convert the first description into a first pressure/velocity signal representation and to convert the second description into a second pressure/velocity signal representation, and

wherein the format combiner is configured to combine the first pressure/velocity signal representation and the second pressure/velocity signal representation by individually combining the individual components of the pressure/velocity signal representations to acquire a combined pressure/velocity signal representation.

5. The apparatus of claim 1 ,

wherein the format converter is configured to convert the first description into a first DirAC parameter representation and to convert the second description into a second DirAC parameter representation, when the second description is different from the DirAC parameter representation, and

wherein the format combiner is configured to combine the first DirAC parameter representation and the second DirAC parameter representation by individually combining the individual components of the first DirAC parameter representation and the second DirAC parameter representations to acquire a combined DirAC parameter representation for the combined audio scene.

6. The apparatus of claim 5 ,

wherein the format combiner is configured to generate, as the combined audio scene metadata,—direction of arrival values for time-frequency tiles or direction of arrival values and diffuseness values for the time-frequency tiles representing the combined audio scene.

7. The apparatus of claim 1 ,

further comprising a DirAC analyzer for analyzing the combined audio scene to derive DirAC parameters for the combined audio scene as the combined audio scene metadata,

wherein the DirAC parameters comprise direction of arrival values for time-frequency tiles or direction of arrival values and diffuseness values for the time-frequency tiles representing the combined audio scene.

8. The apparatus of claim 1 , further comprising:

a metadata encoder

for encoding DirAC metadata as the combined audio scene metadata to acquire encoded DirAC metadata, or

for encoding DirAC metadata derived from the first scene to acquire first encoded DirAC metadata and for encoding DirAC metadata derived from the second scene to acquire second encoded DirAC metadata, wherein the first encoded DirAC metadata and the second encoded DirAC metadata represent the combined audio scene metadata.

9. The apparatus of claim 1 , wherein the combined audio scene metadata comprise encoded DirAC metadata and wherein the encoded transport signal comprises one or more encoded transport channels.

10. The apparatus of claim 1 ,

wherein the metadata for the separate object description for the specific audio object comprises, in addition to the single direction, at least one of a distance, a diffuseness, and any other object attribute, and wherein the combined metadata comprises, as the object metadata for the separate object description for the specific audio object, the single direction, and at least one of the distance, the diffuseness, and the any other object attribute.

11. A method for generating a description of an enhanced audio scene, comprising:

receiving a first description of a first scene in a first format and receiving a second description of a second scene in a second format, wherein the second format is different from the first format;

converting the first description into a common format and converting the second description into the common format, when the second format is different from the common format; and

combining the first description in the common format and the second description in the common format to acquire a description of a combined audio scene comprising combined audio scene metadata and a combined audio scene transport signal;

generating combined metadata, the combined metadata comprising the combined audio scene metadata and object metadata of a separate object description for a specific audio object,

wherein the separate object description for the specific audio object comprises, in the object metadata of the separate object description for the specific audio object, a single direction throughout all frequency bands of the specific audio object, wherein the specific audio object is either static or moving slower than a velocity threshold, and

wherein the separate object description comprises an object waveform signal,

encoding the transport signal and the object waveform signal to obtain an encoded transport signal; and

outputting the description of the enhanced audio scene, wherein the enhanced audio scene comprises the encoded transport signal and the combined metadata.

12. A non-transitory digital storage medium having a computer program stored thereon to perform, when said computer program is run by a computer, the method for generating a description of an enhanced audio scene comprising:

receiving a first description of a first scene in a first format and receiving a second description of a second scene in a second format, wherein the second format is different from the first format;

converting the first description into a common format and converting the second description into the common format, when the second format is different from the common format; and

combining the first description in the common format and the second description in the common format to acquire a description of a combined audio scene comprising combined audio scene metadata and a combined audio scene transport signal;

generating combined metadata, the combined metadata comprising the combined audio scene metadata and object metadata of a separate object description for a specific audio object,

wherein the separate object description for the specific audio object comprises, in the object metadata of the separate object description for the specific audio object, a single direction throughout all frequency bands of the specific audio object, wherein the specific audio object is either static or moving slower than a velocity threshold, and

wherein the separate object description comprises an object waveform signal,

encoding the transport signal and the object waveform signal to obtain an encoded transport signal; and

outputting the description of the enhanced audio scene, wherein the enhanced audio scene comprises the encoded transport signal and the combined metadata.

13. The apparatus of claim 1 , wherein the combined audio scene comprises, as the combined audio scene metadata, direction of arrival data and diffuseness data for each frequency band for a frame, and wherein the separate object description for the specific audio object comprises, as the object metadata, the single direction as direction of arrival data for all frequency bands of the frame, and wherein the combined metadata comprises the direction of arrival data and the diffuseness data for each frequency band for the frame, and the single direction for all frequency bands of the frame.

14. The apparatus of claim 13 , wherein the separate object description for the specific audio object comprises, as the object metadata, a diffuseness of zero or no diffuseness for all the frequency bands of the frame, and wherein the combined metadata comprises the diffuseness of zero or no diffuseness.

15. The apparatus of claim 13 , wherein the object metadata of the separate object description for the specific audio object is updated less frequently than the first audio scene and the second audio scene.

16. The apparatus of claim 13 , wherein the combined audio scene comprises, as the combined audio scene metadata, the direction of arrival data and the diffuseness data for each frequency band and for each frame, and wherein the separate object description for the specific audio object comprises, as the object metadata, the single direction for all the frequency bands of every n-th frame only, wherein n is greater than or equal to two, and wherein the combined metadata comprise the direction of arrival data and the diffuseness data for each frequency band and for each frame and the single direction for all the frequency bands of every n-th frame only, wherein n is greater than or equal to two.

17. The apparatus of claim 1 , further comprising a manipulator for selectively applying a spatial filtering to the specific audio object without affecting the combined audio scene by the spatial filtering.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2022
From: FUCHS, GUILLAUME; HERRE, JÜRGEN; KÜCH, FABIAN; DÖHLA, STEFAN; MULTRUS, MARKUS; THIERGART, OLIVER; WÜBBOLT, OLIVER; GHIDO, FLORIN; BAYER, STEFAN; JAEGERS, WOLFGANG
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 058781/0546 →
Priority Claims (1)
EP 17194816 · Oct 4, 2017 · regional
Continuity (3)
Division 16821069 · Mar 17, 2020
Continuation PCTEP2018076641 · Oct 1, 2018
Related Publication 20220150635A1 · May 12, 2022