IP Library › Granted Patent US 11,743,553
Granted Patent B2
US 11,743,553 · App. 17/664,397 · Granted Aug 29, 2023

Data processor and transport of user control data to audio decoders and renderers

Inventors: Stephan Schreiner (Birgland, DE); Simone Neukam (Kalchreuth, DE); Harald Fuchs (Roettenbach, DE); Jan Plogsties (Fuerth, DE); Stefan Doehla (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Foerderung der angewandten Forschung e.V.
H04N21/8106G10L19/00G10L19/167H04N21/435H04N21/4363H04N21/4394H04N21/44222H04N21/44227H04N21/4852
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,743,553
App. No.
17/664,397
Granted
Aug 29, 2023
Kind
B2
Abstract

Audio data processor, having: a receiver interface for receiving encoded audio data and metadata related to the encoded audio data; a metadata parser for parsing the metadata to determine an audio data manipulation possibility; an interaction interface for receiving an interaction input and for generating, from the interaction input, interaction control data related to the audio data manipulation possibility; and a data stream generator for obtaining the interaction control data and the encoded audio data and the metadata and for generating an output data stream, the output data stream having the encoded audio data, at least a portion of the metadata, and the interaction control data.

Claims (75)

1. An audio data processor for processing packetized audio data, the audio data processor comprising:

a receiver interface for receiving encoded audio data comprising audio elements and metadata related to the audio elements;

a metadata parser for parsing the metadata to determine an audio data manipulation possibility of the audio elements;

an interaction interface configured to receive an interaction input and configured to generate, from the interaction input, interaction control data related to the audio data manipulation possibility for manipulating the audio elements externally from a decoder;

wherein from the interaction interface a user can select and manipulate the audio elements to adapt the audio presentation to his personal preferences; and

a data stream generator configured to acquire the interaction control data and the encoded audio data and the metadata and configured to generate an output data stream, the output data stream being again a valid encoded audio stream comprising the still encoded audio data, the metadata, and the interaction control data.

2. The audio data processor of claim 1 ,

wherein the interaction interface provides two interaction modes,

wherein a first interaction mode is an advanced interaction mode that is signaled for each audio element group being present in an audio scene for enabling the user to freely choose which audio element groups to play back and to interact with all of them, and

wherein a second mode is a basic interaction mode comprising Group Presets, for enabling the user to choose one of the Group Presets as a preset,

wherein audio element groups gather related audio elements that are to be manipulated jointly.

3. The audio data processor of claim 2 ,

wherein the Group Presets comprise at least one of:

an On/Off interactivity, wherein a group of elements is switched on or off;

a position interactivity, wherein the positions of a group of elements are changed;

a gain interactivity, wherein the level or gain of a group of elements is changed;

WIRE interactivity, wherein the audio content of the elements of a group are routed to a WIRE output.

4. The audio data processor of claim 1 ,

wherein the output data stream is sent to the decoder, wherein the presence of the interaction control data in the output data stream enables said decoder to identify that an interaction has happened.

5. The audio data processor of claim 1 ,

wherein the encoded audio data comprises separate encoded audio elements, wherein the metadata is related to the audio elements,

wherein the metadata parser is configured to parse a corresponding portion of the metadata corresponding to the encoded audio elements and to determine, for at least one of the audio elements, the audio manipulation possibility, and

wherein the interaction interface is configured to generate, for the at least one encoded audio element, interaction control data being related to the at least one encoded audio element.

6. The audio data processor of claim 1 ,

wherein the interaction interface is configured to

present, to the user, the audio data manipulation possibility derived from the metadata by the metadata parser, and

to receive, from the user, a user input on the specific data manipulation of the audio data manipulation possibility.

7. The audio data processor of claim 1 ,

wherein the data stream generator is configured to

process a data stream comprising the encoded audio data and the metadata received by the receiver interface without decoding the encoded audio data, or

copy the encoded audio data and the metadata without changes in the output data stream, and

wherein the data stream generator is configured to

add an additional data portion containing the interaction control data to the encoded audio data or to the metadata in the output data stream.

8. The audio data processor of claim 1 ,

wherein the data stream generator is configured to generate, in the output data stream, the interaction control data in a same format as the metadata.

9. The audio data processor of claim 1 ,

wherein the data stream generator is configured to associate, with the interaction control data, an identifier in the output data stream, the identifier being different from an identifier associated with the metadata.

10. The audio data processor of claim 1 ,

wherein the data stream generator is configured to add, to the interaction control data, signature data, the signature data indicating information on an application, a device or a user performing an audio data manipulation or providing the interaction input.

11. The audio data processor of claim 1 ,

wherein the metadata parser is configured to identify a disabling possibility for one or more audio elements,

wherein the interaction interface is configured to receive a disabling information for the one or more audio elements, and

wherein the data stream generator is configured to mark the one or more audio elements as disabled in the interaction control data, or to remove the disabled one or more audio elements from the encoded audio data so that the output data stream does not include encoded audio data for the disabled one or more audio elements.

12. The audio data processor of claim 1 ,

wherein the data stream generator is configured to dynamically generate the output data stream, wherein in response to a new interaction input, the interaction control data is updated to match the new interaction input, and

wherein the data stream generator is configured to include the updated interaction control data in the output data stream.

13. The audio data processor of claim 1 ,

wherein the receiver interface is configured to receive a main audio data stream comprising the encoded audio data and the metadata related to the encoded audio data, and to additionally receive optional audio data comprising an optional audio element,

wherein the metadata related to said optional audio element is contained in said main audio data stream.

14. The audio data processor of claim 1 ,

wherein the metadata parser is configured to determine the audio manipulation possibility for an optional audio element that is not included in the encoded audio data,

wherein the interaction interface is configured to receive an interaction input for the optional audio element, and

wherein the receiver interface is configured to request audio data for the optional audio element from an audio data provider or to receive the audio data for the optional audio element from a different substream contained in a broadcast stream or an internet protocol connection.

15. The audio data processor of claim 1 ,

wherein the data stream generator is configured to assign, in the output data stream, a further packet type to the interaction control data, the further packet type being different from packet types for the encoded audio data and the metadata, or

wherein the data stream generator is configured to add, into the output data stream, fill data in a fill data packet type, wherein an amount of fill data is determined based on a data rate requirement determined by an output interface of the audio data processor.

16. The audio data processor of claim 1 ,

being implemented as a separate first device that is separated from a second device which is configured to receive the processed, but still encoded, audio data from the first device for decoding said audio data,

wherein the receiver interface forms an input to the separate first device via a wired or wireless connection,

wherein the audio data processor further comprises an output interface connected to the data stream generator, the output interface being configured for outputting the output data stream, and

wherein the output interface performs an output of the separate first device and comprises a wireless interface or a wired connector.

17. A method for processing packetized audio data, the method comprising:

receiving encoded audio data comprising audio elements and metadata related to the audio elements;

parsing the metadata to determine an audio data manipulation possibility of the audio elements;

receiving an interaction input and generating, from the interaction input, interaction control data related to the audio data manipulation possibility for manipulating the audio elements externally from a decoder;

wherein by said interaction input a user can select and manipulate the audio elements to adapt the audio presentation to his personal preferences; and

obtaining the interaction control data and the encoded audio data and the metadata, and

generating an output data stream, the output data stream being again a valid encoded audio stream comprising the still encoded audio data, the metadata, and the interaction control data.

18. A computer program for performing, when running on a computer or a processor, a method for processing packetized audio data, the method comprising:

receiving encoded audio data comprising audio elements and metadata related to the audio elements;

parsing the metadata to determine an audio data manipulation possibility of the audio elements;

receiving an interaction input and generating, from the interaction input, interaction control data related to the audio data manipulation possibility for manipulating the audio elements externally from a decoder;

wherein by said interaction input a user can select and manipulate the audio elements to adapt the audio presentation to his personal preferences; and

obtaining the interaction control data and the encoded audio data and the metadata, and

generating an output data stream, the output data stream being again a valid encoded audio stream comprising the still encoded audio data, the metadata, and the interaction control data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2023
From: SCHREINER, STEPHAN; NEUKAM, SIMONE; FUCHS, HARALD; PLOGSTIES, JAN; DOEHLA, STEFAN
To: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 063408/0366 →
Priority Claims (1)
EP 14170416 · May 28, 2014 · regional
Continuity (4)
Continuation 15931422 · May 13, 2020
Continuation 15357640 · Nov 21, 2016
Continuation PCTEP2015056768 · Mar 27, 2015
Related Publication 20220286756A1 · Sep 8, 2022
Cited By (1)
US 12,279,007