Methods and systems for rendering object based audio
Methods for generating an object based audio program, renderable in a personalizable manner, and including a bed of speaker channels renderable in the absence of selection of other program content (e.g., to provide a default full range audio experience). Other embodiments include steps of delivering, decoding, and/or rendering such a program. Rendering of content of the bed, or of a selected mix of other content of the program, may provide an immersive experience. The program may include multiple object channels (e.g., object channels indicative of user-selectable and user-configurable objects), the bed of speaker channels, and other speaker channels. Another aspect is an audio processing unit (e.g., encoder or decoder) configured to perform, or which includes a buffer memory which stores at least one frame (or other segment) of an object based audio program (or bitstream thereof) generated in accordance with, any embodiment of the method.
1 . A method of rendering audio content of an audio program, wherein the audio
program is associated with a mix of the audio content, the method comprising:
receiving the audio program, the audio program further comprising mix metadata indicative of the mix of the audio content, wherein the mix of the audio content comprises: at least one of (1) one or more set(s) of object channels and (2) one or more bed(s) of speaker channels;
parsing, by an audio processing unit, the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels, and the mix metadata;
determining that the mix metadata is associated with the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels;
determining, based on the mix metadata, rendering information for the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels; and
rendering the audio content based on the mix metadata, wherein the rendering comprises rendering the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels based on the rendering information.
2 . The method of claim 1 , further comprising rendering the audio content by selecting and mixing the audio content based on the mix metadata.
3 . The method of claim 1 , further comprising determining rendering parameters based on the rendering information.
4 . The method of claim 1 , wherein the mix metadata indicates a selection of a subset of the one or more of the set(s) object channels and/or the one or more of the bed(s) of speaker channels.
5 . The method of claim 1 , wherein the audio program comprises one or more bitstreams.
6 . The method of claim 1 , wherein the mix metadata indicates a mix of ambient content and non-ambient content.
7 . The method of claim 6 , wherein the ambient content is indicative of ambient sound at a spectator event, and wherein the non-ambient content is indicative of commentary on the spectator event.
8 . The method of claim 1 , wherein the audio program is an un-encoded representation which is indicative of the audio content and the metadata of the program, and the un-encoded representation is a bitstream or at least one file of data stored in a non-transient manner in a memory.
9 . The method of claim 1 , wherein the mix metadata is indicative of a layered mix graph, the layered mix graph is indicative of selectable mixes of the one or more of the set(s) object channels and/or the one or more of the bed(s) of speaker channels, wherein the layered mix graph includes a base layer of metadata and at least one extension layer of metadata.
10 . The method of claim 1 , wherein the audio program is an encoded bitstream comprising frames, and each of the frames of includes mix metadata.
11 . A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performing the method of claim 1 .
12 . A system for rendering audio content of an audio program, wherein the audio program is associated with a mix of the audio content, the system comprising:
a receiver for receiving the audio program, the audio program further comprising mix metadata indicative of the mix of the audio content, wherein the mix of the audio content comprises: at least one of (1) one or more set(s) of object channels and (2) one or more bed(s) of speaker channels;
a parser for parsing, by an audio processing unit, the one or more bed(s) of speaker channels and/or one or more set(s) of object channels and the mix metadata;
a first processor for determining that the mix metadata is associated with the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels;
a second processor for determining, based on the mix metadata, rendering information for the one or more bed(s) of speaker channels; and
a renderer for rendering the audio content based on the mix metadata, wherein the rendering comprises rendering the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels based on the rendering information.
13 . A method of generating an audio program, the method comprising:
determining audio content, wherein the audio content comprises: one or more bed(s) of speaker channels and/or one or more set(s) of object channels;
determining mix metadata indicative of at least a mix of the audio content, wherein the mix metadata is associated with the one or more bed(s) of speaker channels and/or the one or more set(s) of object channels; and
generating the audio program, wherein the audio program comprises the audio content and the mix metadata, wherein the mix of the audio content comprises: at least one of (1) one or more set(s) of object channels and (2) one or more bed(s) of speaker channels.
14 . The method of claim 13 , wherein the mix metadata indicates a selection of a subset of the one or more of the set(s) object channels and/or the one or more of the bed(s) of speaker channels.
15 . The method of claim 13 , wherein the audio program comprises one or more bitstreams.
16 . The method of claim 13 , wherein the mix metadata indicates a mix of ambient content and non-ambient content.
17 . The method of claim 16 , wherein the ambient content is indicative of ambient sound at a spectator event, and wherein the non-ambient content is indicative of commentary on the spectator event.
18 . The method of claim 13 , wherein the audio program is an un-encoded representation which is indicative of the audio content and the metadata of the program, and the un-encoded representation is a bitstream or at least one file of data stored in a non-transient manner in a memory.
19 . The method of claim 13 , wherein the mix metadata is indicative of a layered mix graph, the layered mix graph is indicative of selectable mixes of the one or more of the set(s) object channels and/or the one or more of the bed(s) of speaker channels, wherein the layered mix graph includes a base layer of metadata and at least one extension layer of metadata.
20 . The method of claim 13 , wherein the audio program is an encoded bitstream comprising frames, and each of the frames of includes mix metadata.
21 . A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performing the method of claim 13 .