Decoding audio frames and converted metadata frames from a target encoder
A target decoder may receive a bitstream including an audio frame in a target format associated with a target encoder and a metadata frame associated with the audio frame. The audio frame may be transcoded from an earlier audio frame. The metadata frame may be converted from an earlier metadata frame associated with the earlier audio frame. The target decoder may decode audio data from the audio frame and metadata from the metadata frame.
1 . A method, comprising:
receiving, by a target decoder, a target bitstream including an audio frame in a target format associated with a target encoder, a metadata frame associated with the audio frame, and configuration fields configured to initialize the target decoder to decode metadata transported in the target bitstream, wherein the audio frame is transcoded from an earlier audio frame in a source bitstream from a source decoder that is a type of codec among multiple types of codecs, and the metadata frame is converted from an earlier metadata frame associated with the earlier audio frame; and
decoding audio data from the audio frame and metadata from the metadata frame based on the configuration fields, wherein the configuration fields include a first field indicating a codec identifier that indicates the type of codec used to produce the source bitstream.
2 . The method of claim 1 , further comprising transmitting an entirety of metadata from the earlier metadata frame.
3 . The method of claim 2 , further comprising transmitting metadata from the earlier metadata frame without modification.
4 . The method of claim 1 , further comprising determining a renderer from a plurality of renderers.
5 . The method of claim 1 , wherein the earlier audio frame has a first size or duration, and the audio frame in the target bitstream has a second size or duration.
6 . The method of claim 5 , wherein the first size or duration is greater than the second size or duration.
7 . The method of claim 1 , wherein the earlier metadata frame includes the metadata describing the audio data, and the target bitstream includes the metadata describing the audio data as encoded in the audio frame.
8 . The method of claim 7 , wherein the metadata defines a position of the audio data in a 3D sound environment.
9 . The method of claim 1 , wherein the configuration fields include a second field indicating a version of configuration data in the configuration fields or a version of metadata in the metadata frame.
10 . The method of claim 1 , wherein the target bitstream indicates to the target decoder whether metadata is carried in a frame in the target bitstream.
11 . An apparatus, comprising:
a memory; and
a processor configured to execute instructions stored in the memory to:
receive a target bitstream including an audio frame in a target format associated with a target encoder, a metadata frame associated with the audio frame, and configuration fields configured to initialize a target decoder to decode metadata transported in the target bitstream, wherein the audio frame is transcoded from an earlier audio frame in a source bitstream from a source decoder that is a type of codec among multiple types of codecs, and the metadata frame is converted from an earlier metadata frame associated with the earlier audio frame; and
decode audio data from the audio frame and metadata from the metadata frame based on the configuration fields, wherein the configuration fields include a first field indicating a codec identifier that indicates the type of codec used to produce the source bitstream.
12 . The apparatus of claim 11 , wherein the processor is further configured to execute instructions stored in the memory to transmit all metadata from the earlier metadata frame.
13 . The apparatus of claim 11 , wherein the processor is further configured to execute instructions stored in the memory to transmit metadata from the earlier metadata frame without changing the metadata.
14 . The apparatus of claim 11 , wherein the processor is further configured to execute instructions stored in the memory to determine a renderer from a plurality of renderers.
15 . The apparatus of claim 11 , wherein the earlier audio frame has M samples, and the audio frame in the target bitstream has N samples.
16 . The apparatus of claim 11 , wherein the apparatus is a head unit of a vehicle.
17 . A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:
receiving a target bitstream including an audio frame in a target format associated with a target encoder, a metadata frame associated with the audio frame, and configuration fields configured to initialize a target decoder to decode metadata transported in the target bitstream, wherein the audio frame is transcoded from an earlier audio frame in a source bitstream from a source decoder that is a type of codec among multiple types of codecs, and the metadata frame is converted from an earlier metadata frame associated with the earlier audio frame; and
decoding audio data from the audio frame and metadata from the metadata frame based on the configuration fields, wherein the configuration fields include a first field indicating a codec identifier that indicates the type of codec used to produce the source bitstream.
18 . The non-transitory computer readable medium storing instructions of claim 17 , the operations further comprising determining a renderer from a plurality of renderers.
19 . The non-transitory computer readable medium storing instructions of claim 17 , wherein the earlier audio frame has a first size or duration, and the audio frame has a second size or duration.
20 . The non-transitory computer readable medium storing instructions of claim 19 , wherein the first size or duration is greater than the second size or duration.
21 . The non-transitory computer readable medium storing instructions of claim 17 , wherein the earlier metadata frame includes the metadata describing the audio data, and the target bitstream includes the metadata describing the audio data as encoded in the audio frame.
22 . The non-transitory computer readable medium storing instructions of claim 21 , wherein the metadata defines a position of the audio data among a plurality of speakers in a vehicle.