Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element
Embodiments relate to an audio processing unit that includes a buffer, bitstream payload deformatter, and a decoding subsystem. The buffer stores at least one block of an encoded audio bitstream. The block includes a fill element that begins with an identifier followed by fill data. The fill data includes at least one flag identifying whether enhanced spectral band replication (eSBR) processing is to be performed on audio content of the block. A corresponding method for decoding an encoded audio bitstream is also provided.
1. An audio processing device comprising:
a bitstream payload deformatter configured to demultiplex a block of an encoded audio bitstream;
a decoding subsystem coupled to the bitstream payload deformatter and configured to decode at least a portion of the block of the encoded audio bitstream, wherein the block of the encoded audio bitstream includes:
a fill element with an identifier indicating a start of the fill element and fill data after the identifier, wherein the fill data includes:
at least one flag identifying whether enhanced spectral band replication processing is to be performed on audio content of at least one block of the encoded audio bitstream, and
enhanced spectral band replication metadata which does not include one or more parameters used for both spectral patching and harmonic transposition, wherein the enhanced spectral band replication metadata is metadata configured to enable at least one eSBR tool which is described or mentioned in the MPEG USAC standard and which is not described or mentioned in the MPEG-4 AAC standard,
wherein the enhanced spectral band replication metadata includes a parameter indicating whether to perform pre-flattening, and the decoding subsystem is further configured to perform additional pre-processing to avoid discontinuities in a shape of a spectral envelope of a high frequency signal input to an envelope adjuster if the parameter indicates that pre-flattening is to be performed.
2. The audio processing unit of claim 1 , wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.
3. The audio processing unit of claim 1 , wherein the identifier is a three bit unsigned integer transmitted most significant bit first and having a value of 0×6.
4. The audio processing unit of claim 1 , wherein the fill data includes an extension payload, the extension payload includes spectral band replication extension data, and the extension payload is identified with a four bit unsigned integer transmitted most significant bit first and having a value of ‘1101’ or ‘1110’, and, wherein the spectral band replication extension data includes:
a spectral band replication header,
spectral band replication data after the header, and
an spectral band replication extension element after the spectral band replication data, and wherein the flag is included in the spectral band replication extension element.
5. A method for decoding an encoded audio bitstream, the method comprising:
demultiplexing a block of the encoded audio bitstream;
decoding at least a portion of the block of the encoded audio bitstream,
wherein the block of the encoded audio bitstream includes:
a fill element with an identifier indicating a start of the fill element and fill data after the identifier, wherein the fill data includes:
a flag identifying whether enhanced spectral band replication processing is to be performed on audio content of at least one block of the encoded audio bitstream, and
enhanced spectral band replication metadata which does not include one or more parameters used for both spectral patching and harmonic transposition, wherein the enhanced spectral band replication metadata is metadata configured to enable at least one eSBR tool which is described or mentioned in the MPEG USAC standard and which is not described or mentioned in the MPEG-4 AAC standard; and
wherein the enhanced spectral band replication metadata includes a parameter indicating whether to perform pre-flattening, and performing additional pre-processing to avoid discontinuities in a shape of a spectral envelope of a high frequency signal input to an envelope adjuster if the parameter indicates that pre-flattening is to be performed.
6. The method of claim 5 , wherein the identifier is a three bit unsigned integer transmitted most significant bit first and having a value of 0×6.
7. The method of claim 5 , wherein the fill data includes an extension payload, the extension payload includes spectral band replication extension data, and the extension payload is identified with a four bit unsigned integer transmitted most significant bit first and having a value of ‘1101’ or ‘1110’, and, wherein the spectral band replication extension data includes:
a spectral band replication header,
spectral band replication data after the header, and
an spectral band replication extension element after the spectral band replication data, and wherein the flag is included in the spectral band replication extension element.
8. The method of claim 5 , wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.