Encoding device and encoding method, decoding device and decoding method, and program
There is provided a decoding device including at least one circuit configured to acquire one or more encoded audio signals including a plurality of channels and/or a plurality of objects and priority information for each of the plurality of channels and/or the plurality of objects, and to decode the one or more encoded audio signals according to the priority information.
1 . A decoding device comprising:
circuitry configured to:
acquire one or more encoded audio signals including at least one of a plurality of channels and a plurality of audio objects;
acquire meta-data for each of the plurality of audio objects;
decode the one or more encoded audio signals according to the meta-data; and
render the plurality of audio objects based on the meta-data by using VBAP (Vector Base Amplitude Panning) wherein:
the meta-data is priority information for each of the plurality of channels and/or the plurality of audio objects,
a value range of the priority information is from 0 to 7,
the circuitry is further configured to, prior to the decoding of the one or more encoded audio signals, perform processing with respect to an obtained audio signal such that in the processing:
a high-frequency component of a non-encoded audio signal is generated from an audio signal of:
a low frequency component generated from performing overlapping addition, and
a high frequency power value,
the obtained audio signal of one frame is divided into sections called time slots, and audio of each of the time slots is band-divided into a signal of a plurality of low frequency sub-bands,
a high-frequency sub-band signal and a low-frequency sub-band signal are synthesized,
an audio signal including the high frequency component is generated, and
the audio signal including the high frequency component generated for each of the time slots are combined, resulting in an audio signal of one time frame including the high frequency component.
2 . The decoding device according to claim 1 , wherein
a lowest priority degree is 0 and a highest priority degree is 7.
3 . The decoding device according to claim 1 , wherein
the meta-data is position information indicating each of position of the plurality of audio objects.
4 . The decoding device according to claim 3 , wherein
the position information is expressed by a horizontal angle from a predetermined reference position, a vertical angle from the predetermined reference position, and a distance from the predetermined reference position to a predetermined audio object.
5 . The decoding device according to claim 1 , wherein
the meta-data is a gain of the audio object.
6 . The decoding device according to claim 1 , wherein
a parameter called num_objects indicates a number of objects representing the plurality of audio objects.
7 . The decoding device according to claim 1 , wherein
the circuitry is further configured to perform IMDCT (inverse modified discrete cosine transform) to generate an audio signal.
8 . The decoding device according to claim 1 , wherein
a sum of audio signals from each of the plurality of channels is multiplied by a VBAP gain of corresponding channels in the plurality of channels, producing an audio signal for each channel.
9 . The decoding device according to claim 8 , wherein
gain adjustment of the encoded audio signals of each object is performed through the VBAP gain.
10 . The decoding device according to claim 1 , wherein
a speaker index, s, acts to specify a speaker corresponding to a predetermined channel.
11 . The decoding device according to claim 1 , wherein
the processing comprises SBR (spectral band) processing is performed with respect to and the audio signal comprises an audio signal obtained by overlappingly adding IMDCT signals output from an IMCDT unit.
12 . The decoding device according to claim 1 , wherein
a signal of each of the sub-bands of high frequency is generated based on a signal of the plurality of low frequency sub-bands and a power value of each of the sub-bands of high frequency.
13 . The decoding device according to claim 1 , wherein
a target high frequency sub-band signal is generated by adjusting power of a low frequency sub-band signal of a predetermined sub-band by a power of a target sub-band of high frequency or by shifting the frequency thereof.
14 . A method executed by at least one processing circuit of a decoding device, the method comprising:
acquiring one or more encoded audio signals including at least one of a plurality of channels and a plurality of audio objects;
acquiring meta-data for each of the plurality of audio objects;
decoding the one or more encoded audio signals according to the meta-data; and
rendering the plurality of audio objects based on the meta-data by using VBAP (Vector Base Amplitude Panning) wherein:
the meta-data is priority information for each of the plurality of channels and/or the plurality of audio objects,
a value range of the priority information is from 0 to 7,
the circuitry is further configured to, prior to the decoding of the one or more encoded audio signals, perform processing with respect to an obtained audio signal such that in the processing:
a high-frequency component of a non-encoded audio signal is generated from an audio signal of:
a low frequency component generated from performing overlapping addition, and
a high frequency power value,
the obtained audio signal of one frame is divided into sections called time slots, and audio of each of the time slots is band-divided into a signal of a plurality of low frequency sub-bands,
a high-frequency sub-band signal and a low-frequency sub-band signal are synthesized,
an audio signal including the high frequency component is generated, and
the audio signal including the high frequency component generated for each of the time slots are combined, resulting in an audio signal of one time frame including the high frequency component.
15 . A non-transitory computer readable medium storing instructions that, when executed by at least one processing circuit of a decoding device, causes the at least one processing circuit to perform a method comprising:
acquiring one or more encoded audio signals including at least one of a plurality of channels and a plurality of audio objects;
acquiring meta-data for each of the plurality of audio objects;
decoding the one or more encoded audio signals according to the meta-data; and
rendering the plurality of audio objects based on the meta-data by using VBAP (Vector Base Amplitude Panning) wherein:
the meta-data is priority information for each of the plurality of channels and/or the plurality of audio objects,
a value range of the priority information is from 0 to 7,
the circuitry is further configured to, prior to the decoding of the one or more encoded audio signals, perform processing with respect to an obtained audio signal such that in the processing:
a high-frequency component of a non-encoded audio signal is generated from an audio signal of:
a low frequency component generated from performing overlapping addition, and
a high frequency power value,
the obtained audio signal of one frame is divided into sections called time slots, and audio of each of the time slots is band-divided into a signal of a plurality of low frequency sub-bands,
a high-frequency sub-band signal and a low-frequency sub-band signal are synthesized,
an audio signal including the high frequency component is generated, and
the audio signal including the high frequency component generated for each of the time slots are combined, resulting in an audio signal of one time frame including the high frequency component.