Object audio coding
In one aspect, a computer-implemented method, includes obtaining object audio and metadata that spatially describes the object audio, converting the object audio to Ambisonics audio based on the metadata, encoding, in a first bit stream, the Ambisonics audio, and encoding, in a second bit stream, at least a subset of the metadata.
1 . A computer-implemented method, comprising:
obtaining object audio and metadata that spatially describes the object audio, wherein the object audio and the metadata characterize an original audio scene;
converting the object audio to Ambisonics audio based on the metadata;
encoding, in a first bit stream, the Ambisonics audio;
encoding, in a second bit stream, at least a subset of the metadata; and
transmitting the first and second bit streams to a decode side where the first and second bit streams are then decoded and the Ambisonics audio is rendered into a plurality of audio channels that match the original audio scene characterized by the object audio and the metadata.
2 . The method of claim 1 wherein encoding at least the subset of the metadata includes encoding less than all of the metadata in the second bit stream.
3 . The method of claim 2 wherein the metadata is used by a decoder to convert the Ambisonics audio to the object audio.
4 . The method of claim 3 wherein the Ambisonics audio is time domain Ambisonics audio.
5 . The method of claim 4 wherein the metadata includes at least one of a direction of an object relative to a listening position, a distance of the object relative to the listening position, or position of the object relative to the listening position.
6 . The method of claim 5 wherein converting the object audio to the Ambisonics audio based on the metadata includes mapping a contribution or acoustic energy of each object in the object audio to each component in the Ambisonics audio using spatial information in the metadata.
7 . A computer-implemented method, comprising:
decoding a first bit stream to obtain Ambisonics audio that characterizes an original audio scene;
decoding a second bit stream to obtain a subset of original metadata which spatially describes original object audio, wherein the original object audio also characterizes the original audio scene;
extracting object audio from the Ambisonics audio using the subset of original metadata which spatially describes the original object audio; and
rendering the extracted object audio with the subset of original metadata based on a desired output layout into a plurality of audio channels that also characterize the original audio scene.
8 . The method of claim 7 , wherein the object audio is extracted directly from the Ambisonics audio using the subset of original metadata.
9 . The method of claim 8 wherein extracting the object audio includes reconstructing a quantized version of the Ambisonics audio and extracting the object audio from the quantized version of the Ambisonics audio.
10 . The method of claim 8 wherein extracting the object audio includes reconstructing a quantized version of the subset of original metadata and extracting the object audio from the Ambisonics audio using the quantized version of the subset of original metadata.
11 . The method of claim 8 wherein rendering the extracted object audio with the subset of original metadata includes rendering a quantized version of the extracted object audio based on a quantized version of the subset of original metadata and the desired output layout.
12 . The method of claim 7 wherein the Ambisonics audio is obtained, when decoding the first bit stream, as a time domain Ambisonics audio.
13 . The method of claim 7 wherein the subset of original metadata includes at least one of a direction of an object relative to a listening position, a distance of the object relative to the listening position, or position of the object relative to the listening position.
14 . The method of claim 7 wherein the subset of original metadata was used by an encode side process to convert the original object audio into the Ambisonics audio, before encoding the Ambisonics audio in the first bit stream.
15 . A computer-implemented method, comprising:
converting first object audio to time-frequency domain Ambisonics audio based on metadata that spatially describes the first object audio, wherein the first object audio is associated with a first priority;
converting second object audio to time domain Ambisonics audio wherein the second object audio is associated with a second priority that is different from the first priority;
encoding the time-frequency domain Ambisonics audio as a first bit stream;
encoding the metadata as a second bit stream; and
encoding the time domain Ambisonics audio as a third bit stream.
16 . The method of claim 15 wherein the first priority has a higher priority than the second priority.
17 . The method of claim 16 wherein the time domain Ambisonics audio is encoded with a lower resolution than the time-frequency domain Ambisonics audio.
18 . The method of claim 17 wherein the time-frequency domain Ambisonics audio includes a plurality of time-frequency tiles, each tile of the plurality of time-frequency tiles representing audio in a sub-band of an Ambisonics component.