Masking zone in metadata for spatial audio rendering
The various aspects of the disclosure here enable a content creation side to control how a sound program is spatial audio rendered by a decoding side, so that an audio scene component in a metadata-specified three dimensional acoustic masking zone is not heard while another audio scene component in an un-masked zone of the sound program is heard by a listener of the playback. Other aspects are also described and claimed.
1 . An encoding side method for spatial audio rendering using metadata, the method comprising:
encoding a sound program into a bitstream; and
providing metadata of the sound program, wherein the metadata specifies a masking zone and instructs a process in a decoding side on whether to attenuate the sound program while the process is spatial audio rendering the sound program for playback so that one or more audio scene components of the sound program that are located in the masking zone are not heard while one or more other audio scene components of the sound program that are located outside the masking zone in an un-masked zone are heard by a listener of the playback,
wherein the metadata further specifies a transition zone that is abutting the masking zone and is in between the masking zone and the un-masked zone, in response to which the process in the decoding side is to perform a gain transition when rendering the sound program by applying a low gain to the masking zone, a gradual change of gain in the transition zone, wherein the gradual change of gain is greater than the low gain, and a high gain to the un-masked zone, wherein the high gain is greater than the gradual change of gain.
2 . The method of claim 1 wherein the metadata specifies the masking zone as an index to a dictionary of pre-defined geometrical three dimensional zones, wherein the decoding side is to perform a lookup into the dictionary using the index to obtain one of the pre-defined geometrical three dimensional zones.
3 . The method of claim 2 wherein each of the pre-defined geometrical three dimensional zones comprise a range of Cartesian coordinates in 3D space or a range of Polar coordinates in 3D space.
4 . The method of claim 1 wherein the metadata specifies a distance length of the transition zone.
5 . The method of claim 1 wherein the metadata further specifies that a fade-in of an original audio scene gain starts outside of the masking zone and the fade-in ends inside the masking zone.
6 . The method of claim 1 wherein the metadata specifies the masking zone directly as a geometrical three dimensional zone comprising a range of Cartesian coordinates in 3D space or a range of Polar coordinates in 3D space.
7 . The method of claim 1 wherein the one or more audio scene components of the sound program that are located in the masking zone are in or have a channel-based representation.
8 . The method of claim 1 wherein the one or more audio scene components of the sound program that are located in the masking zone are in a higher order ambisonics representation, an HOA representation.
9 . The method of claim 1 wherein the one or more audio scene components of the sound program that are located in the masking zone are in or have an object representation.
10 . The method of claim 1 wherein the metadata specifies a country or a product market for the masking zone.
11 . An article of manufacture comprising a non-transitory machine-readable having contained therein instructions that when executed by a processor:
encode a sound program into a bitstream; and
provide metadata of the sound program, wherein the metadata specifies a masking zone and instructs a process in a decoding side on whether to attenuate the sound program while the process is spatial audio rendering the sound program for playback so that one or more audio scene components of the sound program that are located in the masking zone are not heard while one or more other audio scene components of the sound program that are located outside the masking zone in an un-masked zone are heard by a listener of the playback,
wherein the metadata further specifies a transition zone that is abutting the masking zone and is in between the masking zone and the un-masked zone, in response to which the process in the decoding side is to perform a gain transition when rendering the sound program by applying a low gain to the masking zone, a gradual change of gain in the transition zone, wherein the gradual change of gain is greater than the low gain, and a high gain to the un-masked zone, wherein the high gain is greater than the gradual change of gain.
12 . The article of manufacture of claim 11 wherein the metadata specifies the masking zone as an index to a dictionary of pre-defined geometrical three dimensional zones, wherein the decoding side is to perform a lookup into the dictionary using the index to obtain one of the pre-defined geometrical three dimensional zones.
13 . The article of manufacture of claim 12 wherein each of the pre-defined geometrical three dimensional zones comprise a range of Cartesian coordinates in 3D space or a range of Polar coordinates in 3D space.
14 . The article of manufacture of claim 11 wherein the metadata specifies a distance length of the transition zone.
15 . The article of manufacture of claim 11 wherein the metadata specifies the masking zone directly as a geometrical three dimensional zone comprising a range of Cartesian coordinates in 3D space or a range of Polar coordinates in 3D space.
16 . The article of manufacture of claim 11 wherein the one or more audio scene components of the sound program that are located in the masking zone are in or have a channel-based representation.
17 . The article of manufacture of claim 11 wherein the one or more audio scene components of the sound program that are located in the masking zone are in a higher order ambisonics representation, an HOA representation.
18 . The article of manufacture of claim 11 wherein the one or more audio scene components of the sound program that are located in the masking zone are in or have an object representation.