Audio rendering system and method and electronic device
The present invention relates to an audio rendering system and method and an electronic apparatus. The audio rendering system comprises: an audio signal encoding module, configured to for an audio signal in a specific audio content format, performing spatial encoding on the audio signal in the specific audio content format on the basis of metadata related information associated with the audio signal in the specific audio content format to obtain an encoded audio signal; and an audio signal decoding module, configured to performing spatial decoding on the encoded audio signal to obtain a decoded audio signal for audio rendering.
1 . An audio rendering system, comprising:
an audio information processing module configured to acquire relevant parameters of an audio signal in a specific audio content format based on metadata associated with the audio signal in the specific audio content format, wherein the audio signal in the specific audio content format comprises at least one of an object-based audio representation signal, a scene-based audio representation signal, or a channel-based audio representation signal;
an audio signal spatial encoding module configured to spatially encode the audio signal in the specific audio content format based on the relevant parameters of the audio signal in the specific audio content format to obtain a spatially encoded audio signal in a common spatial format, wherein the spatially encoded audio signal is an Ambisonics type of audio signal,
wherein the audio signal spatial encoding module is further configured to in response to the audio signal in the specific audio content format comprising the object-based audio representation signal, spatially encode the object-based audio representation signal based on spatial attribute information in relevant parameters of the object-based audio representation signal, wherein the spatial attribute information of the object-based audio representation signal comprises information related to a spatial propagation path of a sound object of the audio signal to a listener, which comprises at least one of propagation duration, propagation distance, azimuth information, path energy intensity and nodes along the way of the spatial propagation path of the sound object to the listener; and
an audio signal decoding module configured to spatially decode the encoded audio signal to obtain a decoded audio signal for audio rendering.
2 . The audio rendering system of claim 1 ,
wherein the spatially encoded audio signal comprises at least one of First Order Ambisonics (FOA), Higher Order Ambisonics (HOA) and Mixed-Order Ambisonics (MOA).
3 . The audio rendering system of claim 1 ,
wherein the audio signal encoding module is further configured to, in response to the audio signal in the specific audio content format comprising an object-based audio representation signal, acquire a reverberation relevant signal of the object-based audio representation signal based on reverberation parameters in the relevant parameters of the object-based audio representation signal, or,
wherein the audio signal encoding module is further configured to, in response to the audio signal in the specific audio content format comprising a scene-based audio representation signal, weight the scene-based audio representation signal based on weight information in the relevant parameters of the scene-based audio representation signal, or
wherein the audio signal encoding module is further configured to, in response to the audio signal in the specific audio content format comprising a scene-based audio representation signal, perform a sound field rotation operation on the scene-based audio representation signal based on the rotation information indicated in the relevant parameters of the scene-based audio representation signal, or,
wherein the audio signal encoding module is further configured to, in response to the audio signal in the specific audio content format comprising a specific type of channel signal in the channel-based audio representation signal, convert the specific type of channel signal into an object-based audio representation signal and then encode it, or,
wherein the audio signal encoding module is further configured to, in response to the audio signal in the specific audio content format comprising a specific type of channel signal in the channel-based audio representation signal, split the specific type of channel signal into audio elements by channel and convert them into metadata for encoding.
4 . The audio rendering system of claim 1 , wherein the audio signal decoding module is further configured to spatially decode an audio signal that has not spatially encoded, wherein the audio signal that has not spatially encoded comprises at least one of a scene-based audio representation signal, a specific type of channel signal in a channel-based audio representation signal, and a reverberated audio signal.
5 . The audio rendering system of claim 1 , wherein the audio signal decoding module is configured to, in a speaker playback mode, spatially decode the audio signal to be decoded by using a decoding matrix corresponding to speaker configuration,
wherein the decoding matrix comprises at least one of the following:
in response to playbacking by a predetermined speaker array, the decoding matrix is a decoding matrix built in the audio rendering system or the audio signal decoding module or a decoding matrix received from the outside and corresponding to the predetermined speaker array, or
in response to playbacking by a custom speaker array, the decoding matrix is calculated according to arrangement manner of the custom speaker array.
6 . The audio rendering system of claim 5 , wherein the decoding matrix is calculated according to azimuth angle and pitch angle of each speaker in the speaker array or three-dimensional coordinate values of the speaker.
7 . The audio rendering system of claim 1 , wherein the audio signal decoding module is configured to, in a binaural playback mode, directly decode an audio signal into a binaural signal as a decoded audio signal or perform speaker virtualization to obtain a decoded signal as a decoded audio signal, or,
wherein the audio signal decoding module is configured to, a binaural playback mode, convert the audio signal to be decoded by using a rotation matrix based on the listener's posture, and perform frequency domain convolution on each signal channel to obtain a decoded audio signal, or,
wherein the audio signal decoding module is configured to perform a sound field rotation operation on the audio signal based on rotation information in the relevant parameters.
8 . The audio rendering system of claim 1 , further comprising a signal post-processing module configured to post-process the decoded audio signal.
9 . An audio rendering method, comprising:
an audio information processing step of acquiring relevant parameters of an audio signal in a specific audio content format based on metadata associated with the audio signal in the specific audio content format, wherein the audio signal in the specific audio content format comprises at least one of an object-based audio representation signal, a scene-based audio representation signal, or a channel-based audio representation signal;
an audio signal encoding step of spatially encoding the audio signal in the specific audio content format based on the relevant parameters of the audio signal in the specific audio content format to obtain a spatially encoded audio signal in a common spatial format, wherein the spatially encoded audio signal is an Ambisonics type of audio signal,
wherein the audio signal encoding step further comprises, in response to the audio signal in the specific audio content format comprising the object-based audio representation signal, spatially encoding the object-based audio representation signal based on spatial attribute information in relevant parameters of the object-based audio representation signal, wherein the spatial attribute information of the object-based audio representation signal comprises information related to a spatial propagation path of a sound object of the audio signal to a listener, which comprises at least one of propagation duration, propagation distance, azimuth information, path energy intensity and nodes along the way of the spatial propagation path of the sound object to the listener; and
an audio signal decoding step of spatially decoding the encoded audio signal to obtain a decoded audio signal for audio rendering.
10 . The audio rendering method of claim 9 ,
wherein the spatially encoded audio signal comprises at least one of First Order Ambisonics (FOA), Higher Order Ambisonics (HOA) and Mixed-Order Ambisonics (MOA).
11 . The audio rendering method of claim 9 , wherein,
wherein the audio signal encoding step further comprises in response to the audio signal in the specific audio content format comprises an object-based audio representation signal, acquiring a reverberation relevant signal of the object-based audio representation signal based on reverberation parameters in the relevant parameters of the object-based audio representation signal, or,
wherein the audio signal encoding step further comprises, in response to the audio signal in the specific audio content format comprises a scene-based audio representation signal, weighting the scene-based audio representation signal based on weight information in the relevant parameters of the scene-based audio representation signal, or,
wherein the audio signal encoding step further comprises, in response to the audio signal in the specific audio content format comprises a scene-based audio representation signal, performing a sound field rotation operation on the scene-based audio representation signal based on the rotation information indicated in the relevant parameters of the scene-based audio representation signal, or,
wherein the audio signal encoding step further comprises, in response to the audio signal in the specific audio content format comprises a specific type of channel signal in the channel-based audio representation signal, converting the specific type of channel signal into an object-based audio representation signal and then encoding it, or,
wherein the audio signal encoding step further comprises, in response to the audio signal in the specific audio content format comprises a specific type of channel signal in the channel-based audio representation signal, splitting the specific type of channel signal into audio elements by channel and converting them into metadata for encoding.
12 . The audio rendering method of claim 9 , wherein the audio signal decoding step further comprises spatially decoding an audio signal that has not spatially encoded, wherein the audio signal that has not spatially encoded comprises at least one of a scene-based audio representation signal, a specific type of channel signal in a channel-based audio representation signal, and a reverberated audio signal.
13 . The audio rendering method of claim 9 , wherein the audio signal decoding step further comprises, in a speaker playback mode, spatially decoding the audio signal to be decoded by using a decoding matrix corresponding to speaker configuration,
wherein, the decoding matrix comprises at least one of the following:
in response to playbacking by a predetermined speaker array, the decoding matrix is a decoding matrix built in an audio rendering system or audio signal decoding module or a decoding matrix received from the outside and corresponding to the predetermined speaker array, or
in response to playbacking by a custom speaker array, the decoding matrix is calculated according to arrangement manner of the custom speaker array.
14 . The audio rendering method of claim 13 , wherein the decoding matrix is calculated according to azimuth angle and pitch angle of each speaker in the speaker array or three-dimensional coordinate values of the speaker.
15 . The audio rendering method of claim 9 , wherein the audio signal decoding step further comprises, in a binaural playback mode, directly decoding an audio signal into a binaural signal as a decoded audio signal or performing speaker virtualization to obtain a decoded signal as a decoded audio signal, or,
wherein the audio signal decoding step further comprises, in a binaural playback mode, converting the audio signal to be decoded by using a rotation matrix based on the listener's posture, and performing frequency domain convolution on each signal channel to obtain a decoded audio signal, or,
wherein the audio signal decoding step further comprises performing a sound field rotation operation on the audio signal based on rotation information in the relevant parameters.
16 . The audio rendering method of claim 9 , further comprising a signal post-processing step of post-processing the decoded audio signal.
17 . An electronic apparatus, comprising:
a memory, and
a processor coupled to the memory; the processor is configured to, based on instructions stored in the memory, execute the following steps:
an audio information processing step of acquiring relevant parameters of an audio signal in a specific audio content format based on metadata associated with the audio signal in the specific audio content format, wherein the audio signal in the specific audio content format comprises at least one of an object-based audio representation signal, a scene-based audio representation signal, or a channel-based audio representation signal;
an audio signal encoding step of spatially encoding the audio signal in the specific audio content format based on the relevant parameters of the audio signal in the specific audio content format to obtain a spatially encoded audio signal in a common spatial format, wherein the spatially encoded audio signal is an Ambisonics type of audio signal,
wherein the audio signal encoding step further comprises, in response to the audio signal in the specific audio content format comprising the object-based audio representation signal, spatially encoding the object-based audio representation signal based on spatial attribute information in relevant parameters of the object-based audio representation signal, wherein the spatial attribute information of the object-based audio representation signal comprises information related to a spatial propagation path of a sound object of the audio signal to a listener, which comprises at least one of propagation duration, propagation distance, azimuth information, path energy intensity and nodes along the way of the spatial propagation path of the sound object to the listener; and
an audio signal decoding step of spatially decoding the encoded audio signal to obtain a decoded audio signal for audio rendering.
18 . The electronic apparatus of claim 17 ,
wherein the spatially encoded audio signal comprises at least one of First Order Ambisonics (FOA), Higher Order Ambisonics (HOA) and Mixed-Order Ambisonics (MOA).