Method and apparatus to determine encoding mode of audio signal and method and apparatus to encode and/or decode audio signal using the encoding mode determination method and apparatus
A method and apparatus to determine an encoding mode of an audio signal, and a method and apparatus to encode an audio signal according to the encoding mode. In the encoding mode determination method, a mode determination threshold for the current frame that is subject to encoding mode determination is adaptively adjusted according to a long-term feature of the audio signal for a frame (the current frame) that is subject to encoding mode determination, thereby improving the hit rate of encoding mode determination and signal classification, suppressing frequent oscillation of an encoding mode in frame units, improving noise tolerance, and improving smoothness of a reconstructed audio signal.
1. An apparatus to determine an encoding mode to encode an audio signal, comprising:
at least one processor; and
a determining unit to generate a short-term feature and a long-term feature for a current frame, to generate a parameter by using the long-term feature, to adaptively adjust a mode determination threshold of the short-term feature in a unit of frames by using the parameter and to determine one of an audio encoding mode and a music encoding mode of the current frame based on a comparison of the short-term feature and the adjusted mode determination threshold so that the current frame of the audio signal is encoded according to the determined encoding mode.
2. The apparatus of claim 1 , further comprising:
a time-domain coding unit to encode the audio signal according to the encoding mode in a time-domain; and
a frequency-domain coding unit to encode the audio signal according to the encoding mode in a frequency-domain.
3. The apparatus of claim 1 , further comprising:
a speech coding unit to encode the audio signal as a speech signal according to the encoding mode; and
a music coding unit to encode the audio signal as a music signal according to the encoding mode.
4. The apparatus of claim 1 , further comprising:
a speech coding unit to receive the audio signal and the encoding mode from the determining unit to encode the audio signal when the encoding mode is the speech encoding mode; and
a music coding unit to receive the audio signal and the encoding mode from the determining unit to encode the audio signal when the encoding mode is music encoding mode.
5. The apparatus of claim 1 , further comprising:
a coding unit to encode the audio signal according to the encoding mode; and
a bitstream generation unit to generate a bitstream according to the encoded audio signal and information on the encoding mode.
6. The apparatus of claim 1 , wherein the determining unit comprises:
a short term feature generation unit to generate the short-term feature from the current frame of the audio signal; and
a long-term feature generation unit to generate the long-term feature from the current frame and at least one previous frame.
7. The apparatus of claim 6 , wherein the determining unit further comprises:
a mode determination threshold adjustment unit to adjust the mode determination threshold using the short term feature and the long-term feature; and
an encoding determination unit to determine the encoding mode according to the adjusted mode determination threshold and the short-term feature.
8. The apparatus of claim 6 , wherein the mode determination threshold adjustment unit adjusts the mode determination threshold using the short term feature, the long-term feature, and an encoding mode of the at least one the previous frame.
9. The apparatus of claim 6 , wherein the encoding determination unit determines the encoding mode according to the adjusted mode determination threshold, the short-term feature, and an encoding mode of the at least one previous frame.
10. The apparatus of claim 6 , wherein the long-term feature generation unit comprises:
a first long-term feature generation unit to generate a first long-term feature according to the short-term feature of the current frame and a short-term feature of the at least one previous frame; and
a second long-term feature generation unit to generate a second long-term feature as the long-term feature according to the first long-term feature and a variation feature of at least one of the current frame and the at least one previous frame.
11. The apparatus of claim 10 , wherein the determining unit further comprises:
a mode determination threshold adjustment unit to adjust a mode determination threshold according to the short term feature and the second long-term feature; and
an encoding determination unit to determine the encoding mode according to the adjusted mode determination threshold and the short-term feature.
12. The apparatus of claim 1 , wherein the determining unit determines the encoding mode of the first frame of the audio signal according to the short-term feature, the long-term feature, and an encoding mode of a previous frame.
13. The apparatus of claim 1 , wherein the determining unit comprises:
an LP-LTP gain generation unit to generate an LP-LTP gain as the short-term feature of the current frame; and
a long-term feature generation unit to generate the long-term feature according to the LP-LTP gain of the current frame and a second LP-LTP gain of at least one previous frames.
14. The apparatus of claim 1 , wherein the determining unit comprises:
a spectrum tilt generation unit to generate a spectrum tilt as the short-term feature of the current frame; and
a long-term feature generation unit to generate the long-term feature according to the spectrum tilt of the current frame and a second spectrum tilt of at least one previous frames.
15. The apparatus of claim 1 , wherein the determining unit comprises:
a zero crossing rate generation unit to generate a zero crossing rate as the short-term feature of the current frame; and
a long-term feature generation unit to generate the long-term feature according to the zero crossing rate of the current frame and a second zero crossing rate of at least one previous frames.
16. The apparatus of claim 1 , wherein the determining unit comprises:
a short-term feature generation unit having one or a combination of an LP-LTP gain generation unit to generate an LP-LTP gain as the short-term feature of the current frame, a spectrum tilt generation unit to generate a spectrum tilt as the short-term feature of the current frame, and a zero crossing rate generation unit to generate a zero crossing rate as the short-term feature of the current frame; and
a long-term feature generation unit to generate the long-term feature according to the short-term feature of the current frame and a second short-term feature of at least one frame.
17. The apparatus of claim 1 , wherein the determination unit comprises a memory to store the short-term and long-term features of the current frame and previous frames.
18. The apparatus of claim 1 , wherein:
the long-term feature is determined according to the short-term feature of the current frame and short-term features of a plurality of previous frames.
19. The apparatus of claim 1 , wherein:
the long-term feature is determined according to a variation feature between the current frame and a previous frame.
20. The apparatus of claim 1 , wherein:
the long-term feature is determined according to a variation feature of a second encoding mode of the previous frame.
21. An apparatus to decode a signal of a bitstream, comprising:
at least one processor; and
a determining unit to determine one of an audio encoding mode and a music encoding mode from a bitstream having an encoded signal and information on the audio encoding mode or the music encoding mode of the encoded signal, so that the encoded signal of the bitstream is decoded according to the audio encoding mode or the music encoding mode,
wherein the audio encoding mode or the music encoding mode of the encoded signal has been determined using a mode determination threshold of the short-term feature for a current frame which is adaptively adjusted in a unit of frames by using a parameter generated from a long-term feature for the current frame.
22. An apparatus to encode and/or decode an audio signal, comprising:
at least one processor;
a first determining unit to generate a short-term feature and a long-term feature for a current frame, to generate a parameter by using the long-term feature, to adaptively adjust a mode determination threshold of the short-term feature in a unit of frames by using the parameter and to determine one of an audio encoding mode and a music encoding mode of the current frame based on a comparison of the short-term feature and the adjusted mode determination threshold, so that the current frame of the audio signal is encoded according to the audio encoding mode or the music encoding mode; and
a second determining unit to determine one of the audio encoding mode and the music encoding mode from a bitstream having the encoded signal and information on the encoding mode, so that the encoded signal of the bitstream is decoded according to the audio encoding mode or the music encoding mode.
23. An encoding apparatus for an audio signal, comprising:
at least one processor; and
a short-term feature generation unit to analyze the audio signal for a current frame and generating a short-term feature;
a long-term feature generation unit to generate a long-term feature for the current frame using the short-term feature;
a mode determination threshold adjustment unit to adaptively adjust a mode determination threshold for the current frame using a parameter obtained from the generated long-term feature;
an encoding mode determination unit to determine one of an audio encoding mode and a speech encoding mode for the current frame using the adaptively adjusted mode determination threshold; and
an encoding unit to perform frequency-domain encoding or time-domain encoding for the current frame in the determined encoding mode.