Signal classifying method and device, and audio encoding method and device using same
The present invention relates to an audio encoding and, more particularly, to a signal classifying method and device, and an audio encoding method and device using the same, which can reduce a delay caused by an encoding mode switching while improving the quality of reconstructed sound. The signal classifying method may comprise the operations of: classifying a current frame into one of a speech signal and a music signal; determining, on the basis of a characteristic parameter obtained from multiple frames, whether a result of the classifying of the current frame includes an error; and correcting the result of the classifying of the current frame in accordance with a result of the determination. By correcting an initial classification result of an audio signal on the basis of a correction parameter, the present invention can determine an optimum coding mode for the characteristic of an audio signal and can prevent frequent coding mode switching between frames.
1. A signal classification method in an encoding device for an audio signal, the signal classification method comprising:
classifying a current frame as one from among a plurality of classes including a speech class and a music class, based on a signal characteristic of an audio signal;
evaluating a condition, based on one or more parameter among a plurality of parameters, wherein the plurality of parameters include a parameter obtained from a plurality of frames;
first determining whether the condition corresponds to a first threshold value;
second determining whether a hangover parameter corresponds to a second threshold value; and
correcting a classification result of the current frame, based on a first result of the first determining and a second result of the second determining,
wherein the plurality of parameters include tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.
2. The signal classification method of claim 1 , wherein the second plurality of signal characteristics are obtained from the current frame and a plurality of previous frames.
3. The signal classification method of claim 1 , wherein the hangover parameter is used to prevent frequent transitions between states.
4. The signal classification method of claim 1 , wherein the correcting comprises correcting the classification result of the current frame from the music class to the speech class when some of the plurality of conditions are satisfied and a first hangover parameter reaches a reference value.
5. The signal classification method of claim 1 , wherein the correcting comprises correcting the classification result of the current frame from the speech class to the music class when some of the plurality of conditions are satisfied and a second hangover parameter reaches a reference value.
6. An audio encoding method in an encoding device for an audio signal, the audio encoding method comprising:
classifying, performed by at least one processor, a current frame as one from among a plurality of classes including a speech class and a music class, based on a signal characteristic of an audio signal;
evaluating a condition, based on one or more parameter among a plurality of parameters, wherein the plurality of parameters include a parameter obtained from a plurality of frames;
first determining whether one of the plurality of conditions corresponds to a first threshold value;
second determining whether a hangover parameter corresponds to a second threshold value; and
correcting a classification result of the current frame, based on a first result of the first determining and a second result of the second determining; and
encoding the current frame based on the classification result or the corrected classification result,
wherein the plurality of parameters include tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.
7. The audio encoding method of claim 6 , wherein the encoding is performed using one of a CELP-type coder, a transform coder and a CELP/transform hybrid coder.