Method and apparatus for encoding/decoding audio signal
A method and apparatus for encoding/decoding audio signal are provided. The encoding method includes transforming an input audio signal in a time domain into an audio signal in a frequency domain, quantizing energy of a frequency band of the audio signal in the frequency domain, generating a normal signal by normalizing the audio signal in the frequency domain according to quantized energy, obtaining a feature vector including information on the energy of the frequency band based on the normal signal and the input audio signal, quantizing the feature vector, obtaining a scale factor used to scale the normal signal based on the quantized feature vector, quantizing an adjustment signal into which the normal signal has been scaled based on the scale factor, and outputting bitstreams based on the quantized energy, the quantized feature vector, and the quantized adjustment signal.
1 . An encoding method, comprising:
transforming an input audio signal in a time domain into an audio signal in a frequency domain;
quantizing energy of a frequency band of the audio signal in the frequency domain;
generating a normal signal by normalizing the audio signal in the frequency domain according to quantized energy;
obtaining a feature vector including information on the energy of the frequency band based on the normal signal and the input audio signal;
quantizing the feature vector;
obtaining a scale factor used to scale the normal signal based on a quantized feature vector;
quantizing an adjustment signal into which the normal signal has been scaled based on the scale factor; and
outputting bitstreams based on the quantized energy, the quantized feature vector, and a quantized adjustment signal,
wherein the obtaining of the scale factor comprises obtaining the scale factor for each frequency band based on the quantized feature vector,
wherein a number of dimensions of the scale factor matches a total number of bands in the frequency band, and
wherein a number of dimensions of the quantized feature vector matches a number of dimensions of the feature vector.
2 . The encoding method of claim 1 , wherein the obtaining of the feature vector comprises:
obtaining a magnitude spectrum of the input audio signal in a frequency domain based on the input audio signal; and
obtaining the feature vector based on the magnitude spectrum and the normal signal.
3 . The encoding method of claim 2 , wherein the obtaining of the feature vector further comprises:
generating a latent representation for extracting the information on the energy of the frequency band based on the magnitude spectrum and the normal signal; and
calculating the feature vector based on the latent representation.
4 . The encoding method of claim 1 , wherein the quantizing of the adjustment signal comprises:
generating the adjustment signal by scaling the normal signal according to the scale factor.
5 . The encoding method of claim 1 , wherein the outputting of the bitstreams comprises:
outputting a first bitstream by encoding the quantized feature vector;
outputting a second bitstream by encoding the quantized adjustment signal; and
outputting a third bitstream by encoding the quantized energy.
6 . A decoding method, comprising:
receiving bitstreams from an encoder;
obtaining a scale factor used to inversely scale a restored adjustment signal based on a first bitstream into which a quantized feature vector is encoded;
generating a restored normal signal based on the scale factor and a second bitstream into which a quantized adjustment signal is encoded;
obtaining a restored audio signal in a frequency domain based on a third bitstream, into which quantized energy is encoded, and the restored normal signal; and
outputting a restored audio signal in a time domain based on the restored audio signal in the frequency domain,
wherein a number of dimensions of the scale factor matches a total number of bands in the frequency band, and
wherein a number of dimensions of the quantized feature vector matches a number of dimensions of the feature vector.
7 . The decoding method of claim 6 , wherein the obtaining of the scale factor comprises:
obtaining a quantized feature vector by decoding the first bitstream; and
calculating the scale factor from the quantized feature vector.
8 . The decoding method of claim 6 , wherein the generating of the restored normal signal comprises:
generating a restored adjustment signal by decoding the second bitstream; and
inversely scaling the restored adjustment signal according to the scale factor.
9 . The decoding method of claim 6 , wherein the obtaining of the restored audio signal in the frequency domain comprises:
outputting restored energy of a frequency band of the restored audio signal in the frequency domain, based on the third bitstream; and
denormalizing the restored normal signal according to the restored energy.
10 . An encoding device, comprising:
a memory configured to store one or more instructions; and
a processor configured to execute the instructions,
wherein, when the instructions are executed, the processor is configured to perform a plurality of operations, and
wherein the plurality of operations comprises:
transforming an input audio signal in a time domain into an audio signal in a frequency domain;
quantizing energy of a frequency band of the audio signal in the frequency domain;
generating a normal signal by normalizing the audio signal in the frequency domain according to quantized energy;
obtaining a feature vector including information on the energy of the frequency band based on the normal signal and the input audio signal;
quantizing the feature vector;
obtaining a scale factor used to scale the normal signal based on a quantized feature vector;
quantizing an adjustment signal into which the normal signal has been scaled based on the scale factor; and
outputting bitstreams based on the quantized energy, the quantized feature vector, and a quantized adjustment signal,
wherein the obtaining of the scale factor comprises obtaining the scale factor for each frequency band based on the quantized feature vector,
wherein a number of dimensions of the scale factor matches a total number of bands in the frequency band, and
wherein a number of dimensions of the quantized feature vector matches a number of dimensions of the feature vector.
11 . The encoding device of claim 10 , wherein the obtaining of the feature vector comprises:
obtaining a magnitude spectrum of the input audio signal in a frequency domain based on the input audio signal; and
obtaining the feature vector based on the magnitude spectrum and the normal signal.
12 . The encoding device of claim 11 , wherein the obtaining of the feature vector further comprises:
generating a latent representation for extracting the information on the energy of the frequency band based on the magnitude spectrum and the normal signal; and
calculating the feature vector based on the latent representation.
13 . The encoding device of claim 10 , wherein the quantizing of the adjustment signal comprises:
generating the adjustment signal by scaling the normal signal according to the scale factor.
14 . The encoding device of claim 10 , wherein the outputting of the bitstreams comprises:
outputting a first bitstream by encoding the quantized feature vector;
outputting a second bitstream by encoding the quantized adjustment signal; and
outputting a third bitstream by encoding the quantized energy.