IP Library Granted Patent US 12706107
Granted Patent B2
US 12706107 · App. 18/653,233 · Granted Aug 11, 2026

Method and apparatus for encoding/decoding audio signal

Inventors: Inseon Jang (Daejeon, KR); Seung Kwon Beack (Daejeon, KR); Jongmo Sung (Daejeon, KR); Tae Jin Lee (Daejeon, KR); Woo-taek Lim (Daejeon, KR); Byeongho Cho (Daejeon, KR); Hong-Goo Kang (Seoul, KR); Byeong Hyeon Kim (Seoul, KR); Jihyun Lee (Seoul, KR); Hyungseob Lim (Seoul, KR)
Assignees: Electronics and Telecommunications Research Institute; UIF (University Industry Foundation), Yonsei University
G10L19/038G10L19/0204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12706107
App. No.
18/653,233
Granted
Aug 11, 2026
Kind
B2
Abstract

A method and apparatus for encoding/decoding audio signal are provided. The encoding method includes transforming an input audio signal in a time domain into an audio signal in a frequency domain, quantizing energy of a frequency band of the audio signal in the frequency domain, generating a normal signal by normalizing the audio signal in the frequency domain according to quantized energy, obtaining a feature vector including information on the energy of the frequency band based on the normal signal and the input audio signal, quantizing the feature vector, obtaining a scale factor used to scale the normal signal based on the quantized feature vector, quantizing an adjustment signal into which the normal signal has been scaled based on the scale factor, and outputting bitstreams based on the quantized energy, the quantized feature vector, and the quantized adjustment signal.

Claims (69)

1 . An encoding method, comprising:

transforming an input audio signal in a time domain into an audio signal in a frequency domain;

quantizing energy of a frequency band of the audio signal in the frequency domain;

generating a normal signal by normalizing the audio signal in the frequency domain according to quantized energy;

obtaining a feature vector including information on the energy of the frequency band based on the normal signal and the input audio signal;

quantizing the feature vector;

obtaining a scale factor used to scale the normal signal based on a quantized feature vector;

quantizing an adjustment signal into which the normal signal has been scaled based on the scale factor; and

outputting bitstreams based on the quantized energy, the quantized feature vector, and a quantized adjustment signal,

wherein the obtaining of the scale factor comprises obtaining the scale factor for each frequency band based on the quantized feature vector,

wherein a number of dimensions of the scale factor matches a total number of bands in the frequency band, and

wherein a number of dimensions of the quantized feature vector matches a number of dimensions of the feature vector.

2 . The encoding method of claim 1 , wherein the obtaining of the feature vector comprises:

obtaining a magnitude spectrum of the input audio signal in a frequency domain based on the input audio signal; and

obtaining the feature vector based on the magnitude spectrum and the normal signal.

3 . The encoding method of claim 2 , wherein the obtaining of the feature vector further comprises:

generating a latent representation for extracting the information on the energy of the frequency band based on the magnitude spectrum and the normal signal; and

calculating the feature vector based on the latent representation.

4 . The encoding method of claim 1 , wherein the quantizing of the adjustment signal comprises:

generating the adjustment signal by scaling the normal signal according to the scale factor.

5 . The encoding method of claim 1 , wherein the outputting of the bitstreams comprises:

outputting a first bitstream by encoding the quantized feature vector;

outputting a second bitstream by encoding the quantized adjustment signal; and

outputting a third bitstream by encoding the quantized energy.

6 . A decoding method, comprising:

receiving bitstreams from an encoder;

obtaining a scale factor used to inversely scale a restored adjustment signal based on a first bitstream into which a quantized feature vector is encoded;

generating a restored normal signal based on the scale factor and a second bitstream into which a quantized adjustment signal is encoded;

obtaining a restored audio signal in a frequency domain based on a third bitstream, into which quantized energy is encoded, and the restored normal signal; and

outputting a restored audio signal in a time domain based on the restored audio signal in the frequency domain,

wherein a number of dimensions of the scale factor matches a total number of bands in the frequency band, and

wherein a number of dimensions of the quantized feature vector matches a number of dimensions of the feature vector.

7 . The decoding method of claim 6 , wherein the obtaining of the scale factor comprises:

obtaining a quantized feature vector by decoding the first bitstream; and

calculating the scale factor from the quantized feature vector.

8 . The decoding method of claim 6 , wherein the generating of the restored normal signal comprises:

generating a restored adjustment signal by decoding the second bitstream; and

inversely scaling the restored adjustment signal according to the scale factor.

9 . The decoding method of claim 6 , wherein the obtaining of the restored audio signal in the frequency domain comprises:

outputting restored energy of a frequency band of the restored audio signal in the frequency domain, based on the third bitstream; and

denormalizing the restored normal signal according to the restored energy.

10 . An encoding device, comprising:

a memory configured to store one or more instructions; and

a processor configured to execute the instructions,

wherein, when the instructions are executed, the processor is configured to perform a plurality of operations, and

wherein the plurality of operations comprises:

transforming an input audio signal in a time domain into an audio signal in a frequency domain;

quantizing energy of a frequency band of the audio signal in the frequency domain;

generating a normal signal by normalizing the audio signal in the frequency domain according to quantized energy;

obtaining a feature vector including information on the energy of the frequency band based on the normal signal and the input audio signal;

quantizing the feature vector;

obtaining a scale factor used to scale the normal signal based on a quantized feature vector;

quantizing an adjustment signal into which the normal signal has been scaled based on the scale factor; and

outputting bitstreams based on the quantized energy, the quantized feature vector, and a quantized adjustment signal,

wherein the obtaining of the scale factor comprises obtaining the scale factor for each frequency band based on the quantized feature vector,

wherein a number of dimensions of the scale factor matches a total number of bands in the frequency band, and

wherein a number of dimensions of the quantized feature vector matches a number of dimensions of the feature vector.

11 . The encoding device of claim 10 , wherein the obtaining of the feature vector comprises:

obtaining a magnitude spectrum of the input audio signal in a frequency domain based on the input audio signal; and

obtaining the feature vector based on the magnitude spectrum and the normal signal.

12 . The encoding device of claim 11 , wherein the obtaining of the feature vector further comprises:

generating a latent representation for extracting the information on the energy of the frequency band based on the magnitude spectrum and the normal signal; and

calculating the feature vector based on the latent representation.

13 . The encoding device of claim 10 , wherein the quantizing of the adjustment signal comprises:

generating the adjustment signal by scaling the normal signal according to the scale factor.

14 . The encoding device of claim 10 , wherein the outputting of the bitstreams comprises:

outputting a first bitstream by encoding the quantized feature vector;

outputting a second bitstream by encoding the quantized adjustment signal; and

outputting a third bitstream by encoding the quantized energy.