IP Library Granted Patent US 10,468,046
Granted Patent B2
US 10,468,046 · App. 16/039,110 · Granted Nov 5, 2019

Coding mode determination method and apparatus, audio encoding method and apparatus, and audio decoding method and apparatus

Inventors: Ki-hyun Choo (Seoul, KR); Anton Victorovich Porov (Saint-Petersburg, RU); Konstantin Sergeevich Osipov (Saint-Petersburg, RU); Nam-suk Lee (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L19/12G10L19/22G10L19/00G10L19/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,468,046
App. No.
16/039,110
Granted
Nov 5, 2019
Kind
B2
Abstract

Provided are a method and an apparatus for determining an encoding mode for improving the quality of a reconstructed audio signal. A method of determining an encoding mode includes determining one from among a plurality of encoding modes including a first encoding mode and a second encoding mode as an initial encoding mode in correspondence to characteristics of an audio signal, and if there is an error in the determination of the initial encoding mode, generating a modified encoding mode by modifying the initial encoding mode to a third encoding mode.

Claims (23)

1. A method of encoding an audio signal, the method comprising:

receiving the audio signal;

obtaining, performed by at least one processor, first parameters of a current frame of the audio signal;

selecting, performed by the at least one processor, a class of the current frame in the audio signal from among a plurality of classes including a music class and a speech class, based on first parameters of the current frame by using a Gaussian mixture model (GMM);

obtaining second parameters including first tonality, second tonality and third tonality;

generating a plurality of conditions, where each of the plurality of conditions is generated based on a combination of the obtained second parameters;

determining, performed by the at least one processor, whether an error occurs in the selected class of the current frame based on whether at least one of the plurality of conditions is met;

when the error occurs in the selected class of the current frame, correcting, performed by the at least one processor, the selected class of the current frame;

encoding, performed by the at least one processor, the current frame, based on either the corrected class or the selected class of the current frame; and

generating a bitstream based on the encoded current frame,

wherein the first tonality is obtained from a subband of 0 to 1 kHz, the second tonality is obtained from a subband of 1 to 2 kHz and the third tonality is obtained from a subband of 2 to 4 kHz, and

wherein the correcting comprises:

when the error occurs in the selected class of the current frame and the selected class of the current frame is the speech class, correcting the selected class of the current frame from the speech class to the music class; and

when the error occurs in the selected class of the current frame and the selected class of the current frame is the music class, correcting the selected class of the current frame from the music class to the speech class.

2. The method of claim 1 , wherein the correcting is performed based on at least two independent states.

3. The method of claim 1 , wherein the second parameters further comprise a difference between a voicing parameter and a correlation parameter.

4. The method of claim 1 , wherein the determining of whether the error occurs in the selected class of the current frame occurs comprises:

determining whether the current frame has speech characteristics when the current frame is classified as the music class; and

determining whether the current frame has music characteristics when the current frame is classified as the speech class.

5. The method of claim 1 , wherein the correcting comprises:

correcting a classification of the current frame, when the current frame is classified as the music class and has speech characteristics; and

correcting the classification of the current frame, when the current frame is classified as the speech class and has music characteristics.

6. The method of claim 1 , wherein the determining is performed further based on a hangover parameter which is used to prevent frequent switching between coding modes.

Continuity (3)
Continuation 14079090 · Nov 13, 2013
Provisional Application 61725694 · Nov 13, 2012
Related Publication 20180322887A1 · Nov 8, 2018