IP Library Granted Patent US 10,090,004
Granted Patent B2
US 10,090,004 · App. 15/121,257 · Granted Oct 2, 2018

Signal classifying method and device, and audio encoding method and device using same

Inventors: Ki-hyun Choo (Seoul, KR); Anton Viktorovich Porov (Saint-Petersburg, RU); Konstantin Sergeevich Osipov (Moscow, RU)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L25/81G10L19/005G10L19/022G10L19/0212G10L19/125G10L19/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,090,004
App. No.
15/121,257
Granted
Oct 2, 2018
Kind
B2
Abstract

The present invention relates to an audio encoding and, more particularly, to a signal classifying method and device, and an audio encoding method and device using the same, which can reduce a delay caused by an encoding mode switching while improving the quality of reconstructed sound. The signal classifying method may comprise the operations of: classifying a current frame into one of a speech signal and a music signal; determining, on the basis of a characteristic parameter obtained from multiple frames, whether a result of the classifying of the current frame includes an error; and correcting the result of the classifying of the current frame in accordance with a result of the determination. By correcting an initial classification result of an audio signal on the basis of a correction parameter, the present invention can determine an optimum coding mode for the characteristic of an audio signal and can prevent frequent coding mode switching between frames.

Claims (30)

1. A signal classification method in an encoding device, the signal classification method comprising:

classifying, performed by at least one processor, a current frame as one from among a plurality of classes including a speech class and a music class, based on a first plurality of signal characteristics;

generating a plurality of conditions, based on one or more of a second plurality of signal characteristics obtained from a plurality of frames including the current frame;

first comparing one of the plurality of conditions with a first threshold value and second comparing a hangover parameter with a second threshold value; and

correcting a classification result of the current frame, based on a result of the first comparing and second comparing,

wherein the second plurality of signal characteristics includes tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.

2. The signal classification method of claim 1 , wherein the second plurality of signal characteristics are obtained from the current frame and a plurality of previous frames.

3. The signal classification method of claim 1 , wherein the hangover parameter is used to prevent frequent transitions between states.

4. The signal classification method of claim 1 , wherein the correcting comprises correcting the classification result of the current frame from the music class to the speech class when some of the plurality of conditions are satisfied and a first hangover parameter reaches a reference value.

5. The signal classification method of claim 1 , wherein the correcting comprises correcting the classification result of the current frame from the speech class to the music class when some of the plurality of conditions are satisfied and a second hangover parameter reaches a reference value.

6. A non-transitory computer-readable recording medium having recorded thereon a program for executing:

classifying a current frame as one from among a plurality of classes including a speech class and a music class, based on a first plurality of signal characteristics;

generating a plurality of conditions, based on one or more of a second plurality of signal characteristics obtained from a plurality of frames including the current frame;

first comparing one of the plurality of conditions with a first threshold value and second comparing a hangover parameter with a second threshold value; and

correcting a classification result of the current frame, based on a result of the first comparing and second comparing,

wherein the second plurality of signal characteristics includes tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.

7. An audio encoding method in an encoding device, the audio encoding method comprising:

classifying, performed by at least one processor, a current frame as one from among a plurality of classes including a speech class and a music class, based on a first plurality of signal characteristics;

generating a plurality of conditions, based on a second plurality of signal characteristics obtained from a plurality of frames including the current frame;

first comparing one of the plurality of conditions with a first threshold value and second comparing a hangover parameter with a second threshold value;

correcting a classification result of the current frame, based on a result of the first comparing and second comparing; and

encoding the current frame based on the classification result or the corrected classification result,

wherein the second plurality of signal characteristics includes tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.

8. The audio encoding method of claim 7 , wherein the encoding is performed using one of a CELP-type coder and a transform coder.

9. The audio encoding method of claim 8 , wherein the encoding is performed using one of the CELP-type coder, the transform coder and a CELP/transform hybrid coder.

10. A signal classification apparatus implemented in an encoding device, the signal classification apparatus comprising at least one processor configured to:

classify a current frame as one from among a plurality of classes including a speech class and a music class, based on a first plurality of signal characteristics, generate a plurality of conditions, based on one or more of a second plurality of signal characteristics obtained from a plurality of frames including the current frame, first compare one of the plurality of conditions with a first threshold value, second compare a hangover parameter with a second threshold value and correct a classification result of the current frame, based on a result of the first comparing and second comparing, wherein the second plurality of signal characteristics includes tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.

11. An audio encoding apparatus implemented in an encoding device, the audio encoding apparatus comprising at least one processor configured to:

classify a current frame as one from among a plurality of classes including a speech class and a music class, based on a first plurality of signal characteristics, generate a plurality of conditions, based on one or more of a second plurality of signal characteristics obtained from a plurality of frames including the current frame, first compare one of the plurality of conditions with a first threshold value, second compare a hangover parameter with a second threshold value, correct a classification result of the current frame, based on a result of the first comparing and second comparing, and encode the current frame based on the classification result or the corrected classification result,

wherein the second plurality of signal characteristics includes tonalities in a plurality of frequency regions, a long term tonality in a low band, a difference between the tonalities in the plurality of frequency regions, a linear prediction error, and a difference between a scaled voicing feature and a scaled correlation map feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2016
From: CHOO, KI-HYUN; POROV, ANTON VIKTOROVICH; OSIPOV, KONSTANTIN SERGEEVICH
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 039529/0133 →
Continuity (3)
Provisional Application 62029672 · Jul 28, 2014
Provisional Application 61943638 · Feb 24, 2014
Related Publication 20170011754A1 · Jan 12, 2017
Cited By (1)
US 12,512,107