IP Library Granted Patent US 10,529,361
Granted Patent B2
US 10,529,361 · App. 16/108,668 · Granted Jan 7, 2020

Audio signal classification method and apparatus

Inventor: Zhe Wang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L25/81G10L19/06G10L19/12G10L25/18G10L25/78G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,529,361
App. No.
16/108,668
Granted
Jan 7, 2020
Kind
B2
Abstract

An audio signal classification method and apparatus includes determining, according to voice activity of a current audio frame, whether to obtain a frequency spectrum fluctuation of the current audio frame and store the frequency spectrum fluctuation in a frequency spectrum fluctuation memory, updating, according to whether the audio frame is percussive music or activity of a historical audio frame, the frequency spectrum fluctuations stored in the frequency spectrum fluctuation memory, and classifying the current audio frame as a speech frame or a music frame according to statistics of a part or all of effective data of the frequency spectrum fluctuations that is stored in the frequency spectrum fluctuation memory.

Claims (72)

1. An audio signal classification method, comprising:

storing, based on at least one condition, data of a frequency spectrum fluctuation parameter of a current audio frame of an audio signal into a memory where data of frequency spectrum fluctuation parameters of a plurality of audio frames are stored, wherein the at least one condition is the current audio frame being an active frame, wherein the frequency spectrum fluctuation parameter denotes an energy fluctuation of a frequency spectrum of the audio signal;

modifying data of frequency spectrum fluctuation parameters of audio frames preceding the current audio frame stored in the memory into ineffective data when the current audio frame is an active frame and an audio frame immediately preceding the current audio frame is an inactive frame, wherein data of frequency spectrum fluctuation parameters in the memory not having been modified into ineffective data are effective data;

modifying the effective data stored in the memory into a value that is less than or equal to a music threshold when a current signal is percussive music, wherein the current signal comprises the current audio frame and a plurality of audio frames precede the current audio frame;

obtain statistics of a part or all of the effective data stored in the memory;

classifying the current audio frame as a speech frame or a music frame according to the statistics of a part or all of the effective data stored in the memory.

2. The method of claim 1 , wherein the current audio frame and a historical frame of the current audio frame belong to a group of multiple consecutive frames, and wherein the at least one condition further comprises none of the group of multiple consecutive frames belongs to an energy attack.

3. The method of claim 1 , wherein classifying the current audio frame as the speech frame or the music frame according to statistics of the part or all of effective data comprises:

obtaining an average value of the part or all of the effective data of the frequency spectrum fluctuation parameters that are stored; and

either classifying the current audio frame as the music frame based on a condition that the average value satisfies a music classification condition or classifying the current audio frame as the speech frame based on a condition that the average value satisfies a speech classification condition.

4. The method of claim 1 , wherein classifying the current audio frame as the speech frame or the music frame comprises:

obtaining a first group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame;

obtaining a second group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame, wherein a quantity of data in the first group and a quantity of data in the second group are different;

obtaining a first statistics according to the quantity of the data in the first group and a second statistics according to the quantity of the data in the second group; and

classifying the current audio frame as the music frame or the speech frame according to the first statistics or the second statistics.

5. The method of claim 1 , wherein the current signal is determined as the percussive music when a relatively acute energy protrusion occurs in the current signal in both a short time period and a long time period, the current signal has no obvious voiced sound characteristic, and several historical frames before the current audio frame are mainly music frames.

6. The method of claim 1 , wherein the current signal is determined as the percussive music when none of subframes of the current signal has an obvious voiced sound characteristic and a relatively obvious increase also occurs in a time domain envelope of the current signal relative to a long-time average of the time domain envelope.

7. An audio signal classification apparatus configured to classify an input audio signal, comprising:

a memory comprising instructions; and

one or more processors in communication with the memory, wherein the one or more processors execute the instructions to:

store, based on at least one condition, data of a frequency spectrum fluctuation parameter of a current audio frame of an audio signal into the memory where data of frequency spectrum fluctuation parameters of a plurality of audio frames are stored, wherein the at least one condition comprises the current audio frame is an active frame, the frequency spectrum fluctuation parameter denotes an energy fluctuation of a frequency spectrum of the audio signal;

modify data of frequency spectrum fluctuation parameters of audio frames preceding the current audio frame stored in the memory into ineffective data when the current audio frame is an active frame and an audio frame immediately preceding the current audio frame is an inactive frame, wherein data of frequency spectrum fluctuation parameters in the memory not having been modified into ineffective data are effective data;

modify the effective data stored in the memory into a value that is less than or equal to a music threshold when a current signal is percussive music, wherein the current signal comprises the current audio frame and a plurality of audio frames precede the current audio frame;

obtain statistics of a part or all of the effective data stored in the memory;

classify the current audio frame as a speech frame or a music frame according to the statistics of a part or all of the effective data stored in the memory.

8. The audio signal classification apparatus of claim 7 , wherein the current audio frame and a historical frame of the current audio frame belong to a group of multiple consecutive frames, and wherein the at least one condition further comprises none of the group of multiple consecutive frames belongs to an energy attack.

9. The audio signal classification apparatus of claim 7 , wherein to classifying the current audio frame as the speech frame or the music frame, the one or more processors are configured to:

obtain an average value of the part or all of the effective data of the frequency spectrum fluctuation parameters that are stored; and

either classify the current audio frame as the music frame based on a condition that the average value satisfies a music classification condition or classify the current audio frame as the speech frame based on a condition that the average value satisfies a speech classification condition.

10. The audio signal classification apparatus of claim 7 , wherein to classify the current audio frame as a speech frame or a music frame, the one or more processors are configured to:

obtain a first group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame;

obtain a second group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame, wherein a quantity of data in the first group and a quantity of data in the second group are different;

obtain a first statistics according to the quantity of the data in the first group and a second statistics according to the quantity of the data in the second group; and

classify the current audio frame as the music frame or the speech frame according to the first statistics or the second statistics.

11. The audio signal classification apparatus of claim 7 , wherein the current signal is determined as the percussive music when a relatively acute energy protrusion occurs in the current signal in both a short time period and a long time period, the current signal has no obvious voiced sound characteristic, and several historical frames before the current audio frame are mainly music frames.

12. The audio signal classification apparatus of claim 7 , wherein the current signal is determined as the percussive music when none of subframes of the current signal has an obvious voiced sound characteristic and a relatively obvious increase also occurs in a time domain envelope of the current signal relative to a long-time average of the time domain envelope.

13. An audio signal classification method, comprising:

storing, based on at least one condition, data of a frequency spectrum fluctuation parameter of a current audio frame of an audio signal into a memory where data of frequency spectrum fluctuation parameters of a plurality of audio frames are stored, wherein the at least one condition comprises the current audio frame is an active frame, the frequency spectrum fluctuation parameter denotes an energy fluctuation of a frequency spectrum of the audio signal;

modifying data of frequency spectrum fluctuation parameters of audio frames preceding the current audio frame stored in the memory into ineffective data when the current audio frame is an active frame and an audio frame immediately preceding the current audio frame is an inactive frame; wherein data of the frequency spectrum fluctuation parameters with negative values is the ineffective data, and data of frequency spectrum fluctuation parameters with a non-negative value is effective data;

modifying the effective data stored in the memory into a value that is less than or equal to a music threshold when a current signal is percussive music, wherein the current signal comprises the current audio frame and a plurality of audio frames precede the current audio frame;

obtaining statistics of a part or all of the effective data stored in the memory; and

classifying the current audio frame as a speech frame or a music frame according to the statistics of a part or all of the effective data stored in the memory.

14. The method of claim 13 , wherein the current audio frame and a historical frame of the current audio frame belong to a group of multiple consecutive frames, and the at least one condition further comprises none of the group of multiple consecutive frames belongs to an energy attack.

15. The method of claim 13 , wherein classifying the current audio frame as the speech frame or the music frame according to statistics of the part or all of effective data comprises:

obtaining an average value of the part or all of the effective data of the frequency spectrum fluctuation parameters that are stored; and

either classifying the current audio frame as the music frame based on a condition that the average value satisfies a music classification condition or classifying the current audio frame as the speech frame based on a condition that the average value satisfies a speech classification condition.

16. The method of claim 13 , wherein classifying the current audio frame as the speech frame or the music frame comprises:

obtaining a first group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame;

obtaining a second group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame, wherein a quantity of data in the first group and a quantity of data in the second group are different;

obtaining a first statistics according to the quantity of the data in the first group and a second statistics according to the quantity of the data in the second group; and

classifying the current audio frame as the music frame or the speech frame according to the first statistics or the second statistics.

17. The method of claim 13 , wherein the current signal is determined as the percussive music when a relatively acute energy protrusion occurs in the current signal in both a short time period and a long time period, the current signal has no obvious voiced sound characteristic, and several historical frames before the current audio frame are mainly music frames.

18. The method of claim 13 , wherein the current signal is determined as the percussive music when none of subframes of the current signal has an obvious voiced sound characteristic and a relatively obvious increase also occurs in a time domain envelope of the current signal relative to a long-time average of the time domain envelope.

19. An audio signal classification apparatus configured to classify an input audio signal, comprising:

a memory comprising instructions; and

one or more processors in communication with the memory, wherein the one or more processors execute the instructions to:

store, based on at least one condition, data of a frequency spectrum fluctuation parameter of a current audio frame of an audio signal into a memory where data of frequency spectrum fluctuation parameters of a plurality of audio frames are stored, wherein the at least one condition comprises the current audio frame is an active frame, the frequency spectrum fluctuation parameter denotes an energy fluctuation of a frequency spectrum of the audio signal;

modify data of frequency spectrum fluctuation parameters of audio frames preceding the current audio frame stored in the memory into ineffective data when the current audio frame is an active frame and an audio frame immediately preceding the current audio frame is an inactive frame; wherein data of the frequency spectrum fluctuation parameters with negative values is the ineffective data, and data of frequency spectrum fluctuation parameters with a non-negative value is effective data;

modify the effective data stored in the memory into a value that is less than or equal to a music threshold when a current signal is percussive music, wherein the current signal comprises the current audio frame and a plurality of audio frames precede the current audio frame;

obtain statistics of a part or all of the effective data stored in the memory; and

classify the current audio frame as a speech frame or a music frame according to the statistics of a part or all of the effective data stored in the memory.

20. The audio signal classification apparatus of claim 19 , wherein the current audio frame and a historical frame of the current audio frame belong to a group of multiple consecutive frames, and wherein the at least one condition further comprises none of the group of multiple consecutive frames belongs to an energy attack.

21. The audio signal classification apparatus of claim 19 , wherein to classifying the current audio frame as the speech frame or the music frame, the one or more processors are configured to:

obtain an average value of the part or all of the effective data of the frequency spectrum fluctuation parameters that are stored; and

either classify the current audio frame as the music frame based on a condition that the average value satisfies a music classification condition or classify the current audio frame as the speech frame based on a condition that the average value satisfies a speech classification condition.

22. The audio signal classification apparatus of claim 19 , wherein to classify the current audio frame as a speech frame or a music frame, the one or more processors are configured to:

obtain a first group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame;

obtain a second group of the effective data comprising data of the frequency spectrum fluctuation parameter of the current audio frame and one or more effective data of frequency spectrum fluctuation parameters of one or more audio frames continuously prior to the current audio frame, wherein a quantity of data in the first group and a quantity of data in the second group are different;

obtain a first statistics according to the quantity of the data in the first group and a second statistics according to the quantity of the data in the second group; and

classify the current audio frame as the music frame or the speech frame according to the first statistics or the second statistics.

23. The audio signal classification apparatus of claim 19 , wherein the current signal is determined as the percussive music when a relatively acute energy protrusion occurs in the current signal in both a short time period and a long time period, the current signal has no obvious voiced sound characteristic, and several historical frames before the current audio frame are mainly music frames.

24. The audio signal classification apparatus of claim 19 , wherein the current signal is determined as the percussive music when none of subframes of the current signal has an obvious voiced sound characteristic and a relatively obvious increase also occurs in a time domain envelope of the current signal relative to a long-time average of the time domain envelope.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2018
From: WANG, ZHE
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 046662/0422 →
Priority Claims (1)
CN 2013 1 0339218 · Aug 6, 2013 · national
Continuity (3)
Continuation 15017075 · Feb 5, 2016
Continuation PCTCN2013084252 · Sep 26, 2013
Related Publication 20180366145A1 · Dec 20, 2018
Cited By (1)
US 12,494,642