IP Library › Granted Patent US 10,818,313
Granted Patent B2
US 10,818,313 · App. 16/391,893 · Granted Oct 27, 2020

Method for detecting audio signal and apparatus

Inventor: Zhe Wang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L25/78G10L25/18G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,818,313
App. No.
16/391,893
Granted
Oct 27, 2020
Kind
B2
Abstract

A method for detecting an audio signal and an apparatus, where the method includes determining an input audio signal as a to-be-determined audio signal, determining an enhanced segmental signal-to-noise ratio (SSNR) of the audio signal, where the enhanced SSNR is greater than a reference SSNR, and comparing the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal. Therefore, the method and the apparatus can accurately distinguish an active voice and an inactive voice.

Claims (76)

1. A method for detecting an active signal, comprising:

determining an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR of the audio signal; and

comparing the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein determining the enhanced SSNR of the audio signal comprises determining the enhanced SSNR according to a signal-to-noise ratio (SNR) of each sub-band and a weight of the SNR of each sub-band in the audio signal,

wherein first weights of SNRs of high-frequency portion sub-bands are greater than a second weight of an SNR of a second sub-band,

wherein the SNRs of the high-frequency portion sub-bands are greater than a first threshold, and

wherein the second sub-band is one of a plurality of sub-bands except the high-frequency portion sub-bands in the audio signal.

2. The method of claim 1 , wherein the audio signal comprises 20 sub-bands.

3. A method for detecting an active signal, comprising:

determining an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR of the audio signal; and

comparing the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein determining the enhanced SSNR of the audio signal comprises:

determining the reference SSNR of the audio signal; and

determining the enhanced SSNR according to the reference SSNR of the audio signal,

wherein the enhanced SSNR is determined using the formula:

SSNR′= x *SSNR+ y , and

wherein SSNR indicates the reference SSNR, SSNR′ indicates the enhanced SSNR, and x and y indicate enhancement parameters.

4. A method for detecting an active signal, comprising:

determining an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR of the audio signal; and

comparing the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein determining the enhanced SSNR of the audio signal comprises:

determining the reference SSNR of the audio signal; and

determining the enhanced SSNR according to the reference SSNR of the audio signal,

wherein the enhanced SSNR is determined using the formula:

SSNR′= f ( x )*SSNR+ h ( y ), and

wherein SSNR indicates the reference SSNR, SSNR′ indicates the enhanced SSNR, f(x) and h(y) indicate enhancement functions, and h(y) is a function related to a Long-term SNR (LSNR) of the audio signal.

5. An apparatus for detecting an active signal, comprising:

a memory storage comprising instructions; and

one or more processors in communication with the memory storage, wherein the one or more processors execute the instructions to:

determine an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR of the audio signal; and

compare the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein the one or more processors further execute the instructions to determine the enhanced SSNR according to a signal-to-noise ratio (SNR) of each sub-band and weight of the SNR of each sub-band in the audio signal,

wherein first weights of SNRs of high-frequency portion sub-bands are greater than a second weight of an SNR of a second sub-band,

wherein the SNRs of the high-frequency portion sub-bands that are greater than a first threshold, and

wherein the second sub-band is one of a plurality of sub-bands except the high-frequency portion sub-bands in the audio signal.

6. The apparatus of claim 5 , wherein the audio signal comprises 20 sub-bands.

7. An apparatus for detecting an active signal, comprising:

a memory storage comprising instructions; and

one or more processors in communication with the memory storage, wherein the one or more processors execute the instructions to:

determine an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR of the audio signal; and

compare the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein the one or more processors further execute the instructions to determine the reference SSNR of the audio signal and determine the enhanced SSNR according to the reference SSNR of the audio signal,

wherein the enhanced SSNR is determined using the formula:

SSNR′= x *SSNR+ y , and

wherein SSNR indicates the reference SSNR, SSNR′ indicates the enhanced SSNR, and x and y indicate enhancement parameters.

8. An apparatus for detecting an active signal, comprising:

a memory storage comprising instructions; and

one or more processors in communication with the memory storage, wherein the one or more processors execute the instructions to:

determine an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR of the audio signal; and

compare the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein the one or more processors further execute the instructions to determine the reference SSNR of the audio signal and determine the enhanced SSNR according to the reference SSNR of the audio signal,

wherein the enhanced SSNR is determined using the formula:

SSNR′= f ( x )*SSNR+ h ( y ), and

wherein SSNR indicates the reference SSNR, SSNR′ indicates the enhanced SSNR, f(x) and h(y) indicate enhancement functions, and h(y) is a function related to a Long-term SNR (LSNR) of the audio signal.

9. A non-transitory computer-readable medium storing computer instructions, that when executed by one or more processors of an apparatus for detecting an active signal, cause the one or more processors to:

determine an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR; and

compare the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein the computer instructions, when executed by the one or more processors, further cause the one or more processors to determine the enhanced SSNR according to a signal-to-noise ratio (SNR) of each sub-band and weight of the SNR of each sub-band in the audio signal,

wherein first weights of SNRs of high-frequency portion sub-bands are greater than a second weight of an SNR of a second sub-band,

wherein the SNRs of the high-frequency portion sub-bands are greater than a first threshold, and

wherein the second sub-band is one of a plurality of sub-bands except the high-frequency portion sub-bands in the audio signal.

10. The non-transitory computer-readable medium of claim 9 , wherein the audio signal comprises 20 sub-bands.

11. A non-transitory computer-readable medium storing computer instructions, that when executed by one or more processors of an apparatus for detecting an active signal, cause the one or more processors to:

determine an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR; and

compare the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein the computer instructions, when executed by the one or more processors, further cause the one or more processors to determine the reference SSNR of the audio signal and determine the enhanced SSNR according to the reference SSNR of the audio signal,

wherein the enhanced SSNR is determined using the formula:

SSNR′= x *SSNR+ y , and

wherein SSNR indicates the reference SSNR, SSNR′ indicates the enhanced SSNR, and x and y indicate enhancement parameters.

12. A non-transitory computer-readable medium storing computer instructions, that when executed by one or more processors of an apparatus for detecting an active signal, cause the one or more processors to:

determine an enhanced segmental signal-to-noise ratio (SSNR) of an audio signal in response to the audio signal being an unvoiced signal, wherein the enhanced SSNR is greater than a reference SSNR; and

compare the enhanced SSNR with a voice activity detection (VAD) decision threshold to determine whether the audio signal is an active signal,

wherein the computer instructions, when executed by the one or more processors, further cause the one or more processors to determine the reference SSNR of the audio signal and determine the enhanced SSNR according to the reference SSNR of the audio signal,

wherein the enhanced SSNR is determined using the following formula:

SSNR′= f ( x )*SSNR+ h ( y ), and

wherein SSNR indicates the reference SSNR, SSNR′ indicates the enhanced SSNR, f(x) and h(y) indicate enhancement functions, and h(y) is a function related to a Long-term SNR (LSNR) of the audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2019
From: WANG, ZHE
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 048970/0954 →
Priority Claims (1)
CN 2014 1 0090386 · Mar 12, 2014 · national
Continuity (3)
Continuation 15262263 · Sep 12, 2016
Continuation PCTCN2014092694 · Dec 1, 2014
Related Publication 20190279657A1 · Sep 12, 2019
Cited By (1)
US 12,444,430