IP Library › Granted Patent US 12,682,911
Granted Patent B2
US 12,682,911 · App. 18/007,324 · Granted Jul 14, 2026

Automatic detection and attenuation of speech-articulation noise events

Inventors: Chunghsin Yeh (Barcelona, ES); Giulio Cengarle (Barcelona, ES); Mark David De Burgh (Mount Colah, AU)
Assignee: DOLBY INTERNATIONAL AB
G10L21/0264G10L21/0232G10L21/0364G10L25/18G10L25/21G10L25/45G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,911
App. No.
18/007,324
Filed
Jan 30, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2657
USPC
704/233
Abstract

Described is a method of performing automatic audio enhancement on an input audio signal including at least one speech-articulation noise event. The method comprises: segmenting the input audio signal into a number of audio frames; obtaining at least one feature parameter from the audio frames; and determining, based at least in part on the obtained feature parameter, a respective type of the speech-articulation noise event and a respective time-frequency range associated with the speech-articulation noise event within the input audio signal.

Claims (60)

1 . A method of performing automatic audio enhancement on an input audio signal including at least one speech-articulation noise event, the method comprising:

segmenting the input audio signal into a number of audio frames;

obtaining at least one feature parameter from the audio frames, wherein obtaining the at least one feature parameter from the audio frames comprises, for each audio frame, obtaining at least one measure of kurtosis based on time-domain sample amplitudes of the audio frames;

determining, based at least in part on the obtained feature parameter, a respective type of the speech-articulation noise event and a respective time-frequency range associated with the speech-articulation noise event within the input audio signal, wherein determining, based on the obtained feature parameter, the respective type of the speech-articulation noise event and the respective range thereof in the input audio signal comprises:

comparing the obtained measure of kurtosis to a predefined kurtosis threshold; and

if the measure of kurtosis exceeds the predefined kurtosis threshold, determining that the audio frame comprises a mouth click event, and determining start and end boundaries of the mouth click event based on respective positions at which the measure of kurtosis rises above and falls below the predefined kurtosis threshold; and

attenuating the speech-articulation noise event in accordance with a determined type of the speech-articulation noise event and the respective time-frequency range associated with the speech-articulation noise event.

2 . The method according to claim 1 , wherein the determined range comprises at least one boundary of the determined speech-articulation noise event, in the time or spectral domain.

3 . The method according to claim 1 , wherein the speech-articulation noise event comprises at least one of: a mouth click event or a speech plosive event.

4 . The method according to claim 3 , wherein the speech-articulation noise event comprises one or more mouth click events; and wherein the one or more mouth click events comprise at least one of: a non-speech click event, a speech click event, or a lip smack event.

5 . The method according to claim 4 , wherein, after segmenting the input audio signal into a number of audio frames, the method further comprises:

classifying the audio frames as either speech frames or non-speech frames.

6 . The method according to claim 5 , wherein the segmentation is performed by using two different time window sizes, one of the two time window sizes being shorter than the other.

7 . The method according to claim 6 , wherein the shorter time window size is used for detecting speech click events in the speech frames and a longer time window size is used for detecting non-speech click events in the non-speech frames.

8 . The method according to claim 5 , wherein obtaining at least one feature parameter from the audio frames comprises:

for each speech frame, obtaining a respective approximation of residual without speech harmonic components and a respective first measure of kurtosis of sample amplitudes for the approximation of residual, and

wherein determining, based on the obtained feature parameter, a respective type of the speech-articulation noise event and a respective range thereof in the input audio signal comprises:

comparing the obtained first measure of kurtosis to a first predefined kurtosis threshold; and

if the first measure of kurtosis exceeds the first predefined kurtosis threshold, determining that the speech frame comprises a speech click event, and determining start and end boundaries of the speech click event based on respective positions at which the first measure of kurtosis rises above and falls below the first predefined kurtosis threshold.

9 . The method according to claim 8 , wherein the approximation of residual without speech harmonic components is a second-order waveform difference.

10 . The method according to claim 8 , further comprising:

obtaining a second measure of kurtosis from residual sample amplitudes of the speech frame;

wherein the type and range of the speech-articulation noise event are determined based on the second measure of kurtosis relative to the first measure of kurtosis.

11 . The method according to claim 8 , further comprising:

refining the determined range of the speech click event by:

locating a sample position with the largest second-order difference within the determined range of the speech click event; and

determining the refined range of the speech click event by applying a predefined speech click event duration around the located sample position.

12 . The method according to claim 8 , further comprising:

determining the range of the speech click event further based on a min/max change rate calculated from local minima and maxima in the speech frame.

13 . The method according to claim 8 , further comprising:

attenuating the determined one or more mouth click events based on respective spectral gains derived from spectral envelopes of the audio frames containing detected mouth click events and target envelopes calculated based on respective reference frames.

14 . The method according to claim 5 , wherein obtaining at least one feature parameter from the audio frames comprises:

for each non-speech frame, obtaining a respective third measure of kurtosis of time-domain sample amplitudes in the non-speech frame, and

wherein determining, based on the obtained feature parameter, a respective type of the speech-articulation noise event and a respective range thereof in the input audio signal comprises:

comparing the obtained third measure of kurtosis to a second predefined kurtosis threshold; and

if the third measure of kurtosis exceeds the second predefined kurtosis threshold, determining that the non-speech frame comprises a non-speech click event; and determining start and end boundaries of the non-speech click event based on respective positions at which the third measure of kurtosis rises above and falls below the second predefined kurtosis threshold.

15 . The method according to claim 14 , further comprising:

if two neighboring non-speech click events are within a predefined gap threshold, merging the two neighboring non-speech click events into a single speech click event.

16 . The method according to claim 14 , wherein

for a determined non-speech click event in a non-speech frame immediately preceding a speech frame:

calculating a high/low-band peak ratio as an amplitude ratio between the largest peak above a predefined frequency and the largest peak below the predefined frequency; and

if the calculated high/low-band peak ratio is above a predefined ratio threshold, determining the non-speech click event as a lip smack event.

17 . The method according to claim 16 , further comprising:

refining the determined range of the lip smack event based on the high/low-band peak ratio, a spectral slope and an energy envelope.

18 . The method according to claim 5 , further comprising:

determining the speech-articulation noise event further based on the center of gravity, COG, calculated for the speech frames in accordance with a further predefined threshold, for distinguishing mouth click events from speech transients.

19 . An apparatus comprising a processor and a memory coupled to the processor storing instructions, that when executed by the processor, cause the apparatus to carry out a method of performing automatic audio enhancement on an input audio signal including at least one speech-articulation noise event, the method comprising:

segmenting the input audio signal into a number of audio frames;

obtaining at least one feature parameter from the audio frames, wherein obtaining the at least one feature parameter from the audio frames comprises, for each audio frame, obtaining at least one measure of kurtosis based on time-domain sample amplitudes of the audio frames;

determining, based at least in part on the obtained feature parameter, a respective type of the speech-articulation noise event and a respective time-frequency range associated with the speech-articulation noise event within the input audio signal, wherein determining, based on the obtained feature parameter, the respective type of the speech-articulation noise event and the respective range thereof in the input audio signal comprises:

comparing the obtained measure of kurtosis to a predefined kurtosis threshold; and

if the measure of kurtosis exceeds the predefined kurtosis threshold, determining that the audio frame comprises a mouth click event, and determining start and end boundaries of the mouth click event based on respective positions at which the measure of kurtosis rises above and falls below the predefined kurtosis threshold; and

attenuating the speech-articulation noise event in accordance with a determined type of the speech-articulation noise event and the respective time-frequency range associated with the speech-articulation noise event.

20 . A non-transitory computer-readable storage medium storing one or more programs comprising instructions that, when executed by a processor, cause the processor to carry out a method of performing automatic audio enhancement on an input audio signal including at least one speech-articulation noise event, the method comprising:

segmenting the input audio signal into a number of audio frames;

obtaining at least one feature parameter from the audio frames, wherein obtaining the at least one feature parameter from the audio frames comprises, for each audio frame, obtaining at least one measure of kurtosis based on time-domain sample amplitudes of the audio frames;

determining, based at least in part on the obtained feature parameter, a respective type of the speech-articulation noise event and a respective time-frequency range associated with the speech-articulation noise event within the input audio signal, wherein determining, based on the obtained feature parameter, the respective type of the speech-articulation noise event and the respective range thereof in the input audio signal comprises:

comparing the obtained measure of kurtosis to a predefined kurtosis threshold; and

if the measure of kurtosis exceeds the predefined kurtosis threshold, determining that the audio frame comprises a mouth click event, and determining start and end boundaries of the mouth click event based on respective positions at which the measure of kurtosis rises above and falls below the predefined kurtosis threshold; and

attenuating the speech-articulation noise event in accordance with a determined type of the speech-articulation noise event and the respective time-frequency range associated with the speech-articulation noise event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: YEH, CHUNGHSIN; CENGARLE, GIULIO; DE BURGH, MARK DAVID
To: DOLBY INTERNATIONAL AB
Reel/Frame 063631/0483 →
Priority Claims (1)
ES ES202030864 · Aug 12, 2020 · national
Continuity (2)
Provisional Application 63107012 · Oct 29, 2020
Related Publication 20230267945A1 · Aug 24, 2023
References Cited (68)
US 4718096A · Meisel · 1988 [cited by applicant]
US 5572623A · Pastor · 1996 [cited by applicant]
US 5611019A · Nakatoh · 1997 [cited by applicant]
US 6304842B1 · Husain · 2001 [cited by applicant]
US 7089180B2 · Heikkinen · 2006 [cited by applicant]
US 7369990B2 · Nemer · 2008 [cited by applicant]
US 7680653B2 · Yeldener · 2010 [cited by applicant]
US 7742914B2 · Kosek · 2010 [cited by applicant]
US 8190432B2 · Matsumoto · 2012 [cited by examiner]
US 8369556B2 · Power · 2013 [cited by applicant]
US 8756054B2 · Kovesi · 2014 [cited by applicant]
US 8775184B2 · Deshmukh · 2014 [cited by applicant]
US 9208794B1 · Mascaro · 2015 [cited by applicant]
US 9454976B2 · Newman · 2016 [cited by applicant]
US 10170126B2 · Kovesi · 2019 [cited by applicant]
US 10204634B2 · Tada · 2019 [cited by applicant]
US 10242696B2 · Ebenezer · 2019 [cited by examiner]
US 10283122B2 · Kjoerling · 2019 [cited by applicant]
US 11227586B2 · Borgstrom · 2022 [cited by examiner]
US 11771372B2 · Voix · 2023 [cited by examiner]
US 20070033042A1 · Marcheret · 2007 [cited by applicant]
US 20080027717A1 · Rajendran · 2008 [cited by applicant]
US 20100274554A1 · Orr · 2010 [cited by applicant]
US 20100280823A1 · Shlomot · 2010 [cited by applicant]
US 20120245927A1 · Bondy · 2012 [cited by examiner]
US 20140324367A1 · Garvey, III · 2014 [cited by applicant]
US 20160293175A1 · Atti · 2016 [cited by applicant]
US 20170018277A1 · Wang · 2017 [cited by applicant]
US 20170180853A1 · Mehta · 2017 [cited by applicant]
US 20200258529A1 · Schnabel · 2020 [cited by applicant]
CN 101009099A · 2007 [cited by applicant]
CN 101790122A · 2014 [cited by applicant]
CN 110335620A · 2019 [cited by applicant]
CN 111263284A · 2021 [cited by applicant]
CN 108682418A · 2022 [cited by applicant]
EP 3038106A1 · 2017 [cited by applicant]
JP S61077099 · 1986 [cited by applicant]
JP S62067598 · 1987 [cited by applicant]
JP H4130500 · 1992 [cited by applicant]
JP 2017112444A · 2017 [cited by applicant]
KR 20000034465A · 2000 [cited by applicant]
WO 1987003127A1 · 1987 [cited by applicant]
WO 2019079909A1 · 2019 [cited by applicant]
WO 2019232684A1 · 2019 [cited by applicant]
Mitra, Vikramjit, Bengt J. Borgstrom, Carol Y. Espy-Wilson, and Abeer Alwan, “A Noise-type and Level-dependent MPO-based Speech Enhancement Architecture with Variable Frame Analysis for Noise-robust Speech Recognition”,… [cited by examiner]
Rao, K. Sreenivasa, and Anil Kumar Vuppala, “Speech Processing in Mobile Environments”, 2014, Springer International Publishing. (Year: 2014). [cited by examiner]
Kong, Ying-Yee, Ala Mullangi, and Kostas Kokkinakis, “Classification of Fricative Consonants for Speech Enhancement in Hearing Devices”, Apr. 2014, PLoS ONE, vol. 9, No. 4, pp. 1-8. (Year: 2014). [cited by examiner]
Talmon, Ronen, Israel Cohen, and Sharon Gannot, “Transient Noise Reduction Using Nonlocal Diffusion Filters”, Aug. 2011, IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, No. 6, pp. 1584-1599. (Year:… [cited by examiner]
Cole, “Location and classification of plosive consonants using expert knowledge and neural net classifiers,” The Journal of the Acoustical Society of America 84, S60 (1988); doi: 10.1121/1.2026396, Jun. 10, 2005. [cited by applicant]
Haque et al., “Zero-Crossings with Adaptation for Automatic Speech Recognition,” Proceedings of the 11th Australasian International Conference on Speech Science and Technology, Jun. 28, 2005. [cited by applicant]
IZotope RX De-plosive, retrieved on Jun. 22, 2020, <https://www.izotope.com/en/products/rx/features/de-plosive.html>. [cited by applicant]
Keshet et al., “Plosive Spotting with Margin Classifiers,” Interspeech 2001. [cited by applicant]
Kondoz, “Digital Speech: Coding for Low Bit Rate Communication Systems,” Section 9.6 Speech Classification, pp. 311-319, 2nd edition, Wiley 2004. [cited by applicant]
Madsack et al., “Phone-Based Plosive Detection,” retrieved Jun. 22, 2020, <https://www.sfb732.uni-stuttgart.de/documents/files/plos_det_techreport.pdf>. [cited by applicant]
Vikhe, “Voice Activity Detection Algorithm for Speech Recognition Applications,” Conference Paper, Jan. 1, 2011. [cited by applicant]
Weigelt et al., “Plosive/Fricative Distinction: The Voiceless Case,” The Journal of The Acoustical Society of America, American Institute of Physics, 2 Huntington Quadrangle, Melville, NY 11747, vol. 87, No. 6, Jun. 1, … [cited by applicant]
Yanxiong et al., “A detection method of lip-smack in spontaneous speech,” Audio, Language and Image Processing, 2008, ICALIP 2008, International Conference on, IEEE, Piscataway, NJ, USA, Jul. 7, 2008, pp. 292-297, XP031… [cited by applicant]
“LS Levelator 2” Le Sound Levelator as accessed on Jan. 27, 2021, https://lesound.io/product/levelator/, pp. 1-7, 7 pages. [cited by applicant]
Asccusonus Mouth De-Clicker https://accusonus.com/products/audio-repair/era-bundle-standard#mouth-de-clicker-anchor as accessed on Jan. 24, 2021, 32 pages. [cited by applicant]
Charpentier et al., “Diphone synthesis using an overlap-add technique for speech waveforms concatenation,” In Int. Conf. Acoustics, Speech, and Signal Processing (ICASSP). vol. 11. IEEE, 1986, pp. 2015-2018, 4 pages. [cited by applicant]
De Leon, “Short-time kurtosis of speech signals with application to co-channel speech separation” Located via IEEE Xplore, Aug. 2002, pp. 831-833, 3 pages. [cited by applicant]
Esquef et al., “Interpolation of Long Gaps in Audio Signals Using the Warped Burg's Method,” Proceedings of the 6th International Conference on Digital Audio Effects (DAFx-03), Lon don, UK, Sep. 2003, pp. DAFX-1-DAFX-6,… [cited by applicant]
Fink et al., “Comparison of various predictors for audio extrapolation,” Proc. of the 16th Int. Conference on Digital Audio Effects (DAFx-13), Maynooth, Ireland, Sep. 2-6, 2013, pp. DAFX-1-DAFX-8, 8 pages. [cited by applicant]
Godsill et al., “Digital Audio Restoration—A Statistical Model Based Approach,” Springer, Sep. 21, 1988, pp. i-328, 346 pages. [cited by applicant]
Lukin et al., “Parametric Interpolation of Gaps in Audio Signals,” AES 125th Convention Oct. 2-5, 2008 San Francisco, CA, pp. 1-8, 8 pages. [cited by applicant]
Lukin, “The Making of RX^ Mouth De-Click,” 2017 https://www.izotope.com/en/learn/the-making-of-rx-6-mouth-de-click.html as accessed on Apr. 14, 2025, pp. 1-191, 191 pages. [cited by applicant]
Roebel et al., “Efficient spectral envelope estimation and its application to pitch shifting and envelope preservation,” International Conference on Digital Audio Effects 2005, Madrid, Spain, hal-01161334, pp. 30-35, 7 … [cited by applicant]
Roebel, “A new approach to transient processing in the phase vocoder,” Proceedings of the 6th Int. Conference on Digital Audio Effects, DAFx-03, London, UK, Sep. 8-11, 2003, pp. DAFX-1-DAFX-6, 7 pages. [cited by applicant]