IP Library › Granted Patent US 12,731,599
Granted Patent B2
US 12,731,599 · App. 18/044,777 · Granted Sep 8, 2026

Adaptive noise estimation

Inventors: Davide Scaini (Barcelona, ES); Chunghsin Yeh (Barcelona, ES); Giulio Cengarle (Barcelona, ES); Mark David de Burgh (Mount Colah, AU)
Assignees: Dolby Laboratories Licensing Corporation; DOLBY INTERNATIONAL AB
G10L21/0232G10L21/028G10L21/034G10L21/0364G10L25/18G10L25/21G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,599
App. No.
18/044,777
Granted
Sep 8, 2026
Kind
B2
Abstract

In some embodiments, a method, comprises: dividing, using at least one processor, an audio input into speech and non-speech segments; for each frame in each non-speech segment, estimating, using the at least one processor, a time-varying noise spectrum of the non-speech segment; for each frame in each speech segment, estimating, using the at least one processor, speech spectrum of the speech segment; for each frame in each speech segment, identifying one or more non-speech frequency components in the speech spectrum; comparing the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra and selecting the estimated noise spectrum from the plurality of estimated noise spectra based on a result of the comparing.

Claims (44)

1 . A method of adaptive noise estimation, comprising:

dividing, using at least one processor, an audio input into speech and non-speech segments;

for each frame in each non-speech segment, estimating, using the at least one processor, a time-varying noise spectrum of the non-speech segment and reducing noise in the audio input based on the time-varying noise spectrum;

for each frame in each speech segment:

estimating, using the at least one processor, a speech spectrum of the speech segment;

identifying one or more non-speech frequency components in the speech spectrum;

comparing the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra, wherein the plurality of estimated noise spectra comprises a first estimated noise spectrum for a past non-speech segment and a second estimated noise spectrum for a future non-speech segment;

selecting one of the first or second estimated noise spectrums based on a result of the comparing; and

reducing noise in the audio input based on the selected first or second estimated noise spectrum.

2 . The method of claim 1 , further comprising:

obtaining a probability of speech in each frame of the audio input and identifying a frame containing speech based on the probability.

3 . The method of claim 1 , wherein the time-varying noise spectrum is estimated by computing a moving average of power spectra of the non-speech segments, and averaging the power spectra of a current non-speech segment and at least one past non-speech segment.

4 . The method of claim 1 , wherein for each speech segment, a past estimated noise spectrum before the speech segment, a future estimated noise spectrum after the speech segment and a current speech frame, are used to determine the estimated noise spectrum that has a highest likelihood to represent noise in the current speech segment.

5 . The method of claim 4 , wherein determining the estimated noise spectrum that has the highest likelihood to represent the noise of the current speech segment, further comprises:

obtaining an average noise spectrum from past and future noise spectra of past and future non-speech segments before and after the speech segment, respectively;

determining an upper frequency limit for the past and future noise spectra;

determining a cutoff frequency to be the lowest one of the two upper frequency limits;

computing a distance metric between frequency components in the speech spectrum and frequency components in the noise spectra; and

selecting one of the past or future noise spectrum that has the smallest distance metric up to the cutoff frequency as the estimated noise spectrum for the audio input.

6 . The method of claim 5 , wherein the distance metric is averaged over a set of speech frames in a speech segment.

7 . The method of claim 1 , wherein speech components are estimated in the speech segments of the audio signal, and then subtracted from actual speech components to obtain a residual spectrum as the estimated non-speech frequency components.

8 . A non-transitory, computer-readable storage medium having stored thereon instructions that when executed by one or more processors, cause the one or more processors to perform operations of claim 1 .

9 . An audio processor comprising:

a divider unit configured to divide an audio input into speech and non-speech segments;

an averaging unit configured to estimate, for each speech segment speech spectra and for each non-speech segment time-varying noise spectra;

a similarity metric unit configured to:

identify one or more non-speech frequency components in the speech spectra;

compare the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra, wherein the plurality of estimated noise spectra comprises a first estimated noise spectrum for a past non-speech segment and a second estimated noise spectrum for a future non-speech segment;

select one of the first or second estimated noise spectrums based on a result of the comparing; and

a noise reduction unit configured to:

for non-speech segments, reduce noise in the audio input based on the time-varying noise spectra; and

for speech segments, reduce noise in the audio input based on the selected first or second estimated noise spectrum.

10 . The audio processor of claim 9 , wherein the noise reduction unit is configured to reduce noise in the audio input using the selected estimated noise spectrum by comparing the spectrum of the audio input with the selected estimated noise spectrum, and applying gain reduction to frequency bands where an energy of the audio input is less than an energy of the noise spectrum plus a predefined threshold.

11 . The audio processor of claim 9 , wherein a voice activity detector (VAD) is configured to obtain a probability of speech in each frame of the audio input and identify a frame containing speech based on the probability; or

wherein the averaging unit is configured to estimate the time-varying noise spectra by computing a moving average of power spectra of the non-speech segments, and averaging the power spectra of a current non-speech segment and at least one past non-speech segment.

12 . The audio processor of claim 9 , wherein for each speech segment, the similarity metric unit is configured to determine the estimated noise spectrum that has a highest likelihood to represent noise in the current speech segments based on a past estimated noise spectrum before the speech segment, a future estimated noise spectrum after the speech segment and a current speech frame.

13 . The audio processor of claim 12 , wherein the similarity metric unit is configured to determine the estimated noise spectrum that has the highest likelihood to represent the noise of the current speech segment by:

obtaining an average noise spectrum from past and future noise spectra of past and future non-speech segments before and after the speech segment, respectively;

determining an upper frequency limit for the past and future noise spectra;

determining a cutoff frequency to be the lowest one of the two upper frequency limits;

computing a distance metric between frequency components in the speech spectrum and frequency components in the noise spectra; and

selecting one of the past or future noise spectrum that has the smallest distance metric up to the cutoff frequency as the estimated noise spectrum for the audio input.

14 . The audio processor of claim 13 , wherein the similarity metric unit is configured to average the distance metric over a set of speech frames in a speech segment.

15 . The audio processor of claim 9 , wherein the similarity metric unit is configured to estimate the one or more speech components in the speech segments of the audio input, and then subtract the one or more estimated speech components from actual speech components to obtain a residual spectrum as the estimated non-speech frequency spectrum.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2023
From: SCAINI, DAVIDE; YEH, CHUNGHSIN; CENGARLE, GIULIO; DE BURGH, MARK DAVID
To: DOLBY INTERNATIONAL AB; DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 064215/0475 →
Priority Claims (1)
ES ES202030960 · Sep 23, 2020 · national
Continuity (3)
Provisional Application 63168998 · Mar 31, 2021
Provisional Application 63120253 · Dec 2, 2020
Related Publication 20240013799A1 · Jan 11, 2024
References Cited (41)
US 6347297B1 · Asghar · 2002 [cited by examiner]
US 6453289B1 · Ertem · 2002 [cited by applicant]
US 7039181B2 · Marchok · 2006 [cited by applicant]
US 7660714B2 · Furuta · 2010 [cited by examiner]
US 8600073B2 · Sun · 2013 [cited by applicant]
US 8694311B2 · Jung · 2014 [cited by examiner]
US 8909522B2 · Shperling · 2014 [cited by applicant]
US 8990073B2 · Malenovsky · 2015 [cited by applicant]
US 9570093B2 · Gao · 2017 [cited by examiner]
US 9576590B2 · Sjoberg · 2017 [cited by applicant]
US 9721580B2 · Skoglund · 2017 [cited by applicant]
US 9754608B2 · Souden · 2017 [cited by applicant]
US 9812149B2 · Yen · 2017 [cited by applicant]
US 9978394B1 · Su · 2018 [cited by examiner]
US 10199033B1 · Yano · 2019 [cited by examiner]
US 11146607B1 · Tang · 2021 [cited by examiner]
US 20030078772A1 · Wu · 2003 [cited by examiner]
US 20050286664A1 · Chen · 2005 [cited by examiner]
US 20090012786A1 · Zhang · 2009 [cited by applicant]
US 20110099007A1 · Zhang · 2011 [cited by examiner]
US 20120197634A1 · Ishikawa · 2012 [cited by examiner]
US 20170345439A1 · Jensen · 2017 [cited by examiner]
US 20180033447A1 · Ramprashad · 2018 [cited by examiner]
US 20190013036A1 · Graf · 2019 [cited by examiner]
US 20190206420A1 · Kandade Rajan et al. · 2019 [cited by applicant]
US 20200312294A1 · Isberg · 2020 [cited by examiner]
CN 1354871A · 2002 [cited by applicant]
CN 1384960A · 2002 [cited by applicant]
GB 2426167B · 2007 [cited by applicant]
JP 2001318687A · 2001 [cited by applicant]
JP 2006039547A · 2006 [cited by examiner]
JP 4765461B2 · 2011 [cited by examiner]
KR 20070108598A · 2006 [cited by examiner]
KR 20100045933A · 2010 [cited by applicant]
WO 2021148342A1 · 2021 [cited by applicant]
Brueckmann et. al., Adaptive Noise Reduction and Voice Activity Detection for improved Verbal Human-Robot Interaction using Binaural Data, Proceedings 2007 IEEE International Conference on Robotics and Automation. [cited by applicant]
Doblinger, Gerhard “Computationally Efficient Speech Enhancement by Spectral Minima Tracking in Subbands” Proc. Eurospeech, pp. 1513-1516, 1995. [cited by applicant]
Martin, Rainer “Noise Power Spectral Density Estimation Based on Optimal Smoothing and Minimum Statistics” IEEE Transactions on Speech and Audio Processing, vol. 9, Issue 5, pp. 504-512, Jul. 2001. [cited by applicant]
Stylianou, Yannis. “Harmonic plus noise models for speech, combined with statistical methods, for speech and speaker modification.” Ph. D thesis, Ecole Nationale Superieure des Telecommunications (1996). [cited by applicant]
Yeh, C., “Multiple Fundamental Frequency Estimation of Polyphonic Recordings,” Ph.D. thesis, 2008, University Paris 6. [cited by applicant]
Z. Zhang, K. Honda and J. Wei, “Retrieving Vocal-Tract Resonance and anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech a… [cited by applicant]