Adaptive noise estimation
In some embodiments, a method, comprises: dividing, using at least one processor, an audio input into speech and non-speech segments; for each frame in each non-speech segment, estimating, using the at least one processor, a time-varying noise spectrum of the non-speech segment; for each frame in each speech segment, estimating, using the at least one processor, speech spectrum of the speech segment; for each frame in each speech segment, identifying one or more non-speech frequency components in the speech spectrum; comparing the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra and selecting the estimated noise spectrum from the plurality of estimated noise spectra based on a result of the comparing.
1 . A method of adaptive noise estimation, comprising:
dividing, using at least one processor, an audio input into speech and non-speech segments;
for each frame in each non-speech segment, estimating, using the at least one processor, a time-varying noise spectrum of the non-speech segment and reducing noise in the audio input based on the time-varying noise spectrum;
for each frame in each speech segment:
estimating, using the at least one processor, a speech spectrum of the speech segment;
identifying one or more non-speech frequency components in the speech spectrum;
comparing the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra, wherein the plurality of estimated noise spectra comprises a first estimated noise spectrum for a past non-speech segment and a second estimated noise spectrum for a future non-speech segment;
selecting one of the first or second estimated noise spectrums based on a result of the comparing; and
reducing noise in the audio input based on the selected first or second estimated noise spectrum.
2 . The method of claim 1 , further comprising:
obtaining a probability of speech in each frame of the audio input and identifying a frame containing speech based on the probability.
3 . The method of claim 1 , wherein the time-varying noise spectrum is estimated by computing a moving average of power spectra of the non-speech segments, and averaging the power spectra of a current non-speech segment and at least one past non-speech segment.
4 . The method of claim 1 , wherein for each speech segment, a past estimated noise spectrum before the speech segment, a future estimated noise spectrum after the speech segment and a current speech frame, are used to determine the estimated noise spectrum that has a highest likelihood to represent noise in the current speech segment.
5 . The method of claim 4 , wherein determining the estimated noise spectrum that has the highest likelihood to represent the noise of the current speech segment, further comprises:
obtaining an average noise spectrum from past and future noise spectra of past and future non-speech segments before and after the speech segment, respectively;
determining an upper frequency limit for the past and future noise spectra;
determining a cutoff frequency to be the lowest one of the two upper frequency limits;
computing a distance metric between frequency components in the speech spectrum and frequency components in the noise spectra; and
selecting one of the past or future noise spectrum that has the smallest distance metric up to the cutoff frequency as the estimated noise spectrum for the audio input.
6 . The method of claim 5 , wherein the distance metric is averaged over a set of speech frames in a speech segment.
7 . The method of claim 1 , wherein speech components are estimated in the speech segments of the audio signal, and then subtracted from actual speech components to obtain a residual spectrum as the estimated non-speech frequency components.
8 . A non-transitory, computer-readable storage medium having stored thereon instructions that when executed by one or more processors, cause the one or more processors to perform operations of claim 1 .
9 . An audio processor comprising:
a divider unit configured to divide an audio input into speech and non-speech segments;
an averaging unit configured to estimate, for each speech segment speech spectra and for each non-speech segment time-varying noise spectra;
a similarity metric unit configured to:
identify one or more non-speech frequency components in the speech spectra;
compare the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra, wherein the plurality of estimated noise spectra comprises a first estimated noise spectrum for a past non-speech segment and a second estimated noise spectrum for a future non-speech segment;
select one of the first or second estimated noise spectrums based on a result of the comparing; and
a noise reduction unit configured to:
for non-speech segments, reduce noise in the audio input based on the time-varying noise spectra; and
for speech segments, reduce noise in the audio input based on the selected first or second estimated noise spectrum.
10 . The audio processor of claim 9 , wherein the noise reduction unit is configured to reduce noise in the audio input using the selected estimated noise spectrum by comparing the spectrum of the audio input with the selected estimated noise spectrum, and applying gain reduction to frequency bands where an energy of the audio input is less than an energy of the noise spectrum plus a predefined threshold.
11 . The audio processor of claim 9 , wherein a voice activity detector (VAD) is configured to obtain a probability of speech in each frame of the audio input and identify a frame containing speech based on the probability; or
wherein the averaging unit is configured to estimate the time-varying noise spectra by computing a moving average of power spectra of the non-speech segments, and averaging the power spectra of a current non-speech segment and at least one past non-speech segment.
12 . The audio processor of claim 9 , wherein for each speech segment, the similarity metric unit is configured to determine the estimated noise spectrum that has a highest likelihood to represent noise in the current speech segments based on a past estimated noise spectrum before the speech segment, a future estimated noise spectrum after the speech segment and a current speech frame.
13 . The audio processor of claim 12 , wherein the similarity metric unit is configured to determine the estimated noise spectrum that has the highest likelihood to represent the noise of the current speech segment by:
obtaining an average noise spectrum from past and future noise spectra of past and future non-speech segments before and after the speech segment, respectively;
determining an upper frequency limit for the past and future noise spectra;
determining a cutoff frequency to be the lowest one of the two upper frequency limits;
computing a distance metric between frequency components in the speech spectrum and frequency components in the noise spectra; and
selecting one of the past or future noise spectrum that has the smallest distance metric up to the cutoff frequency as the estimated noise spectrum for the audio input.
14 . The audio processor of claim 13 , wherein the similarity metric unit is configured to average the distance metric over a set of speech frames in a speech segment.
15 . The audio processor of claim 9 , wherein the similarity metric unit is configured to estimate the one or more speech components in the speech segments of the audio input, and then subtract the one or more estimated speech components from actual speech components to obtain a residual spectrum as the estimated non-speech frequency spectrum.