IP Library › Granted Patent US 12,725,625
Granted Patent B2
US 12,725,625 · App. 18/844,562 · Granted Sep 1, 2026

Method and audio processing system for wind noise suppression

Inventors: Qingyuan Bin (Beijing, CN); Yuanxing Ma (Beijing, CN); Zhiwei Shuang (Beijing, CN)
Assignee: Dolby Laboratories Licensing Corporation
G10L21/0232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,625
App. No.
18/844,562
Granted
Sep 1, 2026
Kind
B2
Abstract

The present disclosure relates to a method and system ( 1 ) for suppressing wind noise. The method comprises obtaining an input audio signal ( 100, 100 ′) comprising a plurality of consecutive audio signal segments ( 101, 102, 103, 101′, 102′, 103 ′) and suppressing wind noise in the input audio signal with a wind noise suppressor module ( 20 ) to generate a wind noise reduced audio signal. The method further comprises sing a neural network ( 10 ) trained to predict a set of gains for reducing noise in the input audio signal ( 100, 100 ′) given samples of the input audio signal ( 100, 100 ′), wherein a noise reduced audio signal is formed by applying said set of gains to the input audio signal ( 100, 100 ′) and mixing the wind noise reduced audio signal and the noise reduced audio signal with a mixer ( 30 ) to obtain an output audio signal with suppressed wind noise.

Claims (68)

1 . A method for suppressing wind noise comprising:

obtaining an input audio signal comprising a plurality of consecutive audio signal segments;

suppressing wind noise in the input audio signal with a wind noise suppressor module to generate a wind noise reduced audio signal, the wind noise suppressor module comprising a high-pass filter;

using a neural network trained to predict a set of gains for reducing noise in the input audio signal given samples of the input audio signal, wherein a noise reduced audio signal is formed by applying said set of gains to the input audio signal; and

mixing the wind noise reduced audio signal and the noise reduced audio signal with a mixer to obtain an output audio signal with suppressed wind noise.

2 . The method according to claim 1 , further comprising:

determining, with a wind noise detector, a respective wind noise indicator for each segment of the input audio signal, the respective wind noise indicator indicating at least one of a probability and a magnitude of wind noise in each respective segment.

3 . The method according to claim 2 , further comprising:

determining, based on at least one respective wind noise indicator, a wind noise state.

4 . The method according to claim 3 , wherein the filter coefficients of the high-pass filter are based on the wind noise state.

5 . The method according to claim 3 , wherein each of said input audio signal, said wind noise reduced audio signal and said noise reduced audio signal comprises two audio channels, the method further comprising:

providing the wind noise state to a gain steering module;

if the wind noise state for said two channels exceeds a first threshold level or if a difference between the wind noise states of said two channel is below a second threshold level,

determining, a common set of gains based on the predicted set of gains of at least one of the two channels; and

applying the common set of gains to both channels;

else,

applying each set of gains to the corresponding channel.

6 . The method according to claim 5 , wherein the common set of gains is an average, a maximum or a minimum gain across the sets of gains.

7 . The method according to claim 3 , wherein determining, based on at least one respective wind noise indicator, a wind noise state comprises:

providing the wind noise indicator to a state machine with at least two states:

a no-wind-noise state, NWN, and

a wind-noise-hold state, WNH,

wherein the state machine transitions to the WNH state in response to detecting a first number of subsequent segments each having a respective wind noise indicator exceeding a high threshold, and outputs a high wind noise state, at least until the next state change, and

wherein the state machine transitions to the NWN state in response to detecting a second number of subsequent segments each having a respective wind noise indicator below a first low threshold, and outputs a low wind noise state at least until the next state change.

8 . The method according to claim 7 , wherein the state machine has four states,

the no-wind-noise state, NWN state,

the wind-noise-hold state, WNH state,

a wind-noise-attack state, WNA, and

a wind-noise-release state, WNR,

wherein the state machine transitions from the NWN state to the WNA state in response to detecting a segment having a respective wind noise indicator exceeding the high threshold and outputs the low wind noise state at least until the next state change,

wherein the state machine transitions from the WNA state to the WNH state in response to detecting the first number of subsequent segments having a respective wind noise indicator exceeding the high threshold and outputs a high wind noise state at least until the next state change,

wherein the state machine transitions from the WNH state to the WNR state in response to detecting a segment having a wind noise indicator being below the first low threshold and outputs the high wind noise state at least until the next state change,

wherein the state machine transitions from the WNR state to the NWN state in response to detecting the second number of subsequent segments being below the first low threshold and outputs the low wind noise state at least until the next state change.

9 . The method according to claim 8 , wherein the state machine transitions from the WNR state to the WNA state in response to detecting a segment having a respective wind noise indicator exceeding the high threshold and outputs the high wind noise state at least until the next state change.

10 . The method according to claim 3 , wherein determining the wind noise state comprises:

smoothing the wind noise indicator across at least two segments, wherein the wind noise state is based on the smoothed wind noise indicator.

11 . The method according to claim 3 , further comprising:

controlling, for each segment of the input audio signal, a mixing ratio of the mixer based on the wind noise state.

12 . The method according to claim 2 , wherein said input audio signal, wind noise reduced audio signal and noise reduced audio signal comprises two audio channels, and

wherein determining the wind noise indicator is based on monaural features of each individual audio channel and difference features associated with both audio channels.

13 . The method according to claim 2 , wherein said input audio signal, wind noise reduced audio signal and noise reduced audio signal comprises two audio channels, said method further comprising:

determining, with a Dual-Mono detector, whether a corresponding segment of each of the two audio channels are similar Dual-Mono segments or dissimilar non Dual-Mono segments based on a spectral power distribution of the segments in at least one frequency band;

if said segments are non Dual-Mono segments, determining the wind noise indicator based on monaural features and difference features associated with both segments;

else, determining the wind noise indicator based on only monaural features of each individual segments.

14 . The method according to claim 13 , wherein the step of determining whether the two audio segments are similar Dual-Mono segments or dissimilar non Dual-Mono segments comprises:

determining a first sum as the sum of the absolute value of the spectral difference between the two audio segments of the input audio signal;

determining a second sum as the sum of the total spectral energy for each of the segments;

calculating a ratio between said first sum and said second sum;

if said ratio is below a predetermined ratio threshold value and the second sum exceeds a predetermined sum threshold value, the segments are determined to be similar Dual-Mono segments; and

else, the segments are determined to be dissimilar non Dual-Mono segments.

15 . The method according to claim 2 , further comprising:

forming a non-real-time global wind noise indicator by aggregating the respective wind noise indicators for two or more segments of the plurality of audio segments of the input audio signal;

wherein the filter coefficients of the high-pass filter are based on the non-real-time global wind noise indicator.

16 . The method according to claim 2 , further comprising:

forming a non-real time global wind noise indicator by aggregating the respective wind noise indicators for two or more segments of the plurality of audio segments of the input audio signal; and

controlling, for each segment of the input audio signal, the mixing ratio of the mixer based on the non-real time global wind noise indicator.

17 . The method according to claim 1 , wherein each set of gains comprises a plurality of gains associated with a respective frequency band of the audio signal.

18 . The method according to claim 1 , wherein said mixing of the wind noise reduced audio signal and the noise reduced audio signal is performed with a fixed mixing ratio.

19 . A wind noise suppression system, comprising a processor and a memory coupled to the processor, the memory storing one or more programs including instructions for:

obtaining an input audio signal comprising a plurality of consecutive audio signal segments;

suppressing wind noise in the input audio signal with a wind noise suppressor module to generate a wind noise reduced audio signal, the wind noise suppressor module comprising a high-pass filter;

using a neural network trained to predict a set of gains for reducing noise in the input audio signal given samples of the input audio signal, wherein a noise reduced audio signal is formed by applying said set of gains to the input audio signal; and

mixing the wind noise reduced audio signal and the noise reduced audio signal with a mixer to obtain an output audio signal with suppressed wind noise.

20 . A non-transitory computer-readable storage medium storing one or more computer programs including instructions for:

obtaining an input audio signal comprising a plurality of consecutive audio signal segments;

suppressing wind noise in the input audio signal with a wind noise suppressor module to generate a wind noise reduced audio signal, the wind noise suppressor module comprising a high-pass filter;

using a neural network trained to predict a set of gains for reducing noise in the input audio signal given samples of the input audio signal, wherein a noise reduced audio signal is formed by applying said set of gains to the input audio signal; and

mixing the wind noise reduced audio signal and the noise reduced audio signal with a mixer to obtain an output audio signal with suppressed wind noise.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2024
From: BIN, QINGYUAN; MA, YUANXING; SHUANG, ZHIWEI
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 069689/0248 →
Priority Claims (1)
WO PCT/CN2022/080242 · Mar 10, 2022 · international
Continuity (3)
Provisional Application 63432996 · Dec 15, 2022
Provisional Application 63327030 · Apr 4, 2022
Related Publication 20250191601A1 · Jun 12, 2025
References Cited (20)
US 8781137B1 · Goodwin · 2014 [cited by applicant]
US 9313597B2 · Dickins et al. · 2016 [cited by applicant]
US 9343056B1 · Goodwin · 2016 [cited by applicant]
US 10721562B1 · Rui · 2020 [cited by applicant]
US 11217264B1 · Yang · 2022 [cited by examiner]
US 20070021958A1 · Visser · 2007 [cited by applicant]
US 20080310646A1 · Amada · 2008 [cited by applicant]
US 20120207325A1 · Taenzer · 2012 [cited by applicant]
US 20130308784A1 · Dickins · 2013 [cited by applicant]
US 20200380945A1 · Woodruff · 2020 [cited by applicant]
US 20200382859A1 · Woodruff · 2020 [cited by applicant]
US 20210110840A1 · Chu · 2021 [cited by examiner]
US 20210151069A1 · Hijazi · 2021 [cited by applicant]
US 20210195343A1 · Aubreville · 2021 [cited by applicant]
CN 112309417A · 2021 [cited by applicant]
DE 102010012941A1 · 2011 [cited by applicant]
JP 2008227595A · 2008 [cited by applicant]
JP 2014102317A · 2014 [cited by applicant]
WO 2022026948A1 · 2022 [cited by applicant]
Mitsuya, Takashi et al., Auditory feedback and articulatory timing, Program abstracts of the 158th meeting of the Acoustical Society of America, Journal of the Acoustical Society of America, 2009. 39 pages. [cited by applicant]