IP Library Granted Patent US 12701359
Granted Patent B2
US 12701359 · App. 18/571,119 · Granted Aug 4, 2026

Audio denoising method and device, apparatus and storage medium

Inventors: Xiaofeng Shu (Beijing, CN); Yehang Zhu (Beijing, CN); Chuxiang Shang (Beijing, CN); Yanjie Chen (Beijing, CN)
Assignee: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
H04R3/04G06N3/08H04R2430/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12701359
App. No.
18/571,119
Granted
Aug 4, 2026
Kind
B2
Abstract

An audio denoising method and device, an apparatus, a computer-readable storage medium, and a program product. The method includes: obtaining audio data to be denoised; estimating amplitude time-frequency mask of the audio data to be denoised by using a preset real-valued network model to obtain a first-order enhanced amplitude spectrum corresponding to the audio data to be denoised; estimating complex time-frequency masking of the audio data to be denoised by using a preset complex-valued network model; and determining denoising resulted audio data corresponding to the audio data to be denoised by combining the first-order enhanced amplitude spectrum with the complex time-frequency mask.

Claims (57)

1 . An audio denoising method, comprising:

acquiring audio data to be denoised;

estimating an amplitude time-frequency mask of the audio data to be denoised by using a preset real-valued network model, wherein the amplitude time-frequency mask is configured to determine a first-order enhanced amplitude spectrum corresponding to the audio data to be denoised;

estimating a complex time-frequency mask of the audio data to be denoised by using a preset complex-valued network model; and

determining denoising-resulted audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised;

wherein the amplitude time-frequency mask is configured to determine the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised, which comprises the case that:

the amplitude time-frequency mask is configured to be multiplied with an original amplitude spectrum of the audio data to be denoised to obtain the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised.

2 . The audio denoising method according to claim 1 , wherein the estimating the complex time-frequency mask of the audio data to be denoised by using the preset complex-valued network model comprises:

determining a complex frequency spectrum to be denoised; wherein the complex frequency spectrum to be denoised comprises a complex frequency spectrum determined based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and an original phase spectrum of the audio data to be denoised, or, a complex frequency spectrum determined based on an original frequency spectrum of the audio data to be denoised and the original phase spectrum of the audio data to be denoised; and

inputting the complex frequency spectrum to be denoised into the preset complex-valued network model, and outputting the complex time-frequency mask corresponding to the audio data to be denoised after a process of the preset complex-valued network model.

3 . The audio denoising method according to claim 2 , wherein the determining the denoising-resulted audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised comprises:

determining an amplitude gain and a phase gain based on the complex time-frequency mask;

determining an enhanced phase spectrum corresponding to the audio data to be denoised based on the phase gain and an original phase spectrum corresponding to the audio data to be denoised;

determining a second-order enhanced amplitude spectrum corresponding to the audio data to be denoised based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised; and

determining the denoising-resulted audio data corresponding to the audio data to be denoised based on the second-order enhanced amplitude spectrum and the enhanced phase spectrum.

4 . The audio denoising method according to claim 1 , wherein the determining the denoising-resulted audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised comprises:

determining an amplitude gain and a phase gain based on the complex time-frequency mask;

determining an enhanced phase spectrum corresponding to the audio data to be denoised based on the phase gain and an original phase spectrum corresponding to the audio data to be denoised;

determining a second-order enhanced amplitude spectrum corresponding to the audio data to be denoised based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised; and

determining the denoising-resulted audio data corresponding to the audio data to be denoised based on the second-order enhanced amplitude spectrum and the enhanced phase spectrum.

5 . The audio denoising method according to claim 1 , wherein the preset real-valued network model and the preset complex-valued network model are configured to form a two-stage time-domain convolutional network (TCN) model.

6 . The audio denoising method according to claim 5 , wherein before the estimating the amplitude time-frequency mask of the audio data to be denoised by using the preset real-valued network model, further comprising:

training the two-stage TCN model by using an audio training sample having a sampling rate higher than a preset sampling rate threshold.

7 . The audio denoising method according to claim 6 , wherein before the training the two-stage TCN model by using the audio training sample having a sampling rate higher than the preset sampling rate threshold, further comprising:

performing a preset data augmentation processing on the audio training sample to obtain an augmented audio training sample; and

correspondingly, the training the two-stage TCN model by using the audio training sample having a sampling rate higher than the preset sampling rate threshold comprises:

training the two-stage TCN model by using the augmented audio training sample, wherein a sampling rate of the augmented audio training sample is higher than the preset sampling rate threshold.

8 . The audio denoising method according to claim 7 , wherein the preset data augmentation processing comprises a high-passing process, a low-passing process, a band-passing process, a setting of different volumes and/or an equalizing process performed on the audio training sample according to a preset probability.

9 . The audio denoising method according to claim 6 , wherein the training the two-stage TCN model comprises:

training the two-stage TCN model by using a time-domain loss function scale-invariant signal-to-noise ratio (SISNR).

10 . An audio denoising device, comprising:

an acquisition module, configured to acquire audio data to be denoised;

a first estimation module, configured to estimate an amplitude time-frequency mask of the audio data to be denoised by using a preset real-valued network model; wherein the amplitude time-frequency mask is configured to determine a first-order enhanced amplitude spectrum corresponding to the audio data to be denoised;

a second estimation module, configured to estimate a complex time-frequency mask of the audio data to be denoised by using a preset complex-valued network model;

a determination module, configured to determine denoising-resulted audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised;

wherein the amplitude time-frequency mask is configured to determine the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised, which comprises the case that:

the amplitude time-frequency mask is configured to be multiplied with an original amplitude spectrum of the audio data to be denoised to obtain the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised.

11 . The audio denoising device according to claim 10 , wherein the second estimation module comprises:

a first determination submodule, configured to determine a complex frequency spectrum to be denoised; wherein the complex frequency spectrum to be denoised comprises a complex frequency spectrum determined based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and an original phase spectrum of the audio data to be denoised, or a complex frequency spectrum determined based on an original frequency spectrum of the audio data to be denoised and the original phase spectrum of the audio data to be denoised; and

a first processing submodule, configured to input the complex frequency spectrum to be denoised into the preset complex-valued network model, and output the complex time-frequency mask corresponding to the audio data to be denoised after a process of the preset complex-valued network model.

12 . The audio denoising device according to claim 11 , wherein the determination module comprises:

a second determination submodule, configured to determine an amplitude gain and a phase gain based on the complex time-frequency mask;

a third determination submodule, configured to determine an enhanced phase spectrum corresponding to the audio data to be denoised based on the phase gain and the original phase spectrum corresponding to the audio data to be denoised;

a fourth determination submodule, configured to determine a second-order enhanced amplitude spectrum corresponding to the audio data to be denoised based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised; and

a fifth determination submodule, configured to determine the denoising-resulted audio data corresponding to the audio data to be denoised based on the second-order enhanced amplitude spectrum and the enhanced phase spectrum.

13 . The audio denoising device according to claim 10 , wherein the determination module comprises:

a second determination submodule, configured to determine an amplitude gain and a phase gain based on the complex time-frequency mask;

a third determination submodule, configured to determine an enhanced phase spectrum corresponding to the audio data to be denoised based on the phase gain and an original phase spectrum corresponding to the audio data to be denoised;

a fourth determination submodule, configured to determine a second-order enhanced amplitude spectrum corresponding to the audio data to be denoised based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised; and

a fifth determination submodule, configured to determine the denoising-resulted audio data corresponding to the audio data to be denoised based on the second-order enhanced amplitude spectrum and the enhanced phase spectrum.

14 . The audio denoising device according to claim 10 , wherein the preset real-valued network model and the preset complex-valued network model are configured to form a two-stage time-domain convolutional network (TCN) model.

15 . The audio denoising device according to claim 14 , further comprising: a training module, configured to train the two-stage TCN model by using an audio training sample having a sampling rate higher than a preset sampling rate threshold.

16 . The audio denoising device according to claim 15 , further comprising:

an augmentation module, configured to perform a preset data augmentation processing on the audio training sample to obtain an augmented audio training sample; and

the training module is configured to train the two-stage TCN model by using the augmented audio training sample; wherein the augmented audio training sample has a sampling rate higher than the preset sampling rate threshold.

17 . The audio denoising device according to claim 16 , wherein the preset data augmentation processing comprises a high-passing process, a low-passing process, a band-passing process, a setting of different volumes and/or an equalizing process performed on the audio training sample according to a preset probability.

18 . The audio denoising device according to claim 16 , wherein the training module is configured to train the two-stage TCN model by using a time-domain loss function scale-invariant signal-to-noise ratio (SISNR).