IP Library Granted Patent US 12,456,476
Granted Patent B2
US 12,456,476 · App. 18/081,492 · Granted Oct 28, 2025

Noise suppression for speech data with reduced power consumption

Inventors: Chandan Karadagur Ananda Reddy (Cupertino, CA); Navin Chatlani (Palo Alto, CA)
Assignee: Google LLC
G10L21/0232G10L21/0224G10L25/60G10L2021/02163
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,476
App. No.
18/081,492
Granted
Oct 28, 2025
Kind
B2
Abstract

Implementations described herein relate to providing noise suppression for speech data with reduced power consumption. In some implementations, a computer-implemented method includes receiving a current time frame of speech data, e.g., after receiving a previous time frame associated with a previous noise suppression mask. The current time frame is transformed to a current frequency frame in the frequency domain. A noise classifier is used to determine whether to create a current noise suppression mask for the current frame. If it is determined to create the mask, the mask is created and multiplied by the current frequency frame to obtain a noise-suppressed frequency frame. If it is determined to not create the current mask, the previous noise suppression mask is multiplied with the current frequency frame to obtain the noise-suppressed frequency frame, without creating a mask. The noise-suppressed frequency frame is transformed to a time frame and output.

Claims (72)

1. A computer-implemented method comprising:

receiving, by one or more processors, a current time frame of speech data in a time domain;

transforming the current time frame to a current frequency frame of the speech data in a frequency domain;

determining, using a noise classifier implemented by the one or more processors, a classification of noise content in the current frequency frame based at least on noise content determined in the current frequency frame, wherein the classification is selected from at least two classifications detectable by the noise classifier, wherein the two classifications include a first classification associated with creating and applying a current noise suppression mask to the current frequency frame, and a second classification associated with applying a previously-created noise suppression mask to the current frequency frame without creating the current noise suppression mask;

in response to determining the first classification:

creating, by the one or more processors, the current noise suppression mask for the current frequency frame based on the noise content in the current frequency frame, wherein creating the current noise suppression mask includes determining one or more gain functions associated with one or more frequency bands of the current frequency frame; and

applying, by the one or more processors, the current noise suppression mask to the current frequency frame to suppress the noise content in the current frequency frame and obtain a current noise-suppressed frequency frame of the speech data;

in response to determining the second classification:

selecting, by the one or more processors, the previously-created noise suppression mask, wherein the previously-created noise suppression mask was created and applied based on a prior frequency frame of the speech data; and

applying, by the one or more processors, the previously-created noise suppression mask to the current frequency frame to suppress the noise content in the current frequency frame and obtain the current noise-suppressed frequency frame of the speech data, without creating the current noise suppression mask;

transforming the current noise-suppressed frequency frame of the speech data to a current noise-suppressed time frame of the speech data that is in the time domain; and

outputting the current noise-suppressed time frame of the speech data that provides audio output having reduced noise.

2. The method of claim 1 , wherein the noise classifier uses a machine-learning model that is trained to determine the classification of noise content in frequency frames provided as input to the machine-learning model.

3. The method of claim 2 , wherein the machine-learning model is trained based on speech data that does not include noise and noise data that includes background noise.

4. The method of claim 2 , wherein the machine-learning model is trained using a speech quality predictor that estimates a quality of the speech data.

5. The method of claim 1 , wherein determining the classification of the noise content in the current frequency frame is based on a determined magnitude of the noise content determined in the current frequency frame.

6. The method of claim 1 , wherein determining the classification of the noise content in the current frequency frame is based on a determined rate of change of the noise content in the current time frame relative to determined noise content in one or more previous time frames.

7. The method of claim 1 , further comprising:

receiving a second time frame of the speech data in the time domain;

transforming the second time frame to a second frequency frame of the speech data in the frequency domain;

determining, using the noise classifier, whether to create a second noise suppression mask for the second frequency frame based at least on noise content determined in the second frequency frame;

in response to determining to create the second noise suppression mask:

creating a second noise suppression mask based on the second frequency frame, wherein creating the second noise suppression mask includes determining one or more gain functions associated with one or more frequency bands of the second frequency frame; and

applying the second noise suppression mask to the second frequency frame to suppress the noise content in the second frequency frame and obtain a second noise-suppressed frequency frame;

in response to determining to not create the second noise suppression mask:

applying the current noise suppression mask to the second frequency frame to suppress the noise content in the second frequency frame and obtain the second noise-suppressed frequency frame of the speech data, without the creating and the applying of the second noise suppression mask;

transforming the second noise-suppressed frequency frame to a second noise-suppressed time frame; and

outputting the second noise-suppressed time frame of the speech data.

8. The method of claim 1 , wherein determining the classification of the noise content in the current frequency frame is based on the determined noise content in the current frequency frame and determined noise content in one or more previous frequency frames.

9. The method of claim 1 , wherein applying the current noise suppression mask to the current frequency frame is performed without applying the previously-created noise suppression mask.

10. The method of claim 1 , wherein selecting the previously-created noise suppression mask includes selecting a most recent previously-created noise suppression mask created for a most recent prior frequency frame of the speech data.

11. The method of claim 1 , wherein the noise classifier includes a statistical classifier, wherein the statistical classifier:

compares data in the current frequency frame to previous data in one or more previous frequency frames using a similarity function or a distance function; and

based on the comparison, determines whether a rate of change of noise in the one or more previous frequency frames and the current frequency frame meets a threshold that indicates the classification of the noise content in the current frequency frame.

12. A computing device comprising:

a processor; and

a memory coupled to the processor, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:

receiving, by the processor, a current time frame of speech data in a time domain after receiving a first time frame of the speech data in the time domain, wherein the current time frame and the first time frame include speech content, and wherein a first noise suppression mask is associated with the first time frame and created to suppress noise in the first time frame of the speech data;

transforming the current time frame to a current frequency frame of the speech data in a frequency domain;

determining, using a noise classifier, a classification of noise content in the current frequency frame based at least on noise content determined in the current frequency frame, wherein the classification is selected from at least two classifications detectable by the noise classifier, wherein the two classifications include a first classification associated with creating and applying a current noise suppression mask to the current frequency frame, and a second classification associated with applying a previously-created noise suppression mask to the current frequency frame without creating the current noise suppression mask;

in response to determining the first classification:

creating the current noise suppression mask for the current frequency frame based on the noise content in the current frequency frame, wherein creating the current noise suppression mask includes determining one or more gain functions associated with one or more frequency bands of the current frequency frame; and

applying the current noise suppression mask to the current frequency frame to suppress the noise content in the current frequency frame and obtain a current noise-suppressed frequency frame of the speech data;

in response to determining the second classification:

selecting the first noise suppression mask; and

applying the first noise suppression mask to the current frequency frame to suppress the noise content in the current frequency frame and obtain the current noise-suppressed frequency frame of the speech data, without the creating and the applying of the current noise suppression mask;

transforming the current noise-suppressed frequency frame of the speech data to a current noise-suppressed time frame of the speech data that is in the time domain; and

outputting the current noise-suppressed time frame of the speech data.

13. The computing device of claim 12 , wherein the noise classifier uses a machine-learning model that is trained to determine the classification of noise content in frequency frames provided as input to the machine-learning model.

14. The computing device of claim 12 , wherein the noise classifier includes a statistical classifier, wherein the statistical classifier:

compares data in the current frequency frame to previous data in one or more previous frames using a similarity function or a distance function; and

based on the comparison, determines whether a rate of change of noise in the current frequency frame and the one or more previous frames meets a threshold that indicates the classification of the noise content in the current frequency frame.

15. A device comprising:

at least one battery;

at least one microphone;

a communication circuit coupled to the battery and the microphone;

a processor coupled to the communication circuit, the battery and the microphone; and

a memory coupled to the processor, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:

receiving, by the processor via the at least one microphone, a current time frame of speech data in a time domain;

transforming the current time frame of speech data to a current frequency frame of the speech data in a frequency domain;

determining, using a noise classifier, a classification of noise content in the current frequency frame based at least on noise content determined in the current frequency frame wherein the classification is selected from at least two classifications detectable by the noise classifier, wherein the two classifications include a first classification associated with creating and applying a current noise suppression mask to the current frequency frame, and a second classification associated with applying a previously-created noise suppression mask to the current frequency frame without creating the current noise suppression mask;

in response to determining the first classification:

creating the current noise suppression mask for the current frequency frame based on the noise content in the current frequency frame, wherein creating the current noise suppression mask includes determining one or more gain functions associated with one or more frequency bands of the current frequency frame; and

applying the current noise suppression mask to the current frequency frame to suppress the noise content in the current frequency frame and obtain a current noise-suppressed frequency frame of the speech data;

in response to determining the second classification:

selecting the previously-created noise suppression mask, wherein the previously-created noise suppression mask was created based on a prior frequency frame of the speech data; and

applying the previously-created noise suppression mask to the current frequency frame to suppress the noise content in the current frequency frame and obtain the current noise-suppressed frequency frame of the speech data, without creating the current noise suppression mask; and

transforming the current noise-suppressed frequency frame of the speech data to a current noise-suppressed time frame of the speech data that is in the time domain.

16. The device of claim 15 , wherein the noise classifier uses a machine-learning model that is trained to determine the classification of noise content in frequency frames provided as input to the machine-learning model.

17. The device of claim 15 , wherein the noise classifier includes a statistical classifier, wherein the statistical classifier:

compares data in the current frequency frame to previous data in one or more previous frequency frames using a similarity function or a distance function; and

based on the comparison, determines whether a rate of change of noise in the current frequency frame and the one or more previous frequency frames meets a threshold that indicates the classification of the noise content in the current frequency frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: KARADAGUR ANANDA REDDY, CHANDAN; CHATLANI, NAVIN
To: GOOGLE LLC
Reel/Frame 062095/0077 →
Continuity (1)
Related Publication 20240203438A1 · Jun 20, 2024
References Cited (26)
US 6785339B1 · Tahernezhaadi et al. · 2004 [cited by applicant]
US 9640194B1 · Nemala · 2017 [cited by examiner]
US 20050278172A1 · Koishida et al. · 2005 [cited by applicant]
US 20130304461A1 · Taleb et al. · 2013 [cited by applicant]
US 20220124433A1 · Kupryjanow · 2022 [cited by examiner]
US 20230125150A1 · Cutler · 2023 [cited by examiner]
US 20240071356A1 · Gu · 2024 [cited by examiner]
US 20240296856A1 · Huang · 2024 [cited by examiner]
Dubey, Harishchandra et al., “ICASSP 2022 Deep Noise Suppression Challenge”, arXiv:2202.13288v1, Feb. 27, 2022, 5 pages. [cited by applicant]
EPO, International Search Report for International Patent Application No. PCT/US2023/033592, Feb. 21, 2024, 4 pages. [cited by applicant]
EPO, Written Opinion for International Patent Application No. PCT/US2023/033592, Feb. 21, 2024, 9 pages. [cited by applicant]
Esquef, Paulo A. et al., “A Double-threshold Based Approach to Impulsive Noise Detection in Audio Signals”, 2000 10th European Signal Processing Conference, 2000, 4 pages. [cited by applicant]
Moattar, M. H. et al., “A Simple but Efficient Real-Time Voice Activity Detection Algorithm”, 17th European Signal Processing Conference (EUSIPCO 2009), 2009, 5 pages. [cited by applicant]
Oudre, Laurent, “Automatic Detection and Removal of Impulsive Noise in Audio Signals”, Image Processing On Line, Nov. 21, 2015, 15 pages. [cited by applicant]
Rix, Antony et al., “Perceptual Evaluation of Speech Quality (PESQ)—A New Method for Speech Quality Assessment of Telephone Networks and Codecs”, 2001 IEEE international conference on acoustics, speech, and signal proce… [cited by applicant]
Tagliasacchi, Marco et al., “SEANet: a Multi-modal Speech Enhancement Network”, INTERSPEECH 2020, Oct. 2020, 5 pages. [cited by applicant]
Reddy, et al., “Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors”, ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)… [cited by applicant]
Reddy, et al., “Dnsmos: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (htt… [cited by applicant]
Lo, et al., “MOSNet: Deep Learning based Objective Assessment for Voice Conversion”, Institute of Information Science (https://arxiv.org/pdf/1904.08352.pdf), Jul. 14, 2021, 5 pages. [cited by applicant]
Zhang, et al., “Multi-Scale Temporal Frequency Convolutional Network with Axial Attention for Multi-Channel Speech Enhancement”, ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing … [cited by applicant]
Kong, et al., “Speech enhancement with weakly labelled data from AudioSet”, (https://arxiv.org/abs/2102.09971), Feb. 19, 2021, 5 Pages. [cited by applicant]
Zorila, et al., “A fast algorithm for improved intelligibility of speech-in-noise based on frequency and time domain energy reallocation”, INTERSPEECH, Sep. 2015, p. 60-64. [cited by applicant]
Zezario, et al., “Deep Learning-Based Non-Intrusive Multi-Objective Speech Assessment Model With Cross-Domain Features”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31 (https://ieeexplore.ieee.… [cited by applicant]
Braun, et al., “Towards Efficient Models for Real-Time Deep Noise Suppression”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (https://arxiv.org/pdf/2101.09249.pdf),… [cited by applicant]
Abdulatif, et al., “CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement”, (https://arxiv.org/pdf/2209.11112.pdf), Sep. 23, 2022, 16 pages. [cited by applicant]
Ju, et al., “TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS Challenge”, ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP),… [cited by applicant]