IP Library Granted Patent US 12,412,590
Granted Patent B1
US 12,412,590 · App. 16/721,881 · Granted Sep 9, 2025

Audio noise removal using one or more neural networks

Inventors: Ambrish Dantrey (Pune, IN); Angshuman Ghosh (Kolkata, IN); Mihir Nyayate (Pune, IN); Abhijit Patait (Pune, IN)
Assignee: NVIDIA Corporation
G10L21/0232G06N3/044G06N3/045G10L15/16G10L15/22G10L25/18G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,590
App. No.
16/721,881
Granted
Sep 9, 2025
Kind
B1
Abstract

Apparatuses, systems, and techniques are presented to reduce noise in audio. In at least one embodiment, a sequence of neural networks is used to remove foreground and background noise from audio including a primary audio signal.

Claims (55)

1. One or more processors, comprising:

one or more circuits to:

cause one or more digital signals to be filtered, where one or more portions of speech are removed from the one or more filtered digital signals;

generate, by one or more neural networks, the one or more portions of speech filtered from the one or more digital signals based, at least in part, on the one or more filtered digital signals; and

add the one or more portions of speech, generated by the one or more neural networks, to the one or more filtered digital signals.

2. The one or more processors of claim 1 , wherein the one or more circuits are to cause one or more second neural networks to filter noise from the one or more digital signals, the one or more second neural networks comprising a frequency-domain deep neural network (DNN) including one or more gated recurrent unit (GRU) layers.

3. The one or more processors of claim 1 , wherein the one or more neural networks comprise a flow-based, generative network to perform audio generation, the flow-based, generative network caused to generate the one or more portions of speech at least by restoring primary audio modified in the one or more filtered digital signals.

4. The one or more processors of claim 1 , wherein the one or more circuits are further to use an auto-encoder network to enhance the one or more portions of speech generated by the one or more neural networks.

5. The one or more processors of claim 1 , wherein the one or more digital signals include a primary signal, corresponding to speech of a determined person, and noise including speech uttered by at least one other person.

6. The one or more processors of claim 1 , wherein the one or more circuits are further to provide the one or more filtered digital signals as input to the one or more neural networks as one or more spectrograms.

7. A system, comprising:

one or more processors to:

cause one or more digital signals to be filtered, where one or more portions of speech are removed from the one or more filtered digital signals;

generate, by one or more neural networks, the one or more portions of speech filtered from the one or more digital signals based, at least in part, on the one or more filtered digital signals; and

add the one or more portions of speech, generated by the one or more neural networks, to the one or more filtered digital signals.

8. The system of claim 7 , wherein the one or more processors are to cause one or more second neural networks to filter noise from the one or more digital signals, the one or more second neural networks comprising a frequency-domain deep neural network (DNN) including one or more gated recurrent unit (GRU) layers.

9. The system of claim 7 , wherein the one or more neural networks comprise a flow-based, generative network to perform audio generation, the flow-based, generative network caused to generate the one or more portions of speech at least by restoring primary audio modified in the one or more filtered digital signals.

10. The system of claim 7 , wherein the one or more processors are further to use an auto-encoder network to enhance the one or more portions of speech generated by the one or more neural networks.

11. The system of claim 7 , wherein the one or more digital signals include a primary signal, corresponding to speech of a determined person, and noise including speech uttered by at least one other person.

12. The system of claim 7 , wherein the one or more processors are further to provide the one or more filtered digital signals as input to the one or more neural networks as one or more spectrograms.

13. A method, comprising:

causing one or more digital signals to be filtered, where one or more portions of speech are removed from the one or more filtered digital signals;

generating, by one or more neural networks, the one or more portions of speech filtered from the one or more digital signals based, at least in part, on the one or more filtered digital signals; and

adding the one or more portions of speech, generated by the one or more neural networks, to the one or more filtered digital signals.

14. The method of claim 13 , wherein causing the one or more digital signals to be filtered comprises causing one or more second neural networks to filter noise from the one or more digital signals, the one or more second neural networks comprising a frequency-domain deep neural network (DNN) including one or more gated recurrent unit (GRU) layers.

15. The method of claim 13 , wherein the one or more neural networks comprise a flow-based, generative network to perform audio generation, and wherein generating, by the one or more neural networks, the one or more portions of speech comprises causing the flow-based, generative network to restore primary audio modified in the one or more filtered digital signals.

16. The method of claim 13 , further comprising using an auto-encoder network to enhance the one or more portions of speech generated by the one or more neural networks.

17. The method of claim 13 , wherein the one or more digital signals include a primary signal, corresponding to speech of a determined person, and noise including speech uttered by at least one other person.

18. The method of claim 13 , further comprising:

providing the one or more filtered digital signals as input to the one or more neural networks as one or more spectrograms.

19. A non-transitory computer-readable storage medium having stored thereon a set of instructions which, if performed by one or more processors, cause the one or more processors to at least:

cause one or more digital signals to be filtered, where one or more portions of speech are removed from the one or more filtered digital signals;

generate, by one or more neural networks, the one or more portions of speech filtered from the one or more digital signals based, at least in part, on the one or more filtered digital signals; and

add the one or more portions of speech, generated by the one or more neural networks, to the one or more filtered digital signals.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the set of instructions, if performed by the one or more processors, cause the one or more processors to cause one or more second neural networks to filter noise from the one or more digital signals, the one or more second neural networks comprising a frequency-domain deep neural network (DNN) including one or more gated recurrent unit (GRU) layers.

21. The non-transitory computer-readable storage medium of claim 19 , wherein the one or more neural networks comprise a flow-based, generative network to perform audio generation, and wherein the set of instructions which, if performed by the one or more processors, cause the one or more processors to generate, by the one or more neural networks, the one or more portions of speech further cause the one or more processors to cause the flow-based, generative network to restore primary audio modified in the one or more filtered digital signals.

22. The non-transitory computer-readable storage medium of claim 19 , wherein the set of instructions, if performed by the one or more processors, further cause the one or more processors to:

use an auto-encoder network to enhance the one or more portions of speech generated by the one or more neural networks.

23. The non-transitory computer-readable storage medium of claim 19 , wherein the one or more digital signals include a primary signal, corresponding to speech of a determined person, and noise including speech uttered by at least one other person.

24. The non-transitory computer-readable storage medium of claim 19 , wherein the set of instructions, if performed by the one or more processors, further cause the one or more processors to:

provide the one or more filtered digital signals as input to the one or more neural networks as one or more spectrograms.

25. An audio de-noising system, comprising:

one or more processors to:

cause one or more digital signals to be filtered, where one or more portions of speech are removed from the one or more filtered digital signals;

generate, by one or more neural networks, the one or more portions of speech filtered from the one or more digital signals based, at least in part, on the one or more filtered digital signals; and

add the one or more portions of speech, generated by the one or more neural networks, to the one or more filtered digital signals; and

memory for storing network parameters to be used by the one or more neural networks.

26. The audio de-noising system of claim 25 , wherein the one or more processors are to cause one or more second neural networks to filter noise from the one or more digital signals, the one or more second neural networks comprising a frequency-domain deep neural network (DNN) including one or more gated recurrent unit (GRU) layers.

27. The audio de-noising system of claim 25 , wherein the one or more neural networks comprise a flow-based, generative network to perform audio generation, the flow-based, generative network caused to generate the one or more portions of speech at least by restoring primary audio modified in the one or more filtered digital signals.

28. The audio de-noising system of claim 25 , wherein the one or more processors are further to use an auto-encoder network for enhancing the one or more portions of speech generated by the one or more neural networks.

29. The audio de-noising system of claim 25 , wherein the one or more digital signals include a primary signal, corresponding to speech of a determined person, and noise including speech uttered by at least one other person.

30. The audio de-noising system of claim 25 , wherein the one or more processors are further to provide the one or more filtered digital signals as input to the one or more neural networks as one or more spectrograms.

31. The one or more processors of claim 1 , wherein the one or more circuits are further to cause one or more neural networks to filter noise from the one or more filtered digital signals.

32. The one or more processors of claim 1 , wherein the one or more circuits are to cause one or more second neural networks to filter the one or more digital signals such that the one or more portions of speech are removed.

33. The one or more processors of claim 1 , wherein the one or more circuits are to cause the one or more neural networks to add the one or more portions of speech, generated by the one or more neural networks, to the one or more filtered digital signals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2019
From: DANTREY, AMBRISH; GHOSH, ANGSHUMAN; NYAYATE, MIHIR; PATAIT, ABHIJIT
To: NVIDIA CORPORATION
Reel/Frame 051357/0824 →
References Cited (15)
US 11678120B2 · Nyayate · 2023 [cited by examiner]
US 12192720B1 · Nyayate · 2025 [cited by examiner]
US 20170270919A1 · Parthasarathi · 2017 [cited by examiner]
US 20180122403A1 · Koretzky · 2018 [cited by examiner]
US 20180233127A1 · Visser · 2018 [cited by examiner]
US 20210360349A1 · Nyayate · 2021 [cited by examiner]
Vincent, “Multichannel Audio Source Separation With Deep Neural Networks,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, No. 9, pp. 1652-1664, Sep. 2016, doi: 10.1109/TASLP.2016.2580946 (Y… [cited by examiner]
Grais et al., “Two-Stage Single-Channel Audio Source Separation Using Deep Neural Networks,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, No. 9, pp. 1773-1783, Sep. 2017, doi: 10.1109/TAS… [cited by examiner]
Ghaffarzadegan et al., “Bosch Rare Sound Events Detection Systems for DCASE2017 Challenge,” Detection and Classification of Acoustic Scenes and Events 2017, Nov. 16, 2017. (Year: 2017). [cited by examiner]
Prenger et al., “Waveglow: A Flow-Based Generative Network for Speech Synthesis,” arXiv:1811.00002v1 [cs.SD], https://doi.org/10.48550/arXiv.1811.00002, Oct. 21, 2018 (Year: 2018). [cited by examiner]
Purwins et al., “Deep Learning for Audio Signal Processing,” in IEEE Journal of Selected Topics in Signal Processing, vol. 13, No. 2, pp. 206-219, May 2019, doi: 10.1109/JSTSP.2019.2908700 (Year: 2019). [cited by examiner]
Xu et al., “Convolutional Gated Recurrent Neural Network Incorporating Spatial Features for Audio Tagging,” arXiv:1702.07787v1 [cs.SD], https://doi.org/10.48550/arXiv.1702.07787, Feb. 24, 2017 (Year: 2017). [cited by examiner]
Lu et al., “Speech Enhancement Based on Deep Denoising Autoencoder,” Proc. Interspeech 2013, 436-440, doi: 10.21437/Interspeech.2013-130 (Year: 2013). [cited by examiner]
Grzywalski, T., & Drgas, S. (Sep. 2018). Application of recurrent U-net architecture to speech enhancement. In 2018 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA) (pp. 82-87). IEEE. (… [cited by examiner]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]