IP Library › Granted Patent US 12,192,720
Granted Patent B1
US 12,192,720 · App. 18/141,964 · Granted Jan 7, 2025

Audio noise determination using one or more neural networks

Inventors: Mihir Nyayate (Pune, IN); Angshuman Ghosh (Kolkata, IN); Revanth Reddy Nalla (Pune, IN); Ambrish Dantrey (Pune, IN)
Assignee: NVIDIA Corporation
H04R5/04G06N3/084G06N5/046G10L21/0232G10L25/84H04R3/04H04R5/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,192,720
App. No.
18/141,964
Filed
May 1, 2023
Granted
Jan 7, 2025
Kind
B1
Art Unit
2691
USPC
381/94.1
Abstract

Apparatuses, systems, and techniques are presented to reduce noise in audio. In at least one embodiment, one or more neural networks are used to determine a noise signal in one or more speech signals.

Claims (26)

1. A processor, comprising:

one or more circuits to use one or more neural networks to identify one or more noise signals in one or more signals based, at least in part, on two or more portions of the one or more neural networks to identify different portions of the one or more signals in parallel.

2. The processor of claim 1 , wherein the one or more circuits are further to generate an audio spectrogram corresponding to one or more features extracted from the one or more signals.

3. The processor of claim 1 , wherein the one or more circuits are further to provide an audio spectrogram as input to the one or more neural networks, wherein the one or more neural networks generate an audio mask corresponding to the noise signal identified in the one or more signals.

4. The processor of claim 2 , wherein the audio spectrogram comprises a mel spectrogram.

5. The processor of claim 1 , wherein the one or more neural networks include two parallel paths to determine patterns in an audio spectrogram, the two parallel paths including a first path comprising a sequence of convolutional layers and a second path comprising one or more gated recurrent unit (GRU) layers.

6. The processor of claim 5 , wherein the one or more circuits are further to concatenate the patterns determined by the two parallel paths and process those concatenated patterns using a sequence of GRU layers to identify important noise patterns in the one or more signals to generate the audio mask.

7. The processor of claim 6 , wherein the one or more circuits are further to invert the audio mask and apply the inverted audio mask to the audio spectrogram to generate an output audio signal with the noise signal removed from the one or more signals.

8. A system comprising:

one or more processors to use one or more neural networks to identify one or more noise signals in one or more signals based, at least in part, on two or more portions of the one or more neural networks to identify different portions of the one or more signals in parallel.

9. The system of claim 8 , wherein the one or more processors are further to generate an audio spectrogram corresponding to one or more features extracted from the one or more signals.

10. The system of claim 8 , wherein the one or more processors are further to provide an audio spectrogram as input to the one or more neural networks, wherein the one or more neural networks generate an audio mask corresponding to the noise signal identified in the one or more signals.

11. The system of claim 9 , wherein the audio spectrogram comprises a mel spectrogram.

12. The system of claim 8 , wherein the one or more neural networks include two parallel paths to determine patterns in an audio spectrogram, the two parallel paths including a first path comprising sequence of convolutional layers and a second path comprising a gated recurrent unit (GRU) layer.

13. The system of claim 12 , wherein the one or more processors are further to concatenate the patterns determined by the two parallel paths and process those concatenated patterns using a sequence of GRU layers to identify important noise patterns in the one or more signals to generate an audio mask.

14. The system of claim 13 , wherein the one or more processors are further to invert the audio mask and apply the inverted audio mask to the audio spectrogram to generate an output audio signal with the noise signal removed from the one or more signals.

15. A method comprising:

using one or more neural networks to identify one or more noise signals in one or more signals based, at least in part, on two or more portions of the one or more neural networks to identify different portions of the two or more signals in parallel.

16. The method of claim 15 , further comprising:

generating an audio spectrogram corresponding to one or more features extracted from the one or more signals.

17. The method of claim 15 , further comprising:

providing an audio spectrogram as input to the one or more neural networks, wherein the one or more neural networks generate an audio mask corresponding to the noise signal identified in the one or more signals.

18. The method of claim 15 , wherein the one or more neural networks include two parallel paths for determining patterns in an audio spectrogram, the two parallel paths including a first path comprising a sequence of convolutional layers and a second path comprising a gated recurrent unit (GRU) layer.

19. The method of claim 18 , further comprising concatenating the patterns determined by the two parallel paths and process those concatenated patterns using a sequence of GRU layers to identify important noise patterns in the one or more signals to generate the audio mask.

20. The method of claim 19 , further comprising:

inverting the audio mask and apply the inverted audio mask to the audio spectrogram in order to generate an output audio signal with the noise signal removed from the one or more signals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2023
From: NYAYATE, MIHIR; GHOSH, ANGSHUMAN; NALLA, REVANTH REDDY; DANTREY, AMBRISH
To: NVIDIA CORPORATION
Reel/Frame 063499/0274 →
Continuity (1)
Continuation 16874171 · May 14, 2020
References Cited (27)
US 10614827B1 · Korjani · 2020 [cited by applicant]
US 20170092268A1 · Kristjansson · 2017 [cited by applicant]
US 20180190268A1 · Lee · 2018 [cited by examiner]
US 20180358003A1 · Calle · 2018 [cited by examiner]
US 20190318755A1 · Tashev et al. · 2019 [cited by applicant]
US 20200066296A1 · Sargsyan · 2020 [cited by examiner]
US 20200243102A1 · Schmidt · 2020 [cited by examiner]
US 20210118462A1 · Tommy · 2021 [cited by examiner]
US 20210350796A1 · Kim et al. · 2021 [cited by applicant]
CN 103337248A · 2013 [cited by applicant]
CN 107452389A · 2017 [cited by applicant]
CN 107564538A · 2018 [cited by applicant]
CN 107609488A · 2018 [cited by applicant]
CN 109036460A · 2018 [cited by applicant]
CN 109410974A · 2019 [cited by applicant]
CN 110503972A · 2019 [cited by applicant]
CN 110600018A · 2019 [cited by applicant]
Das et al., “Das-Vector Taylor Series Expansion with Auditory Masking for Noise Robust Speech Recognition,” Oct. 17-20, 2016, 5 pages. [cited by applicant]
Grzywalski et al., “Application of Recuurent U-net Architecture to Speech Enhancement,” IEEE The Institute of Electrical and Electronics Engineering Inc, Sep. 19-21, 2018, 6 pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
Mandel et al., “Analysis-by-Synthesis Feature Estimation for Robust Automatic Speech Recognition using Spectral Masks,” Jul. 14, 2014, 5 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Patent Application No. 2106937.2 dated Mar. 10, 2022, 9 pages. [cited by applicant]
Office Action for German Application No. 102021112250.3, mailed Sep. 29, 2023, 16 pages. [cited by applicant]
Office Action for Chinese Application No. 202110523004.8, mailed Dec. 5, 2023, 11 pages. [cited by applicant]
United Kingdom Search Report for Application No. GB2306445.4, mailed Mar. 11, 2024, 4 pages. [cited by applicant]
Office Action for Chinese Application No. 202110523004.8, mailed Aug. 21, 2024, 25 pages. [cited by applicant]
Office Action for Chinese Application No. 202110523004.8, mailed Nov. 15, 2024, 33 pages. [cited by applicant]
Cited By (1)
US 12,412,590