IP Library › Granted Patent US 12,057,131
Granted Patent B2
US 12,057,131 · App. 17/662,418 · Granted Aug 6, 2024

Deep learning segmentation of audio using magnitude spectrogram

Inventor: Luke Miner (San Francisco, CA)
Assignee: AUDIOSHAKE, INC.
G10L19/0216G06N3/08G10L21/0272G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,057,131
App. No.
17/662,418
Granted
Aug 6, 2024
Kind
B2
Abstract

A method, system, and computer readable medium for decomposing an audio signal into different isolated sources. The techniques and mechanisms convert an audio signal into K input spectrogram fragments. The fragments are sent into a deep neural network to isolate for different sources. The isolated fragments are then combined to form full isolated source audio signals.

Claims (40)

1. A method for isolating a source from an audio signal, the method comprising:

transforming an audio file into a complex spectrogram;

decomposing the complex spectrogram into a magnitude spectrogram and a phase spectrogram;

splitting the magnitude spectrogram into K small fragments along the time dimension;

sending each fragment in the K small fragments through a convolutional deep neural network to produce K small fragment outputs, wherein the convolutional deep neural network includes convolutional layers augmented with an attention layer to help the convolutional deep neural network find the most informative areas and reduce artifacts in each of the K small fragments, wherein the attention layer informs the convolutional deep neural network of important pixels;

producing a source mask based on the K small fragment outputs, wherein the source mask corresponds to an audio source targeted to be isolated;

multiplying the source mask with the magnitude spectrogram to create a new magnitude spectrogram corresponding to the source.

2. The method of claim 1 , further comprising combining the new magnitude spectrograms with the phase spectrogram in order to produce a new complex spectrogram.

3. The method of claim 1 , wherein transforming the audio file into the complex spectrogram is done via a short-time fourier transform.

4. The method of claim 1 , wherein a second deep neural network is used for isolating a second audio source.

5. The method of claim 1 , wherein the deep neural network includes an input scale layer before a series of downsample layers and an output scale layer following a series of upsample layers.

6. The method of claim 1 , wherein the deep neural network includes a bridge layer comprising a first convolutional 2D layer and a second convolutional 2D layer.

7. The method of claim 1 , further comprising constructing a new phase spectrogram from a new source using a generative adversarial neural network.

8. The method of claim 1 , further comprising constructing a new phase spectrogram from a new source using the Griffin-Lim algorithm.

9. The method of claim 1 , further comprising applying a multi-channel wiener filter to the new magnitude spectrogram.

10. The method of claim 1 , wherein the source mask is produced by concatenating the K small fragment outputs.

11. A system for isolating a source from an audio signal, the system comprising:

a processor; and

memory storing instructions to cause the processor to execute a method, the method comprising:

transforming an audio file into a complex spectrogram;

decomposing the complex spectrogram into a magnitude spectrogram and a phase spectrogram;

splitting the magnitude spectrogram into K small fragments along the time dimension;

sending each fragment in the K small fragments through a convolutional deep neural network to produce K small fragment outputs, wherein the convolutional deep neural network includes convolutional layers augmented with an attention layer to help the convolutional deep neural network find the most informative areas and reduce artifacts in each of the K small fragments, wherein the attention layer informs the convolutional deep neural network of important pixels;

producing a source mask based on the K small fragment outputs, wherein the source mask corresponds to an audio source targeted to be isolated;

multiplying the source mask with the magnitude spectrogram to create a new magnitude spectrogram corresponding to the source.

12. The system of claim 11 , wherein the method further comprises combining the new magnitude spectrograms with the phase spectrogram in order to produce a new complex spectrogram.

13. The system of claim 11 , wherein transforming the audio file into the complex spectrogram is done via a short-time fourier transform.

14. The system of claim 11 , wherein a second deep neural network is used for isolating a second audio source.

15. The system of claim 11 , wherein the deep neural network includes an input scale layer before a series of downsample layers and an output scale layer following a series of upsample layers.

16. The system of claim 11 , wherein the deep neural network includes a bridge layer comprising a first convolutional 2D layer and a second convolutional 2D layer.

17. The system of claim 11 , wherein the method further comprises constructing a new phase spectrogram from a new source using a generative adversarial neural network.

18. The system of claim 11 , wherein the method further comprises constructing a new phase spectrogram from a new source using the Griffin-Lim algorithm.

19. The system of claim 11 , wherein the method further comprises applying a multi-channel wiener filter is applied to the new magnitude spectrogram.

20. A non-transitory computer readable medium storing instructions to be executed by a processor, the instructions comprising:

transforming an audio file into a complex spectrogram;

decomposing the complex spectrogram into a magnitude spectrogram and a phase spectrogram;

splitting the magnitude spectrogram into K small fragments along the time dimension;

sending each fragment in the K small fragments through a convolutional deep neural network to produce K small fragment outputs, wherein the convolutional deep neural network includes convolutional layers augmented with an attention layer to help the convolutional deep neural network find the most informative areas and reduce artifacts in each of the K small fragments, wherein the attention layer informs the convolutional deep neural network of important pixels;

producing a source mask based on the K small fragment outputs, wherein the source mask corresponds to an audio source targeted to be isolated;

multiplying the source mask with the magnitude spectrogram to create a new magnitude spectrogram corresponding to the source.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2022
From: MINER, LUKE
To: AUDIOSHAKE, INC.
Reel/Frame 060038/0071 →
Continuity (4)
Continuation 17062253 · Oct 2, 2020
Continuation 17061799 · Oct 2, 2020
Provisional Application 62882317 · Aug 2, 2019
Related Publication 20220262375A1 · Aug 18, 2022