IP Library Granted Patent US 10,014,002
Granted Patent B2
US 10,014,002 · App. 15/792,566 · Granted Jul 3, 2018

Real-time audio source separation using deep neural networks

Inventors: Alejandro Koretzky (Venice, CA); Karthiek Reddy Bokka (Los Angeles, CA); Naveen Sasalu Rajashekharappa (Los Angeles, CA)
Assignee: Red Pill VR, Inc.
G10L25/18G10L21/028G10L25/30G06F3/165
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,014,002
App. No.
15/792,566
Granted
Jul 3, 2018
Kind
B2
Abstract

Methods and systems for audio source separation in real-time are described. In an embodiment, the present disclosure describes reading and decoding an audio source into PCM samples, fragmenting Pulse Code Modulation (PCM) samples into fragments, transforming fragments into spectrograms, performing audio source separation using a deep neural network (DNN) to generate an estimated magnitude spectrogram of the component(s) of the audio source, reconstructing the estimated time domain component signals, and streaming the component signals to a playback engine. In an embodiment, a semantic equalizer graphical user allows for real-time mixing of individual component signals.

Claims (37)

1. A method comprising:

reading an audio source, wherein the audio source represents a combined audio signal that comprises a plurality of component signals;

generating, based on the audio source, a plurality of pulse-code modulation (PCM) samples;

fragmenting the PCM samples into a plurality of PCM sample fragments;

for each particular PCM sample fragment of the plurality of PCM sample fragments:

transforming the particular PCM sample fragment into a spectrogram;

processing the spectrogram into one or more spectrogram fragments;

inputting the one or more spectrogram fragments into a deep neural network;

generating, by the deep neural network, one or more binary mask fragments wherein each binary mask fragment of the one or more binary mask fragments corresponds to one or more component signals of the plurality of component signals;

using the one or more binary mask fragments to estimate the plurality of component signals;

wherein the method is performed using one or more processors.

2. The method of claim 1 , wherein the deep neural network is a stem model deep neural network that corresponds to a single component signal of the plurality of component signals.

3. The method of claim 1 , wherein the deep neural network is a stacked model deep neural network that corresponds to the plurality of component signals.

4. The method of claim 1 , wherein the deep neural network is a merged model deep neural network that corresponds to the plurality of component signals.

5. The method of claim 1 , wherein each spectrogram fragment of the one or more spectrogram fragments comprises 25 frames of data.

6. The method of claim 1 , wherein the deep neural network is trained using a backpropagation algorithm on a training dataset.

7. The method of claim 1 , wherein transforming the particular PCM sample fragment into a spectrogram comprises applying a Short-Time Fast Fourier Transform to the particular PCM sample fragment.

8. The method of claim 1 , wherein using the one or more binary mask fragments to estimate the plurality of component signals comprises multiplying a middle frame of the one or more spectrogram fragments with the one or more binary mask fragments.

9. The method of claim 1 , wherein transforming the particular PCM sample fragment into a spectrogram comprises applying a Short-Time Fast Fourier Transform to the particular PCM sample fragment.

10. One or more non-transitory computer-readable media storing instructions, which when executed by one or more processors cause:

reading an audio source, wherein the audio source represents a combined audio signal that comprises a plurality of component signals;

generating, based on the audio source, a plurality of pulse-code modulation (PCM) samples;

fragmenting the PCM samples into a plurality of PCM sample fragments;

for each particular PCM sample fragment of the plurality of PCM sample fragments:

transforming the particular PCM sample fragment into a spectrogram;

processing the spectrogram into one or more spectrogram fragments;

inputting the one or more spectrogram fragments into a deep neural network;

generating, by the deep neural network, one or more binary mask fragments corresponding to the plurality of component signals; and

using the one or more binary mask fragments to estimate the plurality of component signals.

11. The one or more non-transitory computer-readable media of claim 10 , wherein the deep neural network is a stem model deep neural network that corresponds to a single component signal of the plurality of component signals.

12. The one or more non-transitory computer-readable media of claim 10 , wherein the deep neural network is a stacked model deep neural network that corresponds to the plurality of component signals.

13. The one or more non-transitory computer-readable media of claim 10 , wherein the deep neural network is a merged model deep neural network that corresponds to the plurality of component signals.

14. The one or more non-transitory computer-readable media of claim 10 , wherein each spectrogram fragment of the one or more spectrogram fragments comprises 25 frames of data.

15. The one or more non-transitory computer-readable media of claim 10 , wherein the deep neural network is trained using a backpropagation algorithm on a training dataset.

16. The one or more non-transitory computer-readable media of claim 10 , wherein transforming the particular PCM sample fragment into a spectrogram comprises applying a Short-Time Fast Fourier Transform to the particular PCM sample fragment.

17. The one or more non-transitory computer-readable media of claim 10 , wherein using the one or more binary mask fragments to estimate the plurality of component signals comprises multiplying a middle frame of the one or more spectrogram fragments with the one or more binary mask fragments.

18. The one or more non-transitory computer-readable media of claim 10 , wherein transforming the particular PCM sample fragment into a spectrogram comprises applying a Short-Time Fast Fourier Transform to the particular PCM sample fragment.

Assignments (2)
CHANGE OF NAME Recorded Apr 30, 2019
From: RED PILL VR, INC.
To: RED PILL VR, INC
Reel/Frame 049043/0093 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2017
From: KORETZKY, ALEJANDRO; BOKKA, KARTHIEK REDDY; RAJASHEKHARAPPA, NAVEEN SASALU
To: RED PILL VR, INC.
Reel/Frame 043940/0206 →
Continuity (3)
Continuation In Part 15434419 · Feb 16, 2017
Provisional Application 62295497 · Feb 16, 2016
Related Publication 20180122403A1 · May 3, 2018
Cited By (7)
US 12,254,892 US 12,254,893 US 12,266,378 US 12,431,159 US 12,499,902 US 12,531,042 US 12,609,127