IP Library Granted Patent US 12,254,892
Granted Patent B2
US 12,254,892 · App. 17/973,482 · Granted Mar 18, 2025

Audio source separation processing workflow systems and methods

Inventors: Emile de la Rey (Wellington, NZ); Paris Smaragdis (Urbana, IL)
Assignee: WingNut Films Productions Limited
G10L21/0308G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,892
App. No.
17/973,482
Granted
Mar 18, 2025
Kind
B2
Abstract

Systems and methods includes receiving a single-track audio input stream having a mixture of audio signals generated from a plurality of sources, training an audio source separation model using, at least in part, the received single-track audio input stream, and separating audio sources, using the audio source separation model, from the audio input stream in accordance with one or more processing recipes to generate a plurality of source separated output stems. The audio separation model is trained to receive the single-track audio input stream and generate a plurality of audio stems corresponding to one or more audio sources of the plurality of sources.

Claims (77)

1. A method comprising:

receiving a single-track audio input stream comprising a mixture of audio signals generated from a plurality of sources;

training an audio source separation model using, at least in part, the single-track audio input stream, wherein the audio source separation model is trained to receive the single-track audio input stream and generate a plurality of audio stems corresponding to one or more of the plurality of sources; and

separating audio sources, using the audio source separation model, from the single-track audio input stream in accordance with one or more processing recipes to generate a plurality of source separated output stems;

wherein training the audio source separation model comprises training a neural network to output one or more classes of source signals and a remaining complement signal through a training process comprising:

(a) training a first neural network at a first audio sample rate to learn a first set of parameters;

(b) training a second neural network at a second audio sample rate that is higher than the first audio sample rate, the second neural network including a second set of parameters inherited from at least one untrained layer of the first neural network, and wherein training the second neural network comprises:

(1) training the second set of parameters while the first set of parameters remains fixed; and

(2) retraining the second neural network to finetune the first set of parameters and the second set of parameters.

2. The method of claim 1 , wherein the audio source separation model comprises a self-iterative training process that repeatedly updates the second training dataset and retrains the audio source separation model on the single-track audio input stream using one or more of the source separation output stems generated from a prior iteration of the trained audio source separation model;

wherein an iteration of the self-iterative training process further comprises:

training the audio source separation model using the second training dataset, the second training dataset further comprising a plurality of labeled speech samples and/or a plurality of labeled music and/or noise data samples;

processing the single-track audio input stream through the trained audio source separation model to generate source separated output stems;

updating the second training dataset to include training data generated from one or more of the source separated output stems; and

retraining the audio source separation model using the updated second training dataset.

3. A system comprising:

a memory component storing machine-readable instructions; and

a logic device configured to execute the machine-readable instructions to:

(A) receive a single-track audio input stream comprising a mixture of audio signals generated from a plurality of audio sources;

(B) perform a first separation of the audio sources, using an audio source separation model, from the single-track audio input stream in accordance with one or more processing recipes to generate a first plurality of source separated output stems corresponding to one or more of the plurality of audio sources;

(C) retrain the audio source separation model using, at least in part, a training dataset comprising at least one of the first plurality of source separated output stems from the single-track audio input stream, wherein the audio source separation model is trained to generate a plurality of audio stems corresponding to one or more of the plurality of audio sources by:

(1) training a first neural network at a first audio sample rate to learn a first set of parameters; and

(2) training a second neural network at a second audio sample rate that is higher than the first audio sample rate, the second neural network including the first set of parameters inherited from the first neural network, and a second set of parameters inherited from at least one untrained layer of the first neural network, and wherein training the second neural network comprises:

(a) training second set of parameters while the first set of parameters remain fixed; and

(b) retraining the second neural network to finetune the first of parameters and the second set of parameters; and

(D) perform a second separation of the audio sources, using the retrained audio source separation model, from the single-track audio input stream in accordance with the one or more processing recipes to generate a second plurality of source separated output stems.

4. The system of claim 3 wherein the one or more processing recipes comprising a plurality of processing branches configured to output one or more audio stems on a first processing branch and a remaining complement signal mixture on a second branch.

5. The system of claim 3 , wherein the audio source separation model comprises a plurality of neural networks, each neural network trained to separate audio signals corresponding to at least one source from the mixture of audio signals in the single-track audio input stream;

wherein the plurality of neural networks are configured to apply a windowing function;

wherein the plurality of neural networks are configured to perform an overlap-add process to smooth banding artefacts; and/or

wherein the plurality of neural networks are configured to perform source separation without applying a mask.

6. The system of claim 3 , wherein the logic device is further configured to train the audio source separation model by implementing a self-iterative training process that repeatedly updates the training dataset and retrains the audio source separation model on the single-track audio input stream using one or more of the source separation output stems generated from a prior iteration of the trained audio source separation model.

7. The system of claim 6 , wherein the logic device is further configured to perform an iteration of the self-iterative training process by:

training the audio source separation model using the training dataset, the training dataset further comprising a plurality of labeled speech samples and/or a plurality of labeled music and/or noise data samples;

processing the single-track audio input stream through the trained audio source separation model to generate source separated output stems;

updating the training dataset to an updated training dataset that includes one or more of the source separated output stems; and

retraining the audio source separation model using the updated training dataset.

8. The system of claim 3 , wherein the logic device is further configured to post-process the second plurality of source separated output stems to remove artefacts introduced during source separation, including clicks, harmonic distortion, ghosting and/or broadband noise.

9. The system of claim 3 , wherein each processing recipe comprises a plurality of processing branches, each of the processing branches configured to process the single-track audio input stream, a separated audio source and/or a complement signal mixture through a trained neural network configured to output both a separated audio source stem and a remaining complement signal mixture.

10. The system of claim 3 , wherein the logic device is further configured to train the audio source separation model by:

obtaining a user evaluation of one or more of the plurality of audio stems; and

fine tuning the audio source separation model by:

(a) adding one more of the plurality of audio stems to the training dataset to retrain the audio source separation model;

(b) adding one or more audio samples to the training dataset to retrain the audio source separation model; and/or

(c) adjusting one or more hyperparameters to fine-tune source separation results.

11. A method comprising:

training an audio source separation model with a first training dataset to generate a plurality of audio stems from an audio input sample;

receiving a single-track audio input stream comprising a mixture of audio signals generated from a plurality of audio sources;

first separating the audio sources, using the audio source separation model, from the single-track audio input stream in accordance with one or more processing recipes to generate a first plurality of source separated output stems corresponding to one or more of the plurality of audio sources;

retraining the audio source separation model to separate the plurality of audio sources from the single-track audio input stream using, at least in part, a second training dataset comprising at least one of the first plurality of source separated output stems generated from the single-track audio input stream; and

second separating the audio sources, using the retrained audio source separation model, from the single-track audio input stream in accordance with the one or more processing recipes to generate a second plurality of source separated output stems corresponding to the one or more of the plurality of audio sources;

wherein retraining the audio source separation model comprises:

(a) training a first neural network at a first audio sample rate to learn a first set of parameters;

(b) training a second neural network at a second audio sample rate that is higher than the first audio sample rate, the second neural network including the first set of parameters inherited from the first neural network, and a second set of parameters inherited from at least one untrained layer of the first neural network, and wherein training the second neural network comprises:

(1) training the second set of parameters while the first set of parameters remain fixed; and

(2) retraining the second neural network to finetune the first set of parameters and the second set of parameters.

12. A method comprising

training an audio source separation model with a first training dataset to generate a plurality of audio stems from an audio input sample;

receiving a single-track audio input stream comprising a mixture of audio signals generated from a plurality of audio sources;

first separating the audio sources, using the audio source separation model, from the single-track audio input stream in accordance with one or more processing recipes to generate a first plurality of source separated output stems corresponding to one or more of the plurality of audio sources;

retraining the audio source separation model to separate the plurality of audio sources from the single-track audio input stream using, at least in part, a second training dataset comprising at least one of the first plurality of source separated output stems generated from the single-track audio input stream; and

second separating the audio sources, using the retrained audio source separation model, from the single-track audio input stream in accordance with the one or more processing recipes to generate a second plurality of source separated output stems corresponding to the one or more of the plurality of audio sources;

wherein retraining the audio source separation model comprises training a first neural network of the audio source separation model at a first audio sample rate, and a second neural network at a second audio sample rate that is higher than the first audio sample rate.

13. The method of claim 12 , wherein the one or more processing recipes comprises a plurality of processing branches, each processing branch trained to output one or more audio stems; and wherein each of the processing branches is further configured to output the one or more audio stems on a first processing branch and a remaining complement signal mixture on a second processing branch.

14. The method of claim 12 , wherein training the audio source separation model comprises training a general audio source separation model; and

wherein retraining the audio source separation model comprises a self-iterative training process that repeatedly updates the training dataset using one or more of the source separation output stems generated from a prior iteration of the retrained audio source separation model such that the training dataset and audio source separation model become increasingly specific to the mixture of audio signals in the single-track audio input stream.

15. The method of claim 12 , further comprising post-processing the second plurality of source separated output stems to remove artefacts introduced during source separation, including clicks, harmonic distortion, ghosting and/or broadband noise.

16. The method of claim 12 , wherein the first plurality of source separated output stems corresponds to one of the plurality of audio sources; and

wherein the second training dataset comprises the first plurality of source separated output stems which correspond to the one of the plurality of audio sources.

17. A method comprising:

training an audio source separation model with a first training dataset to generate a plurality of audio stems from an audio input sample;

receiving a single-track audio input stream comprising a mixture of audio signals generated from a plurality of audio sources;

first separating the audio sources, using the audio source separation model, from the single-track audio input stream in accordance with one or more processing recipes to generate a first plurality of source separated output stems corresponding to one or more of the plurality of audio sources;

retraining the audio source separation model to separate the plurality of audio sources from the single-track audio input stream using, at least in part, a second training dataset comprising at least one of the first plurality of source separated output stems generated from the single-track audio input stream; and

second separating the audio sources, using the retrained audio source separation model, from the single-track audio input stream in accordance with the one or more processing recipes to generate a second plurality of source separated output stems corresponding to the one or more of the plurality of audio sources;

wherein training the audio source separation model comprises feeding the first training dataset through a first neural network model, adjusting first model parameters based on a first loss function, and validating performance of the audio source separation model using a validation dataset; and

wherein retraining the audio source separation model further comprises feeding the second training dataset through a second neural network model, adjusting second model parameters based on a second loss function, and validating performance of the retrained audio source separation model to confirm improved results over the audio source separation model in separating sources from the single-track audio input stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: DE LA REY, EMILE; SMARAGDIS, PARIS
To: WINGNUT FILMS PRODUCTIONS LIMITED
Reel/Frame 063828/0368 →
Continuity (3)
Continuation In Part 17848341 · Jun 23, 2022
Provisional Application 63272650 · Oct 27, 2021
Related Publication 20230130844A1 · Apr 27, 2023
References Cited (25)
US 10014002B2 · Koretzky et al. · 2018 [cited by applicant]
US 10410115B2 · Lewis et al. · 2019 [cited by applicant]
US 20170178664A1 · Wingate · 2017 [cited by examiner]
US 20170251319A1 · Jeong · 2017 [cited by examiner]
US 20170345185A1 · Byron et al. · 2017 [cited by applicant]
US 20180122403A1 · Koretzky · 2018 [cited by examiner]
US 20190066713A1 · Mesgarani · 2019 [cited by examiner]
US 20210090536A1 · Pachet · 2021 [cited by examiner]
US 20210104256A1 · Jansson · 2021 [cited by examiner]
US 20210208842A1 · Cassidy · 2021 [cited by examiner]
US 20210312939A1 · Uhle et al. · 2021 [cited by applicant]
US 20210358513A1 · Narisetty et al. · 2021 [cited by applicant]
US 20220095061A1 · Diehl et al. · 2022 [cited by applicant]
US 20220337952A1 · Neoran · 2022 [cited by examiner]
WO WO2020193929A1 · 2020 [cited by applicant]
WO WO2021159775A1 · 2021 [cited by applicant]
Ruder, Sebastian et al., “Strong Baselines for Neural Semi-supervised Learning under Domain Shift”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 25, 2018. [cited by applicant]
Manilow Ethan et al., “Hierarchical Musical Instrument Separation”, Oct. 11, 2020, pp. 376-383. [URL: https://program.ismir2020.net/static/final_papers/105.pdf]. [cited by applicant]
Brunner, Gino et al., “Monaural Music Source Separation using a ResNet Latent Separator Network”, 2019 IEEE 31st International Conference on Tools With Artificial Intelligence (ICTAI), IEEE, Nov. 4, 2019. [cited by applicant]
Wang, Zhepei et al., “Semi-Supervised Singing Voice Separation With Noisy Self-Training”, ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Jun. 6, 2021 (Jun. 6, 2… [cited by applicant]
International Search Report mailed on Jan. 16, 2023, for PCT Application No. PCT/IB2022/060319. [cited by applicant]
International Search Report mailed on Jan. 27, 2023, for PCT Application No. PCT/IB2022/060320. [cited by applicant]
International Search Report mailed on Jan. 31, 2023, for PCT Application No. PCT/IB2022/060322. [cited by applicant]
International Search Report mailed on Jan. 23, 2023, for PCT Application No. PCT/IB2022/060321. [cited by applicant]
Luo, Yi et al., “Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech Separation”, Mar. 27, 2020. pp. 1-5, arXiv:1910.06379v2. [cited by applicant]
Cited By (1)
US 12,732,739