IP Library Granted Patent US 12,431,159
Granted Patent B2
US 12,431,159 · App. 17/848,341 · Granted Sep 30, 2025

Audio source separation systems and methods

Inventors: Emile de la Rey (Wellington, NZ); Paris Smaragdis (Urbana, IL)
Assignee: WingNut Films Productions Limited
G10L25/51G10L15/063G10L21/0272G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,431,159
App. No.
17/848,341
Filed
Jun 23, 2022
Granted
Sep 30, 2025
Kind
B2
Art Unit
2659
USPC
704/501
Abstract

Systems and methods for audio source separation include receiving an audio input stream including a mixture of audio signals generated from a plurality of audio sources; processing, through a trained audio source separation model, the audio input stream to generate a plurality of audio stems corresponding to one or more of the plurality of audio sources; updating, using a self-iterative processing and training system, the audio source separation model based at least in part on the plurality of audio stems; and re-processing, using the updated trained audio source separation model, the audio input stream to generate a plurality of enhanced audio stems.

Claims (35)

1. A system comprising:

a memory component storing machine-readable instructions; and

a logic device configured to execute the machine-readable instructions to:

a trained audio source separation model configured to receive an audio input sample comprising a single-track mixture of audio signals generated from a plurality of audio sources and generate a plurality of audio stems, the plurality of audio stems corresponding to one or more audio source of the plurality of audio sources; and

a self-iterative training system configured to perform a plurality of training iterations, a training iteration comprising generating a new audio source separation model based at least in part on a training dataset comprising a subset of the generated plurality of audio stems from the preceding training iteration, wherein a subset of the generated plurality of audio stems comprises one or more of: a part of one stem, multiple parts of one stem, one complete stem, multiple stems, or a combination of one complete stem and one or more parts of another stem, and wherein the new audio source separation model generated in each iteration is increasingly specific to the mixture of audio signals in the audio input sample,

wherein the self-iterative training system is further configured to determine whether the new audio source separation model is increasingly specific to the mixture of audio signals in the audio input sample by calculating a first quality metric associated with the generated plurality of audio stems, the first quality metric providing a first performance measure of the audio source separation model of the prior iteration, calculating a second quality metric associated with the audio stems generated in the present iteration, the second quality metric providing a second performance measure of the new audio source separation model, and wherein the second quality metric is greater than the first quality metric; and

wherein the new audio source separation model is configured to re-process the audio input stream to generate a plurality of enhanced audio stems.

2. The system of claim 1 , wherein the trained audio source separation model comprises a neural network trained to separate one or more audio source signals from the single-track mixture of audio signals.

3. The system of claim 2 , wherein the neural network is configured to perform audio source separation without applying a mask.

4. The system of claim 1 , further comprising a training dataset comprising labeled source audio data and labeled noise audio data, and wherein the trained audio source separation model is initially trained using the training dataset to generate a general source separation model.

5. The system of claim 4 , wherein during each training iteration at least a subset of the plurality of audio stems are culled based on a threshold metric and added to the training dataset from the preceding training iteration to form a culled dynamically evolving dataset, and wherein the culled dynamically evolving dataset is used to train the new audio source separation model.

6. The system of claim 1 , wherein the trained audio source separation model is trained using a training dataset comprising a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the system to address source separation associated with an identified source.

7. The system of claim 6 , wherein the plurality of datasets comprises a speech training dataset comprising a plurality of labeled speech samples, and/or a non-speech training dataset comprising a plurality of labeled music and/or noise data samples.

8. The system of claim 1 , wherein the self-iterative training system further comprises a self-iterative dataset generation module configured to generate labeled audio samples from the generated plurality of audio stems.

9. The system of claim 1 , wherein the plurality of enhanced audio stems is generated using a hierarchical branching sequence comprising separating a source signal and a remaining complement signal.

10. A method comprising:

receiving an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources;

generating, using a trained audio source separation model configured to receive the audio input stream, a generated plurality of audio stems corresponding to one or more audio sources of the plurality of audio sources;

updating, the trained audio source separation model through a plurality of training iterations, where a training iteration comprises generating a new audio source separation model based at least in part on a subset of the generated plurality of audio stems derived from a preceding training iteration, wherein a subset of the generated plurality of audio stems comprises one or more of: a part of one stem, multiple parts of one stem, one complete stem, multiple stems, or a combination of one complete stem and one or more parts of another stem;

wherein updating the trained audio source separation model further comprises:

calculating a first quality metric associated with the generated plurality of audio stems, the first quality metric providing a first performance measure of the trained audio source separation model; calculating a second quality metric associated with the enhanced audio stems, the second quality metric providing a performance measure of the new audio source separation model; and comparing the second quality metric to the first quality metric to confirm the second quality metric is greater than the first quality metric; and

re-processing the audio input stream using the new audio source separation model to generate a plurality of enhanced audio stems.

11. The method of claim 10 , wherein the audio input stream comprises one or more single-track audio mixtures, and wherein the trained audio source separation model comprises a neural network trained to separate one or more audio source signals from the one or more single-track audio mixtures.

12. The method of claim 11 , wherein the neural network is configured to perform audio source separation without applying a mask.

13. The method of claim 10 , further comprising:

providing a training dataset comprising labeled source audio data and labeled noise audio data; and

training the trained audio source separation model using the training dataset to generate a general source separation model.

14. The method of claim 13 , further comprising:

adding at least a subset of the generated plurality of audio stems to the training dataset to produce a dynamically evolving dataset;

culling the dynamically evolving dataset based on a threshold metric; and

training the new audio source separation model using the culled, dynamically evolving dataset.

15. The method of claim 10 , wherein the trained audio source separation model is trained using training dataset comprising a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the audio source separation model to address a different source separation problem.

16. The method of claim 15 , wherein the plurality of datasets comprises a speech training dataset comprising a plurality of labeled speech samples, and/or a non-speech training dataset comprising a plurality of labeled music and/or noise data samples.

17. The method of claim 16 , wherein updating the trained audio source separation model further comprises generating labeled audio samples from the generated plurality of audio stems for a self-iterative dataset.

18. The method of claim 10 , further comprising generating the plurality of enhanced audio stems using a hierarchical branching sequence comprising separating a source signal and a remaining complement signal.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: DE LA REY, EMILE; SMARAGDIS, PARIS
To: WINGNUT FILMS PRODUCTIONS LIMITED
Reel/Frame 063828/0368 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2022
From: DE LA REY, EMILE; SMARAGDIS, PARIS
To: WINGNUT FILMS PRODUCTIONS LIMITED
Reel/Frame 060416/0428 →
Continuity (2)
Provisional Application 63272650 · Oct 27, 2021
Related Publication 20230126779A1 · Apr 27, 2023
References Cited (28)
US 10014002B2 · Koretzky · 2018 [cited by examiner]
US 10410115B2 · Lewis et al. · 2019 [cited by applicant]
US 20170178664A1 · Wingate · 2017 [cited by examiner]
US 20170251319A1 · Jeong et al. · 2017 [cited by applicant]
US 20170316792A1 · Chaudhuri · 2017 [cited by examiner]
US 20170345185A1 · Byron et al. · 2017 [cited by applicant]
US 20180122403A1 · Koretzky et al. · 2018 [cited by applicant]
US 20190066713A1 · Mesgarani et al. · 2019 [cited by applicant]
US 20210090536A1 · Pachet et al. · 2021 [cited by applicant]
US 20210104256A1 · Jansson et al. · 2021 [cited by applicant]
US 20210208842A1 · Cassidy et al. · 2021 [cited by applicant]
US 20210312939A1 · Uhle · 2021 [cited by examiner]
US 20210358513A1 · Narisetty et al. · 2021 [cited by applicant]
US 20220095061A1 · Diehl et al. · 2022 [cited by applicant]
US 20220337952A1 · Neoran et al. · 2022 [cited by applicant]
WO WO2020193929A1 · 2020 [cited by applicant]
WO WO2021159775A1 · 2021 [cited by applicant]
Manilow, Hierarchical Musical Instrument Separation, Oct. 11-16, 2020, ISMIR Conference (Year: 2020). [cited by examiner]
Brunner, Monaural Music Source Separation Using a ResNet Latent Separator Network, 2019, IEEE 31st International Conference on Tools with Artificial Intelligence (Year: 2019). [cited by examiner]
Ruder, Sebastian et al., “Strong Baselines for Neural Semi-supervised Learning under Domain Shift”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 25, 2018. [cited by applicant]
Manilow Ethan et al., “Hierarchical Musical Instrument Separation”, Oct. 11, 2020, pp. 376-383. [URL: https://program.ismir2020.net/static/final_papers/105.pdf]. [cited by applicant]
Brunner, Gino et al., “Monaural Music Source Separation using a ResNet Latent Separator Network”, 2019 IEEE 31st International Conference on Tools With Artificial Intelligence (ICTAI), IEEE, Nov. 4, 2019. [cited by applicant]
Wang, Zhepei et al., “Semi-Supervised Singing Voice Separation With Noisy Self-Training”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Jun. 6, 2021 (Jun. 6, 2… [cited by applicant]
International Search Report mailed on Jan. 16, 2023, for PCT Application No. PCT/IB2022/060319. [cited by applicant]
International Search Report mailed on Jan. 27, 2023, for PCT Application No. PCT/IB2022/060320. [cited by applicant]
International Search Report mailed on Jan. 31, 2023, for PCT Application No. PCT/IB2022/060322. [cited by applicant]
International Search Report mailed on Jan. 23, 2023, for PCT Application No. PCT/IB2022/060321. [cited by applicant]
Luo, Yi et al., “Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech Separation”, Mar. 27, 2020. pp. 1-5, arXiv:1910.06379v2. [cited by applicant]