IP Library › Granted Patent US 12,456,493
Granted Patent B2
US 12,456,493 · App. 18/012,245 · Granted Oct 28, 2025

System for automated multitrack mixing

Inventors: Christian James Steinmetz (Barcelona, ES); Joan Serra (Barcelona, ES)
Assignee: Dolby Laboratories Licensing Corporation
G11B27/038G06N3/045H04S3/008H04S2400/01H04S2400/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,493
App. No.
18/012,245
Granted
Oct 28, 2025
Kind
B2
Abstract

A deep-learning-based system for performing automated multitrack mixing based on a plurality of input audio tracks is described herein. The system comprises one or more instances of a deep-learning-based first network and one or more instances of a deep-learning-based second network. Particularly, the first network is configured to, based on the 5 input audio tracks, generate parameters for use in the automated multitrack mixing. The second network is configured to, based on the parameters, apply signal processing and at least one mixing gain to the input audio tracks, for generating an output mix of the audio tracks.

Claims (52)

1. A deep-learning-based system for performing automated multitrack mixing based on a plurality of input audio tracks, wherein the system comprises:

one or more instances of a deep-learning-based controller network; and

one or more instances of a deep-learning-based transformation network,

wherein the controller network is configured to, based on the input audio tracks, generate parameters for use in the automated multitrack mixing;

wherein the transformation network is configured to, based on the parameters, apply signal processing and at least one mixing gain to the input audio tracks, for generating an output mix of the audio tracks;

wherein the controller network and transformation network are trained separately, and

wherein the controller network is trained based on the pre-trained transformation network.

2. The system according to claim 1 , wherein the output mix is a stereo mix.

3. The system according to claim 1 , wherein the controller network comprises:

a first stage; and

a second stage; and

wherein generating the parameters by the controller network comprises:

mapping, by the first stage, each of the input audio tracks into a respective feature space representation; and

generating, by the second stage, parameters for use by the transformation network, based on the feature space representations.

4. The system according to claim 3 , wherein the generating, by the second stage, the parameters for use by the transformation network comprises:

generating a combined representation based on the feature space representations of the input audio tracks; and

generating parameters for use by the transformation network based on the combined representation.

5. The system according to claim 4 , wherein generating the combined representation involves an averaging process on the feature space representations of the input audio tracks.

6. The system according to claim 1 , wherein the controller network is trained based on at least one loss function that indicates differences between predetermined mixes of audio tracks and respective predictions thereof.

7. The system according to claim 1 , wherein the controller network is trained by:

obtaining, as input, at least one first training set, wherein the first training set comprises a plurality of subsets of audio tracks, and, for each subset, a respective predetermined mix of the audio tracks in the subset;

inputting the first training set to the controller network; and

iteratively training the controller network to predict respective mixes of the audio tracks of the subsets in the training set,

wherein the training is based on at least one first loss function that indicates differences between the predetermined mixes of the audio tracks and respective predictions thereof.

8. The system according to claim 7 , wherein the predicted mixes of the audio tracks are stereo mixes, and wherein the first loss function is a stereo loss function and is constructed in such a manner that it is invariant under re-assignment of left and right channels.

9. The system according to claim 7 , wherein the training of the controller network to predict the mixes of the audio tracks comprises, for each subset of audio tracks:

generating, by the controller network, a plurality of predicted parameters in accordance with the subset of audio tracks;

feeding the predicted parameters to the transformation network; and

generating, by the transformation network, the prediction of the mix of the subset of audio tracks, based on the predicted parameters and on the subset of audio tracks.

10. The system according to claim 1 , wherein a number of instances of the transformation network equals a number of the input audio tracks, wherein the transformation network is configured to, based on at least part of the parameters, perform signal processing on a respective input audio track to generate a respective processed output, wherein the processed output comprises left and right channels, and wherein the output mix is generated based on the processed outputs.

11. The system according to claim 10 , wherein the system further comprises a routing component, wherein the routing component is configured to generate a number of bus-level mixes based on the processed outputs, and wherein the output mix is generated based on the bus-level mixes.

12. The system according to claim 11 , wherein the controller network is configured to further generate parameters for the routing component.

13. The system according to claim 11 , wherein the one or more instances of the transformation network is a first set of one or more instances of the transformation network, wherein the system further comprises a second set of one or more instances of the transformation network, and wherein a number of instances of the second set of one or more instances of the transformation network is determined in accordance with the number of the bus-level mixes.

14. The system according to claim 13 , wherein the system is configured to further generate a left mastering mix and a right mastering mix based on the bus-level mixes, wherein the system further comprises a pair of instances of the transformation network, and wherein the pair of instances of transformation network are configured to generate the output mix based on the left and right mastering mixes.

15. The system according to claim 1 , wherein the transformation network is trained by:

obtaining, as input, at least one second training set, wherein the second training set comprises a plurality of audio signals, and, for each audio signal, at least one transformation parameter for signal processing of the audio signal and a respective predetermined processed audio signal;

inputting the second training set to the transformation network; and

iteratively training the transformation network to predict respective processed audio signals based on the audio signals and the transformation parameters,

wherein the training is based on at least one second loss function that indicates differences between the predetermined processed audio signals and the respective predictions thereof.

16. The system according to claim 1 , wherein the parameters generated by the controller network comprise at least one of human parameters, machine interpretable parameters, control parameters, and/or panning parameters.

17. The system according to claim 1 , wherein the controller and/or transformation network comprises at least one neural network, the neural network comprising a linear layer and/or a multilayer perceptron, MLP.

18. A deep-learning-based system for performing automated multitrack mixing based on a plurality of input audio tracks, the system comprising:

a controller network, and

a transformation network,

wherein the input audio tracks are provided in parallel to the controller network and the transformation network, and

wherein the transformation network is configured to apply signal processing and at least one mixing gain to the input audio tracks for generating an output mix of the audio tracks, based on one or more parameters generated by the controller network.

19. The system according to claim 18 , wherein the parameters are human interpretable parameters.

20. The system according to claim 18 , wherein the system comprises a plurality of instances of the controller network in a weight-sharing configuration; and/or a plurality of instances of the transformation network in a weight-sharing configuration.

21. A method of operating a deep-learning-based system for performing automated multitrack mixing based on a plurality of input audio tracks, wherein the system comprises one or more instances of a deep-learning-based controller network and one or more instances of a deep-learning-based transformation network, the method comprising:

providing the input audio tracks to the controller network and the transformation network in parallel;

generating, by the controller network, parameters for use in the automated multitrack mixing, based on the input audio tracks; and

applying, by the transformation network, signal processing and at least one mixing gain to the input audio tracks based on the parameters, for generating an output mix of the audio tracks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2023
From: STEINMETZ, CHRISTIAN JAMES; SERRA, JOAN
To: DOLBY INTERNATIONAL AB
Reel/Frame 064514/0335 →
Priority Claims (2)
ES P202030604 · Jun 22, 2020 · national
EP 20203276 · Oct 22, 2020 · regional
Continuity (3)
Provisional Application 63092310 · Oct 15, 2020
Provisional Application 63072762 · Aug 31, 2020
Related Publication 20230352058A1 · Nov 2, 2023
References Cited (22)
US 8367923B2 · Humphrey · 2013 [cited by applicant]
US 9304988B2 · Terrell · 2016 [cited by applicant]
US 9665823B2 · Saon · 2017 [cited by applicant]
US 9697826B2 · Sainath · 2017 [cited by applicant]
US 9716948B2 · Tang · 2017 [cited by applicant]
US 9952826B2 · Rowe · 2018 [cited by applicant]
US 10564923B2 · Cardinaux · 2020 [cited by applicant]
US 11552611B2 · Veselinovic · 2023 [cited by examiner]
US 20070083365A1 · Shmunk · 2007 [cited by applicant]
US 20080199027A1 · Kleczkowski · 2008 [cited by applicant]
US 20200321975A1 · Milot · 2020 [cited by examiner]
US 20210037287A1 · Ha · 2021 [cited by examiner]
JP 2016521925A · 2016 [cited by applicant]
WO 2014183879A1 · 2014 [cited by applicant]
WO 2019121574A1 · 2019 [cited by applicant]
Hawley, S. et al “Signal Train: Profiling Audio Compressors with Deep Neural Networks” arxiv.org, May 28, 2019, Olin Library Cornell University Ithaca, NY. [cited by applicant]
Martinez, M. et al, Intelligent Audio Mixing Using Deep Learning, DMRN+11: Digital Music Research Network, Dec. 20, 2016, Queen Mary University of London, UK. [cited by applicant]
Moffat, D. et al, Approaches in Intelligent Music Production, MDPI Arts Journal, Sep. 25, 2019, arts8040125, MDPI, Basel, Switzerland. [cited by applicant]
Purwins, H. et al “Deep Learning for Audio Signal Processing” IEEE Journal of Selected Topics in Signal Processing, vol. 13, No. 2, May 2019, pp. 206-219. [cited by applicant]
Scott, J. et al “Automatic Multi-Track Mixing Using Linear Dynamical Systems” Jul. 6, 2011, retrieved from the Internet: SMC papers. [cited by applicant]
Steinmetz, C. “Towards end-to-end Multirack Mixing with Deep Learning” Master Thesis on Sound and Music Computing, Jul. 2020, pp. 1-29. [cited by applicant]
Xu, K. et al, Mixup-Based Acoustic Scene Classification Using Multi-Channel Convolutional Neural Network, Association for Computing Machinery (ACM), May 18, 2018; 1805.07319v1, Cornell University, Ithaca, New York. [cited by applicant]