IP Library › Granted Patent US 12,620,407
Granted Patent B2
US 12,620,407 · App. 18/067,069 · Granted May 5, 2026

Neural modeler of audio systems

Inventors: Douglas Andres Castro Borquez (Helsinki, FI); Eero-Pekka Damskägg (Helsinki, FI); Athanasios Gotsopoulos (Helsinki, FI); Lauri Juvela (Helsinki, FI); Thomas William Sherson (Utrecht, NL)
Assignee: Neural DSP Technologies Oy
G10L25/30G06F3/165G10L15/16H04R3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,407
App. No.
18/067,069
Granted
May 5, 2026
Kind
B2
Abstract

A process is provided for training a neural network that digitally models an audio system. A sound source is utilized to electrically couple a test signal into an input of a reference audio system. The output of the reference audio system is collected into an audio interface coupled to a computer. A neural network is then trained using the test signal and the captured information to derive a set of weight vectors with appropriate values such that the overall output of the neural network converges towards an output representative of the reference audio system, and a signal in the time domain from a musical instrument is processed through the trained neural network with a latency under 20 milliseconds. A graphical user interface then outputs a graphical representation of the trained neural network, where the graphical representation visually displays at least one virtual control for interaction by a user.

Claims (59)

1 . A process for training a neural network that digitally models an audio system, comprising:

utilizing a sound source to electrically couple a test signal into an input of a reference audio system;

collecting an output of the reference audio system into an audio interface coupled to a computer to store captured information;

training a neural network using the test signal and the captured information to derive a set of weight vectors with appropriate values such that:

the overall output of the neural network converges towards an output representative of the reference audio system; and

a digital signal representing an analog musical instrument signal is processed through the trained neural network in the time domain with an algorithmic latency under 20 milliseconds;

outputting to a graphical user interface, a graphical representation of the trained neural network, the graphical representation visually displaying at least one virtual control; and

enabling a user to interact with the virtual control of the graphical representation of the trained neural network via the graphical user interface to define a virtualization of the reference audio system.

2 . The process of claim 1 , wherein the training is carried out on the computer.

3 . The process of claim 1 , wherein the training is carried out on a separate, remote computer.

4 . The process of claim 1 , wherein training the neural network further comprises:

setting a stopping condition that determines when training ends by performing at least one of:

setting a user-initiated stopping condition; or

processing a number of iterations of the training data as the stopping condition.

5 . The process of claim 1 , wherein training the neural network further comprises:

setting a stopping condition that determines when training ends by processing a perceptual loss function where the perceptual loss function serves as an indicator of the stopping condition.

6 . The process of claim 1 , wherein:

outputting to the graphical user interface, the graphical representation of the trained neural network further comprises outputting to the graphical user interface, a graphical representation of an effects processor that is not within a native capability of the reference audio system.

7 . The process of claim 6 , wherein:

outputting to the graphical user interface, the graphical representation of the effects processor that is not within the native capability of the reference audio system comprises outputting to the graphical user interface, a graphical representation of an equalizer for equalization, wherein the equalizer is not part of the reference audio system.

8 . The process of claim 6 , wherein:

outputting to the graphical user interface, the graphical representation of the effects processor that is not within the native capability of the reference audio system comprises outputting to the graphical user interface, a graphical representation of a dynamics processor, wherein the dynamics processor is not part of the reference audio system.

9 . The process of claim 1 , wherein:

outputting to the graphical user interface, the graphical representation of the effects processor that is not within the native capability of the reference audio system comprises outputting to the graphical user interface, a graphical representation of a time-based processor, wherein the time-based processor is not part of the reference audio system.

10 . A process for training a neural network that digitally models an audio system, comprising:

utilizing a sound source to electrically couple a test signal into an input of a reference audio system;

collecting an output of the reference audio system into an audio interface coupled to a computer to store captured information;

training a neural network using the test signal and the captured information to derive a set of weight vectors with appropriate values such that the overall output of the neural network converges towards an output representative of the reference audio system, and a signal in the time domain from a musical instrument is processed through the trained neural network with a latency under 20 milliseconds;

outputting to a graphical user interface, a graphical representation of the trained neural network, the graphical representation visually displaying at least one virtual control; and

enabling a user to interact with the virtual control of the graphical representation of the trained neural network via the graphical user interface to define a virtualization of the reference audio system;

wherein training the neural network further comprises modeling a non-linear behavior of the reference audio system, modeling a first linear aspect of the reference audio system, and modeling a second linear aspect of the reference audio system.

11 . The process of claim 1 , wherein:

outputting to the graphical user interface, the graphical representation of the trained neural network further comprises outputting to the graphical user interface, a graphical representation of at least one of an emulation of a speaker and an emulation of a speaker cabinet, wherein the emulation is not part of the reference audio system.

12 . The process of claim 1 , wherein:

training the neural network models a non-linear behavior and a linear aspect of the reference audio system.

13 . The process of claim 1 , wherein:

training the neural network further comprises creating a model file that includes sufficient data such that that when read out and processed by a modeling audio system, a functioning model of the reference audio system is realized; wherein:

the virtualization comprises a framework that enables the modeling audio system to model the reference audio system by loading the model file into the modeling audio system.

14 . A process for creating digital audio systems, comprising:

utilizing a sound source to electrically couple a test signal into an input of a reference audio system;

collecting an output of the reference audio system into an audio interface coupled to a computer to store captured information;

training a neural network using the test signal and the captured information to derive a set of weight vectors with appropriate values such that the overall output of the neural network converges towards an output representative of the reference audio system, wherein the training digitally models a non-linear behavior of the reference audio system, models a first linear aspect of the reference audio system, and models a second linear aspect of the reference audio system, the training carried out by repeatedly performing operations comprising:

predicting by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system;

computing an error in the prediction; and

adjusting the weight vectors to minimize the computed error;

outputting a neural network model file upon training;

outputting to a graphical user interface, a graphical representation of the trained neural network in the neural network model file, the graphical representation visually displaying at least one virtual control; and

enabling a user to interact with the virtual control of the graphical representation of the trained neural network via the graphical user interface to define a virtualization of the reference audio system.

15 . The process of claim 14 , wherein predicting by the neural network, the model output comprises carrying out the prediction in the time domain.

16 . The process of claim 14 , wherein:

modeling the non-linear behavior of the reference audio system, modeling the first linear aspect of the reference audio system, and modeling the second linear aspect of the reference audio system are arranged in series, parallel, or a combination thereof.

17 . The process of claim 16 further comprising modeling a temporal dependency of the reference audio system in addition to, or in lieu of modeling the first linear aspect of the reference audio system.

18 . The process of claim 14 , wherein the training further comprises:

applying a perceptual loss function to the neural network based upon a determined psychoacoustic property, wherein the perceptual loss function is applied in the frequency domain; and

adjusting the neural network responsive to the output of the perceptual loss function.

19 . The process of claim 18 , wherein applying the perceptual loss function to the neural network comprises establishing a loudness threshold of hearing for each of multiple frequency bins such that a signal below the loudness threshold is not optimized further, wherein:

for each frequency bin, a loudness threshold is independently set under which a signal is not optimized further in order to optimize further that particular frequency bin.

20 . The process of claim 18 , wherein applying the perceptual loss function to the neural network comprises:

implementing frequency masking such that a frequency component is not further processed if a computed error is below a masking threshold, where the masking threshold is based upon a target signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2022
From: CASTRO BORQUEZ, DOUGLAS ANDRES; DAMSKAGG, EERO-PEKKA; GOTSOPOULOS, ATHANASIOS; JUVELA, LAURI; SHERSON, THOMAS WILLIAM
To: NEURAL DSP TECHNOLOGIES OY
Reel/Frame 062136/0984 →
Continuity (3)
Continuation 16738512 · Jan 9, 2020
Provisional Application 62941986 · Nov 29, 2019
Related Publication 20230119557A1 · Apr 20, 2023
References Cited (34)
US 6292791B1 · Su et al. · 2001 [cited by applicant]
US 6476308B1 · Zhang · 2002 [cited by applicant]
US 8600068B2 · DeBoer et al. · 2013 [cited by applicant]
US 8796530B2 · Kemper · 2014 [cited by applicant]
US 9099066B2 · Welch · 2015 [cited by applicant]
US 9626949B2 · Wang et al. · 2017 [cited by applicant]
US 20070227344A1 · Ryle · 2007 [cited by examiner]
US 20080037804A1 · Shmunk · 2008 [cited by applicant]
US 20080267419A1 · DeBoer · 2008 [cited by examiner]
US 20140216235A1 · Alt · 2014 [cited by applicant]
US 20140260906A1 · Welch · 2014 [cited by applicant]
US 20150278686A1 · Cardinaux et al. · 2015 [cited by applicant]
US 20160179458A1 · Chase · 2016 [cited by applicant]
US 20160328501A1 · Chase · 2016 [cited by applicant]
US 20190164052A1 · Sung et al. · 2019 [cited by applicant]
WO 2017102972A1 · 2017 [cited by applicant]
WO 2019199995A1 · 2019 [cited by applicant]
Julian Parker; Fabian Esqueda; André Bergner, “Modelling of Nonlinear State-Space Systems Using a Deep Neural Network”, Proceedings of the 22nd International Conference on Digital Audio Effects (DAFx-19), Birmingham, UK… [cited by examiner]
Fichas, Felix et al.; “Virtual Analog Modeling of Guitar Amplifiers with Wiener-Hammerstein Models”; Department of Signal Processing and Communications; Hamburg, Germany; Mar. 2018. [cited by applicant]
Yao, J. et al.; “Coarse-to-fine Optimization for Speech Enhancement”; Apple Inc.; arXiv:1908.08044v1; Aug. 21, 2019. [cited by applicant]
Martin-Donas, Juan Manuel et al.; “A Deep Learning Loss Function Based on the Perceptual Evaluation of the Speech Quality”; IEEE Signal Processing Letters; vol. 25, No. 11; Nov. 2018. [cited by applicant]
Damskagg, Eero-Pekka et al.; “Deep Learning for Tube Amplifier Emulation”; Acoustics Lab, Dept. of Signal Processing and Acoustics; Aalto University; Espoo, Finland; arXiv:1811.00334v1; Nov. 2018. [cited by applicant]
Feng, Berthy et al.; “Learning Bandwidth Expansion Perceptually-Motivated Loss”; May 2019. [cited by applicant]
Schmitz, Thomas et al.; “Real Time Emulation of Parametric Guitar Tube Amplifier with Long Short Term Memory Neural Network”; Department of Electrical Engineering and Computer Science; Liege University; Montefiore Insti… [cited by applicant]
Damskagg, Eero-Pekka et al.; “Real-Time Modeling of Audio Distortion Circuits with Deep Learning”; Acoustics Lab, Department of Signal Processing and Acoustics; Aalto University; Espoo, Finland; Jan. 2019. [cited by applicant]
Su, Jiaqi et al.; “Perceptually-Motivated Environment-Specific Speech Enhancement”; 2019. [cited by applicant]
Crama, Philippe et al.; “Computing an Initial Estimate of a Wiener-Hammerstein System With a Random Phase Multisine Excitation”; IEEE Transactions on Instrumentation and Measurement; vol. 54, No. 1,; Feb. 2005. [cited by applicant]
Extended European Search Report for European Patent Application No. 20151065.8; European Patent Office; Munich, Germany; May 29, 2020. [cited by applicant]
Hawley, Scott H et al.; “SignalTrain: Profiling Audio Compressors with Deep Neural Networks”; arXiv:1905.11928v2; ; May 30, 2019. [cited by applicant]
Ramirez, Marco A. Martinez et al.; “End-to-End Equalization with Convolutional Neural Networks”; Proceedings of the 21st International Conference on Digital Audio Effects (DAFx-18); Aveiro, Portugal; Sep. 4-18, 2018. [cited by applicant]
Wright, Alec et al.; “Perceptual Loss Function for Neural Modelling of Audio Systems”; arXiv:1911.08922v1; Nov. 20, 2019. [cited by applicant]
Martinez, Marco et al; “Modeling of nonlinear audio effects with end-to-end deep neural networks”; located at https://www.groundai.com/search/?text=modeling+of+nonlinear+audio+effects+with+end-to-end+deep+neural+network… [cited by applicant]
Mendoza, David Sanchez; “Emulating Electric Guitar Effects with Neural Networks”; Graduation Project, Computer Engineering, Universitat Pompeu Fabra; Sep. 2005. [cited by applicant]
Catala, Omar del Tejo; “Audio Effects Emulation with Neural Networks”; Universitat Politecnica de Valencia; 2016/2017. [cited by applicant]