IP Library › Granted Patent US 11,532,318
Granted Patent B2
US 11,532,318 · App. 16/738,512 · Granted Dec 20, 2022

Neural modeler of audio systems

Inventors: Douglas Andres Castro Borquez (Helsinki, FI); Eero-Pekka Damskägg (Helsinki, FI); Athanasios Gotsopoulos (Helsinki, FI); Lauri Juvela (Helsinki, FI); Thomas William Sherson (Utrecht, NL)
Assignee: Neural DSP Technologies Oy
G10L25/30G10L15/16H04R3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,532,318
App. No.
16/738,512
Granted
Dec 20, 2022
Kind
B2
Abstract

A neural network is trained to digitally model a reference audio system. Training is carried out by repeatedly performing a set of operations. The set of operations includes predicting by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system, and the prediction is carried out in the time domain. The set of operations also includes applying a perceptual loss function to the neural network based upon a determined psychoacoustic property, wherein the perceptual loss function is applied in the frequency domain. Moreover, the set of operations includes adjusting the neural network responsive to the output of the perceptual loss function. A neural model file is output that can be loaded to generate a virtualization of the reference audio system.

Claims (87)

1. A process for creating digital audio systems, comprising:

training a neural network that digitally models a reference audio system by modeling a non-linear behavior of the reference audio system, modeling a first linear aspect of the reference audio system, and modeling a second linear aspect of the reference audio system, and the training carried out by repeatedly performing operations for:

predicting by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system, and the prediction is carried out in the time domain;

applying a perceptual loss function to the neural network based upon a determined psychoacoustic property, wherein the perceptual loss function is applied in the frequency domain; and

adjusting the neural network responsive to the output of the perceptual loss function; and

outputting a neural model file that can be loaded to generate a virtualization of the reference audio system.

2. The process of claim 1 further comprising:

training each role of the neural network occurs at the same time such that all parameters of the model are learned simultaneously.

3. The process of claim 1 , wherein:

modeling the non-linear behavior of the reference audio system, modeling the first linear aspect and/or a temporal dependency of the reference audio system, and modeling the second linear aspect of the reference audio system are arranged in series, parallel, or a combination thereof.

4. The process of claim 1 , wherein applying the perceptual loss function to the neural network comprises establishing a loudness threshold such that a signal below the threshold is not optimized further.

5. The process of claim 4 , wherein establishing the loudness threshold comprises establishing a threshold of hearing for each of multiple frequency bins;

wherein:

for each frequency bin, a loudness threshold is independently set under which the signal is not optimized further in order to optimize further that particular frequency bin.

6. The process of claim 1 , wherein applying the perceptual loss function to the neural network comprises:

implementing frequency masking such that a frequency component is not further processed if a computed error is below a masking threshold, where the masking threshold is based upon a target signal.

7. The process of claim 6 , wherein implementing frequency masking comprises selecting a specific masking threshold for each of multiple frequency bins.

8. The process of claim 1 further comprising:

loading the neural model file into a model audio system to define a virtualization of the reference audio system; and

outputting an audio signal using the virtualization such that the output of the model audio system includes at least one characteristic of an output of the reference audio system, wherein outputting the audio signal is performed upon coupling a musical instrument based on an input from the musical instrument to the model audio system.

9. The process of claim 1 , wherein training the neural network comprises at least one of:

training a convolutional neural network; and

training a recurrent neural network.

10. The process of claim 1 further comprising:

initializing the neural network based on measurements of the reference audio system.

11. The process of claim 1 , wherein initializing the neural network comprises initializing the neural network using measurements based on sine sweeps.

12. The process of claim 1 , wherein the neural network is extended to any combination and/or order of signal-processing waveshapers and signal-processing filters.

13. The process of claim 1 , wherein the operations are differentiable so as to be able to calculate gradients with regard to the predicted signal.

14. A process for creating digital audio systems, comprising:

training a neural network that digitally models a reference audio system by repeatedly performing:

predicting by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system, and the prediction is carried out in the time domain;

applying a perceptual loss function to the neural network based upon a determined psychoacoustic property, wherein the perceptual loss function is applied in the frequency domain;

adjusting the neural network responsive to the output of the perceptual loss function; and

computing an error signal by:

receiving a target signal and an associated predicted signal given from the neural network, and computing therefrom in the time domain, the error signal;

further comprising:

converting the target signal and the error signal to the frequency domain;

applying critical band filtering to the target signal in the frequency domain;

applying critical band filtering to the error signal in the frequency domain; and

thresholding the error signal in the frequency domain according to a frequency-dependent threshold level, where the threshold level is established based upon a predetermined threshold of hearing; and

outputting a neural model file that can be loaded to generate a virtualization of the reference audio system.

15. The process of claim 14 further comprising:

establishing a frequency-dependent mask thresholding level based on the target signal; and

thresholding the error signal in the frequency domain according to the established frequency-dependent mask thresholding level.

16. A process for creating digital audio systems, comprising:

training a neural network that digitally models a reference audio system by repeatedly performing:

predicting by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system, and the prediction is carried out in the time domain;

applying a perceptual loss function to the neural network based upon a determined psychoacoustic property, wherein the perceptual loss function is applied in the frequency domain by:

receiving a target signal and sorting the received target signal into target critical bands;

generating from the target signal in each target critical band, an associated masking threshold;

receiving an error signal generated from the target signal and an associated prediction signal, where the error signal is sorted into error signal critical bands;

applying a threshold of hearing function to the error signal, wherein an error signal below an associated hearing threshold of a corresponding error signal critical band does not contribute to a final error; and

applying a masking function to the error signal, wherein an error signal below the associated masking threshold of a corresponding error signal critical band does not contribute to the final error; and

adjusting the neural network responsive to the output of the perceptual loss function;

changing at least one parameter of the neural network responsive to the final error output of the perceptual loss function; and

outputting a neural model file that can be loaded to generate a virtualization of the reference audio system.

17. A process for creating and using digital audio systems, comprising:

training a neural network that digitally models a reference audio system by repeatedly performing:

predicting by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system, and the prediction is carried out in the time domain;

applying a perceptual loss function to the neural network, where the perceptual loss function is applied in the frequency domain, the perceptual loss function implemented by:

receiving a target signal and sorting the received target signal into target critical bands;

generating from the target signal in each target critical band, an associated masking threshold;

receiving an error signal generated from the target signal and an associated prediction signal, where the error signal is sorted into error signal critical bands;

applying a threshold of hearing function to the error signal, wherein an error signal below an associated hearing threshold of a corresponding error signal critical band does not contribute to a final error; and

applying a masking function to the error signal, wherein an error signal below the associated masking threshold of a corresponding error signal critical band does not contribute to the final error; and

changing at least one parameter of the neural network responsive to the final error output of the perceptual loss function; and

generating a neural model; and

loading the neural model file into a model audio system to define a virtualization of the reference audio system;

wherein:

upon coupling a musical instrument to the model audio system, a user can perform using the virtualization in place of the reference audio system such that an output of the model audio system includes at least one characteristic of an output of the reference audio system.

18. A hardware system, comprising:

an analog to digital converter;

a digital to analog converter; and

processing circuitry that couples to the analog to digital converter and to the digital to analog converter, the processing circuitry having a processor coupled to memory, wherein the processor executes instructions that:

train a neural network that digitally models a reference audio system by repeatedly performing instructions to:

predict by the neural network, a model output based upon an input, where the output approximates an expected output of the reference audio system, and the prediction is carried out in the time domain;

apply a perceptual loss function to the neural network, where the perceptual loss function is applied in the frequency domain, the perceptual loss function implemented to:

receive a target signal and sort the received target signal into target critical bands;

generate from the target signal in each target critical band, an associated masking threshold;

receive an error signal generated from the target signal and an associated prediction signal, where the error signal is sorted into error signal critical bands;

apply a threshold of hearing function to the error signal, wherein an error signal below an associated hearing threshold of a corresponding error signal critical band does not contribute to a final error; and

apply a masking function to the error signal, wherein an error signal below the associated masking threshold of a corresponding error signal critical band does not contribute to the final error; and

change at least one parameter of the neural network responsive to the final error output of the perceptual loss function; and

generate a neural model file; and

load the neural model file into a model audio system to define a virtualization of the reference audio system;

wherein:

upon coupling a musical instrument to the hardware system, a user can perform using the virtualization in place of the reference audio system such that an output of the model audio system includes at least one characteristic of an output of the reference audio system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2020
From: CASTRO BORQUEZ, DOUGLAS ANDRES; DAMSKÄGG, EERO-PEKKA; GOTSOPOULOS, ATHANASIOS; JUVELA, LAURI; SHERSON, THOMAS WILLIAM
To: NEURAL DSP TECHNOLOGIES OY
Reel/Frame 051713/0085 →
Continuity (2)
Provisional Application 62941986 · Nov 29, 2019
Related Publication 20210166718A1 · Jun 3, 2021