IP Library › Granted Patent US 12,334,043
Granted Patent B2
US 12,334,043 · App. 17/924,701 · Granted Jun 17, 2025

Time-varying and nonlinear audio processing using deep neural networks

Inventors: Marco Antonio Martinez Ramirez (London, GB); Joshua Daniel Reiss (London, GB); Emmanouil Benetos (London, GB)
Assignee: WAVESHAPER TECHNOLOGIES INC.
G10H1/0091G06N3/0442G06N3/045G06N3/0499G06N3/08G10H1/16G10H2210/215G10H2210/281G10H2210/311G10H2250/025G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,043
App. No.
17/924,701
Granted
Jun 17, 2025
Kind
B2
Abstract

A computer-implemented method of processing audio data, the method comprising receiving input audio data (x) comprising a time-series of amplitude values; transforming the input audio data (x) into an input frequency band decomposition (X1) of the input audio data (x); transforming the input frequency band decomposition (X1) into a first latent representation (Z); processing the first latent representation (Z) by a first deep neural network to obtain a second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )}); transforming the second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )}) to obtain a discrete approximation (X3{circumflex over ( )}); element-wise multiplying the discrete approximation (X3{circumflex over ( )}) and a residual feature map (R, X5{circumflex over ( )}) to obtain a modified feature map, wherein the residual feature map (R, X5{circumflex over ( )}) is derived from the input frequency band decomposition (X1); processing a pre-shaped frequency band decomposition by a waveshaping unit to obtain a waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}), wherein the pre-shaped frequency band decomposition is derived from the input frequency band decomposition (X1), wherein the waveshaping unit comprises a second deep neural network; summing the waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}) and a modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) to obtain a summation output (X0{circumflex over ( )}), wherein the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) is derived from the modified feature map; and transforming the summation output (X0{circumflex over ( )}) to obtain target audio data (y{circumflex over ( )}).

Claims (113)

1. A computer-implemented method of processing audio data, the method comprising:

receiving input audio data (x) comprising a time-series of amplitude values;

transforming the input audio data (x) into an input frequency band decomposition (X1) of the input audio data (x);

transforming the input frequency band decomposition (X1) into a first latent representation (Z);

processing the first latent representation (Z) by a first deep neural network to obtain a second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )});

transforming the second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )}) to obtain a discrete approximation (X3{circumflex over ( )});

element-wise multiplying the discrete approximation (X3{circumflex over ( )}) and a residual feature map (R, X5{circumflex over ( )}) to obtain a modified feature map, wherein the residual feature map (R, X5{circumflex over ( )}) is derived from the input frequency band decomposition (X1);

processing a pre-shaped frequency band decomposition by a waveshaping unit to obtain a waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}), wherein the pre-shaped frequency band decomposition is derived from the input frequency band decomposition (X1), wherein the waveshaping unit comprises a second deep neural network;

summing the waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}) and a modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) to obtain a summation output (X0{circumflex over ( )}), wherein the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) is derived from the modified feature map; and

transforming the summation output (X0{circumflex over ( )}) to obtain target audio data (y{circumflex over ( )}).

2. The method of claim 1 , wherein transforming the input audio data (x) into the input frequency band decomposition (X1) comprises convolving the input audio data (x) with kernel matrix (W1).

3. The method of claim 1 , wherein transforming the input audio data (x) into the input frequency band decomposition (X1) comprises convolving the input audio data (x) with kernel matrix (W1); wherein transforming the summation output (X0{circumflex over ( )}) to obtain the target audio data (y{circumflex over ( )}) comprises convolving the summation output (X0{circumflex over ( )}) with the transpose of the kernel matrix (W1T).

4. The method of claim 1 , wherein transforming the input frequency band decomposition (X1) into the first latent representation (Z) comprises locally-connected convolving the absolute value (|X1|) of the input frequency band decomposition (X1) with a weight matrix (W2) to obtain a feature map (X2); and max-pooling the feature map (X2) to obtain the first latent representation (Z).

5. The method of claim 1 , wherein the waveshaping unit further comprises a locally connected smooth adaptive activation function layer following the second deep neural network.

6. The method of claim 1 , wherein the waveshaping unit further comprises a locally connected smooth adaptive activation function layer following the second deep neural network; wherein the waveshaping unit further comprises a first squeeze-and-excitation layer following the locally connected smooth adaptive activation function layer.

7. The method of claim 1 , wherein at least one of the waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}) and the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) is scaled by a gain factor (se, se1, se2) before summing to produce the summation output (X0{circumflex over ( )}).

8. The method of claim 1 , wherein the second deep neural network comprises first to fourth dense layers.

9. The method of claim 1 ,

wherein the waveshaping unit further comprises a locally connected smooth adaptive activation function layer following the second deep neural network;

wherein the waveshaping unit further comprises a first squeeze-and-excitation layer following the locally connected smooth adaptive activation function layer;

wherein, in the waveshaping unit, the first squeeze-and-excitation layer comprises an absolute value layer preceding a global average pooling operation.

10. The method of claim 1 , further comprising:

passing on the input frequency band decomposition (X1) as the residual feature map (R);

passing on the modified feature map as the pre-shaped frequency band decomposition; and

passing on the modified feature map as the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}).

11. The method of claim 1 , further comprising:

passing on the input frequency band decomposition (X1) as the residual feature map (R);

passing on the modified feature map as the pre-shaped frequency band decomposition; and

passing on the modified feature map as the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )});

wherein the first deep neural network comprises a plurality of bidirectional long short-term memory layers.

12. The method of claim 1 , further comprising:

passing on the input frequency band decomposition (X1) as the residual feature map (R);

passing on the modified feature map as the pre-shaped frequency band decomposition; and

passing on the modified feature map as the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )});

wherein the first deep neural network comprises a plurality of bidirectional long short-term memory layers;

wherein the plurality of bidirectional long short-term memory layers comprises first, second and third bidirectional long short-term memory layers.

13. The method of claim 1 , further comprising:

passing on the input frequency band decomposition (X1) as the residual feature map (R);

passing on the modified feature map as the pre-shaped frequency band decomposition; and

passing on the modified feature map as the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )});

wherein the first deep neural network comprises a plurality of bidirectional long short-term memory layers;

wherein the plurality of bidirectional long short-term memory layers is followed by a plurality of smooth adaptive activation function layers.

14. The method of claim 1 , further comprising:

passing on the input frequency band decomposition (X1) as the residual feature map (R);

passing on the modified feature map as the pre-shaped frequency band decomposition; and

passing on the modified feature map as the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )});

wherein the first deep neural network comprises a plurality of bidirectional long short-term memory layers;

wherein the first deep neural network comprises a feedforward WaveNet comprising a plurality of layers.

15. The method of claim 1 ,

wherein the first deep neural network comprises a plurality of shared bidirectional long short-term memory layers, followed by, in parallel, first and second independent bidirectional long short-term memory layers;

wherein the second latent representation (Z1{circumflex over ( )}) is derived from the output of the first independent bidirectional long short-term memory layer;

wherein, in the waveshaping unit, the first squeeze-and-excitation layer further comprises a long short-term memory layer;

wherein the method further comprises:

passing on the input frequency band decomposition (X1) as the pre-shaped frequency band decomposition;

processing the first latent representation (Z) using the second independent bidirectional long short-term memory layer to obtain a third latent representation (Z2{circumflex over ( )});

processing the third latent representation (Z2{circumflex over ( )}) using a sparse finite impulse response layer to obtain a fourth latent representation (Z3{circumflex over ( )});

convolving the frequency band representation (X1) with the fourth latent representation (Z3{circumflex over ( )}) to obtain said residual feature map (X5{circumflex over ( )}); and

processing the modified feature map by a second squeeze-and-excitation layer comprising a long short-term memory layer to obtain said modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}).

16. The method of claim 1 ,

wherein the first deep neural network comprises a plurality of shared bidirectional long short-term memory layers, followed by, in parallel, first and second independent bidirectional long short-term memory layers;

wherein the second latent representation (Z1{circumflex over ( )}) is derived from the output of the first independent bidirectional long short-term memory layer;

wherein, in the waveshaping unit, the first squeeze-and-excitation layer further comprises a long short-term memory layer;

wherein the method further comprises:

passing on the input frequency band decomposition (X1) as the pre-shaped frequency band decomposition;

processing the first latent representation (Z) using the second independent bidirectional long short-term memory layer to obtain a third latent representation (Z2{circumflex over ( )});

processing the third latent representation (Z2{circumflex over ( )}) using a sparse finite impulse response layer to obtain a fourth latent representation (Z3{circumflex over ( )});

convolving the frequency band representation (X1) with the fourth latent representation (Z3{circumflex over ( )}) to obtain said residual feature map (X5{circumflex over ( )}); and

processing the modified feature map by a second squeeze-and-excitation layer comprising a long short-term memory layer to obtain said modified frequency band decomposition (X2{circumflex over ( )}, X1.1 {circumflex over ( )});

wherein the plurality of shared bidirectional long short-term memory layers comprises first and second shared bidirectional long short-term memory layers.

17. The method of claim 1 ,

wherein the first deep neural network comprises a plurality of shared bidirectional long short-term memory layers, followed by, in parallel, first and second independent bidirectional long short-term memory layers;

wherein the second latent representation (Z1{circumflex over ( )}) is derived from the output of the first independent bidirectional long short-term memory layer;

wherein, in the waveshaping unit, the first squeeze-and-excitation layer further comprises a long short-term memory layer;

wherein the method further comprises:

passing on the input frequency band decomposition (X1) as the pre-shaped frequency band decomposition;

processing the first latent representation (Z) using the second independent bidirectional long short-term memory layer to obtain a third latent representation (Z2{circumflex over ( )});

processing the third latent representation (Z2{circumflex over ( )}) using a sparse finite impulse response layer to obtain a fourth latent representation (Z3{circumflex over ( )});

convolving the frequency band representation (X1) with the fourth latent representation (Z3{circumflex over ( )}) to obtain said residual feature map (X5{circumflex over ( )}); and

processing the modified feature map by a second squeeze-and-excitation layer comprising a long short-term memory layer to obtain said modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )})

wherein each of the first and second independent bidirectional long short-term memory layers comprises 16 units.

18. The method of claim 1 ,

wherein the first deep neural network comprises a plurality of shared bidirectional long short-term memory layers, followed by, in parallel, first and second independent bidirectional long short-term memory layers;

wherein the second latent representation (Z1{circumflex over ( )}) is derived from the output of the first independent bidirectional long short-term memory layer;

wherein, in the waveshaping unit, the first squeeze-and-excitation layer further comprises a long short-term memory layer;

wherein the method further comprises:

passing on the input frequency band decomposition (X1) as the pre-shaped frequency band decomposition;

processing the first latent representation (Z) using the second independent bidirectional long short-term memory layer to obtain a third latent representation (Z2{circumflex over ( )});

processing the third latent representation (Z2{circumflex over ( )}) using a sparse finite impulse response layer to obtain a fourth latent representation (Z3{circumflex over ( )});

convolving the frequency band representation (X1) with the fourth latent representation (Z3 {circumflex over ( )}) to obtain said residual feature map (X5{circumflex over ( )}); and

processing the modified feature map by a second squeeze-and-excitation layer comprising a long short-term memory layer to obtain said modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )});

wherein the sparse finite impulse response layer comprises:

first and second independent dense layers taking the third latent representation (Z2{circumflex over ( )}) as input; and

a sparse tensor taking the respective output of the first and second independent dense layers as inputs, the output of the sparse tensor being the fourth latent representation (Z3{circumflex over ( )}).

19. A non-transitory computer-readable medium having stored thereon a computer program comprising instructions which, when the computer program is executed by a computer, cause the computer to perform a method of processing audio data, the method comprising:

receiving input audio data (x) comprising a time-series of amplitude values;

transforming the input audio data (x) into an input frequency band decomposition (X1) of the input audio data (x);

transforming the input frequency band decomposition (X1) into a first latent representation (Z);

processing the first latent representation (Z) by a first deep neural network to obtain a second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )});

transforming the second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )}) to obtain a discrete approximation (X3{circumflex over ( )});

element-wise multiplying the discrete approximation (X3{circumflex over ( )}) and a residual feature map (R, X5{circumflex over ( )}) to obtain a modified feature map, wherein the residual feature map (R, X5{circumflex over ( )}) is derived from the input frequency band decomposition (X1);

processing a pre-shaped frequency band decomposition by a waveshaping unit to obtain a waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}), wherein the pre-shaped frequency band decomposition is derived from the input frequency band decomposition (X1), wherein the waveshaping unit comprises a second deep neural network;

summing the waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}) and a modified frequency band decomposition (X2 {circumflex over ( )}, X1.1{circumflex over ( )}) to obtain a summation output (X0{circumflex over ( )}), wherein the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) is derived from the modified feature map; and

transforming the summation output (X0{circumflex over ( )}) to obtain target audio data (y{circumflex over ( )}).

20. An audio data processing device comprising a processor configured to perform the method of processing audio data, the method comprising:

receiving input audio data (x) comprising a time-series of amplitude values;

transforming the input audio data (x) into an input frequency band decomposition (X1) of the input audio data (x);

transforming the input frequency band decomposition (X1) into a first latent representation (Z);

processing the first latent representation (Z) by a first deep neural network to obtain a second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )});

transforming the second latent representation (Z{circumflex over ( )}, Z1{circumflex over ( )}) to obtain a discrete approximation (X3{circumflex over ( )});

element-wise multiplying the discrete approximation (X3{circumflex over ( )}) and a residual feature map (R, X5{circumflex over ( )}) to obtain a modified feature map, wherein the residual feature map (R, X5{circumflex over ( )}) is derived from the input frequency band decomposition (X1);

processing a pre-shaped frequency band decomposition by a waveshaping unit to obtain a waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}), wherein the pre-shaped frequency band decomposition is derived from the input frequency band decomposition (X1), wherein the waveshaping unit comprises a second deep neural network;

summing the waveshaped frequency band decomposition (X1{circumflex over ( )}, X1.2{circumflex over ( )}) and a modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) to obtain a summation output (X0{circumflex over ( )}), wherein the modified frequency band decomposition (X2{circumflex over ( )}, X1.1{circumflex over ( )}) is derived from the modified feature map; and

transforming the summation output (X0{circumflex over ( )}) to obtain target audio data (y{circumflex over ( )}).

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2025
From: QUEEN MARY UNIVERSITY OF LONDON
To: WAVESHAPER TECHNOLOGIES INC.
Reel/Frame 070314/0715 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2022
From: MARTINEZ RAMIREZ, MARCO ANTONIO; REISS, JOSHUA DANIEL; BENETOS, EMMANOUIL
To: QUEEN MARY UNIVERSITY OF LONDON
Reel/Frame 062155/0462 →
Continuity (1)
Related Publication 20230197043A1 · Jun 22, 2023
References Cited (122)
US 10068557B1 · Engel · 2018 [cited by examiner]
JP 202027245A · 2020 [cited by applicant]
Martinez et al.; Modeling Plate and Spring Reverberation Using a DSP-Informed Deep Neural Network; (Year: 2019). [cited by examiner]
Martinez et al.; A General-Purpose Deep Learning Approach to Model Time-Varying Audio Efects (Year: 2019). [cited by examiner]
Martinez et al.; Deep Leaning for Black-Box Modeling of Audio Efects (Year: 2020). [cited by examiner]
Stephan Möller, Martin Gromowski, and Udo Zölzer. A measurement technique for highly nonlinear transfer functions. In 5th International Conference on Digital Audio Effects (DAFx-02), Sep. 26-28, 2002. [cited by applicant]
Filip Korzeniowski and Gerhard Widmer. Feature learning for chord recognition: The deep chroma extractor. In 17th International Society for Music Information Retrieval Conference (ISMIR), 2016. [cited by applicant]
Oliver Kröning, Kristjan Dempwolf, and Udo Zölzer. Analysis and simulation of an analog guitar compressor. In 14th International Conference on Digital Audio Effects (DAFx-11), 2011. [cited by applicant]
Walter Kuhl. The acoustical and technological properties of the reverberation plate. E. B. U. Review, 49, 1958. [cited by applicant]
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller. Efficient backprop. Neural networks: Tricks of the trade, pp. 9-48, 2012. [cited by applicant]
Honglak Lee, Peter Pham, Yan Largman, and Andrew Y Ng. Unsupervised feature learning for audio classification using convolutional deep belief networks. In Advances in neural information processing systems, pp. 1096-1104… [cited by applicant]
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam. SampleCNN: End-to-end deep convolutional neural networks using very small filters for music classification. Applied Sciences, 8(1):150, 2018. [cited by applicant]
Keun Sup Lee, Nicholas J Bryan, and Jonathan S Abel. Approximating measured reverberation using a hybrid fixed/switched convolution structure. In 13th International Conference on Digital Audio Effects (DAFx-10), 2010. [cited by applicant]
Teck Yian Lim, Raymond A Yeh, Yijia Xu, Minh N Do, and Mark Hasegawa-Johnson. Time-frequency networks for audio super-resolution. In IEEE Inter-national Conference on Acoustics, Speech and Signal Processing (ICASSP), 20… [cited by applicant]
Jaromír Macák. Simulation of analog flanger effect using BBD circuit. In 19th International Conference on Digital Audio Effects (DAFx-16), 2016. [cited by applicant]
Jacob A Maddams, Saoirse Finn, and Joshua D Reiss. An autonomous method for multi-track dynamic range compression. In 15th International Conference on Digital Audio Effects (DAFx-12), 2012. [cited by applicant]
EP Matthew Davies and Sebastian Bock. Temporal convolutional networks for musical audio beat tracking. In 27th IEEE European Signal Processing Conference (EUSIPCO), 2019. [cited by applicant]
Daniel Matz, Estefanía Cano, and Jakob Abeßer. New sonorities for early jazz recordings using sound source separation and automatic mixing tools. In 16th International Society for Music Information Retrieval Conference … [cited by applicant]
Josh H McDermott and Eero P Simoncelli. Sound texture perception via statistics of the auditory periphery: evidence from sound synthesis. Neuron, 71, 2011. [cited by applicant]
Martin McKinney and Jeroen Breebaart. Features for audio and music classification. In 4th International Society for Music Information Retrieval Conference (ISMIR), 2003. [cited by applicant]
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio. SampleRNN: An unconditional end-to-end neural audio generation model. In 5th International Con… [cited by applicant]
Stephan Möller, Martin Gromowski, and Udo Zölzer. A measurement technique for highly nonlinear transfer functions. In 5th International Conference on Digital Audio Effects (DAFx-02), 2002. [cited by applicant]
James A Moorer. About this reverberation business. Computer music journal, pp. 13-28, 1979. [cited by applicant]
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio. In CoRR abs/1609.03499, 2016. [cited by applicant]
Jyri Pakarinen and David T Yeh. A review of digital techniques for modeling vacuum-tube guitar amplifiers. Computer Music Journal, 33(2): 85-100, 2009. [cited by applicant]
Bryan Pardo, David Little, and Darren Gergle. Building a personalized audio equalizer interface with transfer learning and active learning. In 2nd International ACM Workshop on Music Information Retrieval with User-Cent… [cited by applicant]
Julian Parker. Efficient dispersion generation structures for spring reverb emulation. EURASIP Journal on Advances in Signal Processing, 2011a. [cited by applicant]
Julian Parker. A simple digital model of the diode-based ring-modulator. In 14th International Conference on Digital Audio Effects (DAFx-11), 2011b. [cited by applicant]
Julian Parker and Stefan Bilbao. Spring reverberation: A physical perspective. In 12th International Conference on Digital Audio Effects (DAFx-09), 2009. [cited by applicant]
Julian Parker and Fabian Esqueda. Modelling of nonlinear state-space systems using a deep neural network. In 22nd International Conference on Digital Audio Effects (DAFx-19), 2019. [cited by applicant]
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International Conference on Machine Learning, 2013. [cited by applicant]
Jussi Pekonen, Tapani Pihlajamaki, and Vesa Välimäki. Computationally efficient hammond organ synthesis. In 14th International Conference on Digital Audio Effects (DAFx-11), 2011. [cited by applicant]
Enrique Perez-Gonzalez and Joshua D. Reiss. Automatic equalization of multi-channel audio using crossadaptive methods. In 127th Audio Engineering Society Convention, 2009. [cited by applicant]
Jordi Pons, Oriol Nieto, Matthew Prockup, Erik Schmidt, Andreas Ehmann, and Xavier Serra. End-to-end learning for music audio tagging at scale. In 31st Conference on Neural Information Processing Systems, 2017. [cited by applicant]
Colin Raffel and Julius O Smith. Practical modeling of bucket-brigade device circuits. In 13th International Conference on Digital Audio Effects (DAFx-10), 2010. [cited by applicant]
Jussi Rämö and Vesa Välimäki. Neural third-octave graphic equalizer. In 22nd International Conference on Digital Audio Effects (DAFx-19), 2019. [cited by applicant]
Dale Reed. A perceptual assistant to do sound equalization. In 5th International Conference on Intelligent User Interfaces, pp. 212-218. ACM, 2000. [cited by applicant]
Dario Rethage, Jordi Pons, and Xavier Serra. A wavenet for speech denoising. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018. [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 2015. [cited by applicant]
Andrew T Sabin and Bryan Pardo. A method for rapid personalization of audio equalization parameters. In 17th ACM International Conference on Multimedia, 2009. [cited by applicant]
Jan Schlüter and Sebastian Böck. Musical onset detection with convolutional neural networks. In 6th International Workshop on Machine Learning and Music, 2013. [cited by applicant]
Jan Schlüter and Sebastian Böck. Improved musical onset detection with convolutional neural networks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014. [cited by applicant]
Manfred R Schroeder and Benjamin F Logan. “Colorless” artificial reverberation. IRE Transactions on Audio, (6):209-214, 1961. [cited by applicant]
Mike Schuster and Kuldip K Paliwal. Bidirectional recurrent neural networks. IEEE transactions on Signal Processing, 45(11):2673-2681, 1997. [cited by applicant]
Di Sheng and György Fazekas. Automatic control of the dynamic range compressor using a regression model and a reference sound. In 20th International Conference on Digital Audio Effects (DAFx-17), 2017. [cited by applicant]
Di Sheng and György Fazekas. A feature learning siamese model for intelligent control of the dynamic range compressor. In International Joint Conference on Neural Networks (IJCNN), 2019. [cited by applicant]
Siddharth Sigtia and Simon Dixon. Improved music feature learning with deep neural networks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014. [cited by applicant]
Siddharth Sigtia, Emmanouil Benetos, Nicolas Boulanger-Lewandowski, Tillman Weyde, Artur S d'Avila Garcez, and Simon Dixon. A hybrid recurrent neural network for music transcription. In IEEE international conference on … [cited by applicant]
Siddharth Sigtia, Emmanouil Benetos, and Simon Dixon. An end-to-end neural network for polyphonic piano music transcription. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 24(5):927-939, 2016. [cited by applicant]
Julius O Smith and Jonathan S Abel. Bark and ERB bilinear transforms. IEEE Transactions on Speech and Audio Processing, 7(6):697-708, 1999. [cited by applicant]
Julius O Smith, Stefania Serafin, Jonathan Abel, and David Berners. Doppler simulation and the leslie. In 5th International Conference on Digital Audio Effects (DAFx-02), 2002. [cited by applicant]
Mirko Solazzi and Aurelio Uncini. Artificial neural networks with adaptive multi-dimensional spline activation functions. In IEEE International Joint Conference on Neural Networks (IJCNN), 2000. [cited by applicant]
Karl Steinberg. Steinberg virtual studio technology (VST) plug-in specification 2.0 software development kit. Hamburg: Steinberg Soft-und Hardware GMBH, 1999. [cited by applicant]
Dan Stowell and Mark D Plumbley. Automatic large-scale classification of bird sounds is strongly improved by unsupervised feature learning. PeerJ, 2:e488, 2014. [cited by applicant]
Bob L Sturm, Joao Felipe Santos, Oded Ben-Tal, and Iryna Korshunova. Music transcription modelling and composition using deep learning. In 1st Conference on Computer Simulation of Musical Creativity, 2016. [cited by applicant]
Somsak Sukittanon, Les E Atlas, and James W Pitton. Modulation-scale analysis for content identification. IEEE Transactions on Signal Processing, 52, 2004. [cited by applicant]
Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples. In International Conference on Machine Learning, 2018. [cited by applicant]
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. Convolutional sequence modeling revisited. In 6th International Conference on Learning Representations (ICLR), 2018. [cited by applicant]
Stefan Bilbao. Numerical simulation of spring reverberation. In 16th International Conference on Digital Audio Effects (DAFx-13), 2013. [cited by applicant]
Stefan Bilbao and Julian Parker. A virtual model of spring reverberation. IEEE Transactions on Audio, Speech and Language Processing, 18(4):799-808, 2009. [cited by applicant]
Stefan Bilbao, Kevin Arcas, and Antoine Chaigne. A physical model for plate reverberation. In IEEE International Conference on Acoustics, Speech, and Signal Processing, 2006. [cited by applicant]
Merlijn Blaauw and Jordi Bonada. A neural parametric singing synthesizer. In Interspeech, 2017. [cited by applicant]
Ólafur Bogason and Kurt James Werner. Modeling circuits with operational transconductance amplifiers using wave digital filters. In 20th International Conference on Digital Audio Effects (DAFx-17), 2017. [cited by applicant]
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. cuDNN: Efficient primitives for deep learning. CoRR, abs / 1410.0759, 2014. [cited by applicant]
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase repre-sentations using RNN encoder-decoder for statistical machine translation. … [cited by applicant]
Eero-Pekka Damskägg, Lauri Juvela, Etienne Thuillier, and Vesa Välimäki. Deep learning for tube amplifier emulation. In IEEE International Conference on Acous-tics, Speech, and Signal Processing (ICASSP), 2019. [cited by applicant]
Brecht De Man, Joshua D Reiss, and Ryan Stables. Ten years of automatic mixing. In Proceedings of the 3rd Workshop on Intelligent Music Production, 2017. [cited by applicant]
Michele Ducceschi and Craig J Webb. Plate reverberation: Towards the development of a real-time physical model for the working musician. In International Congress on Acoustics (ICA), 2016. [cited by applicant]
John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of machine learning research, (Jul. 12):2121-2159, 2011. [cited by applicant]
Simon Durand, Juan P Bello, Bertrand David, and Gaël Richard. Downbeat tracking with multiple features and deep neural networks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. [cited by applicant]
Douglas Eck and Juergen Schmidhuber. A first look at music composition using Istm recurrent neural networks. Istituto Dalle Molle Di Studi Sull Intelligenza Artificiale, 103, 2002. [cited by applicant]
Felix Eichas and Udo Zolzer. Black-box modeling of distortion circuits with block- oriented models. In 19th International Conference on Digital Audio Effects (DAFx-16), 2016. [cited by applicant]
Felix Eichas and Udo Zölzer. Virtual analog modeling of guitar amplifiers with wiener-hammerstein models. In 44th Annual Convention on Acoustics, 2018. [cited by applicant]
Felix Eichas, Marco Fink, Martin Holters, and Udo Zölzer. Physical modeling of the mxr phase 90 guitar effect pedal. In 17th International Conference on Digital Audio Effects (DAFx-14), 2014. [cited by applicant]
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan. Neural audio synthesis of musical notes with wavenet autoencoders. 34th International Conference on Machine … [cited by applicant]
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts. DDSP: Differentiable digital signal processing. In 8th International Conference on Learning Representations (ICLR), 2020. [cited by applicant]
Angelo Farina. Simultaneous measurement of impulse response and distortion with a swept-sine technique. In 108th Audio Engineering Society Convention, 2000. [cited by applicant]
Xue Feng, Yaodong Zhang, and James Glass. Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition. In IEEE International Conference on Acoustics, Speech, and Signal Pr… [cited by applicant]
Todor Ganchev, Nikos Fakotakis, and George Kokkinakis. Comparative evaluation of various mfcc implementations on the speaker verification task. In International Conference on Speech and Computer, 2005. [cited by applicant]
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins. Learning to forget: Continual prediction with LSTM. IET, 1999. [cited by applicant]
Dimitrios Giannoulis, Michael Massberg, and Joshua D Reiss. Parameter automation in a dynamic range compressor. Journal of the Audio Engineering Society, 61 (10):716-726, 2013. [cited by applicant]
Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In the 13th International Conference on Artificial Intelligence and Statistics, 2010. [cited by applicant]
Luke B Godfrey and Michael S Gashler. A continuum among logarithmic, linear, and exponential functions, and its potential to improve generalization in neural networks. In 7th IEEE International Joint Conference on Knowl… [cited by applicant]
Alex Graves and Jürgen Schmidhuber. Framewise phoneme classification with bidirectional lstm and other neural network architectures. Neural Networks, 18 (5-6):602-610, 2005. [cited by applicant]
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In IEEE International Conference on Acous-tics, Speech, and Signal Processing (ICASSP), 2013. [cited by applicant]
Anna Hagenblad. Aspects of the identification of Wiener models. PhD thesis, Linköpings Universitet, 1999. [cited by applicant]
Philippe Hamel, Matthew EP Davies, Kazuyoshi Yoshii, and Masataka Goto. Transfer learning in MIR: Sharing learned latent representations for music audio classification and similarity. In 14th International Society for M… [cited by applicant]
Kun Han, Yuxuan Wang, DeLiang Wang, William S Woods, Ivo Merks, and Tao Zhang. Learning spectral mapping for speech dereverberation and denoising. IEEE Transactions on Audio, Speech and Language Processing, 23(6):982-99… [cited by applicant]
Yoonchang Han, Jaehun Kim, and Kyogu Lee. Deep convolutional neural networks for predominant instrument recognition in polyphonic music. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 25 (1):208-221, 2… [cited by applicant]
Scott H Hawley, Benjamin Colburn, and Stylianos I Mimilakis. SignalTrain: Profiling audio compressors with deep neural networks. In 147th Audio Engineering Society Convention, 2019. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, 2016. [cited by applicant]
Thomas Helie. On the use of volterra series for real-time simulations of weakly nonlinear analog audio devices: Application to the moog ladder filter. In 9th International Conference on Digital Audio Effects (DAFx-06), … [cited by applicant]
Marcel Hilsamer and Stephan Herzog. A statistical approach to automated offline dynamic processing in the audio mastering process. In 17th International Conference on Digital Audio Effects (DAFx-14), 2014. [cited by applicant]
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735-1780, 1997. [cited by applicant]
Martin Holters and Julian D Parker. A combined model for a bucket brigade device and its input and output filters. In 21st International Conference on Digital Audio Effects (DAFx-17), 2018. [cited by applicant]
Martin Holters and Udo Zölzer. Physical modelling of a wah-wah effect pedal as a case study for application of the nodal dk method to circuits with variable parts. In 14th International Conference on Digital Audio Effec… [cited by applicant]
Le Hou, Dimitris Samaras, Tahsin M Kurc, Yi Gao, and Joel H Saltz. Neural networks with smooth adaptive activation functions for regression. arXiv preprint arXiv:1608.06557, 2016. [cited by applicant]
Le Hou, Dimitris Samaras, Tahsin M Kurc, Yi Gao, and Joel H Saltz. Convnets with smooth adaptive activation functions for regression. In 20th International Conference on Artificial Intelligence and Statistics (AISTATS),… [cited by applicant]
Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Allen Huang and Raymond Wu. Deep learning for music. CoRR, abs / 1606.04930, 2016. [cited by applicant]
Eric J Humphrey and Juan P Bello. From music audio to chord tablature: Teaching deep convolutional networks to play guitar. In IEEE international conference on acoustics, speech and signal processing (ICASSP), 2014. [cited by applicant]
Antti Huovilainen. Enhanced digital models for analog modulation effects. In 8th International Conference on Digital Audio Effects (DAFx-05), 2005. [cited by applicant]
Nicholas Jillings, Brecht De Man, David Moffat, and Joshua D Reiss. Web Audio Evaluation Tool: A browser-based listening test environment. In 12th Sound and Music Computing Conference, 2015. [cited by applicant]
Roope Kiiski, Fabian Esqueda, and Vesa Välimäki. Time-variant gray-box modeling of a phaser pedal. In 19th International Conference on Digital Audio Effects (DAFx-16), 2016. [cited by applicant]
Taejun Kim, Jongpil Lee, and Juhan Nam. Sample-level CNN architectures for music auto-tagging using raw waveforms. In IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2018. [cited by applicant]
Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations (ICLR), 2015. [cited by applicant]
Tijmen Tieleman and Geoffrey Hinton. RMSprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4(2):26-31, 2012. [cited by applicant]
Aurelio Uncini. Audio signal processing by neural networks. Neurocomputing, 55 (3-4):593-625, 2003. [cited by applicant]
International Telecommunication Union. Recommendation ITU-R BS. 1534-1: Method for the subjective assessment of intermediate quality level of coding systems. 2003. [cited by applicant]
Vesa Välimaki and Joshua D. Reiss. All about audio equalization: Solutions and frontiers. Applied Sciences, 6(5):129, 2016. [cited by applicant]
Aaron Van den Oord, Sander Dieleman, and Benjamin Schrauwen. Deep content-based music recommendation. In Advances in Neural Information Processing Sys-tems, pp. 2643-2651, 2013. [cited by applicant]
Shrikant Venkataramani, Jonah Casebeer, and Paris Smaragdis. Adaptive front- ends for end-to-end source separation. In 31st Conference on Neural Information Processing Systems, 2017. [cited by applicant]
Vincent Verfaille, U. Zölzer, and Daniel Arfib. Adaptive digital audio effects (A-DAFx): A new class of sound transformations. IEEE Transactions on Audio, Speech and Language Processing, 14(5):1817-1831, 2006. [cited by applicant]
Kurt J Werner, W Ross Dunkel, and François G Germain. A computational model of the hammond organ vibrato/chorus using wave digital filters. In 19th International Conference on Digital Audio Effects (DAFx-16), 2016. [cited by applicant]
Silvin Willemsen, Stefania Serafin, and Jesper R Jensen. Virtual analog simulation and extensions of plate reverberation. In 14th Sound and Music Computing Conference, 2017. [cited by applicant]
Alec Wright, Eero-Pekka Damskägg, and Vesa Välimäki. Real-time black-box modelling with recurrent neural networks. In 22nd International Conference on Digital Audio Effects (DAFx-19), 2019. [cited by applicant]
David T Yeh and Julius O Smith. Simulating guitar distortion circuits using wave digital and nonlinear statespace formulations. In 11th International Conference on Digital Audio Effects (DAFx-08), 2008. [cited by applicant]
David T Yeh, Jonathan S Abel, and Julius O Smith. Automated physical modeling of nonlinear audio circuits for real-time audio effects part I: Theoretical development. IEEE Transactions on Audio, Speech, and Language Pro… [cited by applicant]
Matthew D Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. In European conference on computer vision. Springer, 2014. [cited by applicant]
Japanese Office Action regarding Application No. 2022-568979, issued Aug. 20, 2024. [cited by applicant]
Marco A. Martinez Ramirez et al. and “Modeling Nonlinear Audio Effects with End-to-end Deep Neural Networks”, [online] and ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS… [cited by applicant]
Eero-Pekka Damskagg et al. and “Deep Learning for Tube Amplifier Emulation”, [online] and ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 12, 2019, pp. 471-4… [cited by applicant]