IP Library › Granted Patent US 12,579,961
Granted Patent B2
US 12,579,961 · App. 17/726,289 · Granted Mar 17, 2026

Music enhancement systems

Inventors: Nikhil Kandpal (Carrboro, NC); Oriol Nieto-Caballero (Oakland, CA); Zeyu Jin (San Francisco, CA)
Assignee: Adobe Inc.
G10H1/0008G10H1/06G10H2210/066G10H2250/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,961
App. No.
17/726,289
Granted
Mar 17, 2026
Kind
B2
Abstract

In implementations of music enhancement systems, a computing device implements an enhancement system to receive input data describing a recorded acoustic waveform of a musical instrument. The recorded acoustic waveform is represented as an input mel spectrogram. The enhancement system generates an enhanced mel spectrogram by processing the input mel spectrogram using a first machine learning model trained on a first type of training data to generate enhanced mel spectrograms based on input mel spectrograms. An acoustic waveform of the musical instrument is generated by processing the enhanced mel spectrogram using a second machine learning model trained on a second type of training data to generate acoustic waveforms based on mel spectrograms. The acoustic waveform of the musical instrument does not include an acoustic artifact that is included in the recorded waveform of the musical instrument.

Claims (33)

1 . A method comprising:

receiving, by a computing device, input data describing a recorded acoustic waveform of a musical instrument;

representing, by the computing device, the recorded acoustic waveform of the musical instrument as an input mel spectrogram;

generating, by the computing device, an enhanced mel spectrogram by processing the input mel spectrogram using a first machine learning model, the first machine learning model including a conditional generative adversarial network trained on a first type of training data to generate enhanced mel spectrograms based on input mel spectrograms, the first type of training data including pairs of recorded acoustic waveforms and perturbed acoustic waveforms that are generated by modifying the recorded acoustic waveforms, the pairs of recorded acoustic waveforms being used to generate high-quality mel spectrograms and low quality spectrograms for training the first machine learning model to generate the enhanced mel spectrograms; and

generating, by the computing device, an acoustic waveform of the musical instrument by processing the enhanced mel spectrogram using a second machine learning model trained to generate acoustic waveforms based on mel spectrograms, the second machine learning model is a denoising diffusion probabilistic model trained on a second type of training data, which is generated by iteratively adding Gaussian noise to the recorded acoustic waveforms used for generating the first type of training data, the second type of training data further including mel spectrograms computed for the recorded acoustic waveforms, the second machine learning model trained to estimate reverse transition distributions from adding the Gaussian noise to the recorded acoustic waveforms conditioned on the mel spectrograms computed for the recorded acoustic waveforms, the acoustic waveform of the musical instrument does not include an acoustic artifact that is included in the recorded acoustic waveform of the musical instrument.

2 . The method as described in claim 1 , wherein the acoustic artifact is noise, a reverberation, or a particular frequency energy.

3 . The method as described in claim 1 , further comprising filtering inaudible frequencies out from the recorded acoustic waveform of the musical instrument.

4 . The method as described in claim 1 , wherein the acoustic waveform of the musical instrument includes an additional acoustic artifact that is not included in the recorded acoustic waveform of the musical instrument.

5 . The method as described in claim 1 , wherein the perturbed acoustic waveforms are generated by convolving the recorded acoustic waveforms with a room impulse response.

6 . The method as described in claim 1 , wherein the perturbed acoustic waveforms are generated by applying additive background noise to the recorded acoustic waveforms or by applying multi-band equalization with randomly sampled gains to the recorded acoustic waveforms.

7 . The method as described in claim 1 , wherein the musical instrument includes a piano.

8 . The method as described in claim 1 , wherein the pairs of recorded acoustic waveforms and perturbed acoustic waveforms are generated by convolving the recorded acoustic waveforms with a room impulse response.

9 . The method as described in claim 1 , wherein the pairs of recorded acoustic waveforms and perturbed acoustic waveforms are generated by applying additive background noise to the recorded acoustic waveforms or by applying multi-band equalization with randomly sampled gains to the recorded acoustic waveforms.

10 . A system comprising:

a mel spectrogram module implemented by one or more processing devices to:

receive input data describing a recorded acoustic waveform of a musical instrument; and

represent the recorded acoustic waveform of the musical instrument as an input mel spectrogram;

a translation module implemented by the one or more processing devices to generate an enhanced mel spectrogram by processing the input mel spectrogram using a first machine learning model, the first machine learning model including a conditional generative adversarial network trained on a first type of training data to generate enhanced mel spectrograms based on input mel spectrograms, the first type of training data including pairs of recorded acoustic waveforms and perturbed acoustic waveforms that are generated by modifying the recorded acoustic waveforms, the pairs of recorded acoustic waveforms being used to generate high-quality mel spectrograms and low quality spectrograms for training the first machine learning model to generate the enhanced mel spectrograms; and

a vocoding module implemented by the one or more processing devices to generate an acoustic waveform of the musical instrument using a second machine learning model trained to generate acoustic waveforms based on mel spectrograms, the second machine learning model is a denoising diffusion probabilistic model trained on a second type of training data, which is generated by iteratively adding Gaussian noise to the recorded acoustic waveforms used for generating the first type of training data, the second type of training data further including mel spectrograms computed for the recorded acoustic waveforms, the second machine learning model trained to estimate reverse transition distributions from adding the Gaussian noise to the recorded acoustic waveforms conditioned on the mel spectrograms computed for the recorded acoustic waveforms, the acoustic waveform of the musical instrument does not include an acoustic artifact that is included in the recorded acoustic waveform of the musical instrument.

11 . The system as described in claim 10 , wherein the acoustic artifact is noise, a reverberation, or a particular frequency energy.

12 . The system as described in claim 10 , wherein the musical instrument includes a piano.

13 . The system as described in claim 10 , wherein the pairs of recorded acoustic waveforms and perturbed acoustic waveforms are generated by convolving the recorded acoustic waveforms with a room impulse response.

14 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving input data describing a recorded acoustic waveform of a musical instrument that includes an acoustic artifact;

representing the recorded acoustic waveform of the musical instrument as an input mel spectrogram;

generating an enhanced mel spectrogram by processing the input mel spectrogram using a first machine learning model, the first machine learning model including a conditional generative adversarial network trained on a first type of training data to generate enhanced mel spectrograms based on input mel spectrograms, the first type of training data including pairs of recorded acoustic waveforms and perturbed acoustic waveforms that are generated by modifying the recorded acoustic waveforms, the pairs of recorded acoustic waveforms being used to generate a high-quality mel spectrograms and low quality spectrograms for training the first machine learning model to generate the enhanced mel spectrograms; and

generate an acoustic waveform of the musical instrument that does not include the acoustic artifact by processing the enhanced mel spectrogram using a second machine learning model trained to generate acoustic waveforms based on mel spectrograms, the second machine learning model is a denoising diffusion probabilistic model trained on a second type of training data, which is generated by iteratively adding Gaussian noise to the recorded acoustic waveforms used for generating the first type of training data, the second type of training data further including mel spectrograms computed for the recorded acoustic waveforms, the second machine learning model trained to estimate reverse transition distributions from adding the Gaussian noise to the recorded acoustic waveforms conditioned on the mel spectrograms computed for the recorded acoustic waveforms.

15 . The non-transitory computer-readable storage medium as described in claim 14 , wherein the acoustic artifact is noise, a reverberation, or a particular frequency energy.

16 . The non-transitory computer-readable storage medium as described in claim 14 , wherein the acoustic waveform of the musical instrument includes an additional acoustic artifact that is not included in the recorded acoustic waveform of the musical instrument.

17 . The non-transitory computer-readable storage medium as described in claim 14 , wherein the operations further comprise filtering inaudible frequencies out from the recorded acoustic waveform of the musical instrument.

18 . The non-transitory computer-readable storage medium as described in claim 14 , wherein the pairs of recorded acoustic waveforms and perturbed acoustic waveforms are generated by convolving the recorded acoustic waveforms with a room impulse response.

19 . The non-transitory computer-readable storage medium as described in claim 14 , wherein the pairs of recorded acoustic waveforms and perturbed acoustic waveforms are generated by applying additive background noise to the recorded acoustic waveforms or by applying multi-band equalization with randomly sampled gains to the recorded acoustic waveforms.

20 . The non-transitory computer-readable storage medium as described in claim 14 , wherein the musical instrument includes a piano.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2022
From: KANDPAL, NIKHIL; NIETO-CABALLERO, ORIOL; JIN, ZEYU
To: ADOBE INC.
Reel/Frame 059670/0758 →
Continuity (1)
Related Publication 20230343312A1 · Oct 26, 2023
References Cited (41)
US 10511908B1 · Fisher · 2019 [cited by examiner]
US 20230230609A1 · Seetharaman · 2023 [cited by examiner]
Jonathan Ho et al., “Denoising Diffusion Probabilistic Models,” (https://arxiv.org/pdf/2006.11239.pdf, arXiv:2006.11239v2, published Dec. 16, 2020, retrieved Feb. 6, 2025) (Year: 2020). [cited by examiner]
Chen et al. (“WaveGrad: Estimating Gradients for Waveform Generation),” https://arxiv.org/abs/2009.00713, Oct. 9, 2020. [cited by examiner]
Mehri et al. , SampleRNN: An Unconditional End-To-End Neural Audio Generation Model, Feb. 11, 2017, pp. 1-11. [cited by examiner]
Chen, Nanxin , et al., “WaveGrad: Estimating Gradients for Waveform Generation”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2009.00713.pdf>., Oct. 9… [cited by applicant]
Defossez, Alexandre , et al., “Real Time Speech Enhancement in the Waveform Domain”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2006.12847.pdf>., Se… [cited by applicant]
Défossez, Alexandre , et al., “Music Source Separation in the Waveform Domain”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1911.13254.pdf>., Apr. 28… [cited by applicant]
Duan, Zhiyao , et al., “Speech Enhancement by Online Non-negative Spectrogram Decomposition in Non-stationary Noise Environments”, Interspeech [retrieved Feb. 24, 2022]. Retrieved from the Internet <http://www2.ece.roch… [cited by applicant]
Eaton, J. , et al., “The ACE challenge—Corpus description and performance evaluation”, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) [retrieved Feb. 24, 2022]. Retrieved from the Int… [cited by applicant]
Griffin, Daniel , et al., “Signal estimation from modified short-time Fourier transform”, IEEE Transactions on Acoustics, Speech, and Signal Processing vol. 32, No. 2 [retrieved Feb. 24, 2022]. Retrieved from the Intern… [cited by applicant]
Han, Kun , et al., “Learning Spectral Mapping for Speech Dereverberation and Denoising”, IEEE/ACM Transactions on Audio, Speech, and Language Processing vol. 23, No. 6 [retrieved Feb. 24, 2022]. Retrieved from the Inter… [cited by applicant]
He, Kaiming , et al., “Deep Residual Learning for Image Recognition”, arXiv Preprint, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1512.03385.pdf>., Dec. 10, 2015, 12 pages. [cited by applicant]
Hennequin, Romain , et al., “Spleeter: a fast and efficient music source separation tool with pre-trained models”, Journal of Open Source Software, vol. 5, No. 50 [retrieved Feb. 24, 2022]. Retrieved from the Internet <… [cited by applicant]
Ho, Jonathan , et al., “Denoising Diffusion Probabilistic Models”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2006.11239.pdf>., Dec. 16, 2020, 25 Pa… [cited by applicant]
Hu, Yi , et al., “Evaluation of Objective Quality Measures for Speech Enhancement”, Audio, Speech, and Language Processing, IEEE Transactions on, vol. 16, No. 1 [retrieved Feb. 24, 2022]. Retrieved from the Internet <ht… [cited by applicant]
Isola, Phillip , et al., “Image-to-Image Translation with Conditional Adversarial Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) [retrieved Feb. 24, 2022]. Retrieved from the Internet … [cited by applicant]
Kagami, Hideaki , et al., “Joint Separation and Dereverberation of Reverberant Mixtures with Determined Multichannel Non-Negative Matrix Factorization”, IEEE International Conference on Acoustics, Speech and Signal Proc… [cited by applicant]
Kandpal, Nikhil , et al., “Music Enhancement via Image Translation and Vocoding”, GitHub, Inc., [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://nkandpa2.github.io/music-enhancement/>., 2 Pages. [cited by applicant]
Kilgour, Kevin , et al., “Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/… [cited by applicant]
Kong, Zhifeng , et al., “DiffWave: A Versatile Diffusion Model for Audio Synthesis”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2009.09761.pdf>., Ma… [cited by applicant]
Kumar, Kundan , et al., “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis”, NIPS'19: Proceedings of the 33rd International Conference on Neural Information Processing Systems [retrieved Feb. 24… [cited by applicant]
Lostanlen, Vincent , et al., “Deep convolutional networks on the pitch spiral for musical instrument recognition”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxi… [cited by applicant]
Luo, Yi , et al., “Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation”, IEEE/ACM Transactions on Audio, Speech, and Language Processing vol. 27, No. 8 [retrieved Feb. 24, 2022]. Retriev… [cited by applicant]
Manocha, Pranay , et al., “CDPAM: Contrastive learning for perceptual audio similarity”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2102.05109.pdf>.… [cited by applicant]
Mehri, Soroush , et al., “SampleRNN: An Unconditional End-to-End Neural Audio Generation Model”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1612.078… [cited by applicant]
Michelsanti, Daniel , et al., “Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker Verification”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the… [cited by applicant]
Pascual, Santiago , et al., “SEGAN: Speech Enhancement Generative Adversarial Network”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1703.09452.pdf>.,… [cited by applicant]
Pascual, Santiago , et al., “Towards Generalized Speech Enhancement with Generative Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pd… [cited by applicant]
Polyak, Adam , et al., “High Fidelity Speech Regeneration with Application to Speech Enhancement”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2102.0… [cited by applicant]
Reddy, Chandan K A K A, et al., “DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to evaluate Noise Suppressors”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Inter… [cited by applicant]
Reddy, Chandan KA, et al., “Interspeech 2021 Deep Noise Suppression Challenge”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2101.01902.pdf>., Apr. 5,… [cited by applicant]
Scalart, Pascal , et al., “Speech Enhancement Based on a Priori Signal to Noise Estimation”, Conference Proceedings of 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing; vol. 2 [retrieved Fe… [cited by applicant]
Su, Jiaqi , et al., “Bandwidth Extension is All You Need”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://pixl.cs.prince… [cited by applicant]
Su, Jiaqi , et al., “HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Intern… [cited by applicant]
Su, Jiaqi , et al., “HiFi-GAN-2: Studio-Quality Speech Enhancement via Generative Adversarial Networks Conditioned on Acoustic Features”, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA… [cited by applicant]
Ulyanov, Dmitry , et al., “Instance Normalization: The Missing Ingredient for Fast Stylization”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1607.080… [cited by applicant]
Van Den Oord, Aaron , et al., “Wavenet: A Generative Model for Raw Audio”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1609.03499.pdf>., Sep. 19, 201… [cited by applicant]
Williamson, Donald S, et al., “Speech dereverberation and denoising using complex ratio masks”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) [retrieved Feb. 24, 2022]. Retrieved from… [cited by applicant]
Yamamoto, Ryuichi , et al., “Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022].… [cited by applicant]
You, Jaeseong , et al., “GAN Vocoder: Multi-Resolution Discriminator Is All You Need”, Cornell University arXiv, arXiv.org [retrieved Feb. 24, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.05236.pdf>., … [cited by applicant]