IP Library Granted Patent US 12,354,576
Granted Patent B2
US 12,354,576 · App. 18/796,182 · Granted Jul 8, 2025

Artificial intelligence music generation model and method for configuring the same

Inventors: Yijun Wang (Auckland, NZ); Yao Yao (Shenzhen, CN); Peike Li (Sydney, AU); Boyu Chen (Sydney, AU); David McDonald (Auckland, NZ); Nicolas Fourrier (Auckland, NZ); Erin Zink (Phoenix, AZ); Aaron McDonald (Auckland, NZ); Yilun Wang (Auckland, NZ)
Assignee: Futureverse IP Limited
G10H1/0025G10H2210/111G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,576
App. No.
18/796,182
Granted
Jul 8, 2025
Kind
B2
Abstract

The present disclosure provides a method for configuring a learning model for music generation and the corresponding learning model. The method includes training a masked autoencoder with training data comprising a combination of a reconstruction loss over time and frequency domains and a patch-based adversarial objective operating at different resolutions. An omnidirectional latent diffusion model is trained based on music data represented in a latent space to obtain a pretrained diffusion model. The pretrained diffusion model is fine-tuned based on text-guided music generation, bidirectional music in-painting, and unidirectional music continuation. The method enables high-fidelity music generation conditioned on text or music representations while maintaining computational efficiency.

Claims (24)

1. A method for configuring a learning model for music generation, the method comprising:

providing a masked autoencoder which executes on a computing device;

providing an omnidirectional latent diffusion model which executes on a computing device, and which is operatively coupled to the masked autoencoder to process latent embeddings produced by the masked autoencoder;

training the masked autoencoder with training data, the training data including a combination of a reconstruction loss over time and frequency domains, and a patch-based adversarial objective operating at different resolutions, including processing of first training data with the masked autoencoder, applying a first loss function to the results of the processing of the first training data by the masked autoencoder, and adjusting parameters of the masked autoencoder in accordance with the loss function;

configuring a pretrained diffusion model by training the omnidirectional latent diffusion model based on music data represented in a latent space to obtain a pretrained diffusion model, including processing of second training data with the omnidirectional latent diffusion model, applying a second loss function to the results of the processing of the second training data by the omnidirectional latent diffusion model, and adjusting parameters of the omnidirectional latent diffusion model in accordance with the loss function;

fine-tuning the pretrained diffusion model based on text-guided music generation;

fine-tuning the pretrained diffusion model based on bidirectional music in-painting; and

fine-tuning the pretrained diffusion model based on unidirectional music continuation.

2. The method of claim 1 , wherein a data masking percentage of the masked autoencoder is 5 percent.

3. The method of claim 1 , wherein fine-tuning the pretrained diffusion model based on text-guided music generation includes a bidirectional mode and a unidirectional mode, wherein the bidirectional mode allows the latent embeddings to attend to one another during a denoising process, and wherein the unidirectional mode restricts the latent embeddings to attend solely to previous time counterparts thereof.

4. The method of claim 1 , wherein fine-tuning the pretrained diffusion model based on bidirectional music in-painting comprises simulating a music inpainting process by randomly generating audio masks and applying the audio mask to obtain corresponding masked audio, wherein the masked audio serves as conditional in-context learning inputs into the diffusion model in the process of fine-tuning the pretrained diffusion model.

5. The method of claim 1 , wherein fine-tuning the pretrained diffusion model based on unidirectional music continuation comprises simulating a music continuation process through the random generation of exclusive right-only masks.

6. The method of claim 1 , wherein the omnidirectional latent diffusion model includes at least one convolutional block and at least one transformer block.

7. The method of claim 6 , wherein the at least one convolutional block includes causal padding in a unidirectional mode to restrict the latent embeddings to attend solely to previous time counterparts thereof.

8. A system for music generation, comprising:

a masked autoencoder, executed on a computing device and trained with training data including a combination of a reconstruction loss over time and frequency domains, and a patch-based adversarial objective operating at different resolutions, wherein the training of the masked autoencoder includes processing of first training data with the masked autoencoder, applying a first loss function to the results to the processing of the first training data by the masked autoencoder, and adjusting parameters of the masked autoencoder in accordance with the loss function;

a pretrained omnidirectional latent diffusion model operatively coupled to the masked autoencoder to process latent embeddings produced by the masked autoencoder and which is trained based on music data represented in a latent space to obtain a pretrained diffusion model, wherein the training of the pretrained omnidirectional latent diffusion model includes processing of second training data with an omnidirectional latent diffusion model, applying a second loss function to the results of the processing of the second training data by the omnidirectional latent diffusion model, and adjusting parameters of the omnidirectional latent diffusion model in accordance with the loss function; and

wherein the pretrained omnidirectional latent diffusion model is fine-tuned based on text-guided music generation, bidirectional music in-painting, and unidirectional music continuation.

9. The system of claim 8 , wherein a masking percentage of the masked autoencoder is 5 percent.

10. The system of claim 8 , wherein fine-tuning the pretrained diffusion model based on text-guided music generation includes a bidirectional mode and a unidirectional mode, wherein the bidirectional mode allows the latent embeddings to attend to one another during a denoising process, and wherein the unidirectional mode restricts the latent embeddings to attend solely to previous time counterparts thereof.

11. The system of claim 8 , wherein fine-tuning the pretrained diffusion model based on bidirectional music in-painting comprises simulating a music inpainting process by randomly generating audio masks and applying the audio masks to obtain corresponding masked audio into the diffusion model in the process of fine-tuning the pretrained diffusion model.

12. The system of claim 11 , wherein the masked audio serves as conditional in-context learning inputs.

13. The system of claim 8 , wherein fine-tuning the pretrained diffusion model based on unidirectional music continuation comprises simulating a music continuation process through random generation of exclusive right-only masks.

14. The system of claim 8 , wherein the pretrained omnidirectional latent diffusion model includes at least one convolutional block and at least one transformer block, and wherein the at least one convolutional block includes causal padding in a unidirectional mode to restrict latent embeddings to attend solely to their previous time counterparts thereof.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2025
From: WANG, YILUN
To: FUTUREVERSE IP LIMITED
Reel/Frame 071042/0270 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2024
From: WANG, YIJUN; YAO, YAO; LI, PEIKE; CHEN, BOYU; MCDONALD, DAVID; FOURRIER, NICOLAS; ZINK, ERIN; MCDONALD, AARON
To: FUTUREVERSE IP LIMITED
Reel/Frame 068333/0209 →
Continuity (2)
Provisional Application 63531693 · Aug 9, 2023
Related Publication 20250054473A1 · Feb 13, 2025
References Cited (87)
US 11080591B2 · van den Oord · 2021 [cited by examiner]
US 11164109B2 · Browne et al. · 2021 [cited by applicant]
US 11429762B2 · Mallya Kasaragod et al. · 2022 [cited by applicant]
US 11710027B2 · Zhu et al. · 2023 [cited by applicant]
US 11836640B2 · Ji et al. · 2023 [cited by applicant]
US 11853724B2 · Hunter · 2023 [cited by applicant]
US 11868896B2 · Brown et al. · 2024 [cited by applicant]
US 11915689B1 · Agostinelli · 2024 [cited by examiner]
US 20150023345A1 · Schechner · 2015 [cited by examiner]
US 20180357047A1 · Brown et al. · 2018 [cited by applicant]
US 20200042879A1 · Jansson · 2020 [cited by examiner]
US 20200043518A1 · Jansson · 2020 [cited by examiner]
US 20210125398A1 · Bradley et al. · 2021 [cited by applicant]
US 20210149958A1 · Hunter · 2021 [cited by applicant]
US 20210247954A1 · Balassanian et al. · 2021 [cited by applicant]
US 20210279957A1 · Eder et al. · 2021 [cited by applicant]
US 20210357780A1 · Ji et al. · 2021 [cited by applicant]
US 20220157294A1 · Li et al. · 2022 [cited by applicant]
US 20220188810A1 · Doney · 2022 [cited by applicant]
US 20220391635A1 · Lian et al. · 2022 [cited by applicant]
US 20230009454A1 · Paciello · 2023 [cited by applicant]
US 20230075884A1 · Jakobsson et al. · 2023 [cited by applicant]
US 20230154090A1 · Bradley et al. · 2023 [cited by applicant]
US 20230169080A1 · Iyer et al. · 2023 [cited by applicant]
US 20230222777A1 · Jain et al. · 2023 [cited by applicant]
US 20230281601A9 · Doney · 2023 [cited by applicant]
US 20230282202A1 · Ahmed et al. · 2023 [cited by applicant]
US 20230350936A1 · Alayrac et al. · 2023 [cited by applicant]
US 20230385085A1 · Singh · 2023 [cited by applicant]
US 20240096017A1 · Gao et al. · 2024 [cited by applicant]
US 20240127775A1 · Vechtomova · 2024 [cited by examiner]
US 20240161470A1 · Sminchisescu et al. · 2024 [cited by applicant]
US 20240161761A1 · Islam · 2024 [cited by examiner]
US 20240394511A1 · Thevenin et al. · 2024 [cited by applicant]
US 20240419949A1 · Aykut · 2024 [cited by examiner]
US 20250054473A1 · Wang · 2025 [cited by examiner]
CA 3150262A1 · 2021 [cited by applicant]
CN 116072098A · 2023 [cited by examiner]
CN 116343723A · 2023 [cited by examiner]
EP 3270379A1 · 2018 [cited by examiner]
EP 4383133A1 · 2024 [cited by examiner]
WO 2021046541A1 · 2021 [cited by applicant]
WO 2021097259A1 · 2021 [cited by applicant]
WO 2022160054A1 · 2022 [cited by applicant]
WO 2024118464A1 · 2024 [cited by applicant]
WO 2024129331A1 · 2024 [cited by applicant]
WO WO2024184745A1 · 2024 [cited by examiner]
Agostinelli, A. et al.: “MusicLm: Generating music from text”, arXiv preprint arXiv:2301.11325, 2023. [cited by applicant]
Borsos Z. et al.: “AudioLm: A language modeling approach to audio generation”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023. [cited by applicant]
Chung, H.W. et al.: “Scaling instruction-finetuned language models”, arXiv preprint arXiv:2210.11416, 2022. [cited by applicant]
Copet, J. et al.: “Simple and controllable music generation”, arXiv preprint arXiv:2306.05284, 2023. [cited by applicant]
Creswell, A. et al.: “Generative adversarial networks: An overview”, IEEE signal processing magazine, 35(1):53-65, 2018. [cited by applicant]
Defossez, A. et al.: “High fidelity neural audio compression”, arXiv preprint arXiv:2210.13438, 2022. [cited by applicant]
Dhariwal, P. et al.: “Jukebox: A generative model for music”, arXiv preprint arXiv:2005.00341, 2020. [cited by applicant]
Elizalde, B. et al.: “Clap learning audio concepts from natural language supervision”, In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1-5. IEEE, 2023. [cited by applicant]
Garbacea, C. et al.: “Low bit-rate speech coding with vq-vae and a wavenet decoder”, In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 735-739. IEEE, 2019. [cited by applicant]
Gemmeke, J.F., et al.: “Audio set: An ontology and human-labeled dataset for audio events”, In 20/7 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 776-780. IEEE, 2017. [cited by applicant]
Ghosal, D. et al.: “. Text-to-audio gen.eration using instruction-tuned Ilm and latent diffusion model”, arXiv preprint arXiv:2304.1373, 2023. [cited by applicant]
Hawthorne, C. et al.: “General- purpose, long-context autoregressive modeling with perceiver ar”, In International Conference on Machine Learning, pp. 8535-8558. PMLR, 2022. [cited by applicant]
Hertz, A. et al.: “Prompt-to-prompt image editing with cross attention control”, arXiv preprint arXiv:2208.01626, 2022. [cited by applicant]
Ho, J. et al.: “Classifier-free diffusion guidance”, arXiv preprint, arXiv:2207.12598, 2022. [cited by applicant]
Ho, J. et al.: “Denoising diffusion probabilistic models”, Advances in neural information processing systems, 33:6840-6851, 2020. [cited by applicant]
Huang, Q. et al.: “Noise2music: Text-conditioned music generation with diffusion models”, arXiv preprint arXiv:2302.03917, 2023. [cited by applicant]
Huang, R. et al.: “Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models”, arXiv preprint arXiv:2301.12661, 2023. [cited by applicant]
International Search Report and Written Opinion PCT/IB2024/059047 dated Jan. 14, 2025; 10 pages. [cited by applicant]
Kilgour, K. et al.: “Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms”, In INTERSPEECH, pp. 2350-2354, 2019. [cited by applicant]
Kong, J. et al.: “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis”, Advances in Neural Information Processing Systems, 33: 17022-17033, 2020. [cited by applicant]
Kong, Z. et al.: Diffwave: A versatile diffusion model for audio synthesis. arXiv preprint arXiv:2009.09761, 2020. [cited by applicant]
Kreuk, F. et al.: “Audiogen: Textually guided audio generation”, arXiv preprint arXiv:2209.15352, 2022. [cited by applicant]
Liu, H. et al.: “Audioldm: Text-to-audio generation with latent diffusion models”, arXiv preprint arXiv:2301.12503, 2023. [cited by applicant]
Loshchilov, I. et al.: “ Decoupled weight decay regularization”, arXiv preprint arXiv: 1711.05101, 2017. [cited by applicant]
Marafioti, A. et al.: “A context encoder for audio inpainting”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(12): 2362-2372, 2019. [cited by applicant]
Muhamed, A. et al.: “Symbolic music generation with transformer-gans”, In Proceedings of the AAA! conference on artificial intelligence, vol. 35, pp. 408-417, 2021. [cited by applicant]
Rombach, R. et al.: “High resolution image synthesis with latent diffusion models”, In Proceedings of the IEEEICVF conference on computer vision and pattern recognition, pp. 10684-10695, 2022. [cited by applicant]
Rubenstein, P.K. et al.: “. Audiopalm: A large language model that can speak and listen”, arXiv preprint arXiv:2306.12925, 2023. [cited by applicant]
Saharia, C. et al.: “Photorealistic text-to-image diffusion models with deep language understanding”, Advances in Neural Information Processing Systems, 35:36479-36494, 2022. [cited by applicant]
Schneider, F. et al.: “Mousai Text-to-music generation with long-context latent diffusion”, arXiv preprint arXiv:2301.11757, 2023. [cited by applicant]
Steinwold: “AI + NFTs: What is an iNFT?”, Apr. 6, 2021, Available at: https://andrewsteinwold.substack. com/p/ai-nfts-what-is-an-inft-. [cited by applicant]
Van Den Oord, A. et al.: “Wavenet: A generative model for raw audio”, arXiv preprint arXiv: 1609.03499, 2016. [cited by applicant]
Van Der Oord, A. et al.: “Neural discrete representation learning”, Advances in neural information processing systems, 30, 2017. [cited by applicant]
Van Erven, T. et al.: “Renyi divergence and kullback-leibler divergence”, IEEE Transactions on Information Theory, 60(7):3797-3820, 2014. [cited by applicant]
Vaswani, A. et al.: “Attention is all you need”, Advances in neural information processing systems, 30, 2017. [cited by applicant]
WO Publication 2022160054 Aug. 2022 (Year: 2022). [cited by applicant]
Yang, D. et al.: “Diffsound: Discrete diffusion model for text-to-sound generation”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023. [cited by applicant]
Yu, Y. et al.: “Conditional Istm-gan for melody generation from lyrics”, ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 17(1): 1-20, 2021. [cited by applicant]
Zeghidour, N. et al.: “Soundstream: An end-to-end neural audio codec”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495-507, 2021. [cited by applicant]
Zhu, Hongyuan, et al., “Pop Music Generation: From Melody to Multi-style Arrangement”, ACM Transactions on Knowledge Discovery from Data (TKDD), 14(5): 1-31, 2020. [cited by applicant]
Cited By (1)
US 12,475,869