IP Library Granted Patent US 12,730,862
Granted Patent B2
US 12,730,862 · App. 19/468,947 · Granted Sep 8, 2026

Watermarking of AI-generated derivative works

Inventor: Daniel A. Drolet (Charleston, SC)
Assignee: Music IP Holdings (MIH), Inc.
G06F21/106G06F16/632G06F21/1084G06F40/205G06F2221/2137
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,862
App. No.
19/468,947
Granted
Sep 8, 2026
Kind
B2
Abstract

Herein disclosed is a method for generating and managing derivative works using generative artificial intelligence. A request to generate a derivative work based on predetermined content is received, the request comprising a requested theme. The derivative work is created as a function of the predetermined content and the requested theme using generative artificial intelligence. A watermark is embedded into the derivative work by distributing the watermark across multiple frequency bands of the derivative work, the frequency bands being less perceptible to a human ear. The watermark may comprise a unique identifier associated with the derivative work that references rights information stored in a database. Use of the derivative work may be tracked based on the unique identifier. An authorization server may be configured to govern use of the derivative work based on the watermark, including validating user requests for access based on the rights information associated with the derivative work.

Claims (33)

1 . A method comprising:

receiving, at a generative artificial intelligence system, a request to generate a derivative audio work based on predetermined content, the request comprising a requested theme for the derivative audio work;

creating the derivative audio work as a function of the predetermined content and the requested theme by the generative artificial intelligence system; and

after creating the derivative audio work, embedding a watermark into the derivative audio work by distributing the watermark across multiple frequency bands of the derivative audio work and across a plurality of temporal segments of the derivative audio work, wherein distributing the watermark across multiple frequency bands of the derivative audio work comprises modifying frequency components of the multiple frequency bands, the multiple frequency bands being less perceptible to a human ear,

wherein the watermark comprises a payload, wherein presence of the payload in the derivative audio work is determined by the generative artificial intelligence system and indicates that the derivative audio work was generated by the generative artificial intelligence system, and wherein the payload comprises metadata that facilitates tracking of the derivative audio work and management of access to the derivative audio work.

2 . The method of claim 1 , wherein distributing the watermark across multiple frequency bands comprises a spread spectrum technique.

3 . The method of claim 1 , further comprising detecting the watermark in the derivative audio work using a detection algorithm configured to identify the watermark.

4 . The method of claim 3 , wherein the detection algorithm is configured to identify the watermark after at least one of: audio compression, format conversion, or digital editing of the derivative audio work.

5 . The method of claim 1 , wherein the watermark comprises metadata identifying at least one of: a source of the derivative audio work, a content owner identity, a date associated with the derivative audio work, usage rights associated with the derivative audio work, a model of the generative artificial intelligence system, a transformation parameter used to create the derivative audio work, an indicator that the derivative audio work is AI-generated, or provenance data documenting creation of the derivative audio work.

6 . The method of claim 1 , wherein the requested theme is converted to a text embedding in a shared latent space, and wherein creating the derivative audio work is based at least in part on the text embedding.

7 . The method of claim 1 , wherein the generative artificial intelligence system comprises a diffusion model.

8 . The method of claim 1 , wherein the predetermined content comprises audio content.

9 . A method comprising:

receiving, at a generative artificial intelligence system, a request to generate a derivative audio work based on predetermined content, the request comprising a requested theme for the derivative audio work;

creating the derivative audio work as a function of the predetermined content and the requested theme by the generative artificial intelligence system;

after creating the derivative audio work, applying a digital watermark to the derivative audio work by distributing the digital watermark across multiple frequency bands of the derivative audio work and across a plurality of temporal segments of the derivative audio work, wherein distributing the digital watermark across multiple frequency bands of the derivative audio work comprises modifying frequency components of the multiple frequency bands, the multiple frequency bands being less perceptible to a human ear, and wherein the digital watermark comprises a unique identifier associated with the derivative audio work and a payload, wherein presence of the payload in the derivative audio work is determined by the generative artificial intelligence system and indicates that the derivative audio work was generated by the generative artificial intelligence system, and wherein the payload comprises metadata that facilitates tracking of the derivative audio work and management of access to the derivative audio work;

storing rights information associated with the derivative audio work in a database, the rights information being referenced by the unique identifier; and

tracking use of the derivative audio work based on the unique identifier.

10 . The method of claim 9 , wherein the digital watermark comprises an audio watermark embedded in the derivative audio work.

11 . The method of claim 9 , wherein the digital watermark is applied using a dynamic watermarking technique, wherein the dynamic watermarking technique comprises periodically updating the digital watermark.

12 . The method of claim 9 , further comprising encrypting watermark data before applying the digital watermark to the derivative audio work.

13 . The method of claim 9 , further comprising configuring an authorization server to govern use of the derivative audio work based on the digital watermark.

14 . The method of claim 13 , wherein governing use of the derivative audio work further comprises validating user requests for access to the derivative audio work based on the rights information stored in the database.

15 . The method of claim 13 , wherein governing use of the derivative audio work further comprises revoking access to the derivative audio work upon determining that usage rights associated with the derivative audio work have expired.

16 . The method of claim 13 , wherein governing use of the derivative audio work further comprises automatically requesting an automated payment via a smart contract execution triggered based at least in part on use of the derivative audio work detected as a function of the digital watermark.

17 . The method of claim 9 , wherein the digital watermark comprises metadata identifying at least one of: a content owner identity, a date of approval, usage rights, a model of the generative artificial intelligence system, a transformation parameter used to create the derivative audio work, an indicator that the derivative audio work is AI-generated, or provenance data documenting creation of the derivative audio work.

18 . The method of claim 9 , wherein the digital watermark is imperceptible to a human observer and detectable by a detection algorithm.

19 . An article of manufacture comprising:

a non-transitory computer-readable memory storing instructions that, when executed by at least one processor, cause the at least one processor to implement a generative artificial intelligence system and perform operations comprising:

receiving, by the generative artificial intelligence system, a request to generate a derivative audio work based on predetermined content, the request comprising a requested theme for the derivative audio work;

creating the derivative audio work as a function of the predetermined content and the requested theme by the generative artificial intelligence system; and

after creating the derivative audio work, embedding a watermark into the derivative audio work by distributing the watermark across multiple frequency bands of the derivative audio work and across a plurality of temporal segments of the derivative audio work, wherein distributing the watermark across multiple frequency bands of the derivative audio work comprises modifying frequency components of the multiple frequency bands, the multiple frequency bands being less perceptible to a human ear,

wherein the watermark comprises a payload, wherein presence of the payload in the derivative audio work is determined by the generative artificial intelligence system and indicates that the derivative audio work was generated by the generative artificial intelligence system, and wherein the payload comprises metadata that facilitates tracking of the derivative audio work and management of access to the derivative audio work.

Continuity (5)
Continuation 19312216 · Aug 27, 2025
Continuation 19197818 · May 2, 2025
Continuation In Part 18926097 · Oct 24, 2024
Provisional Application 63592741 · Oct 24, 2023
Related Publication 20260195422A1 · Jul 9, 2026
References Cited (152)
US 6700989B1 · Itoh et al. · 2004 [cited by applicant]
US 6810388B1 · Sato · 2004 [cited by applicant]
US 10217181B2 · Butler · 2019 [cited by examiner]
US 12019982B2 · Veyseh et al. · 2024 [cited by applicant]
US 12080046B2 · Saraee et al. · 2024 [cited by applicant]
US 12086857B2 · Kharbanda et al. · 2024 [cited by applicant]
US 12105729B1 · Haq et al. · 2024 [cited by applicant]
US 12106318B1 · Chiang et al. · 2024 [cited by applicant]
US 12106548B1 · Brudalla et al. · 2024 [cited by applicant]
US 12118325B2 · Gray et al. · 2024 [cited by applicant]
US 12118976B1 · Chen et al. · 2024 [cited by applicant]
US 12165655B1 · Sandrew · 2024 [cited by applicant]
US 12204627B2 · Wexler · 2025 [cited by applicant]
US 20040024588A1 · Watson et al. · 2004 [cited by applicant]
US 20060004669A1 · Ito · 2006 [cited by examiner]
US 20060190970A1 · Hellman · 2006 [cited by applicant]
US 20060271494A1 · Ito · 2006 [cited by examiner]
US 20070140318A1 · Hellman · 2007 [cited by applicant]
US 20070266252A1 · Davis et al. · 2007 [cited by applicant]
US 20210233204A1 · Alattar et al. · 2021 [cited by applicant]
US 20220059063A1 · Balassanian et al. · 2022 [cited by applicant]
US 20220092267A1 · Hou et al. · 2022 [cited by applicant]
US 20220134914A1 · Jung · 2022 [cited by applicant]
US 20230095092A1 · Xiao et al. · 2023 [cited by applicant]
US 20230100289A1 · Kare et al. · 2023 [cited by applicant]
US 20230377099A1 · Kreis et al. · 2023 [cited by applicant]
US 20230377214A1 · Kansy et al. · 2023 [cited by applicant]
US 20240005604A1 · Kreis et al. · 2024 [cited by applicant]
US 20240095987A1 · Piramutha et al. · 2024 [cited by applicant]
US 20240152544A1 · Aykut et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240185396A1 · Hatamizadeh et al. · 2024 [cited by applicant]
US 20240202795A1 · Kharbanda et al. · 2024 [cited by applicant]
US 20240253217A1 · Vahdat et al. · 2024 [cited by applicant]
US 20240282079A1 · Saraee et al. · 2024 [cited by applicant]
US 20240289407A1 · Rofouei et al. · 2024 [cited by applicant]
US 20240304177A1 · Wu et al. · 2024 [cited by applicant]
US 20240312087A1 · Agrawal et al. · 2024 [cited by applicant]
US 20240346629A1 · Harikumar et al. · 2024 [cited by applicant]
US 20250131928A1 · Drolet · 2025 [cited by examiner]
US 20250139375A1 · Bright et al. · 2025 [cited by applicant]
AU 2004258523A1 · 2005 [cited by examiner]
AU 2004258523B2 · 2005 [cited by applicant]
CA 2065641A1 · 2006 [cited by applicant]
CA 2605641A1 · 2006 [cited by examiner]
CA 2605646A1 · 2006 [cited by examiner]
CN 1525363A · 2004 [cited by examiner]
EP 1146411B1 · 2005 [cited by applicant]
JP 2004013493A · 2004 [cited by examiner]
JP 2004506947A · 2004 [cited by examiner]
JP 2004193843A · 2004 [cited by examiner]
JP 2006244075A · 2006 [cited by applicant]
JP 3990853B2 · 2007 [cited by examiner]
JP 4353651B2 · 2009 [cited by examiner]
JP 4456185B2 · 2010 [cited by examiner]
KR 100865247B1 · 2008 [cited by applicant]
TR 2024005874 · 2024 [cited by applicant]
TR 2024006991 · 2024 [cited by applicant]
WO WO2006116394A2 · 2006 [cited by examiner]
WO WO2024097380A1 · 2024 [cited by examiner]
WO 2024158853A1 · 2024 [cited by applicant]
WO 2024220450A1 · 2024 [cited by applicant]
WO 2024243183A2 · 2024 [cited by applicant]
Ramponi, Marco, “Recent developments in Generative AI for Audio”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/recent-developments-i n-generative-ai-for-audio/, 34 pages. [cited by applicant]
Weng, Lilian, “What are Diffusion Models?”, GitHub, Jul. 11, 2021, https://liliamweng.github.io/posts/2021-07-11-diffusion-models/#reverse-diffusion-process, 25 pages. [cited by applicant]
O'Connor, Ryan, “Automatic summarization with LLMs in Python”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/automatic-summarization-llms-python/, 12 pages. [cited by applicant]
“Apply LLMs to audio files, Learn how to leverage LLMs for speech using LeMUR”, Assembly AI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/getting-started/apply-llm-to-audio-files, 5 page… [cited by applicant]
“Building In-Video Search”, Netflix Technology Blog, Nov. 6, 2023, 12 pages. [cited by applicant]
Stevens, Ingrid, “Chat with Your Audio Locally: A guide to RAG with Whisper, Ollama, and FAISS”, Medium, Nov. 19, 2023, https://medi um. com/@ingridstevens/chat-with-your-audio-locally-a-gui de-to-rag-with-whisperollama… [cited by applicant]
Anderson, Brian, “Reverse-Time Diffusion Equation Models”, Stochastic Processes and their Applications 12 (1982) 313-326, North-Holland Publishing Company, 14 pages. [cited by applicant]
“Content Moderation”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assembyai.com/docs/audio-intelligence/content-moderation, 11 pages. [cited by applicant]
Muthukumar, “Detecting Voiced, Unvoiced and Silent parts of a speech signal”, Medium, Mar. 19, 2024, https://muthuku37.medium.com/detecting-voiced-unvoiced-and-silent-parts-of-a-speech-signal-?4e6fbf5e 75, 26 pages. [cited by applicant]
“Diffusion Models: A Comprehensive High-Level Understanding”, Research Graph, Medium, May 21, 2024, https://medium.com/@researchgraph/diffusion-model-compreshensive-high-level-understanding-55d6ecad2cba, 22 pages. [cited by applicant]
Larcher, Mario, “Diffusion Transformer Explained”, Towards Data Science, Feb. 28, 2024, https://medium.com/towards-data-sciene/diffusion-transforer-explai ned-e603c4770f7 e, 19 pages. [cited by applicant]
O'Connor, Ryan, “Introduction to Diffusion Models for Machine Learning” AssemblyAI, May 12, 2022, https://www. assemblyai.com/blog/diffusion-models-for-machine-learning-introduction/, 34 pages. [cited by applicant]
Andreas, et al., “DRCap_Zeroshot_Audio-Captioning”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/drcap_zeroshot_aac/README.md, 3 pages. [cited by applicant]
Di Pietro, Mauro, “GenAI with Python: Build Agents from Scratch (Complete Tutorial)”, Towards Data Science, Sep. 29, 2024, https://towardsdatascience.com/genai-with-python-build-agents-from-scratch-com pletetutorial-4fc… [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, Cornell Univ., arXiv:2204.06125v1 [cs.CV], Apr. 13, 2022, 27 pages. [cited by applicant]
Ramirez, et al., “Voice Activity Detection. Fundamentals and Speech Recognition System Robustness.” InTech Open Science Open Minds, 2007, 24 pages. [cited by applicant]
Swimberghe, Niels, “How to integrate spoken audio into LangChain.js using AssemblyAI”, AssemblyAI, Aug. 15, 2023, https://www.assemblyai.com/blog/integrate-audio-langchainjs/, 12 pages. [cited by applicant]
“A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, F5-TTS, retrieved from the internet on Oct. 20, 2024, https:/swivid.github.io.F5-TTS/, 17 pages. [cited by applicant]
“CLIP: Connecting text and images”, OpenAI, Jan. 5, 2021, https://openai.com/index/clip/, 16 pages. [cited by applicant]
“CLIP”, Hugging Face, retrieved from the internet on Oct. 22, 2024, https://huggingface.co/docs/transformers/model_doc/clip, 49 pages. [cited by applicant]
Rustamy, Fahim, Phd., “CLIP Model and The Importance of Multimodal Embeddings”, Towards Data Science, Dec. 11, 2023, https://towardsdatascience.com/clip-model-and-the-importance-of-multimodalembeddings-1c8f6b13bf72, 20 … [cited by applicant]
“Diffusion Models from Scratch”, Hugging Face Diffusion Course, retrieved from the internet on Oct. 22, 2024, https://huggingface. co/learn/diffusion-course/en/unit1 /3, 31 pages. [cited by applicant]
“MC_MusicCaps”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/L-LANCES/SLAM-LLM/blob/main/examples/mc_musiccaps/README. md, 2 pages. [cited by applicant]
Briggs, James, “Quick-fire Guide to Multi-Modal ML With OpenAI's CLIP”, Towards Data Science, Aug. 11, 2022, https://towardsdatascience. com/quick-tire-guide-to-multi-modal-ml-with-openais-cli p. 2dad7 e398ac0, 21 pages. [cited by applicant]
Bouchard, Louis-Francois, “Stable Diffusion for Videos Explained”, Towards AI, Nov. 29, 2023, https://pub.towardsai.net/stable-diffusion-for-videos-explai ned-fawf0b6af3b0, 15 pages. [cited by applicant]
Erdem, Kemal, “Step by Step visual introduction to Diffusion Models”, published Nov. 1, 2023, https://erdem.pl/2023/11/step-by-step-visual-introduction-to-diffusion-models, 15 pages. [cited by applicant]
“Stable Diffusion: Training Your Own Model in 3 Simple Steps”, run:ai, https://www.run.ai.guides/generative-ai/stablediffusion-training, 10 pages. [cited by applicant]
Stevens, Ingrid, “Uncovering Insights in Audio: An Exploration”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/ingridstevens/whisper-audio-transcriber/tree/main, 6 pages. [cited by applicant]
Palucha, Szymon, “Understanding OpenAI's CLIP model”, Medium, Feb. 24, 2024 https://medium.com/@paluchasz/understanding-openais-cl i p-m odel-6b52bade3fa3, 23 pages. [cited by applicant]
Assembly AI, “Summarization”, retrieved from the internet on Oct. 20, 2024, https://assemblyai.com/docs/audiointelligence/summarization, 6 pages. [cited by applicant]
GitHub, “SLAM-MC”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 5 pages. [cited by applicant]
GitHub, “SLAM-LLM”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 4 pages. [cited by applicant]
Huggingface.co Blog, “The Annotated Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/blog/annotated-diffusion, 38 pages. [cited by applicant]
Huggingface.co Blog, “Train a Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/docs/diffusers/tutorials/basic_training, 12 pages. [cited by applicant]
IBM, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 18 pages. [cited by applicant]
Lil Log, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 25 pages. [cited by applicant]
Assembly AI, “Topic Detection”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audiointelligence/topic-detection, 6 pages. [cited by applicant]
Assembly AI, “Key Phrases”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audiointelligence/key-phrases, 5 pages. [cited by applicant]
Assembly All, “Sentiment Analysis”, retrieved from the internet Oct. 20, 2024, https://www.assemblyai.com/docs/audio-intelligence/sentiment-analysis, 4 pages. [cited by applicant]
Atal et al., “A pattern recognition approach to voiced-unvoiced-silence classification with applications to speech recognition”, IEEE Transactions on Acoustics, Speech, and Signal Processing (vol. 24, Issue: 3, Jun. 197… [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, Cornell Univ., arXiv:1312.6114v11 [stat. ML]—, Dec. 10, 2022, pgs. [cited by applicant]
Sohl-Dickstein et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics” Cornell Univ., arXiv:1503.03585v8 [cs.LG], Nov. 18, 2015. 18 pgs. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, Cornell Univ., arXiv:1505.04597v1 [cs.CV], May 18, 2015, 8 pgs. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, Cornell Univ., arXiv:2006.11239v2 [cs.LG], Dec. 16, 2020, 25 pgs. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Cornell Univ., arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, 48 pgs. [cited by applicant]
Dhariwal et al., “Diffusion Models Beat GANs on Image Synthesis”, Cornell Univ., arXiv:2105.05233v4 [cs.LG], Jun. 2021, 44 pgs. [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, Cornell Univ., arXiv:2112.10752v2 [cs.CV], Apr. 13, 2022, 45 pgs. [cited by applicant]
Blattmann et al., “Semi-Parametric Neural Image Synthesis”, Cornell Univ., arXiv:2204.11824v3 [cs.CV], Oct. 24, 2022, 34 pgs. [cited by applicant]
Karras et al.,.. “Elucidating the Design Space of Diffusion-Based Generative Models”, Cornell Univ., arXiv:2206.00364v2 [cs.CV], Oct. 11, 2022, 47 pgs. [cited by applicant]
Graikos et al., “Diffusion models as plug-and-play priors” Cornell Univ., arXiv:2206.09012v3 [cs.LG], Jan. 8, 2023, 22 pgs. [cited by applicant]
Luo, “Understanding Diffusion Models: A Unified Perspective”, Cornell Univ., arXiv:2208.11970v1 [cs.LG], Aug. 25, 2022, 23 pgs. [cited by applicant]
Ruiz et al., “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, Cornell Univ., arXiv:2208.12242v2 [cs.CV], Mar. 15, 2023, 25 pgs. [cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v9 [cs.LG], Oct. 24, 2022, 39 pgs. [cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v13 [cs.LG], Oct. 24, 2022, 39 pgs. [cited by applicant]
Lipman et al., “Flow Matching for Generative Modeling”, Cornell Univ., arXiv:2210.02747v2 [cs.LG], Feb. 8, 2023, 28 pgs. [cited by applicant]
Peebles et al., “Scalable Diffusion Models with Transformers”, Cornell Univ., arXiv:2212.09748v2 [cs. CV], Mar. 2023, 25 pgs. [cited by applicant]
Schneider et al., “Mo0sai: Text-to-Music Generation with Long-Context Latent Diffusion”, Cornell Univ., arXiv:2301.11757v2 [cs.CL], Jan. 30, 2023, 13 pgs. [cited by applicant]
Fei, et al., “Generative Diffusion Prior for Unified Image Restoration and Enhancement”, Cornell Univ., arXiv:2304.01247v1 [cs.CV], Apr. 3, 2023, 46 pages. [cited by applicant]
Ghosal, et al., “Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model”, Cornell Univ., arXiv:2304.13731v2 [eess.AS], May 29, 2023, 15 pages. [cited by applicant]
Liu, et al., AudioLDM: Text-to-Audio Generation with Latent Diffusion Models, Cornell Univ., arXiv:2301.12503v3 [cs.SD], Sep. 9, 2023, 25 pages. [cited by applicant]
Wei, et al., “ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation”, Cornell Univ., arXiv:2302.13848v2 [cs.CV], Aug. 18, 2023, 16 pages. [cited by applicant]
Zhang, et al., “A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI”, Cornell Univ., arXiv:2303.13336v2 [cs.SD], Apr. 2, 2023, 18 pages. [cited by applicant]
Nikkiran, et al., “Step-by-Step Diffusion: An Elementary Tutorial”, Cornell Univ., arXiv:2406.08929v2 [cs.LG], Jun. 23, 2024, 51 pages. [cited by applicant]
Copet, et al., “Simple and Controllable Music Generation”, Cornell Univ., arXiv:2306.05284v3 [cs. SD], Jan. 30, 2024, 17 pages. [cited by applicant]
Li, et al., “JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models”, Cornell Univ., arXiv:2308.04729v1 [cs. SD], Aug. 9, 2023, 12 pages. [cited by applicant]
Yao, et al., “JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation”, Cornell Univ., arXiv:2310.19180v2 [cs.SD], Nov. 3, 2023, 12 pages. [cited by applicant]
Xue, et al., “Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation”, Cornell Univ., arXiv:2401.01044v1 [cs.SD], Jan. 2, 2024, 17 pages. [cited by applicant]
Zheng, et al., “BAT: Learning to Reason about Spatial Sounds with Large Language Models”, Cornell Univ., arXiv:2402.01591v2 [eess.AS], May 25, 2024, 16 pages. [cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v2 [cs.CV], Oct. 20, 2024, 28 pages. [cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v1 [cs.CV], Oct. 16, 2024, 28 pages. [cited by applicant]
Chen, et al., “F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, Cornell Univ., arXiv:2410.06885v2 [eess.AS], Oct. 15, 2024, 18 pages. [cited by applicant]
Chen, et al., “Contrastive Localized Language-Image Pre-Training”, Cornell Univ., arXiv:2410.02746v1 [cs.CV], Oct. 3, 2024, 20 pages. [cited by applicant]
Xin, et al., “DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval”, Cornell Univ., arXiv:2409.10025v2 [cs.SD], Oct. 17, 2024, 5 pages. [cited by applicant]
Chen, et al., “JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning”, Cornell Univ., arXiv:2406.12292v1 [cs.SD], Jun. 28, 2024, 13 pages. [cited by applicant]
Chan, Stanley, H., “Tutorial on Diffusion Models for Imaging and Vision”, Cornell Univ., arXiv:2403.18103v2 [cs.LG], Sep. 6, 2024, 89 pages. [cited by applicant]
Tu, et al., “A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)”, Cornell Univ., arXiv:2402.07410v1 [cs.CV], Feb. 12, 2024, 14 pages. [cited by applicant]
Esser, et al., “Scaling Rectified Flow Transformers for High-Resolution Image Synthesis”, Cornell Univ., arXiv:2403.03206v1 [cs.CV], Mar. 5, 2024, 28 pages. [cited by applicant]
Bansal, et al., “Universal Guidance for Diffusion Models”, Cornell Univ., arXiv:2302.07121v1 [cs.CV], Feb. 14, 2023, 10 pages. [cited by applicant]
Young, Mike, “A Complete Guide to Turning Text into Audio with Audio-LDM”, retrieved from the internet on Oct. 20, 2024, https://notes.airmodels.fyi/audio-ldm-ai-text-to-audio-generation-with-latent-diffusion-models/, 9… [cited by applicant]
“Lecture 14: LPC speech synthesis and autocorrelation-based pitch tracking”, ECE 417, Multimedia Signal Processing, Oct. 10, 2019, 37 pages. [cited by applicant]
Ribeiro, Andre, “Linking Images and Text with OpenAI CLIP”, Tawards Data Scoemce.retrieved from the internet on Oct. 22, 2024, https://towardsdatascience.com/linking-images-and-text-wth-openai-clip-abb4bdf5dbd2, 23 page… [cited by applicant]
Copet, Jade, “MusicGen: Simple and Controllable Music Generation”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/facebookresearch/audiocraft/blob/main/docs/MUSICGEN.md, 10 pages. [cited by applicant]
Agostinelli, et al., “MusicLM: Generating Music From Text”, Cornell Univ., arXiv:2301.11325v1 [cs. SD], Jan. 26, 2023, 15 pages. [cited by applicant]
Microsoft Copilot: Your AI Companion, retrieved from the internet on Oct. 23, 2024, https://copilot. microsoft.com/?FORM=hpcodx&showconv=1, 3 pages. [cited by applicant]
Graf et al., “Features for voice activity detection: a comparative analysis”, EURASIP Journal on Advances in Signal Processing, a SpringerOpen Journal, 2015, 15 pages. [cited by applicant]
“SELD SpatialSoundQA”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/seld_spatialsoundqa/README.md, 4 pages. [cited by applicant]
Radford, et al., “Contrastive Language-Image Pretraining”, CLIP Explained, Papers With Code, retrieved from the internet on Oct. 22, 2024, https://paperswithcode.com/method/clip, 5 pages. [cited by applicant]
Huang, et al., “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models”, Cornell Univ., arXiv:2301.12661v1 [cs. SD], Jan. 30, 2023, 16 pages. [cited by applicant]
Aguilera, Frank Morales, “Open AI CLIP: Bridging Text and Images”, The Deep Hub, https://medium.com/thedeephub/openai-clip-bridging-text-and-images-aaf3cd20299e, Apr. 11, 2024, 15 pages. [cited by applicant]