IP Library Granted Patent US 12,632,516
Granted Patent B2
US 12,632,516 · App. 19/312,216 · Granted May 19, 2026

AI-generated music derivative works

Inventor: Daniel A. Drolet (Charleston, SC)
Assignee: Music IP Holdings, Inc.
G06F21/106G06F16/632G06F21/1084G06F40/205G06F2221/2137
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,516
App. No.
19/312,216
Granted
May 19, 2026
Kind
B2
Abstract

A system for generating derivative works comprises a content derivation platform with at least one processor configured to receive predetermined content and a request to transform it into a derivative work with a requested theme. The system uses generative artificial intelligence to create the derivative work comprising audio, video, or images based on the requested theme. A content approval machine learning model with a neural network determines an approval score based on content owner preferences and text embeddings identifying items in the derivative work. When the score exceeds a predetermined minimum, the system applies a digital watermark, configures an authorization server to govern use based on the digital watermark, and provides access to the approved derivative work. The requested theme may be determined through a chatbot interview using a Large Language Model. The generative artificial intelligence may comprise a diffusion model. The content may comprise music or audio.

Claims (39)

1 . A system comprising:

a content derivation platform comprising at least one processor configured to:

receive predetermined content using a database server;

receive a request to transform the predetermined content into a derivative work;

receive a requested theme for the derivative work;

create the derivative work as a function of the predetermined content and the requested theme, using generative artificial intelligence, wherein to create the derivative work further comprises to transform the predetermined content into the derivative work comprising audio, video or images, based on an embedding for the requested theme;

determine if the generated derivative work is approved based on a content approval machine learning model configured to determine a content approval score as a function of at least one content owner preference and the generated derivative work, wherein the content approval machine learning model further comprises a neural network configured to determine a score identifying a degree of like or dislike by the content owner for the derivative work, the score determined as a function of a text embedding identifying an item in the derivative work; and

in response to the content approval score being greater than a predetermined minimum:

apply a digital watermark to the approved derivative work;

configure an authorization server to govern use of the approved derivative work based on the digital watermark; and

provide access to the approved derivative work.

2 . The system of claim 1 , wherein the content comprises music.

3 . The system of claim 1 , wherein the content comprises audio.

4 . The system of claim 3 , wherein the audio further comprises a human voice sound.

5 . The system of claim 4 , wherein the at least one processor is further configured to detect the human voice based on a technique comprising autocorrelation.

6 . The system of claim 5 , wherein the autocorrelation further comprises frequency domain autocorrelation.

7 . The system of claim 3 , wherein the audio further comprises a musical instrument sound.

8 . The system of claim 1 , wherein the requested theme is determined based on an interview with a user.

9 . The system of claim 8 , wherein the interview with the user is performed by a chatbot.

10 . The system of claim 8 , wherein the requested theme is determined based on a response from the user that is matched with a semantically similar predetermined theme identified by a Large Language Model (LLM) as a function of the response from the user.

11 . The system of claim 10 , wherein the predetermined theme is pre-approved by the content owner.

12 . The system of claim 1 , wherein the generative artificial intelligence comprises a diffusion model.

13 . The system of claim 12 , wherein the diffusion model is a latent diffusion model.

14 . The system of claim 12 , wherein the at least one processor is further configured to encode the content to a latent space, using an encoder network.

15 . The system of claim 14 , wherein the encoder network further comprises a convolutional neural network (CNN) configured to extract mel-frequency cepstral coefficients (MFCCs) from the content.

16 . The system of claim 14 , wherein the encoder network further comprises a CNN configured to extract a spatial or temporal feature from the content.

17 . The system of claim 14 , wherein the encoder network further comprises a recurrent neural network (RNN) or a transformer, configured to extract a word embedding from the content.

18 . The system of claim 14 , wherein the at least one processor is further configured to decode the content from the latent space, using a decoder network.

19 . The system of claim 1 , wherein the at least one processor is further configured to determine a text embedding identifying an item, using a CLIP model.

20 . The system of claim 1 , wherein the at least one processor is further configured to convert the requested theme to a text embedding in a shared latent space.

21 . The system of claim 1 , wherein to apply the digital watermark further comprises to embed the digital watermark in the derivative work.

22 . The system of claim 21 , wherein the at least one processor is further configured to embed the digital watermark using frequency domain embedding.

23 . The system of claim 1 , wherein the at least one processor is further configured to:

update the derivative work with a new digital watermark that is valid for a limited time; and

provide access to the updated derivative work.

24 . The system of claim 1 , wherein to govern use of the approved derivative work further comprises to configure a tracking system to determine authenticity of the derivative work, verified as a function of the digital watermark by the tracking system.

25 . The system of claim 1 , wherein to govern use of the approved derivative work further comprises to validate user requests for access to the derivative work authorized as a function of the digital watermark.

26 . The system of claim 1 , wherein to govern use of the approved derivative work further comprises to automatically request an automated payment via a smart contract execution triggered based on use of the derivative work detected as a function of the digital watermark.

27 . The system of claim 1 , wherein to govern use of the approved derivative work further comprises to revoke access to the derivative work upon determining a time-sensitive watermark has expired.

Continuity (4)
Continuation 19197818 · May 2, 2025
Continuation In Part 18926097 · Oct 24, 2024
Provisional Application 63592741 · Oct 24, 2023
Related Publication 20250390559A1 · Dec 25, 2025
References Cited (150)
US 6700989B1 · Itoh et al. · 2004 [cited by applicant]
US 6810388B1 · Sato · 2004 [cited by applicant]
US 12019982B2 · Veyseh et al. · 2024 [cited by applicant]
US 12080046B2 · Saraee et al. · 2024 [cited by applicant]
US 12086857B2 · Kharbanda et al. · 2024 [cited by applicant]
US 12105729B1 · Haq et al. · 2024 [cited by applicant]
US 12106318B1 · Chiang et al. · 2024 [cited by applicant]
US 12106548B1 · Brudalla et al. · 2024 [cited by applicant]
US 12118325B2 · Gray et al. · 2024 [cited by applicant]
US 12118976B1 · Chen et al. · 2024 [cited by applicant]
US 12165655B1 · Sandrew · 2024 [cited by examiner]
US 12204627B2 · Wexler · 2025 [cited by examiner]
US 20040024588A1 · Watson et al. · 2004 [cited by applicant]
US 20060004669A1 · Ito · 2006 [cited by examiner]
US 20060190970A1 · Hellman · 2006 [cited by applicant]
US 20060271494A1 · Ito · 2006 [cited by examiner]
US 20070140318A1 · Hellman · 2007 [cited by applicant]
US 20070266252A1 · Davis et al. · 2007 [cited by applicant]
US 20210233204A1 · Alattar et al. · 2021 [cited by applicant]
US 20220059063A1 · Balassanian et al. · 2022 [cited by applicant]
US 20220092267A1 · Hou et al. · 2022 [cited by applicant]
US 20220134914A1 · Jung · 2022 [cited by applicant]
US 20230095092A1 · Xiao et al. · 2023 [cited by applicant]
US 20230100289A1 · Kare et al. · 2023 [cited by applicant]
US 20230377099A1 · Kreis et al. · 2023 [cited by applicant]
US 20230377214A1 · Kansy et al. · 2023 [cited by applicant]
US 20240005604A1 · Kreis et al. · 2024 [cited by applicant]
US 20240095987A1 · Piramutha et al. · 2024 [cited by applicant]
US 20240152544A1 · Aykut et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240185396A1 · Hatamizadeh et al. · 2024 [cited by applicant]
US 20240202795A1 · Kharbanda et al. · 2024 [cited by applicant]
US 20240253217A1 · Vahdat et al. · 2024 [cited by applicant]
US 20240282079A1 · Saraee et al. · 2024 [cited by applicant]
US 20240289407A1 · Rofouei et al. · 2024 [cited by applicant]
US 20240304177A1 · Wu et al. · 2024 [cited by applicant]
US 20240312087A1 · Agrawal et al. · 2024 [cited by applicant]
US 20240346629A1 · Harikumar et al. · 2024 [cited by applicant]
US 20250131928A1 · Drolet · 2025 [cited by applicant]
US 20250139375A1 · Bright et al. · 2025 [cited by applicant]
AU 2004258523A1 · 2005 [cited by examiner]
CA 2065641A1 · 2006 [cited by applicant]
CA 2605641A1 · 2006 [cited by examiner]
CA 2605646A1 · 2006 [cited by examiner]
CN 1525363A · 2004 [cited by examiner]
EP 1146411B1 · 2005 [cited by applicant]
JP 2004013493A · 2004 [cited by examiner]
JP 2004506947A · 2004 [cited by examiner]
JP 2004193843A · 2004 [cited by examiner]
JP 2006244075A · 2006 [cited by applicant]
JP 3990853B2 · 2007 [cited by examiner]
JP 4353651B2 · 2009 [cited by examiner]
JP 4456185B2 · 2010 [cited by examiner]
KR 100865247B1 · 2008 [cited by examiner]
TR 2024005874 · 2024 [cited by applicant]
TR 2024006991 · 2024 [cited by applicant]
WO WO2024097380A1 · 2024 [cited by examiner]
WO 2024158853A1 · 2024 [cited by applicant]
WO 2024220450A1 · 2024 [cited by applicant]
WO 2024243183A2 · 2024 [cited by applicant]
Ramponi, Marco, “Recent developments in Generative AI for Audio”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/recent-developments-i n-generative-ai-for-audio/, 34 pages. [cited by applicant]
Weng, Lilian, “What are Diffusion Models?”, GitHub, Jul. 11, 2021, https://liliamweng.github.io/posts/2021-07-11-diffusion-models/#reverse-diffusion-process, 25 pages. [cited by applicant]
O'Connor, Ryan, “Automatic summarization with LLMs in Python”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/automatic-summarization-llms-python/, 12 pages. [cited by applicant]
“Apply LLMs to audio files, Learn how to leverage LLMs for speech using LeMUR”, Assembly AI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/getting-started/apply-llm-to-audio-files, 5 page… [cited by applicant]
Chen et al., “Buildin In-Video Search”, Medium, Nov. 6, 2023, https://netflixtechblog.com/building-in-videosearch-936766f0017c, 12 pages. [cited by applicant]
Stevens, Ingrid, “Chat with Your Audio Locally: A guide to RAG with Whisper, Ollama, and FAISS”, Medium, Nov. 19, 2023, https://medi um. com/@ingridstevens/chat-with-your-audio-locally-a-gui de-to-rag-with-whisperollama… [cited by applicant]
Anderson, Brian, “Reverse-Time Diffusion Equation Models”, Stochastic Processes and their Applications 12 (1982) 313-326, North-Holland Publishing Company, 14 pages. [cited by applicant]
“Content Moderation”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assembyai.com/docs/audio-intelligence/content-moderation, 11 pages. [cited by applicant]
Muthukumar, “Detecting Voiced, Unvoiced and Silent parts of a speech signal”, Medium, Mar. 19, 2024, https://muthuku37.medium.com/detecting-voiced-unvoiced-and-silent-parts-of-a-speech-signal-?4e6fbf5e 75, 26 pages. [cited by applicant]
“Diffusion Models: A Comprehensive High-Level Understanding”, Research Graph, Medium, May 21, 2024, https://medium.com/@researchgraph/diffusion-model-compreshensive-high-level-understanding-55d6ecad2cba, 22 pages. [cited by applicant]
Larcher, Mario, “Diffusion Transformer Explained”, Towards Data Science, Feb. 28, 2024, https://medium.com/towards-data-sciene/diffusion-transforer-explained-e603c4770f7 e, 19 pages. [cited by applicant]
O'Connor, Ryan, “Introduction to Diffusion Models for Machine Learning” AssemblyAI, May 12, 2022, https://www. assemblyai.com/blog/diffusion-models-for-machine-learning-introduction/, 34 pages. [cited by applicant]
Andreas, et al., “DRCap_Zeroshot_Audio-Captioning”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/drcap_zeroshot_aac/README.md, 3 pages. [cited by applicant]
Di Pietro, Mauro, “GenAI with Python: Build Agents from Scratch (Complete Tutorial)”, Towards Data Science, Sep. 29, 2024, https://towardsdatascience.com/genai-with-python-build-agents-from-scratch-com pletetutorial-4fc… [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, Cornell Univ., arXiv:2204.06125v1 [cs.CV], Apr. 13, 2022, 27 pages. [cited by applicant]
Ramirez, et al., “Voice Activity Detection. Fundamentals and Speech Recognition System Robustness.” InTech Open Science Open Minds, 2007, 24 pages. [cited by applicant]
Swimberghe, Niels, “How to integrate spoken audio into LangChain.js using AssemblyAI”, AssemblyAI, Aug. 15, 2023, https://www.assemblyai.com/blog/integrate-audio-langchainjs/, 12 pages. [cited by applicant]
“A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, F5-TTS, retrieved from the internet on Oct. 20, 2024, https:/swivid.github.io.F5-TTS/, 17 pages. [cited by applicant]
“CLIP: Connecting text and images”, OpenAI, Jan. 5, 2021, https://openai.com/index/clip/, 16 pages. [cited by applicant]
“CLIP”, Hugging Face, retrieved from the internet on Oct. 22, 2024, https://huggingface.co/docs/transformers/model_doc/clip, 49 pages. [cited by applicant]
Rustamy, Fahim, Phd., “CLIP Model and The Importance of Multimodal Embeddings”, Towards Data Science, Dec. 11, 2023, https://towardsdatascience.com/clip-model-and-the-importance-of-multimodalembeddings-1c8f6b13bf72, 20 … [cited by applicant]
“Diffusion Models from Scratch”, Hugging Face Diffusion Course, retrieved from the internet on Oct. 22, 2024, https://huggingface. co/learn/diffusion-course/en/unit1 /3, 31 pages. [cited by applicant]
“MC_MusicCaps”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/L-LANCES/SLAM-LLM/blob/main/examples/mc_musiccaps/README. md, 2 pages. [cited by applicant]
Briggs, James, “Quick-fire Guide to Multi-Modal ML With OpenAI's CLIP”, Towards Data Science, Aug. 11, 2022, https://towardsdatascience. com/quick-tire-guide-to-multi-modal-ml-with-openais-cli p 2dad7 e398ac0, 21 pages. [cited by applicant]
Bouchard, Louis-Francois, “Stable Diffusion for Videos Explained”, Towards AI, Nov. 29, 2023, https://pub.towardsai.net/stable-diffusion-for-videos-explained-fawf0b6af3b0, 15 pages. [cited by applicant]
Erdem, Kemal, “Step by Step visual introduction to Diffusion Models”, published Nov. 1, 2023, https://erdem.pl/2023/11/step-by-step-visual-introduction-to-diffusion-models, 15 pages. [cited by applicant]
“Stable Diffusion: Training Your Own Model in 3 Simple Steps”, run:ai, https://www.run.ai.guides/generative-ai/stablediffusion-training, 10 pages. [cited by applicant]
Stevens, Ingrid, “Uncovering Insights in Audio: An Exploration”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/ingridstevens/whisper-audio-transcriber/tree/main, 6 pages. [cited by applicant]
Palucha, Szymon, “Understanding OpenAI's CLIP model”, Medium, Feb. 24, 2024 https://medium.com/@paluchasz/understanding-openais-cl i p-m odel-6b52bade3fa3, 23 pages. [cited by applicant]
Chen et al., JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning, Cornell Univ., arXiv:2406.12292 [cs.SD], Jun. 18, 2024, 13 pages. [cited by applicant]
Github, “SLAM-MC”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 5 pages. [cited by applicant]
Github, “SLAM-LLM”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 4 pages. [cited by applicant]
Huggingface.co Blog, “The Annotated Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/blog/annotated-diffusion, 38 pages. [cited by applicant]
Huggingface.co Blog, “Train a Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/docs/ diffusers/tutorials/basic_training, 12 pages. [cited by applicant]
IBM, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 18 pages. [cited by applicant]
Lil 'Log, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 25 pages. [cited by applicant]
Assembly AI, “Topic Detection”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audiointelligence/topic-detection, 6 pages. [cited by applicant]
Assembly AI, “Key Phrases”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audiointelligence/key-phrases, 5 pages. [cited by applicant]
Assembly All, “Sentiment Analysis”, retrieved from the internet Oct. 20, 2024, https://www.assemblyai.com/docs/audio-intelligence/sentiment-analysis, 4 pages. [cited by applicant]
Assembly AI, “Summarization”, retrieved from the internet on Oct. 20, 2024, https://assemblyai.com/docs/audiointelligence/summarization, 6 pages. [cited by applicant]
Atal et al., “A pattern recognition approach to voiced-unvoiced-silence classification with applications to speech recognition”, IEEE Transactions on Acoustics, Speech, and Signal Processing (vol. 24, Issue: 3, Jun. 197… [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, Cornell Univ., arXiv:1312.6114v11 [stat. ML] -, Dec. 10, 2022, pgs. [cited by applicant]
Sohl-Dickstein et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics” Cornell Univ., arXiv:1503.03585v8 [cs.LG], Nov. 18, 2015. 18 pgs. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, Cornell Univ., arXiv:1505.04597v1 [cs.CV], May 18, 2015, 8 pgs. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, Cornell Univ., arXiv:2006.11239v2 [cs.LG], Dec. 16, 2020, 25 pgs. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Cornell Univ., arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, 48 pgs. [cited by applicant]
Dhariwal et al., “Diffusion Models Beat GANs on Image Synthesis”, Cornell Univ., arXiv:2105.05233v4 [cs.LG], Jun. 2021, 44 pgs. [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, Cornell Univ., arXiv:2112.10752v2 [cs.CV], Apr. 13, 2022, 45 pgs. [cited by applicant]
Blattmann et al., “Semi-Parametric Neural Image Synthesis”, Cornell Univ., arXiv:2204.11824v3 [cs.CV], Oct. 24, 2022, 34 pgs. [cited by applicant]
Karras et al., . . . “Elucidating the Design Space of Diffusion-Based Generative Models”, Cornell Univ., arXiv:2206.00364v2 [cs.CV], Oct. 11, 2022, 47 pgs. [cited by applicant]
Graikos et al., “Diffusion models as plug-and-play priors” Cornell Univ., arXiv:2206.09012v3 [cs.LG], Jan. 8, 2023, 22 pgs. [cited by applicant]
Luo, “Understanding Diffusion Models: A Unified Perspective”, Cornell Univ., arXiv:2208.11970v1 [cs.LG], Aug. 25, 2022, 23 pgs. [cited by applicant]
Ruiz et al., “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, Cornell Univ., arXiv:2208.12242v2 [cs.CV], Mar. 15, 2023, 25 pgs. [cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v9 [cs.LG], Oct. 24, 2022, 39 pgs. [cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v13 [cs.LG], Oct. 24, 2022, 39 pgs. [cited by applicant]
Lipman et al., “Flow Matching for Generative Modeling”, Cornell Univ., arXiv:2210.02747v2 [cs.LG], Feb. 8, 2023, 28 pgs. [cited by applicant]
Peebles et al., “Scalable Diffusion Models with Transformers”, Cornell Univ., arXiv:2212.09748v2 [cs.CV], Mar. 2023, 25 pgs. [cited by applicant]
Schneider et al., “Mo0sai: Text-to-Music Generation with Long-Context Latent Diffusion”, Cornell Univ., arXiv:2301.11757v2 [cs.CL], Jan. 30, 2023, 13 pgs. [cited by applicant]
Fei, et al., “Generative Diffusion Prior for Unified Image Restoration and Enhancement”, Cornell Univ., arXiv:2304.01247v1 [cs.CV], Apr. 3, 2023, 46 pages. [cited by applicant]
Ghosal, et al., “Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model”, Cornell Univ., arXiv:2304.13731v2 [eess.AS], May 29, 2023, 15 pages. [cited by applicant]
Liu, et al., AudioLDM: Text-to-Audio Generation with Latent Diffusion Models, Cornell Univ., arXiv:2301.12503v3 [cs.SD], Sep. 9, 2023, 25 pages. [cited by applicant]
Wei, et al., “ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation”, Cornell Univ., arXiv:2302.13848v2 [cs.CV], Aug. 18, 2023, 16 pages. [cited by applicant]
Zhang, et al., “A Survey on Audio Diffusion Models: Text to Speech Synthesis and Enhancement in Generative AI”, Cornell Univ., arXiv:2303.13336v2 [cs.SD], Apr. 2, 2023, 18 pages. [cited by applicant]
Nikkiran, et al., “Step-by-Step Diffusion: An Elementary Tutorial”, Cornell Univ., arXiv:2406.08929v2 [cs.LG], Jun. 23, 2024, 51 pages. [cited by applicant]
Copet, et al., “Simple and Controllable Music Generation”, Cornell Univ., arXiv:2306.05284v3 [cs.SD], Jan. 30, 2024, 17 pages. [cited by applicant]
Li, et al., “JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models”, Cornell Univ., arXiv:2308.04729v1 [cs.SD], Aug. 9, 2023, 12 pages. [cited by applicant]
Yao, et al., “JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation”, Cornell Univ., arXiv:2310.19180v2 [cs.SD], Nov. 3, 2023, 12 pages. [cited by applicant]
Xue, et al., “Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation”, Cornell Univ., arXiv:2401.01044v1 [cs.SD], Jan. 2, 2024, 17 pages. [cited by applicant]
Zheng, et al., “BAT: Learning to Reason about Spatial Sounds with Large Language Models”, Cornell Univ., arXiv:2402.01591v2 [eess.AS], May 25, 2024, 16 pages. [cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v2 [cs.CV], Oct. 20, 2024, 28 pages. [cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v1 [cs.CV], Oct. 16, 2024, 28 pages. [cited by applicant]
Chen, et al., “F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, Cornell Univ., arXiv:2410.06885v2 [eess.AS], Oct. 15, 2024, 18 pages. [cited by applicant]
Chen, et al., “Contrastive Localized Language-Image Pre-Training”, Cornell Univ., arXiv:2410.02746v1 [cs.CV], Oct. 3, 2024, 20 pages. [cited by applicant]
Xin, et al., “DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval”, Cornell Univ., arXiv:2409.10025v2 [cs.SD], Oct. 17, 2024, 5 pages. [cited by applicant]
Chen, et al., “JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning”, Cornell Univ., arXiv:2406.12292v1 [cs.SD], Jun. 28, 2024, 13 pages. [cited by applicant]
Chan, Stanley, H., “Tutorial on Diffusion Models for Imaging and Vision”, Cornell Univ., arXiv:2403.18103v2 [cs.LG], Sep. 6, 2024, 89 pages. [cited by applicant]
Tu, et al., “A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)”, Cornell Univ., arXiv:2402.07410v1 [cs.CV], Feb. 12, 2024, 14 pages. [cited by applicant]
Esser, et al., “Scaling Rectified Flow Transformers for High-Resolution Image Synthesis”, Cornell Univ., arXiv:2403.03206v1 [cs.CV], Mar. 5, 2024, 28 pages. [cited by applicant]
Bansal, et al., “Universal Guidance for Diffusion Models”, Cornell Univ., arXiv:2302.07121v1 [cs.CV], Feb. 14, 2023, 10 pages. [cited by applicant]
Young, Mike, “A Complete Guide to Turning Text into Audio with Audio-LDM”, retrieved from the internet on Oct. 20, 2024, https://notes.airmodels.fyi/audio-ldm-ai-text-to-audio-generation-with-latent-diffusion-models/, 9… [cited by applicant]
“Lecture 14: LPC speech synthesis and autocorrelation-based pitch tracking”, ECE 417, Multimedia Signal Processing, Oct. 10, 2019, 37 pages. [cited by applicant]
Ribeiro, Andre, “Linking Images and Text with OpenAI CLIP”, Tawards Data Scoemce.retrieved from the internet on Oct. 22, 2024, https://towardsdatascience.com/linking-images-and-text-wth-openai-clip-abb4bdf5dbd2, 23 page… [cited by applicant]
Copet, Jade, “MusicGen:Simple and Controllable Music Generation”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/facebookresearch/audiocraft/blob/main/docs/MUSICGEN.md, 10 pages. [cited by applicant]
Agostinelli, et al., “MusicLM: Generating Music From Text”, Cornell Univ., arXiv:2301.11325v1 [cs. SD], Jan. 26, 2023, 15 pages. [cited by applicant]
Microsoft Copilot: Your AI Companion, retrieved from the internet on Oct. 23, 2024, https://copilot. microsoft.com/?FORM=hpcodx&showconv=1, 3 pages. [cited by applicant]
Graf et al., “Features for voice activity detection: a comparative analysis”, EURASIP Journal on Advances in Signal Processing, a SpringerOpen Journal, 2015, 15 pages. [cited by applicant]
“SELD SpatialSoundQA”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/seld_spatialsoundqa/README.md, 4 pages. [cited by applicant]
Radford, et al., “Contrastive Language-Image Pretraining”, CLIP Explained, Papers With Code, retrieved from the internet on Oct. 22, 2024, https://paperswithcode.com/method/clip, 5 pages. [cited by applicant]
Huang, et al., “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models”, Cornell Univ., arXiv:2301.12661v1 [cs. SD], Jan. 30, 2023, 16 pages. [cited by applicant]
Aguilera, Frank Morales, “Open AI CLIP: Bridging Text and Images”, The Deep Hub, https://medium.com/thedeephub/openai-clip-bridging-text-and-images-aaf3cd20299e, Apr. 11, 2024, 15 pages. [cited by applicant]