IP Library Granted Patent US 12,322,402
Granted Patent B2
US 12,322,402 · App. 18/926,097 · Granted Jun 3, 2025

AI-generated music derivative works

Inventor: Daniel A Drolet (Charleston, SC)
G10L19/018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,402
App. No.
18/926,097
Granted
Jun 3, 2025
Kind
B2
Abstract

Herein disclosed is receiving predetermined content, receiving a request to transform the predetermined content into a derivative work, receiving a requested theme for the derivative work, using generative artificial intelligence to create the derivative work generated as a function of the predetermined content and the requested theme, determining if the generated derivative work is approved based on a machine learning model configured to determine a content approval score as a function of content owner preferences, in response to determining the generated derivative work is approved, applying a digital watermark to the approved derivative work, configuring an authorization server to govern use of the approved derivative work based on the digital watermark and providing user access to the authorized derivative work. The requested theme may be determined using a Large Language Model (LLM) and a chatbot interview. The generative artificial intelligence may comprise a diffusion model. The content may comprise music.

Claims (28)

1. A method comprising: receiving predetermined content, using a database server; receiving a request to transform the predetermined content into a derivative work, using a content derivation platform comprising at least one processor; receiving a requested theme for the derivative work, using the at least one processor; creating the derivative work generated as a function of the predetermined content and the requested theme, using generative artificial intelligence and the at least one processor, wherein the creating further comprises transforming the predetermined content into the derivative work comprising audio, video or images, based on an embedding for the requested theme, using the at least one processor; determining if the generated derivative work is approved based on a content approval machine learning model configured to determine a content approval score as a function of at least one content owner preference and the generated derivative work, using the at least one processor; wherein the content approval machine learning model further comprises a neural network configured to determine a score identifying a degree of like or dislike by the content owner for the derivative work, the score determined as a function of a text embedding identifying an item in the derivative work, using the at least one processor; and in response to determining the content approval score is greater than a predetermined minimum: applying a digital watermark to the approved derivative work, using the at least one processor; configuring an authorization server to govern use of the approved derivative work based on the digital watermark, using the at least one processor; and providing access to the approved derivative work.

2. The method of claim 1 , wherein the content comprises music.

3. The method of claim 1 , wherein the content comprises audio.

4. The method of claim 3 , wherein the audio further comprises a human voice sound.

5. The method of claim 4 , wherein the method further comprises detecting the human voice based on a technique comprising autocorrelation, using the at least one processor.

6. The method of claim 5 , wherein the autocorrelation further comprises frequency domain autocorrelation, using the at least one processor.

7. The method of claim 3 , wherein the audio further comprises a musical instrument sound.

8. The method of claim 1 , wherein the requested theme is determined based on an interview with a user, using the at least one processor.

9. The method of claim 8 , wherein the interview with the user is performed by a chatbot, using the at least one processor.

10. The method of claim 8 , wherein the requested theme is determined based on matching a response from the user with a semantically similar predetermined theme identified by a Large Language Model (LLM) as a function of the response from the user, using the at least one processor.

11. The method of claim 10 , wherein the predetermined theme is pre-approved by the content owner.

12. The method of claim 1 , wherein the generative artificial intelligence comprises a diffusion model.

13. The method of claim 12 , wherein the diffusion model is a latent diffusion model.

14. The method of claim 12 , wherein the method further comprises encoding the content to a latent space, using a encoder network.

15. The method of claim 14 , wherein the encoder network further comprises a convolutional neural network (CNN) configured to extract mel-frequency cepstral coefficients (MFCCs) from the content.

16. The method of claim 14 , wherein the encoder network further comprises a CNN configured to extract a spatial or temporal feature from the content.

17. The method of claim 14 , wherein the encoder network further comprises a recurrent neural network (RNN) or a transformer, configured to extract a word embedding from the content.

18. The method of claim 14 , wherein the method further comprises decoding the content from the latent space, using a decoder network.

19. The method of claim 1 , wherein the method further comprises determining a text embedding identifying an item, using a CLIP model.

20. The method of claim 1 , wherein the method further comprises converting the requested theme to a text embedding in a shared latent space, using the at least one processor.

21. The method of claim 1 , wherein applying the digital watermark further comprises embedding the digital watermark in the derivative work.

22. The method of claim 21 , wherein the method further comprises frequency domain embedding of the digital watermark.

23. The method of claim 1 , wherein the method further comprises updating the derivative work with a new digital watermark that is valid for a limited time and providing access to the updated derivative work.

24. The method of claim 1 , wherein governing use of the approved derivative work further comprises configuring a tracking system to determine authenticity of the derivative work, verified as a function of the digital watermark by the tracking system.

25. The method of claim 1 , wherein governing use of the approved derivative work further comprises validating user requests for access to the derivative work authorized as a function of the digital watermark.

26. The method of claim 1 , wherein governing use of the approved derivative work further comprises automatically request an automated payment via a smart contract execution triggered based on use of the derivative work detected as a function of the digital watermark.

27. The method of claim 1 , wherein governing use of the approved derivative work further comprises revoke access to the derivative work upon determining a time-sensitive watermark has expired.

28. An article of manufacture comprising: a memory that is not a transitory propagating signal, wherein the memory further comprises computer readable instructions configured that when executed by at least one processor the computer readable instructions cause the at least one processor to perform operations comprising: receive predetermined content; receive a request to transform the predetermined content into a derivative work; receive a requested theme for the derivative work; create the derivative work generated as a function of the predetermined content and the requested theme, using generative artificial intelligence, wherein the predetermined content is transformed into the derivative work comprising audio, video or images, based on an embedding for the requested theme, using the at least one processor; determine if the generated derivative work is approved based on a content approval machine learning model configured to determine a content approval score as a function of a content owner preference and the generated derivative work; wherein the content approval machine learning model further comprises a neural network configured to determine a score identifying a degree of like or dislike by the content owner for the derivative work, the score determined as a function of a text embedding identifying an item in the derivative work, using the at least one processor; and in response to determining the content approval score is greater than a predetermined minimum: apply a digital watermark to the approved derivative work; configure an authorization server to govern use of the approved derivative work based on the digital watermark; and provide access to the approved derivative work.

Continuity (2)
Provisional Application 63592741 · Oct 24, 2023
Related Publication 20250131928A1 · Apr 24, 2025
References Cited (132)
US 12019982B2 · Pouran Ben Veyseh et al. · 2024 [cited by applicant]
US 12080046B2 · Saraee et al. · 2024 [cited by applicant]
US 12086857B2 · Kharbanda et al. · 2024 [cited by applicant]
US 12106318B1 · Chiang et al. · 2024 [cited by applicant]
US 12106548B1 · Brudalla et al. · 2024 [cited by applicant]
US 12118325B2 · Gray et al. · 2024 [cited by applicant]
US 12118976B1 · Chen et al. · 2024 [cited by applicant]
US 12165655B1 · Sandrew · 2024 [cited by examiner]
US 20060004669A1 · Ito · 2006 [cited by examiner]
US 20060271494A1 · Ito · 2006 [cited by examiner]
US 20220092267A1 · Hou et al. · 2022 [cited by applicant]
US 20230095092A1 · Xiao et al. · 2023 [cited by applicant]
US 20230377099A1 · Kreis et al. · 2023 [cited by applicant]
US 20230377214A1 · Kansy et al. · 2023 [cited by applicant]
US 20240005604A1 · Kreis et al. · 2024 [cited by applicant]
US 20240095987A1 · Piramutha et al. · 2024 [cited by applicant]
US 20240152544A1 · Aykut et al. · 2024 [cited by applicant]
US 20240185396A1 · Hatamizadeh et al. · 2024 [cited by applicant]
US 20240253217A1 · Vahdat et al. · 2024 [cited by applicant]
US 20240282079A1 · Saraee et al. · 2024 [cited by applicant]
US 20240289407A1 · Rofouei et al. · 2024 [cited by applicant]
US 20240304177A1 · Wu et al. · 2024 [cited by applicant]
US 20240312087A1 · Agrawal et al. · 2024 [cited by applicant]
US 20240346629A1 · Harikumar et al. · 2024 [cited by applicant]
AU 2004258523A1 · 2006 [cited by examiner]
CA 2605641A1 · 2006 [cited by examiner]
CA 2605646A1 · 2006 [cited by examiner]
CN 1525363A · 2004 [cited by examiner]
JP 2004013493A · 2004 [cited by examiner]
JP 2004506947A · 2004 [cited by examiner]
JP 2004193843A · 2004 [cited by examiner]
JP 2006244075A · 2006 [cited by examiner]
JP 3990853B2 · 2007 [cited by examiner]
JP 4353651B2 · 2009 [cited by examiner]
JP 4456185B2 · 2010 [cited by examiner]
KR 100865247B1 · 2008 [cited by examiner]
WO WO2024097380A1 · 2024 [cited by examiner]
WO WO2024158853A1 · 2024 [cited by examiner]
WO WO2024220450A1 · 2024 [cited by examiner]
WO WO2024243183A2 · 2024 [cited by examiner]
Ramponi, Marco, “Recent developments in Generative AI for Audio”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/recent-developments-in-generative-ai-for-audio/, 34 pages. [cited by applicant]
Weng, Lilian, “What are Diffusion Models?”, GitHub, Jul. 11, 2021, https://liliamweng.github.io/posts/2021-07-11-diffusion-models/#reverse-diffusion-process, 25 pages. [cited by applicant]
O'Connor, Ryan, “Automatic summarization with LLMs in Python”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/automatic-summarization-Ilms-python/, 12 pages. [cited by applicant]
“Apply LLMs to audio files, Learn how to leverage LLMs for speech using LeMUR”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/getting-started/apply-llm-to-audio-files, 5 pages. [cited by applicant]
“Building In-Video Search”, Netflix Technology Blog, Nov. 6, 2023, 12 pages. [cited by applicant]
Stevens, Ingrid, “Chat with Your Audio Locally: A guide to RAG with Whisper, Ollama, and FAISS”, Medium, Nov. 19, 2023, https://medium.com/@ingridstevens/chat-with-your-audio-locally-a-guide-to-rag-with-whisper-ollama-a… [cited by applicant]
Anderson, Brian, “Reverse-Time Diffusion Equation Models”, Stochastic Processes and their Applications 12 (1982) 313-326, North-Holland Publishing Company, 14 pages. [cited by applicant]
“Content Moderation”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assembyai.com/docs/audio-intelligence/content-moderation, 11 pages. [cited by applicant]
Muthukumar, “Detecting Voiced, Unvoiced and Silent parts of a speech signal”, Medium, Mar. 19, 2024, https://muthuku37.medium.com/detecting-voiced-unvoiced-and-silent-parts-of-a-speech-signal-74e6fbf5e75, 26 pages. [cited by applicant]
“Diffusion Models: A Comprehensive High-Level Understanding”, Research Graph, Medium, May 21, 2024, https://medium.com/@researchgraph/diffusion-model-compreshensive-high-level-understanding-55d6ecad2cba, 22 pages. [cited by applicant]
Larcher, Mario, “Diffusion Transformer Explained”, Towards Data Science, Feb. 28, 2024, https://medium.com/towards-data-sciene/diffusion-transforer-explained-e603c4770f7e, 19 pages. [cited by applicant]
O'Connor, Ryan, “Introduction to Diffusion Models for Machine Learning” AssemblyAI, May 12, 2022, https://www.assemblyai.com/blog/diffusion-models-for-machine-learning-introduction/, 34 pages. [cited by applicant]
Andreas, et al., “DRCap_Zeroshot_Audio-Captioning”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/drcap_zeroshot_aac/README.md, 3 pages. [cited by applicant]
Di Pietro, Mauro, “GenAI with Python: Build Agents from Scratch (Complete Tutorial)”, Towards Data Science, Sep. 29, 2024, https://towardsdatascience.com/genai-with-python-build-agents-from-scratch-complete-tutorial-4fc… [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, Cornell Univ., arXiv:2204.06125v1 [cs.CV], Apr. 13, 2022, 27 pages. [cited by applicant]
Ramírez, et al., “Voice Activity Detection. Fundamentals and Speech Recognition System Robustness.” InTech Open Science Open Minds, 2007, 24 pages. [cited by applicant]
Swimberghe, Niels, “How to integrate spoken audio into LangChain.js using AssemblyAI”, AssemblyAI, Aug. 15, 2023, https://www.assemblyai.com/blog/integrate-audio-langchainjs/, 12 pages. [cited by applicant]
“A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, F5-TTS, retrieved from the internet on Oct. 20, 2024, https:/swivid.github.io.F5-TTS/, 17 pages. [cited by applicant]
“CLIP: Connecting text and images”, OpenAI, Jan. 5, 2021, https://openai.com/index/clip/, 16 pages. [cited by applicant]
“CLIP”, Hugging Face, retrieved from the internet on Oct. 22, 2024, https://huggingface.co/docs/transformers/model_doc/clip, 49 pages. [cited by applicant]
Rustamy, Fahim, Phd., “CLIP Model and The Importance of Multimodal Embeddings”, Towards Data Science, Dec. 11, 2023, https://towardsdatascience.com/clip-model-and-the-importance-of-multimodal-embeddings-1c8f6b13bf72, 20… [cited by applicant]
“Diffusion Models from Scratch”, Hugging Face Diffusion Course, retrieved from the internet on Oct. 22, 2024, https://huggingface.co/learn/diffusion-course/en/unit1/3, 31 pages. [cited by applicant]
“MC_MusicCaps”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/L-LANCES/SLAM-LLM/blob/main/examples/mc_musiccaps/README.md, 2 pages. [cited by applicant]
Briggs, James, “Quick-fire Guide to Multi-Modal ML With OpenAI's CLIP”, Towards Data Science, Aug. 11, 2022, https://towardsdatascience.com/quick-fire-guide-to-multi-modal-ml-with-openais-clip-2dad7e398ac0, 21 pages. [cited by applicant]
Bouchard, Louis-Francois, “Stable Diffusion for Videos Explained”, Towards AI, Nov. 29, 2023, https://pub.towardsai.net/stable-diffusion-for-videos-explained-fawf0b6af3b0, 15 pages. [cited by applicant]
Erdem, Kemal, “Step by Step visual introduction to Diffusion Models”, published Nov. 1, 2023, https://erdem.pl/2023/11/step-by-step-visual-introduction-to-diffusion-models, 15 pages. [cited by applicant]
“Stable Diffusion: Training Your Own Model in 3 Simple Steps”, run:ai, https://www.run.ai.guides/generative-ai/stable-diffusion-training, 10 pages. [cited by applicant]
Stevens, Ingrid, “Uncovering Insights in Audio: An Exploration”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/ingridstevens/whisper-audio-transcriber/tree/main, 6 pages. [cited by applicant]
Palucha, Szymon, “Understanding OpenAI's CLIP model”, Medium, Feb. 24, 2024, https://medium.com/@paluchasz/understanding-openais-clip-model-6b52bade3fa3, 23 pages. [cited by applicant]
Chen et al., JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning, Cornell Univ., arXiv:2406.12292 [cs.SD], Jun. 18, 2024, 13 pages. [cited by applicant]
Github, “SLAM-AAC”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 5 pages. [cited by applicant]
Github, “SLAM-LLM”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 4 pages. [cited by applicant]
Huggingface.co Blog, “The Annotated Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/blog/annotated-diffusion, 38 pages. [cited by applicant]
Huggingface.co Blog, “Train a Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/docs/diffusers/tutorials/basic_training, 12 pages. [cited by applicant]
IBM, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 18 pages. [cited by applicant]
Lil 'Log, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 25 pages. [cited by applicant]
Assembly AI, “Topic Detection”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audio-intelligence/topic-detection, 6 pages. [cited by applicant]
Assembly AI, “Key Phrases”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audio-intelligence/key-phrases, 5 pages. [cited by applicant]
Assembly All, “Sentiment Analysis”, retrieved from the internet Oct. 20, 2024, https://www.assemblyai.com/docs/audio-intelligence/sentiment-analysis, 4 pages. [cited by applicant]
Assembly AI, “Summarization”, retrieved from the internet on Oct. 20, 2024, https://assemblyai.com/docs/audio-intelligence/summarization, 6 pages. [cited by applicant]
Chen et al., “Buildin In-Video Search”, Medium, Nov. 6, 2023, https://netflixtechblog.com/building-in-video-search-936766f0017c, 12 pages. [cited by applicant]
Atal et al., “A pattern recognition approach to voiced-unvoiced-silence classification with applications to speech recognition”, IEEE Transactions on Acoustics, Speech, and Signal Processing (vol. 24, Issue: 3, Jun. 197… [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, Cornell Univ., arXiv:1312.6114v11 [stat.ML] - , Dec. 10, 2022, 14 pgs. [cited by applicant]
Sohl-Dickstein et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics” Cornell Univ., arXiv:1503.03585v8 [cs.LG], Nov. 18, 2015. 18 pgs. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, Cornell Univ., arXiv:1505.04597v1 [cs.CV], May 18, 2015, 8 pgs. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, Cornell Univ., arXiv:2006.11239v2 [cs.LG], Dec. 16, 2020, 25 pgs. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Cornell Univ., arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, 48 pgs. [cited by applicant]
Dhariwal et al., “Diffusion Models Beat GANs on Image Synthesis”, Cornell Univ., arXiv:2105.05233v4 [cs.LG], Jun. 1, 2021, 44 pgs. [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, Cornell Univ., arXiv:2112.10752v2 [cs.CV], Apr. 13, 2022, 45 pgs. [cited by applicant]
Blattmann et al., “Semi-Parametric Neural Image Synthesis”, Cornell Univ., arXiv:2204.11824v3 [cs.CV], Oct. 24, 2022, 34 pgs. [cited by applicant]
Karras et al., “Elucidating the Design Space of Diffusion-Based Generative Models”, Cornell Univ., arXiv:2206.00364v2 [cs.CV], Oct. 11, 2022, 47 pgs. [cited by applicant]
Graikos et al., “Diffusion models as plug-and-play priors” Cornell Univ., arXiv:2206.09012v3 [cs.LG], Jan. 8, 2023, 22 pgs. [cited by applicant]
Luo, “Understanding Diffusion Models: A Unified Perspective”, Cornell Univ., arXiv:2208.11970v1 [cs.LG], Aug. 25, 2022, 23 pgs. [cited by applicant]
Ruiz et al., “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, Cornell Univ., arXiv:2208.12242v2 [cs.CV], Mar. 15, 2023, 25 pgs. [cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v9 [cs.LG], Oct. 24, 2022, 39 pgs. [cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v13 [cs.LG], Oct. 24, 2022, 39 pgs. [cited by applicant]
Lipman et al., “Flow Matching for Generative Modeling”, Cornell Univ., arXiv:2210.02747v2 [cs.LG], Feb. 8, 2023, 28 pgs. [cited by applicant]
Peebles et al., “Scalable Diffusion Models with Transformers”, Cornell Univ., arXiv:2212.09748v2 [cs.CV], Mar. 2, 2023, 25 pgs. [cited by applicant]
Schneider et al., “Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion”, Cornell Univ., arXiv:2301.11757v2 [cs.CL], Jan. 30, 2023, 13 pgs. [cited by applicant]
Fei, et al., “Generative Diffusion Prior for Unified Image Restoration and Enhancement”, Cornell Univ., arXiv:2304.01247v1 [cs.CV], Apr. 3, 2023, 46 pages. [cited by applicant]
Ghosal, et al., “Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model”, Cornell Univ., arXiv:2304.13731v2 [eess.AS], May 29, 2023, 15 pages. [cited by applicant]
Liu, et al., AudioLDM: Text-to-Audio Generation with Latent Diffusion Models, Cornell Univ., arXiv:2301.12503v3 [cs.SD], Sep. 9, 2023, 25 pages. [cited by applicant]
Wei, et al., “ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation”, Cornell Univ., arXiv:2302.13848v2 [cs.CV], Aug. 18, 2023, 16 pages. [cited by applicant]
Zhang, et al., “A Survey on Audio Diffusion Models: Text to Speech Synthesis and Enhancement in Generative AI”, Cornell Univ., arXiv:2303.13336v2 [cs.SD], Apr. 2, 2023, 18 pages. [cited by applicant]
Nikkiran, et al., “Step-by-Step Diffusion: An Elementary Tutorial”, Cornell Univ., arXiv:2406.08929v2 [cs.LG], Jun. 23, 2024, 51 pages. [cited by applicant]
Copet, et al., “Simple and Controllable Music Generation”, Cornell Univ., arXiv:2306.05284v3 [cs.SD], Jan. 30, 2024, 17 pages. [cited by applicant]
Li, et al., “JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models”, Cornell Univ., arXiv:2308.04729v1 [cs.SD], Aug. 9, 2023, 12 pages. [cited by applicant]
Yao, et al., “JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation”, Cornell Univ., arXiv:2310.19180v2 [cs.SD], Nov. 3, 2023, 12 pages. [cited by applicant]
Xue, et al., “Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation”, Cornell Univ., arXiv:2401.01044v1 [cs.SD], Jan. 2, 2024, 17 pages. [cited by applicant]
Zheng, et al., “BAT: Learning to Reason about Spatial Sounds with Large Language Models”, Cornell Univ., arXiv:2402.01591v2 [eess.AS], May 25, 2024, 16 pages. [cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v2 [cs.CV], Oct. 20, 2024, 28 pages. [cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v1 [cs.CV], Oct. 16, 2024, 28 pages. [cited by applicant]
Chen, et al., “F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, Cornell Univ., arXiv:2410.06885v2 [eess.AS], Oct. 15, 2024, 18 pages. [cited by applicant]
Chen, et al., “Contrastive Localized Language-Image Pre-Training”, Cornell Univ., arXiv:2410.02746v1 [cs.CV], Oct. 3, 2024, 20 pages. [cited by applicant]
Xin, et al., “DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval”, Cornell Univ., arXiv:2409.10025v2 [cs.SD], Oct. 17, 2024, 5 pages. [cited by applicant]
Chen, et al., “JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning”, Cornell Univ., arXiv:2406.12292v1 [cs.SD], Jun. 28, 2024, 13 pages. [cited by applicant]
Chan, Stanley, H., “Tutorial on Diffusion Models for Imaging and Vision”, Cornell Univ., arXiv:2403.18103v2 [cs.LG], Sep. 6, 2024, 89 pages. [cited by applicant]
Tu, et al., “A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)”, Cornell Univ., arXiv:2402.07410v1 [cs.CV], Feb. 12, 2024, 14 pages. [cited by applicant]
Esser, et al., “Scaling Rectified Flow Transformers for High-Resolution Image Synthesis”, Cornell Univ., arXiv:2403.03206v1 [cs.CV], Mar. 5, 2024, 28 pages. [cited by applicant]
Bansal, et al., “Universal Guidance for Diffusion Models”, Cornell Univ., arXiv:2302.07121v1 [cs.CV], Feb. 14, 2023, 10 pages. [cited by applicant]
Young, Mike, “A Complete Guide to Turning Text into Audio with Audio-LDM”, retrieved from the internet on Oct. 20, 2024, https://notes.airmodels.fyi/audio-Idm-ai-text-to-audio-generation-with-latent-diffusion-models/, 9… [cited by applicant]
“Lecture 14: LPC speech synthesis and autocorrelation-based pitch tracking”, ECE 417, Multimedia Signal Processing, Oct. 10, 2019, 37 pages. [cited by applicant]
Ribeiro, Andre, “Linking Images and Text with OpenAI CLIP”, Tawards Data Scoemce.retrieved from the internet on Oct. 22, 2024, https://towardsdatascience.com/linking-images-and-text-wth-openai-clip-abb4bdf5dbd2, 23 page… [cited by applicant]
Copet, Jade, “MusicGen:Simple and Controllable Music Generation”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/facebookresearch/audiocraft/blob/main/docs/MUSICGEN.md, 10 pages. [cited by applicant]
Agostinelli, et al., “MusicLM: Generating Music From Text”, Cornell Univ., arXiv:2301.11325v1 [cs.SD], Jan. 26, 2023, 15 pages. [cited by applicant]
Microsoft Copilot: Your Al Companion, retrieved from the internet on Oct. 23, 2024, https://copilot.microsoft.com/?FORM=hpcodx&showconv=1, 3 pages. [cited by applicant]
Graf et al., “Features for voice activity detection: a comparative analysis”, EURASIP Journal on Advances in Signal Processing, a SpringerOpen Journal, 2015, 15 pages. [cited by applicant]
“SELD_SpatialSoundQA”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/seld_spatialsoundqa/README.md, 4 pages. [cited by applicant]
Radford, et al., “Contrastive Language-Image Pretraining”, CLIP Explained, Papers With Code, retrieved from the internet on Oct. 22, 2024, https://paperswithcode.com/method/clip, 5 pages. [cited by applicant]
Huang, et al., “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models”, Cornell Univ., arXiv:2301.12661v1 [cs.SD], Jan. 30, 2023, 16 pages. [cited by applicant]
Aguilera, Frank Morales, “Open AI Clip: Bridging Text and Images”, The Deep Hub, https://medium.com/thedeephub/openai-clip-bridging-text-and-images-aaf3cd20299e, Apr. 11, 2024, 15 pages. [cited by applicant]
Atal, Bishnu S.; and Rabiner, Lawrence R.; “A Pattern Recognition Approach to Voiced-Unvoiced-Silence Classification with Applications to Speech Recognition” IEEE Transactions on Acoustics, Speech, and Signal Processing… [cited by applicant]
Cited By (2)
US 12,574,213 US 12,700,998