US 6700989B1
· Itoh et al.
· 2004
[cited by applicant]
US 6810388B1
· Sato
· 2004
[cited by applicant]
US 20060271494A1
· Ito
· 2006
[cited by examiner]
US 20220134914A1
· Jung
· 2022
[cited by applicant]
US 20240185396A1
· Hatamizadeh et al.
· 2024
[cited by applicant]
US 20240253217A1
· Vahdat et al.
· 2024
[cited by applicant]
US 20240289407A1
· Rofouei et al.
· 2024
[cited by applicant]
US 20240312087A1
· Agrawal et al.
· 2024
[cited by applicant]
US 20250139375A1
· Bright et al.
· 2025
[cited by applicant]
AU 2004258523A1
· 2005
[cited by examiner]
CA 2065641A1
· 2006
[cited by applicant]
CA 2605641A1
· 2006
[cited by examiner]
CA 2605646A1
· 2006
[cited by examiner]
CN 1525363A
· 2004
[cited by examiner]
EP 1146411B1
· 2005
[cited by applicant]
JP 2004013493A
· 2004
[cited by examiner]
JP 2004506947A
· 2004
[cited by examiner]
JP 2004193843A
· 2004
[cited by examiner]
JP 2006244075A
· 2006
[cited by applicant]
JP 3990853B2
· 2007
[cited by examiner]
JP 4353651B2
· 2009
[cited by examiner]
JP 4456185B2
· 2010
[cited by examiner]
KR 100865247B1
· 2008
[cited by examiner]
TR 2024005874
· 2024
[cited by applicant]
TR 2024006991
· 2024
[cited by applicant]
WO WO2024097380A1
· 2024
[cited by examiner]
WO 2024158853A1
· 2024
[cited by applicant]
WO 2024220450A1
· 2024
[cited by applicant]
WO 2024243183A2
· 2024
[cited by applicant]
Ramponi, Marco, “Recent developments in Generative AI for Audio”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/recent-developments-i n-generative-ai-for-audio/, 34 pages.
[cited by applicant]
Weng, Lilian, “What are Diffusion Models?”, GitHub, Jul. 11, 2021, https://liliamweng.github.io/posts/2021-07-11-diffusion-models/#reverse-diffusion-process, 25 pages.
[cited by applicant]
O'Connor, Ryan, “Automatic summarization with LLMs in Python”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/blog/automatic-summarization-llms-python/, 12 pages.
[cited by applicant]
“Apply LLMs to audio files, Learn how to leverage LLMs for speech using LeMUR”, Assembly AI, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/getting-started/apply-llm-to-audio-files, 5 page…
[cited by applicant]
Chen et al., “Buildin In-Video Search”, Medium, Nov. 6, 2023, https://netflixtechblog.com/building-in-videosearch-936766f0017c, 12 pages.
[cited by applicant]
Stevens, Ingrid, “Chat with Your Audio Locally: A guide to RAG with Whisper, Ollama, and FAISS”, Medium, Nov. 19, 2023, https://medi um. com/@ingridstevens/chat-with-your-audio-locally-a-gui de-to-rag-with-whisperollama…
[cited by applicant]
Anderson, Brian, “Reverse-Time Diffusion Equation Models”, Stochastic Processes and their Applications 12 (1982) 313-326, North-Holland Publishing Company, 14 pages.
[cited by applicant]
“Content Moderation”, AssemblyAI, retrieved from the internet on Oct. 20, 2024, https://www.assembyai.com/docs/audio-intelligence/content-moderation, 11 pages.
[cited by applicant]
Muthukumar, “Detecting Voiced, Unvoiced and Silent parts of a speech signal”, Medium, Mar. 19, 2024, https://muthuku37.medium.com/detecting-voiced-unvoiced-and-silent-parts-of-a-speech-signal-?4e6fbf5e 75, 26 pages.
[cited by applicant]
“Diffusion Models: A Comprehensive High-Level Understanding”, Research Graph, Medium, May 21, 2024, https://medium.com/@researchgraph/diffusion-model-compreshensive-high-level-understanding-55d6ecad2cba, 22 pages.
[cited by applicant]
Larcher, Mario, “Diffusion Transformer Explained”, Towards Data Science, Feb. 28, 2024, https://medium.com/towards-data-sciene/diffusion-transforer-explained-e603c4770f7 e, 19 pages.
[cited by applicant]
O'Connor, Ryan, “Introduction to Diffusion Models for Machine Learning” AssemblyAI, May 12, 2022, https://www. assemblyai.com/blog/diffusion-models-for-machine-learning-introduction/, 34 pages.
[cited by applicant]
Andreas, et al., “DRCap_Zeroshot_Audio-Captioning”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/drcap_zeroshot_aac/README.md, 3 pages.
[cited by applicant]
Di Pietro, Mauro, “GenAI with Python: Build Agents from Scratch (Complete Tutorial)”, Towards Data Science, Sep. 29, 2024, https://towardsdatascience.com/genai-with-python-build-agents-from-scratch-com pletetutorial-4fc…
[cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, Cornell Univ., arXiv:2204.06125v1 [cs.CV], Apr. 13, 2022, 27 pages.
[cited by applicant]
Ramirez, et al., “Voice Activity Detection. Fundamentals and Speech Recognition System Robustness.” InTech Open Science Open Minds, 2007, 24 pages.
[cited by applicant]
Swimberghe, Niels, “How to integrate spoken audio into LangChain.js using AssemblyAI”, AssemblyAI, Aug. 15, 2023, https://www.assemblyai.com/blog/integrate-audio-langchainjs/, 12 pages.
[cited by applicant]
“A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, F5-TTS, retrieved from the internet on Oct. 20, 2024, https:/swivid.github.io.F5-TTS/, 17 pages.
[cited by applicant]
“CLIP: Connecting text and images”, OpenAI, Jan. 5, 2021, https://openai.com/index/clip/, 16 pages.
[cited by applicant]
“CLIP”, Hugging Face, retrieved from the internet on Oct. 22, 2024, https://huggingface.co/docs/transformers/model_doc/clip, 49 pages.
[cited by applicant]
Rustamy, Fahim, Phd., “CLIP Model and The Importance of Multimodal Embeddings”, Towards Data Science, Dec. 11, 2023, https://towardsdatascience.com/clip-model-and-the-importance-of-multimodalembeddings-1c8f6b13bf72, 20 …
[cited by applicant]
“Diffusion Models from Scratch”, Hugging Face Diffusion Course, retrieved from the internet on Oct. 22, 2024, https://huggingface. co/learn/diffusion-course/en/unit1 /3, 31 pages.
[cited by applicant]
“MC_MusicCaps”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/L-LANCES/SLAM-LLM/blob/main/examples/mc_musiccaps/README. md, 2 pages.
[cited by applicant]
Briggs, James, “Quick-fire Guide to Multi-Modal ML With OpenAI's CLIP”, Towards Data Science, Aug. 11, 2022, https://towardsdatascience. com/quick-tire-guide-to-multi-modal-ml-with-openais-cli p 2dad7 e398ac0, 21 pages.
[cited by applicant]
Bouchard, Louis-Francois, “Stable Diffusion for Videos Explained”, Towards AI, Nov. 29, 2023, https://pub.towardsai.net/stable-diffusion-for-videos-explained-fawf0b6af3b0, 15 pages.
[cited by applicant]
Erdem, Kemal, “Step by Step visual introduction to Diffusion Models”, published Nov. 1, 2023, https://erdem.pl/2023/11/step-by-step-visual-introduction-to-diffusion-models, 15 pages.
[cited by applicant]
“Stable Diffusion: Training Your Own Model in 3 Simple Steps”, run:ai, https://www.run.ai.guides/generative-ai/stablediffusion-training, 10 pages.
[cited by applicant]
Stevens, Ingrid, “Uncovering Insights in Audio: An Exploration”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/ingridstevens/whisper-audio-transcriber/tree/main, 6 pages.
[cited by applicant]
Palucha, Szymon, “Understanding OpenAI's CLIP model”, Medium, Feb. 24, 2024 https://medium.com/@paluchasz/understanding-openais-cl i p-m odel-6b52bade3fa3, 23 pages.
[cited by applicant]
Chen et al., JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning, Cornell Univ., arXiv:2406.12292 [cs.SD], Jun. 18, 2024, 13 pages.
[cited by applicant]
Github, “SLAM-MC”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 5 pages.
[cited by applicant]
Github, “SLAM-LLM”, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM, 4 pages.
[cited by applicant]
Huggingface.co Blog, “The Annotated Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/blog/annotated-diffusion, 38 pages.
[cited by applicant]
Huggingface.co Blog, “Train a Diffusion Model”, retrieved from the internet on Oct. 20, 2024, https://huggingface.co/docs/ diffusers/tutorials/basic_training, 12 pages.
[cited by applicant]
IBM, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 18 pages.
[cited by applicant]
Lil 'Log, “What are Diffusion Models?”, retrieved from the internet on Oct. 22, 2024, 25 pages.
[cited by applicant]
Assembly AI, “Topic Detection”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audiointelligence/topic-detection, 6 pages.
[cited by applicant]
Assembly AI, “Key Phrases”, retrieved from the internet on Oct. 20, 2024, https://www.assemblyai.com/docs/audiointelligence/key-phrases, 5 pages.
[cited by applicant]
Assembly All, “Sentiment Analysis”, retrieved from the internet Oct. 20, 2024, https://www.assemblyai.com/docs/audio-intelligence/sentiment-analysis, 4 pages.
[cited by applicant]
Assembly AI, “Summarization”, retrieved from the internet on Oct. 20, 2024, https://assemblyai.com/docs/audiointelligence/summarization, 6 pages.
[cited by applicant]
Atal et al., “A pattern recognition approach to voiced-unvoiced-silence classification with applications to speech recognition”, IEEE Transactions on Acoustics, Speech, and Signal Processing (vol. 24, Issue: 3, Jun. 197…
[cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, Cornell Univ., arXiv:1312.6114v11 [stat. ML] -, Dec. 10, 2022, pgs.
[cited by applicant]
Sohl-Dickstein et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics” Cornell Univ., arXiv:1503.03585v8 [cs.LG], Nov. 18, 2015. 18 pgs.
[cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, Cornell Univ., arXiv:1505.04597v1 [cs.CV], May 18, 2015, 8 pgs.
[cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, Cornell Univ., arXiv:2006.11239v2 [cs.LG], Dec. 16, 2020, 25 pgs.
[cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Cornell Univ., arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, 48 pgs.
[cited by applicant]
Dhariwal et al., “Diffusion Models Beat GANs on Image Synthesis”, Cornell Univ., arXiv:2105.05233v4 [cs.LG], Jun. 2021, 44 pgs.
[cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, Cornell Univ., arXiv:2112.10752v2 [cs.CV], Apr. 13, 2022, 45 pgs.
[cited by applicant]
Blattmann et al., “Semi-Parametric Neural Image Synthesis”, Cornell Univ., arXiv:2204.11824v3 [cs.CV], Oct. 24, 2022, 34 pgs.
[cited by applicant]
Karras et al., . . . “Elucidating the Design Space of Diffusion-Based Generative Models”, Cornell Univ., arXiv:2206.00364v2 [cs.CV], Oct. 11, 2022, 47 pgs.
[cited by applicant]
Graikos et al., “Diffusion models as plug-and-play priors” Cornell Univ., arXiv:2206.09012v3 [cs.LG], Jan. 8, 2023, 22 pgs.
[cited by applicant]
Luo, “Understanding Diffusion Models: A Unified Perspective”, Cornell Univ., arXiv:2208.11970v1 [cs.LG], Aug. 25, 2022, 23 pgs.
[cited by applicant]
Ruiz et al., “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, Cornell Univ., arXiv:2208.12242v2 [cs.CV], Mar. 15, 2023, 25 pgs.
[cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v9 [cs.LG], Oct. 24, 2022, 39 pgs.
[cited by applicant]
Yang et al., “Diffusion Models: A Comprehensive Survey of Methods and Applications”, Cornell Univ., arXiv:2209.00796v13 [cs.LG], Oct. 24, 2022, 39 pgs.
[cited by applicant]
Lipman et al., “Flow Matching for Generative Modeling”, Cornell Univ., arXiv:2210.02747v2 [cs.LG], Feb. 8, 2023, 28 pgs.
[cited by applicant]
Peebles et al., “Scalable Diffusion Models with Transformers”, Cornell Univ., arXiv:2212.09748v2 [cs.CV], Mar. 2023, 25 pgs.
[cited by applicant]
Schneider et al., “Mo0sai: Text-to-Music Generation with Long-Context Latent Diffusion”, Cornell Univ., arXiv:2301.11757v2 [cs.CL], Jan. 30, 2023, 13 pgs.
[cited by applicant]
Fei, et al., “Generative Diffusion Prior for Unified Image Restoration and Enhancement”, Cornell Univ., arXiv:2304.01247v1 [cs.CV], Apr. 3, 2023, 46 pages.
[cited by applicant]
Ghosal, et al., “Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model”, Cornell Univ., arXiv:2304.13731v2 [eess.AS], May 29, 2023, 15 pages.
[cited by applicant]
Liu, et al., AudioLDM: Text-to-Audio Generation with Latent Diffusion Models, Cornell Univ., arXiv:2301.12503v3 [cs.SD], Sep. 9, 2023, 25 pages.
[cited by applicant]
Wei, et al., “ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation”, Cornell Univ., arXiv:2302.13848v2 [cs.CV], Aug. 18, 2023, 16 pages.
[cited by applicant]
Zhang, et al., “A Survey on Audio Diffusion Models: Text to Speech Synthesis and Enhancement in Generative AI”, Cornell Univ., arXiv:2303.13336v2 [cs.SD], Apr. 2, 2023, 18 pages.
[cited by applicant]
Nikkiran, et al., “Step-by-Step Diffusion: An Elementary Tutorial”, Cornell Univ., arXiv:2406.08929v2 [cs.LG], Jun. 23, 2024, 51 pages.
[cited by applicant]
Copet, et al., “Simple and Controllable Music Generation”, Cornell Univ., arXiv:2306.05284v3 [cs.SD], Jan. 30, 2024, 17 pages.
[cited by applicant]
Li, et al., “JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models”, Cornell Univ., arXiv:2308.04729v1 [cs.SD], Aug. 9, 2023, 12 pages.
[cited by applicant]
Yao, et al., “JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation”, Cornell Univ., arXiv:2310.19180v2 [cs.SD], Nov. 3, 2023, 12 pages.
[cited by applicant]
Xue, et al., “Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation”, Cornell Univ., arXiv:2401.01044v1 [cs.SD], Jan. 2, 2024, 17 pages.
[cited by applicant]
Zheng, et al., “BAT: Learning to Reason about Spatial Sounds with Large Language Models”, Cornell Univ., arXiv:2402.01591v2 [eess.AS], May 25, 2024, 16 pages.
[cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v2 [cs.CV], Oct. 20, 2024, 28 pages.
[cited by applicant]
Sammani, et al., “Interpreting and Analyzing CLIP's Zero-Shot Image Classification via Mutual Knowledge”, Cornell Univ., arXiv:2410.13016v1 [cs.CV], Oct. 16, 2024, 28 pages.
[cited by applicant]
Chen, et al., “F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching”, Cornell Univ., arXiv:2410.06885v2 [eess.AS], Oct. 15, 2024, 18 pages.
[cited by applicant]
Chen, et al., “Contrastive Localized Language-Image Pre-Training”, Cornell Univ., arXiv:2410.02746v1 [cs.CV], Oct. 3, 2024, 20 pages.
[cited by applicant]
Xin, et al., “DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval”, Cornell Univ., arXiv:2409.10025v2 [cs.SD], Oct. 17, 2024, 5 pages.
[cited by applicant]
Chen, et al., “JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning”, Cornell Univ., arXiv:2406.12292v1 [cs.SD], Jun. 28, 2024, 13 pages.
[cited by applicant]
Chan, Stanley, H., “Tutorial on Diffusion Models for Imaging and Vision”, Cornell Univ., arXiv:2403.18103v2 [cs.LG], Sep. 6, 2024, 89 pages.
[cited by applicant]
Tu, et al., “A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)”, Cornell Univ., arXiv:2402.07410v1 [cs.CV], Feb. 12, 2024, 14 pages.
[cited by applicant]
Esser, et al., “Scaling Rectified Flow Transformers for High-Resolution Image Synthesis”, Cornell Univ., arXiv:2403.03206v1 [cs.CV], Mar. 5, 2024, 28 pages.
[cited by applicant]
Bansal, et al., “Universal Guidance for Diffusion Models”, Cornell Univ., arXiv:2302.07121v1 [cs.CV], Feb. 14, 2023, 10 pages.
[cited by applicant]
Young, Mike, “A Complete Guide to Turning Text into Audio with Audio-LDM”, retrieved from the internet on Oct. 20, 2024, https://notes.airmodels.fyi/audio-ldm-ai-text-to-audio-generation-with-latent-diffusion-models/, 9…
[cited by applicant]
“Lecture 14: LPC speech synthesis and autocorrelation-based pitch tracking”, ECE 417, Multimedia Signal Processing, Oct. 10, 2019, 37 pages.
[cited by applicant]
Ribeiro, Andre, “Linking Images and Text with OpenAI CLIP”, Tawards Data Scoemce.retrieved from the internet on Oct. 22, 2024, https://towardsdatascience.com/linking-images-and-text-wth-openai-clip-abb4bdf5dbd2, 23 page…
[cited by applicant]
Copet, Jade, “MusicGen:Simple and Controllable Music Generation”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/facebookresearch/audiocraft/blob/main/docs/MUSICGEN.md, 10 pages.
[cited by applicant]
Agostinelli, et al., “MusicLM: Generating Music From Text”, Cornell Univ., arXiv:2301.11325v1 [cs. SD], Jan. 26, 2023, 15 pages.
[cited by applicant]
Microsoft Copilot: Your AI Companion, retrieved from the internet on Oct. 23, 2024, https://copilot. microsoft.com/?FORM=hpcodx&showconv=1, 3 pages.
[cited by applicant]
Graf et al., “Features for voice activity detection: a comparative analysis”, EURASIP Journal on Advances in Signal Processing, a SpringerOpen Journal, 2015, 15 pages.
[cited by applicant]
“SELD SpatialSoundQA”, GitHub, retrieved from the internet on Oct. 20, 2024, https://github.com/X-LANCE/SLAM-LLM/blob/main/examples/seld_spatialsoundqa/README.md, 4 pages.
[cited by applicant]
Radford, et al., “Contrastive Language-Image Pretraining”, CLIP Explained, Papers With Code, retrieved from the internet on Oct. 22, 2024, https://paperswithcode.com/method/clip, 5 pages.
[cited by applicant]
Huang, et al., “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models”, Cornell Univ., arXiv:2301.12661v1 [cs. SD], Jan. 30, 2023, 16 pages.
[cited by applicant]
Aguilera, Frank Morales, “Open AI CLIP: Bridging Text and Images”, The Deep Hub, https://medium.com/thedeephub/openai-clip-bridging-text-and-images-aaf3cd20299e, Apr. 11, 2024, 15 pages.
[cited by applicant]