IP Library › Granted Patent US 12,711,308
Granted Patent B2
US 12,711,308 · App. 18/646,202 · Granted Aug 18, 2026

System and method for knowledge-based audio-text modeling via automatic multimodal graph construction

Inventors: Wei-Cheng Lin (Pittsburgh, PA); Ho-Hsiang Wu (Morrisville, NC); Shabnam Ghaffarzadegan (Livermore, CA); Luca Bondi (Pittsburgh, PA); Abinaya Kumar (Pittsburgh, PA); Samarjit Das (Wexford, PA)
Assignee: Robert Bosch GmbH
G06F40/20G06F16/65G06F16/685G06F40/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,308
App. No.
18/646,202
Filed
Apr 25, 2024
Granted
Aug 18, 2026
Kind
B2
Art Unit
2658
USPC
704/9
Abstract

Knowledge-based audio-text modeling via automatic multimodal graph construction is performed. An audio dataset is received, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip of the audio data. Graph nodes of interest are identified from a sematic network, the graph nodes being descriptive of semantics of the knowledge domain of the contents of the audio dataset. A large language model (LLM) is utilized for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph. The extracted knowledge graph is validated utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data.

Claims (44)

1 . A method for knowledge-based audio-text modeling via automatic multimodal graph construction, comprising:

receiving an audio dataset, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of audio contents of the respective clip;

identifying graph nodes of interest from a semantic network, the graph nodes being descriptive of semantics of a knowledge domain of the audio dataset;

utilizing a large language model (LLM) for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph;

validating the extracted knowledge graph utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data; and

utilizing the knowledge graph, as validated, for a downstream application, including providing the knowledge graph to an audio generation model as a prompt to generate audio data based on the knowledge graph.

2 . The method of claim 1 , wherein the metadata includes human annotations describing the audio contents of the respective clips of the audio data.

3 . The method of claim 1 , wherein the metadata includes machine-learned labels, attributes, and/or other forms of recognition outcomes inferred from audio and/or speech data using one or more machine learning models.

4 . The method of claim 1 , wherein the graph nodes of interest are one or more of: defined based on user knowledge of the knowledge domain, queried from the semantic network as graph nodes describing semantics of the knowledge domain; extracted from a database of domain knowledge; received from the LLM responsive to a prompt for relevant graph nodes for the knowledge domain.

5 . The method of claim 1 , wherein inferring the supplemental data includes receiving the supplemental data from the LLM responsive to a prompt for requesting the LLM to infer content for names of the graph nodes for which there is no metadata available.

6 . The method of claim 1 , wherein the downstream application includes an audio classification application using the knowledge graph for sound event detection and/or audio tagging.

7 . The method of claim 1 , wherein the downstream application includes an audio captioning application using the knowledge graph for audio retrieval.

8 . The method of claim 1 , wherein the downstream application includes representing the knowledge graph as an adjacency matrix to perform multimodal graph representation learning.

9 . The method of claim 1 , wherein the downstream application includes using the knowledge graph to define knowledge-based clusters for contrastive learning.

10 . The method of claim 1 , wherein the downstream application includes using the knowledge graph to curate controllable prompts, captions, and/or descriptive contents for building knowledge-guided generative models.

11 . The method of claim 1 , further comprising:

converting the knowledge graph into a textual representation; and

providing the textual representation as a textual prompt to a text-to-audio model to generate the audio data.

12 . The method of claim 1 , further comprising:

displaying the knowledge graph via a knowledge graph editor;

receiving user input via the knowledge graph editor to substitute at least one node of the knowledge graph with a different node; and

providing the knowledge graph, as substituted, to the audio generation model, thereby modifying a characteristic of the audio data generated by the audio generation model.

13 . A system for knowledge-based audio-text modeling via automatic multimodal graph construction, comprising:

one or more hardware computing devices configured to:

receive an audio dataset, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip;

identify graph nodes of interest from a semantic network, the graph nodes being descriptive of semantics of a knowledge domain of the audio dataset;

utilize a large language model (LLM) for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph;

validate the extracted knowledge graph utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data; and

utilize the knowledge graph, as validated, for a downstream application, including providing the knowledge graph to an audio generation model as a prompt to generate audio data based on the knowledge graph.

14 . The system of claim 13 , wherein the metadata includes human annotations describing the audio contents of the respective clips of the audio data.

15 . The system of claim 13 , wherein the metadata includes machine-learned labels, attributes, and/or other forms of recognition outcomes inferred from audio and/or speech data using one or more machine learning models.

16 . The system of claim 13 , wherein the graph nodes of interest are one or more of: defined based on user knowledge of the knowledge domain, queried from the semantic network as graph nodes describing semantics of the knowledge domain; extracted from a database of domain knowledge; received from the LLM responsive to a prompt for relevant graph nodes for the knowledge domain.

17 . The system of claim 13 , wherein inferring the supplemental data includes receiving the supplemental data from the LLM responsive to a prompt for requesting the LLM to infer content for names of the graph nodes for which there is no metadata available.

18 . The system of claim 13 , wherein the downstream application includes an audio classification application using the knowledge graph for sound event detection and/or audio tagging.

19 . The system of claim 13 , wherein the downstream application includes an audio captioning application using the knowledge graph for audio retrieval.

20 . The system of claim 13 , wherein the downstream application includes representing the knowledge graph as an adjacency matrix to perform multimodal graph representation learning.

21 . The system of claim 13 , wherein the downstream application includes using the knowledge graph to define knowledge-based clusters for contrastive learning.

22 . The system of claim 13 , wherein the downstream application includes using the knowledge graph to curate controllable prompts/captions/descriptive contents for building knowledge-guided generative models.

23 . A non-transitory computer-readable medium comprising instructions for a knowledge-based audio-text modeling via automatic multimodal graph construction that, when executed by one or more hardware computing devices cause the one or more hardware computing devices to perform operations including to:

receive an audio dataset, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip;

identify graph nodes of interest from a semantic network, the graph nodes being descriptive of semantics of a knowledge domain of the audio dataset;

utilize a large language model (LLM) for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph;

validate the extracted knowledge graph utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data; and

utilize the knowledge graph, as validated, for a downstream application, including providing the knowledge graph to an audio generation model as a prompt to generate audio data based on the knowledge graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2024
From: LIN, WEI-CHENG; WU, HO-HSIANG; GHAFFARZADEGAN, SHABNAM; BONDI, LUCA; KUMAR, ABINAYA; DAS, SAMARJIT
To: ROBERT BOSCH GMBH
Reel/Frame 067237/0168 →
Continuity (1)
Related Publication 20250335705A1 · Oct 30, 2025
References Cited (68)
US 11861320B1 · Gajek · 2024 [cited by examiner]
US 12242503B1 · Kelsey · 2025 [cited by examiner]
US 12243513B2 · Zhu · 2025 [cited by examiner]
US 12340238B1 · Stänescu · 2025 [cited by examiner]
US 12379948B1 · Vlasceanu · 2025 [cited by examiner]
US 12412138B1 · Geene · 2025 [cited by examiner]
US 12437188B2 · Cantrell · 2025 [cited by examiner]
US 12450887B1 · Etemadi · 2025 [cited by examiner]
US 20210287103A1 · Bangalore · 2021 [cited by examiner]
US 20220050864A1 · Osmon · 2022 [cited by examiner]
US 20220377431A1 · Lerner · 2022 [cited by examiner]
US 20230067177A1 · Ding · 2023 [cited by examiner]
US 20230071799A1 · Ramnani · 2023 [cited by examiner]
US 20240037328A1 · Jiang · 2024 [cited by examiner]
US 20240289395A1 · Zhou · 2024 [cited by examiner]
US 20240354586A1 · Cummings · 2024 [cited by examiner]
US 20240370709A1 · Siebel · 2024 [cited by examiner]
US 20240370769A1 · Sheth · 2024 [cited by examiner]
US 20240394176A1 · Pean · 2024 [cited by examiner]
US 20240412720A1 · Vasylyev · 2024 [cited by examiner]
US 20240419912A1 · Somech · 2024 [cited by examiner]
US 20250068667A1 · Blum · 2025 [cited by examiner]
US 20250110975A1 · Shea · 2025 [cited by examiner]
US 20250140246A1 · Lee · 2025 [cited by examiner]
US 20250217209A1 · Pedersen · 2025 [cited by examiner]
US 20250217576A1 · Mann · 2025 [cited by examiner]
US 20250217585A1 · Jain · 2025 [cited by examiner]
US 20250217673A1 · Silver · 2025 [cited by examiner]
US 20250225572A1 · Chiu · 2025 [cited by examiner]
US 20250238629A1 · Malkiel · 2025 [cited by examiner]
US 20250240494A1 · Avendano · 2025 [cited by examiner]
US 20250245872A1 · Cokely · 2025 [cited by examiner]
US 20250259080A1 · Wang · 2025 [cited by examiner]
US 20250284721A1 · Agrawal · 2025 [cited by examiner]
US 20250285348A1 · Cantrell · 2025 [cited by examiner]
US 20250293998A1 · Courcelle · 2025 [cited by examiner]
US 20250307690A1 · Hoffer · 2025 [cited by examiner]
US 20250321976A1 · Isles · 2025 [cited by examiner]
US 20250321977A1 · Kelsey · 2025 [cited by examiner]
US 20250331085A1 · Gilg · 2025 [cited by examiner]
US 20250335705A1 · Lin · 2025 [cited by examiner]
DE 10047212A1 · 2002 [cited by applicant]
DE 10307840B3 · 2004 [cited by applicant]
DE 202010013504U1 · 2010 [cited by applicant]
EP 1269923A2 · 2003 [cited by applicant]
EP 3584024A1 · 2019 [cited by applicant]
Chen et al., “A Survey on Multimodal Knowledge Graphs: Construction, Completion and Applications”, Mathematics 2023, 11, 1815, pp. 1-27. (Year: 2023). [cited by examiner]
Doh et al., “Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to Music Retrieval”, ICASSP-24, 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Apr. 14,… [cited by examiner]
Doh et al., “Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval”, 2024 International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2024, 5 pages. [cited by applicant]
Fan et al., “Graph Machine Learning in the Era of Large Language Models (LLMs)”, IEEE Transactions on Knowledge and Data Engineering Subission, 2023, 21 pages. [cited by applicant]
Zhu et al., “LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities”, arxiv.org, Cornell University Library, 2023, 18 pages. [cited by applicant]
Baker et al., “The Berkeley FrameNet Project”, The 17th International Conference on Computational Linguistics, Coling, 1998, 5 pages, vol. 1. [cited by applicant]
Chia et al., “RelationPrompt: Leveraging Prompts to Generate Synthetic Data for Zero-Shot Relation Triplet Extraction”, Finding of the Association for Computational Linguistics: ACL, May 2022, p. 45-57. [cited by applicant]
Dhuliawala et al., “Chain-Of-Verification Reduces Hallucination in Large Language Models”, arXiv preprint arXiv:2309.11495, 2023, p. 1-19. [cited by applicant]
Liu et al., “Entity-Duet Neural Ranking: Understanding the Role of Knowledge Graph Semantics in Neural Information Retrieval”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Lon… [cited by applicant]
Hsu et al. “Degree: A Data-Efficient Generation-Based Event Extraction Model”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi… [cited by applicant]
Hamilton, “Graph Representation Learning”, Synthesis Lectures on Artificial Intelligence and Machine Learning, 2020, p. 1-159, vol. 14, No. 3. [cited by applicant]
Gong et al., “Listen, Think, and Understand”, Conference Paper at ICLR 2024, 2024, p. 1-29. [cited by applicant]
Jiang et al., “Active Retrieval Augmented Generation”, arXiv preprint arXiv:2305.06983, 2023, 24 pages. [cited by applicant]
Kreuk et al., “AudioGen: Textually Guided Audio Generation”, Conference Paper at ICLR 2023, 2023, p. 1-16. [cited by applicant]
Lan et al., “A Survey on Complex Knowledge Base Question Answering: Methods, Challenges and Solutions”, Proceedings of the Thirtieth Joint Conference on Artificial Intelligence (IJCAI-12) Survey Track, Aug. 2021, p. 448… [cited by applicant]
Liu et al., “AudioLDM: Text-to-Audio Generation with Latent Diffusion Model”, Proceedings of the 40th International Conference on Machine Learning, Jul. 2023, p. 1-25. [cited by applicant]
Sallam et al., “ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns”, Healthcare, 2023, p. 1-20, vol. 11, No. 887. [cited by applicant]
Zheng et al., “PRGC: Potential Relation and Global Correspondence Based Joint Relational Triple Extraction”, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internati… [cited by applicant]
Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Model”, 36th Conference on Neural Information Processing Systems, 2022, 43 pages. [cited by applicant]
White et al., “A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT”, arXiv preprint arXiv:2302.11382, 2023, 19 pages. [cited by applicant]
Yao et al., “Schema-Aware Reference as Prompt Improves Data-Efficient Knowledge Graph Construction”, SIGIR 2023, Jul. 2023, 12 pages. [cited by applicant]
Zhang et al., “Supporing Clustering with Contrastive Learning”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 2021, … [cited by applicant]