IP Library Granted Patent US 12,243,513
Granted Patent B2
US 12,243,513 · App. 17/323,847 · Granted Mar 4, 2025

Generation of optimized spoken language understanding model through joint training with integrated acoustic knowledge-speech module

Inventors: Chenguang Zhu (Sammamish, WA); Nanshan Zeng (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G06F40/279G10L13/08G10L15/02G10L15/1815G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,513
App. No.
17/323,847
Granted
Mar 4, 2025
Kind
B2
Abstract

A speech module is joint trained with a knowledge module by transforming a first knowledge graph into an acoustic knowledge graph. The knowledge module is trained on the acoustic knowledge graph. Then, the knowledge module is integrated with the speech module to generate an integrated knowledge-speech module. In some instances, the speech module included in the integrated knowledge-speech module is aligned with a language module to generate an optimized speech model configured to leverage acoustic information and acoustic-based knowledge information, along with language information.

Claims (63)

1. A computer implemented method for joint training a speech module with a knowledge module for natural language understanding, the method comprising:

obtaining a first knowledge graph comprising a set of entities and a set of relations between two or more entities included in the set of entities;

transforming the first knowledge graph into an acoustic knowledge graph, the acoustic knowledge graph comprising a set of acoustic data corresponding to each entity of the set of entities and each relation of the set of relations between entities;

extracting a first set of acoustic features from the acoustic knowledge graph;

training a knowledge module on the set of acoustic features to generate knowledge-based entity representations for the acoustic knowledge graph;

pre-training a speech module on a training data set comprising unlabeled speech data to understand acoustic information from speech transcriptions;

pre-training a language module on a second training data set comprising unlabeled text-based data, the language module being configured to understand semantic information from speech transcriptions;

aligning the speech module and the language module, the speech module being configured to leverage acoustic information and language information in natural language processing tasks;

obtaining a third training data set comprising paired acoustic data and transcript data;

applying the third training data set to the speech module and language module;

obtaining acoustic output embeddings from the speech module;

obtaining language output embeddings from the language module;

after pre-training the language module, extracting a set of textual features from the first knowledge graph;

training a second knowledge module on the set of textual features; and

integrating the second knowledge module with the language module before the language module and speech module are aligned;

aligning the acoustic output embeddings and the language output embeddings to a shared semantic space; and

generating an integrated knowledge-speech module to perform semantic analysis on audio data by integrating the knowledge module with the speech module, the integrated knowledge-speech module being generated by providing context information from the speech module to the knowledge module and knowledge information from the knowledge module to the speech module.

2. The method of claim 1 , wherein transforming the first knowledge graph into an acoustic knowledge graph comprises:

obtaining electronic content that describes the set of entities and the set of relations between two or more entities included in the set of entities;

extracting a second set of acoustic features from the electronic content;

transcribing the second set of acoustic features from the electronic content into text, the text representing the set of entities and the set of relations included in the first knowledge graph; and

training the knowledge module on the text and the second set of acoustic features, the knowledge module being configured to generate entity representations.

3. The method of claim 1 , wherein transforming the first knowledge graph into an acoustic knowledge graph comprises:

extracting a set of textual representations for each entity of the set of entities and each relation of the set of relations;

accessing a text-to-speech module;

applying the set of textual representations as input to the text-to-speech module; and

obtaining a set of acoustic signals corresponding to the set of textual representations as output from the text-to-speech module, each acoustic signal of the set of acoustic signals representing an entity of the set of entities included in the first knowledge graph or a relation of the set of relations included in the first knowledge graph.

4. The method of claim 1 , wherein the knowledge module of the integrated knowledge-speech module comprises a graph attention network and is configured to provide structure-aware entity embeddings for speech modeling.

5. The method of claim 1 , wherein the speech module of the integrated knowledge-speech module is further configured to produce acoustic language representations as initial embeddings for acoustic knowledge graph entities and relations.

6. The method of claim 1 , wherein integrating the speech module and knowledge module comprises projecting entity and relations output embeddings and language acoustic embeddings into a shared semantic space.

7. The method of claim 1 , wherein the speech module of the integrated knowledge-speech module comprises a first speech module and a second speech module.

8. The method of claim 7 , wherein the first speech module comprise a first set of transformer layers and the second speech module comprises a second set of transformer layers.

9. The method of claim 7 , wherein integrating the speech module and knowledge module further comprises:

obtaining a first set of acoustic language embeddings from the first speech module; and

applying the first set of acoustic language embeddings as input to the second speech module and the knowledge module.

10. The method of claim 9 , further comprising:

obtaining a first set of acoustic entity embeddings from the knowledge module;

applying the first set of acoustic entity embeddings as input to the second speech module; and

obtaining a final representation output embedding from the second speech module based on the first set of acoustic entity embeddings and the first set of acoustic language embeddings, the final representation output embedding including acoustic and knowledge information.

11. The method of claim 1 , wherein each entity of the set of entities is represented by entity description audio data to describe a concept and meaning of the entities, and each relation of the set of relations being represented by relation description audio data to describe each relation between two or more entities.

12. The method of claim 1 , further comprising:

optimizing the integrated knowledge-speech module by performing one or more following training tasks: entity category prediction, relation type prediction, masked token prediction, or masked entity prediction.

13. A computing system configured for joint training a speech module with a knowledge module for natural language understanding, the computing system comprising:

one or more processors; and

one or more hardware storage devices that store computer-executable instructions that are structured to be executed by the one or more processors to cause the computing system to at least:

obtain a first knowledge graph comprising a set of entities and a set of relations between two or more entities included in the set of entities;

transform the first knowledge graph into an acoustic knowledge graph, the acoustic knowledge graph comprising acoustic data corresponding to each entity of the set of entities and each relation of the set of relations between entities;

extract a set of acoustic features from the acoustic knowledge graph;

train a knowledge module on the set of acoustic features to generate knowledge-based entity representations for the acoustic knowledge graph;

pre-train a speech module on a training data set comprising unlabeled speech data to understand acoustic information from speech transcriptions;

pre-train a language module on a second training data set comprising unlabeled text-based data, the language module being configured to understand semantic information from speech transcriptions; and

after pre-training the language module, extracting a set of textual features from the first knowledge graph;

training a second knowledge module on the set of textual features;

integrating the second knowledge module with the language module before the language module and speech module are aligned;

align the speech module and the language module, the speech module being configured to leverage acoustic information and language information in natural language processing tasks; and

generate an integrated knowledge-speech module to perform semantic analysis on audio data by integrating the knowledge module with the speech module, the integrated knowledge-speech module being trained by providing context information from the speech module to the knowledge module and knowledge information from the knowledge module to the speech module.

14. A computer implemented method for joint training a speech module with a knowledge module for natural language understanding, the method comprising:

obtaining a first knowledge graph comprising a set of entities and a set of relations between two or more entities included in the set of entities;

transforming the first knowledge graph into an acoustic knowledge graph, the acoustic knowledge graph comprising a set of acoustic data corresponding to each entity of the set of entities and each relation of the set of relations between entities;

extracting a first set of acoustic features from the acoustic knowledge graph;

training a knowledge module on the set of acoustic features to generate knowledge-based entity representations for the acoustic knowledge graph;

pre-training a speech module on a training data set comprising unlabeled speech data to understand acoustic information from speech transcriptions, the speech module comprising a first set of transformer layers in a first speech module and a second set of transformer layers in a second speech module; and

generating an integrated knowledge-speech module to perform semantic analysis on audio data by integrating the knowledge module with the speech module, the integrated knowledge-speech module being generated by at least obtaining a first set of acoustic language embeddings from the first speech module and applying the first set of acoustic language embeddings as input to the second speech module and the knowledge module, providing context information from the speech module to the knowledge module and knowledge information from the knowledge module to the speech module, obtaining a first set of acoustic entity embeddings from the knowledge module, applying the first set of acoustic entity embeddings as input to the second speech module, and obtaining a final representation output embedding from the second speech module based on the first set of acoustic entity embeddings and the first set of acoustic language embeddings, the final representation output embedding including acoustic and knowledge information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2021
From: ZHU, CHENGUANG; ZENG, NANSHAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056279/0514 →
Continuity (2)
Provisional Application 63205646 · Jan 20, 2021
Related Publication 20220230629A1 · Jul 21, 2022
References Cited (129)
US 10388274B1 · Hoffmeister · 2019 [cited by applicant]
US 20150309992A1 · Visel · 2015 [cited by applicant]
US 20150332672A1 · Akbacak · 2015 [cited by examiner]
US 20190005951A1 · Kang et al. · 2019 [cited by applicant]
US 20190147367A1 · Bellamy · 2019 [cited by examiner]
US 20190244611A1 · Godambe · 2019 [cited by examiner]
US 20200342853A1 · Ji et al. · 2020 [cited by applicant]
US 20210312906A1 · Kuo · 2021 [cited by examiner]
US 20220044671A1 · Kapoor · 2022 [cited by examiner]
US 20220122622A1 · Narayanan · 2022 [cited by examiner]
US 20220230625A1 · Zhu et al. · 2022 [cited by applicant]
US 20220230628A1 · Zhu et al. · 2022 [cited by applicant]
WO 2016196320A1 · 2016 [cited by applicant]
Zhang, Yuyu, et al. “Variational Reasoning for Question Answering with Knowledge Graph.” arXiv preprint arXiv:1709.04071 (2017), pp. 1-22 (Year: 2017). [cited by examiner]
Gemmeke, Jort F., et al. “Audio set: An ontology and human-labeled dataset for audio events.” 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) (2017), pp. 776-780 (Year: 2017). [cited by examiner]
Chuang, Yung-Sung, et al. “SpeechBERT: An audio-and-text jointly learned language model for end-to-end spoken question answering.” arXiv preprint arXiv:1910.11559 (2019), pp. 1-6 (Year: 2019). [cited by examiner]
Sun, Tianxiang, et al. “CoLAKE: Contextualized language and knowledge embedding.” arXiv preprint arXiv:2010.00309 (2020), pp. 1-11 (Year: 2020). [cited by examiner]
Pruksachatkun, Yada, et al. “Intermediate-task transfer learning with pretrained models for natural language understanding: When and why does it work?. ” arXiv preprint arXiv:2005.00628 (2020), pp. 1-17 (Year: 2020). [cited by examiner]
Narayanan, Arun, et al. “Cascaded encoders for unifying streaming and non-streaming ASR.” arXiv preprint arXiv:2010.14606 ( 2020), pp. 1-5 (Year: 2020). [cited by examiner]
Song, Meina, Wet al. “KGAnet: a knowledge graph attention network for enhancing natural language inference.” Neural Computing and Applications 32 (2020), pp. 14963-14973 (Year: 2020). [cited by examiner]
Fu, Xiaoyi, et al. “A speech-to-knowledge-graph construction system.” Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence. (Jan. 7, 2021), pp. 5303-5305… [cited by examiner]
Tang, Zhiyuan, et al. “Collaborative joint training with multitask recurrent model for speech and speaker recognition.” IEEE/ACM Transactions on Audio, Speech, and Language Processing 25.3 (2016): pp. 493-504. (Year: 20… [cited by examiner]
Definition of “module” in the Free On-Line Dictionary of Computing, available at https://foldoc.org/module (last updated Oct. 27, 1997). (Year: 1997). [cited by examiner]
Vegupatti, Mani, et al. “Simple Question Answering Over a Domain-Specific Knowledge Graph using BERT by Transfer Learning.” AICS. 2020, pp. 1-12 (Year: 2020). [cited by examiner]
Kumar, et al., “A Knowledge Graph Based Speech Interface for Question Answering Systems”, In Proceedings of Speech Communication, vol. 92, Sep. 1, 2017, 12 Pages. [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 17/323,854”, Mailed Date: Dec. 13, 2022, 77 Pages. (MS# 409334-US-NP). [cited by applicant]
Zhu, et al., “SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering”, In Repository of arXiv:1812.03593v4, Dec. 19, 2018, pp. 1-8. [cited by applicant]
Ba, et al., “Layer Normalization”, In Repository of arXiv:1607.06450, Jul. 21, 2016, 14 Pages. [cited by applicant]
Baevski, et al., “VQ-WAV2VEC: SELF-Supervised Learning of Discrete Speech Representations”, In Proceedings of Eighth International Conference on Learning Representations, Apr. 26, 2020, 12 pages. [cited by applicant]
Bhargava, et al., “Easy Contextual Intent Prediction and Slot Detection”, In Proceedings of International Conference on Acoustics, Speech and Signal Processing, May 26, 2013, pp. 8337-8341. [cited by applicant]
Bordes, et al., “Translating Embeddings for Modeling Multi-Relational Data”, In Proceedings of the 26th International Conference on Neural Information Processing Systems, vol. 2, Dec. 5, 2013, 9 Pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, In Repository of arXiv:2005.14165, May 28, 2020, 72 Pages. [cited by applicant]
Calhoun, et al., “The NXT-format Switchboard Corpus: A rich resource for investigating the syntax, semantics, pragmatics and prosody of dialogue”, In Language Resources and Evaluation, Dec. 2010, pp. 387-419. [cited by applicant]
Chen, et al., “Spoken language understanding without speech recognition”, In IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 15, 2018, 5 Pages. [cited by applicant]
Chuang, et al., “SpeechBERT: Cross-modal pre-trained language model for end-to-end spoken question answering”, In International Speech Communication Association, Oct. 25, 2020, 5 Pages. [cited by applicant]
Chuang,, et al., “SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering”, In Repository arXiv preprint arXiv:1910.11559, Aug. 11, 2020, 6 Pages. [cited by applicant]
Chuing, et al., “Generative pre-training for speech with autoregressive predictive coding”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 4, 2020, pp. 3497-3501. [cited by applicant]
Chung, et al., “An unsupervised autoregressive model for speech representation learning”, In Proceedings of The 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 146-150. [cited by applicant]
Chung, et al., “Vector-quantized autoregressive predictive coding”, In International Speech Communication Association, Oct. 25, 2020, pp. 3760-3764. [cited by applicant]
Chung,, et al., “Semi-Supervised Speech-Language Joint Pre-Training for Spoken Language Understanding”, In Repository of arXiv:2010.02295v1, Oct. 5, 2020, 10 Pages. [cited by applicant]
Denisov, et al., “Pretrained semantic speech embeddings for end-to-end spoken language understanding via cross-modal teacher-student learning”, In International Speech Communication Association, Oct. 25, 2020, 5 Pages. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human … [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Repository of arXiv:1810.04805v1, Oct. 11, 2018, 14 Pages. [cited by applicant]
Ding, et al., “Cognitive graph for multi-hop reading comprehension at scale”, In Repository of arXiv:1905.05460v2, Jun. 4, 2019, 10 Pages. [cited by applicant]
Dong, et al., “Unified Language Model Pre-training for Natural Language Understanding and Generation”, In Proceedings of 33rd Conference on Neural Information Processing Systems, Dec. 8, 2019, pp. 1-13. [cited by applicant]
Duran, et al., “Probabilistic word association for dialogue act classification with recurrent neural networks”, In Engineering Applications of Neural Networks, Sep. 3, 2018, pp. 1-12. [cited by applicant]
Fevry, et al., “Entitles as experts: Sparse memory access with entity supervision”, In Repository of arXiv:2004.07202v2, Oct. 6, 2020, 15 Pages. [cited by applicant]
Gao, et al., “Fewrel 2.0: Towards more challenging few-shot relation classification”, In Repository of arXiv:1910.07124v1, Oct. 16, 2019, 6 pages. [cited by applicant]
Ghosal, et al., “Contextual inter-modal attention for multi-modal sentiment analysis”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 3454-3466. [cited by applicant]
Gorman, et al., “Prosodylab-aligner: A tool for forced alignment of laboratory speech”, In Journal of the Canadian Acoustical Association, vol. 39, Issue 3, Sep. 2011, 3 Pages. [cited by applicant]
Guu, et al., “Realm: Retrieval-augmented language model pre-training”, In Repository of arXiv:2002.08909v1, Feb. 10, 2020, 12 pages. [cited by applicant]
Haghani, et al., “From audio to semantics: Approaches to end-to-end spoken language understanding”, In Proceedings of IEEE Spoken Language Technology Workshop (SLT), Dec. 18, 2018, pp. 720-726. [cited by applicant]
Hamilton, et al., “Inductive representation learning on large graphs”, In Journal of Advances in neural information processing systems, Dec. 4, 2017, pp. 1-11. [cited by applicant]
Han, et al., “Fewrel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation”, In Repository of arXiv:1810.10147, Oct. 26, 2018, 7 Pages. [cited by applicant]
Kalinowski,, et al., “A Survey of Embedding Space Alignment Methods for Language and Knowledge Graphs”, In Repository of arXiv preprint arXiv:2010.13688, Oct. 26, 2020, 27 Pages. [cited by applicant]
Lee, et al., “ODSQA: Opendomain spoken question answering dataset”, In Proceedings of IEEE Spoken Language Technology Workshop (SLT), Dec. 2018, 8 Pages. [cited by applicant]
Levine, et al., “Sensebert: Driving some sense into bert”, In Repository of arXiv:1908.05646v1, Aug. 15, 2019, 10 Pages. [cited by applicant]
Lewis, et al., “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguist… [cited by applicant]
Li, et al., “Spoken Squad: A study of mitigating the impact of speech recognition errors on listening comprehension”, In 19th Annual Conference of the International Speech Communication Association, Sep. 2, 2018, pp. 34… [cited by applicant]
Ling, et al., “Deep contextualized acoustic representations for semi-supervised speech recognition”, In IEEE International Conference on Acoustics, Speech and Signal Processing, May 4, 2020, pp. 6429-6433. [cited by applicant]
Liu, et al., “K-bert: Enabling language representation with knowledge graph”, In Proceedings of The Thirty-Fourth AAAI Conference on Artificial Intelligence, Feb. 7, 2020, 8 Pages. [cited by applicant]
Liu, et al., “Linguistic knowledge and transferability of contextual representations”, In Repository of arXiv:1903.08855v1, Mar. 21, 2019, 22 Pages. [cited by applicant]
Liu, et al., “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders”, In IEEE International Conference on Acoustics, Speech and Signal Processing, May 4, 2020, pp. 6419-642… [cited by applicant]
Liu, et al., “Roberta: A Robustly Optimized BERT Pretraining Approach”, In Repository of arXiv:1907.11692, Jul. 26, 2019, 13 Pages. [cited by applicant]
Loshchilov, et al., “Decoupled weight decay regularization”, In Proceedings of the 7th International Conference on Learning Representations, Sep. 28, 2018, pp. 1-18. [cited by applicant]
Lugosch, et al., “Speech model pre-training for end-to-end spoken language understanding”, In Proceedings of the 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 814-818. [cited by applicant]
Lv, et al., “Graph-based reasoning over heterogeneous external knowledge for commonsense question answering”, In Proceedings of The Thirty-Fourth AAAI Conference on Artificial Intelligence, Apr. 2020, pp. 8449-8456. [cited by applicant]
Mori, et al., “Spoken Language Understanding: Systems for Extracting Semantic Information from Speech”, In Publication of John Wiley and Sons, Apr. 25, 2011, 3 Pages. [cited by applicant]
Oord, et al., “Representation learning with contrastive predictive coding”, In Repository of arXiv:1807.03748v1, Jul. 10, 2018, pp. 1-13. [cited by applicant]
Panayotov, et al., “Librispeech: an Asr Corpus Based on Public Domain Audio Books”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2015, pp. 5206-5210. [cited by applicant]
Park, et al., “SpecAugment: A simple data augmentation method for automatic speech recognition”, In Proceedings of The 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 261… [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US21/062518”, Mailed Date: Apr. 20, 2022, 10 Pages. (MS# 409308-WO-PCT). [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US21/062723”, Mailed Date: Apr. 4, 2022, 10 Pages. (MS# 409334-WO-PCT). [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US21/063940”, Mailed Date: Apr. 13, 2022, 11 Pages. (MS# 409116-WO-PCT). [cited by applicant]
Peters, et al., “Deep contextualized word representations”, In Repository of arXiv:1802.05365v2, Mar. 22, 2018, 15 Pages. [cited by applicant]
Peters, et al., “Knowledge enhanced contextual word representations”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language P… [cited by applicant]
Poerner, et al., “Bert is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised qa”, In Repository of arXiv:1911.03681v1, Nov. 9, 2019. 7 Pages. [cited by applicant]
Qian, et al., “Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system”, InProceedings of IEEE Automatic Speech Recognition and Understanding Workshop, Dec. 16, 201… [cited by applicant]
Radford, et al., “Improving Language Understanding by Generative Pre-training”, Retrieved From: https://www.cs.ubc.ca/˜amuham01/LING530/papers/radford2018improving.pdf, Jun. 11, 2018, 12 Pages. [cited by applicant]
Radford, et al., “Language Models are Unsupervised Multitask Learners”, In Journal of OpenAI Blog, vol. 1, Issue 8, Feb. 24, 2019, 24 Pages. [cited by applicant]
Rajpurkar, et al., “SQUAD: 100,000+ Questions for Machine Comprehension of Text”, In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Nov. 1, 2016, pp. 2383-2392. [cited by applicant]
Rao, et al., “Speech To Semantics: Improve ASR and NLU Jointly via All-Neural Interfaces”, In Repository of arXiv:2008.06173v1, Aug. 14, 2020, 5 Pages. [cited by applicant]
Ravuri, et al., “Recurrent Neural Network and LSTM Models for Lexical Utterance Classification”, In Proceedings of 16th Annual Conference of the International Speech Communication Association, Sep. 6, 2015, 5 Pages. [cited by applicant]
Ringgaard, et al., “Sling: A framework for frame semantic parsing”, In Repository of arXiv:1710.07032v1, Oct. 19, 2017, 9 Pages. [cited by applicant]
Riviere, et al., “Unsupervised pretraining transfers well across languages”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 4, 2020, pp. 7414-7418. [cited by applicant]
Schneider, et al., “wav2vec: Unsupervised pre-training for speech recognition”, In Proceedings of The 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 3465-3469. [cited by applicant]
Serdyuk, et al., “Towards end-to-end spoken language understanding”, In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 15, 2018, pp. 5754-5758. [cited by applicant]
Shen, et al., “Exploiting structured knowledge in text via graph-guided representation learning”, In Repository of arXiv:2004.14224v1, Apr. 19, 2020, pp. 1-19. [cited by applicant]
Soares, et al., “Matching the blanks: Distributional similarity for relation learning”, In Repository of arXiv:1906.03158v1, Jun. 7, 2019, 10 Pages. [cited by applicant]
Song, et al., “Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks”, In International Speech Communication Association, Oct. 25, 2020, pp. 3765-3769. [cited by applicant]
Sun, et al., “ERNIE: Enhanced Representation through Knowledge Integration”, In Journal of arXiv preprint arXiv:1904.09223v1, Apr. 19, 2019, 8 Pages. [cited by applicant]
Sun, et al., “Open domain question answering using early fusion of knowledge bases and text”, In Repository of arXiv:1809.00782v1, Sep. 4, 2018, 12 pages. [cited by applicant]
Sun, et al., “Pullnet: Open domain question answering with iterative retrieval on knowledge bases and text”, In Repository of arXiv:1904.09537v1, Apr. 21, 2019, 11 Pages. [cited by applicant]
Talmor, et al., “olmpics-on what language model pre-training captures”, In Repository of arXiv:1912.13283v1, Dec. 31, 2019, 15 Pages. [cited by applicant]
Vashishth, et al., “Composition-based multirelational graph convolutional networks”, In Repository of arXiv:1911.03082v1, Nov. 8, 2019, pp. 1-14. [cited by applicant]
Vaswani, et al., “Attention is All You Need”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, pp. 1-11. [cited by applicant]
Veliković, et al., “Graph Attention Networks”, In Repository of arXiv:1710.10903, Oct. 30, 2017, pp. 1-11. [cited by applicant]
Verga, et al., “Facts as experts: Adaptable and interpretable neural memory over symbolic knowledge”, In Repository of arXiv:2007.00849v1, Jul. 2, 2020, 12 pages. [cited by applicant]
Vrandecic, “Wikidata: a free collaborative knowledgebase”, In Journal of Communications of the ACM, vol. 57, Issue 10, Sep. 2014, pp. 78-85. [cited by applicant]
Wang, et al., “Deep graph library: Towards efficient and scalable deep learning on graphs”, In Repository of arXiv:1909.01315v1, Sep. 3, 2019, pp. 1-7. [cited by applicant]
Wang, et al., “Kepler: A unified model for knowledge embedding and pre-trained language representation”, In Repository of arXiv:1911.06136v1, Nov. 13, 2019, 10 Pages. [cited by applicant]
Wang, et al., “Unsupervised pre-training of bidirectional speech encoders via masked reconstruction”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 4, 2020, pp.… [cited by applicant]
Wolf, et al., “Transformers: State-of-the-Art Natural Language Processing”, In Repository of arXiv:1910.03771v1, Oct. 9, 2019, pp. 1-11. [cited by applicant]
Xiong, et al., “Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model”, In Repository of arXiv:1912.09637v1, Dec. 20, 2019, pp. 1-13. [cited by applicant]
Yang, et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, In Proceedings of 33rd Conference on Neural Information Processing Systems, Dec. 8, 2019, 11 Pages. [cited by applicant]
YU,, et al., “JAKET: Joint Pre-training of Knowledge Graph and Language Understanding”, In Repository of arXiv:2010.00796v1, Oct. 2020, 11 Pages. [cited by applicant]
Zadeh, et al., “Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Lo… [cited by applicant]
Zhang, et al., “BERTScore: Evaluating text generation with BERT”, In Proceedings of Eighth International Conference on Learning Representations, Apr. 26, 2020, pp. 1-43. [cited by applicant]
Zhang, et al., “ERNIE: Enhanced Language Representation with Informative Entities”, In Repository of arXiv:1905.07129, Jun. 4, 2019, 11 Pages. [cited by applicant]
Zhang, et al., “Semantics-aware BERT for Language Understanding”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, Issue 05, Apr. 2020, 8 Pages. [cited by applicant]
Zhang, et al., “Variational reasoning for question answering with knowledge graph”, In Repository of arXiv:1709.04071v5, Nov. 27, 2017, pp. 1-22. [cited by applicant]
Zhu, et al., “Make lead bias in your favor: A simple and effective method for news summarization”, In Repository of arXiv:1912.11602v1, Dec. 25, 2019, 7 Pages. [cited by applicant]
“Notice of Allowance Issued in U.S. Appl. No. 17/323,854”, Mailed Date: Jun. 23, 2023, 9 Pages. (MS# 409334-US-NP). [cited by applicant]
“Non Final Office Action issued in U.S. Appl. No. 17/323,561”, Mailed Date: Mar. 15, 2023, 24 Pages. (MS#409116-US-NP). [cited by applicant]
Bhattacharjee, et al., “To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging”, In Repository of arXiv:2010.14042v1, Oct. 27, 2020, 8 Pages. [cited by applicant]
Do, et al., “Cross-Lingual Phone Mapping for Large Vocabulary Speech Recognition of Under-Resourced Languages”, In Journal of IEICE Transactions on Information and Systems, vol. 97, Issue 2, Feb. 1, 2014, pp. 285-295. [cited by applicant]
Hsu, et al., “Unsupervised Domain Adaptation for Robust Speech Recognition via Variational Autoencoder-based Data Augmentation”, In Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop, Dec. 16, 2… [cited by applicant]
Hsuan, et al., “Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension”, In Repository of arXiv:1804.00320v1, Apr. 1, 2018, 5 Pages. [cited by applicant]
Jin, et al., “Relation of the Relations: A New Paradigm of the Relation Extraction Problem”, In Repository of arXiv:2006.03719v1, Jun. 5, 2020, 11 Pages. [cited by applicant]
Nathani, et al., “Learning Attention-based Embeddings for Relation Prediction in Knowledge Graphs”, In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28, 2019, pp. 4710-472… [cited by applicant]
Peters, et al., “Knowledge Enhanced Contextual Word Representations”, In Repository of arXiv:1909.04164v1, Sep. 9, 2019, 14 Pages. [cited by applicant]
Tsunoo, et al., “Towards Online End-to-end Transformer Automatic Speech Recognition”, In Repository of arXiv:1910.11871v1, Oct. 25, 2019, 5 Pages. [cited by applicant]
U.S. Appl. No. 63/205,647, filed Jan. 20, 2021. [cited by applicant]
U.S. Appl. No. 63/205,646, filed Jan. 20, 2021. [cited by applicant]
U.S. Appl. No. 63/205,645, filed Jan. 20, 2021. [cited by applicant]
U.S. Appl. No. 17/323,561, filed May 18, 2021. [cited by applicant]
U.S. Appl. No. 17/323,854, filed May 18, 2021. [cited by applicant]
Final Office Action mailed on Dec. 8, 2023, in U.S. Appl. No. 17/323,561 (MS# 409116-US02), 18 Pages. [cited by applicant]
Non-Final Office Action mailed on Jun. 21, 2024, in U.S. Appl. No. 17/323,561, (MS#409116-US02), 19 pages. [cited by applicant]
Cited By (2)
US 12,518,741 US 12,711,308