IP Library Granted Patent US 12,387,720
Granted Patent B2
US 12,387,720 · App. 17/455,727 · Granted Aug 12, 2025

Neural sentence generator for virtual assistants

Inventors: Pranav Singh (Sunnyvale, CA); Keyvan Mohajer (Los Gatos, CA); Yilun Zhang (Toronto, CA)
Assignee: SoundHound AI IP, LLC.
G10L15/1822G06F40/284G06F40/30G06F40/35G10L15/02G10L15/063G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,720
App. No.
17/455,727
Granted
Aug 12, 2025
Kind
B2
Abstract

Methods and systems for automatically generating sample phrases or sentences that a user can say to invoke a set of defined actions performed by a virtual assistant are disclosed. By enabling finetuned general-purpose natural language models, the system can generate potential and accurate utterance sentences based on extracted keywords or the input utterance sentence. Furthermore, domain-specific datasets can be used to train the pre-trained, general-purpose natural language models via unsupervised learning. These generated sentences can improve the efficiency of configuring a virtual assistant. The system can further optimize the effectiveness of a virtual assistant in understanding the user, which can enhance the user experience of communicating with it.

Claims (51)

1. A computer-implemented method for virtual assistants, comprising:

receiving an utterance sentence corresponding to a customized user intent for a virtual assistant;

extracting one or more keywords from the utterance sentence to represent the customized user intent;

generating, via a sentence generation model, preliminary utterance sentences based on the one or more keywords;

training a binary classifier model with supported utterance sentences that are known to invoke the customized user intent;

computing, via the binary classifier model, correctness scores for the preliminary utterance sentences based on probability of whether a preliminary utterance sentence is grammatically or semantically correct;

selecting a plurality of preliminary utterance sentences with correctness scores higher than a threshold;

generating, via the binary classifier model, sample utterance sentences corresponding to the customized user intent based on the plurality of preliminary utterance sentences by mapping the plurality of preliminary utterance sentences to the supported utterance sentences that are known to invoke the customized user intent; and

configuring a voice interaction model of the virtual assistant with the sample utterance sentences, wherein the sample utterance sentences are supported by the voice interaction model to invoke the customized user intent.

2. The computer-implemented method of claim 1 , wherein the utterance sentence comprises one or more spoken phrases that a user can speak to invoke the customized user intent, and wherein the customized user intent invokes one or more defined actions to be performed by the virtual assistant.

3. The computer-implemented method of claim 1 , wherein extracting one or more keywords from the utterance sentence is based on a keyword extraction model.

4. The computer-implemented method of claim 1 , further comprising:

replacing at least one keyword with a placeholder representing a specific type of word.

5. The computer-implemented method of claim 1 , wherein the sentence generation model is a general-purpose natural language generation model finetuned by associated keywords combined with corresponding utterance sentences.

6. The computer-implemented method of claim 1 , wherein the sentence generation model is a general-purpose natural language generation model finetuned by domain-specific datasets.

7. The computer-implemented method of claim 1 , wherein the sentence generation model is a general-purpose natural language generation model finetuned by domain identifiers.

8. The computer-implemented method of claim 1 , wherein the binary classifier model is trained by at least one of positive datasets, negative datasets, and unlabeled datasets.

9. The computer-implemented method of claim 8 , wherein the positive datasets comprise supported utterance sentences combined with the customized user intent, and wherein the supported utterance sentences are configured to invoke the customized user intent.

10. The computer-implemented method of claim 1 , further comprising:

mapping, via the binary classifier model, the selected plurality of preliminary utterance sentences to the customized user intent to generate the sample utterance sentences.

11. A computer-implemented method for virtual assistants, comprising:

receiving an utterance sentence corresponding to an intent for a virtual assistant;

extracting one or more keywords from the utterance sentence to represent the intent;

generating, via a sentence generation model, preliminary utterance sentences based on the one or more keywords;

training a binary classifier model with supported utterance sentences that are known to invoke the customized user intent;

computing, via the binary classifier model, correctness scores for the preliminary utterance sentences based on probability of whether a preliminary utterance sentence is grammatically or semantically correct;

selecting a plurality of preliminary utterance sentences with correctness scores higher than a threshold;

generating sample utterance sentences based on the plurality of preliminary utterance sentences by mapping the plurality of preliminary utterance sentences to the supported utterance sentences that are known to invoke the customized user intent, wherein the sample utterance sentences are generated by the sentence generation model and selected by the binary classifier model; and

configuring the virtual assistant with the sample utterance sentences, wherein the sample utterance sentences are configured to invoke the intent.

12. The computer-implemented method of claim 11 , wherein the virtual assistant further comprises a voice interaction model that supports the sample utterance sentences to invoke the intent.

13. The computer-implemented method of claim 11 , wherein extracting one or more keywords from the utterance sentence is based on a keyword extraction model.

14. The computer-implemented method of claim 11 , further comprising:

replacing at least one keyword with a placeholder representing a specific type of word.

15. The computer-implemented method of claim 11 , wherein the sentence generation model is a general-purpose natural language generation model finetuned by relevant datasets comprising one or more of associated keywords combined with corresponding utterance sentences, domain-specific datasets, and domain identifiers.

16. The computer-implemented method of claim 11 , wherein the binary classifier model is trained by at least one of positive datasets, negative datasets, and unlabeled datasets.

17. The computer-implemented method of claim 16 , wherein the positive datasets comprise supported utterance sentences combined with the intent, and wherein the supported utterance sentences are known to invoke the intent.

18. The computer-implemented method of claim 11 , further comprising:

mapping, via the binary classifier model, the plurality of preliminary utterance sentences to the intent to generate the sample utterance sentences.

19. A computer-implemented method for virtual assistants, comprising:

obtaining one or more keywords associated with an utterance sentence corresponding to a customized user intent for a virtual assistant;

generating, via a sentence generation model, preliminary utterance sentences based on the one or more keywords;

training a binary classifier model with supported utterance sentences that are known to invoke the customized user intent;

computing, via the binary classifier model, correctness scores for the preliminary utterance sentences based on probability of whether a preliminary utterance sentence is grammatically or semantically correct;

selecting a plurality of preliminary utterance sentences with correctness scores higher than a threshold;

generating sample utterance sentences based on the plurality of preliminary utterance sentences by mapping the plurality of preliminary utterance sentences to the supported utterance sentences that are known to invoke the customized user intent, wherein the sample utterance sentences are generated by the sentence generation model and selected by the binary classifier model; and

configuring the virtual assistant with the sample utterance sentences, wherein the sample utterance sentences are configured to invoke the intent.

20. The computer-implemented method of claim 19 , wherein the one or more keywords are extracted from the utterance sentence that is known to invoke the intent by a keyword extraction model.

21. The computer-implemented method of claim 19 , wherein the one or more keywords are specified by a developer from the utterance sentence that is known to invoke the intent.

22. The computer-implemented method of claim 19 , wherein the sentence generation model is a general-purpose natural language generation model finetuned by relevant datasets comprising one or more of associated keywords combined with corresponding utterance sentences, domain-specific datasets, and domain identifiers.

23. The computer-implemented method of claim 19 , further comprising:

mapping, via the binary classifier model, the plurality of preliminary utterance sentences to the intent to generate the sample utterance sentences.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2021
From: SINGH, PRANAV; MOHAJER, KEYVAN; ZHANG, YILUN
To: SOUNDHOUND, INC.
Reel/Frame 058186/0801 →
Continuity (2)
Provisional Application 63198912 · Nov 20, 2020
Related Publication 20220165257A1 · May 26, 2022
References Cited (100)
US 8812316B1 · Chen · 2014 [cited by applicant]
US 10452782B1 · Kumar · 2019 [cited by examiner]
US 10515155B2 · Bachrach et al. · 2019 [cited by applicant]
US 11394799B2 · Jackson · 2022 [cited by third party]
US 11580959B2 · Freed et al. · 2023 [cited by applicant]
US 20020082833A1 · Marasek et al. · 2002 [cited by applicant]
US 20020133340A1 · Basson et al. · 2002 [cited by applicant]
US 20060129381A1 · Wakita · 2006 [cited by applicant]
US 20060195318A1 · Stanglmayr · 2006 [cited by applicant]
US 20070033026A1 · Bartosik et al. · 2007 [cited by applicant]
US 20070208567A1 · Amento et al. · 2007 [cited by applicant]
US 20100106505A1 · Shu · 2010 [cited by applicant]
US 20110218802A1 · Bouganim et al. · 2011 [cited by applicant]
US 20120059838A1 · Berntson et al. · 2012 [cited by applicant]
US 20130018649A1 · Deshmukh · 2013 [cited by examiner]
US 20150179169A1 · John et al. · 2015 [cited by applicant]
US 20160155436A1 · Choi et al. · 2016 [cited by applicant]
US 20160180242A1 · Byron et al. · 2016 [cited by applicant]
US 20160196820A1 · Williams · 2016 [cited by examiner]
US 20170161373A1 · Goyal et al. · 2017 [cited by applicant]
US 20170200458A1 · Kang et al. · 2017 [cited by applicant]
US 20170286869A1 · Zarosim · 2017 [cited by examiner]
US 20180061408A1 · Andreas · 2018 [cited by examiner]
US 20180068653A1 · Trawick et al. · 2018 [cited by applicant]
US 20180121419A1 · Lee · 2018 [cited by examiner]
US 20180329883A1 · Leidner · 2018 [cited by examiner]
US 20190108257A1 · Lefebure et al. · 2019 [cited by applicant]
US 20190147104A1 · Wu et al. · 2019 [cited by applicant]
US 20190155905A1 · Bachrach et al. · 2019 [cited by applicant]
US 20190251165A1 · Bachrach et al. · 2019 [cited by applicant]
US 20190278841A1 · Pusateri et al. · 2019 [cited by applicant]
US 20190370323A1 · Davidson et al. · 2019 [cited by applicant]
US 20200004787A1 · Gupta · 2020 [cited by examiner]
US 20200034357A1 · Panuganty et al. · 2020 [cited by applicant]
US 20200065334A1 · Rodriguez et al. · 2020 [cited by applicant]
US 20200142888A1 · Alakuijala · 2020 [cited by examiner]
US 20200142959A1 · Mallinar · 2020 [cited by examiner]
US 20200143247A1 · Jonnalagadda et al. · 2020 [cited by applicant]
US 20200167379A1 · Faruqui et al. · 2020 [cited by applicant]
US 20200334334A1 · Keskar · 2020 [cited by examiner]
US 20210042614A1 · Walters et al. · 2021 [cited by applicant]
US 20210056266A1 · Ma · 2021 [cited by examiner]
US 20210110816A1 · Choi · 2021 [cited by examiner]
US 20210118436A1 · Kim et al. · 2021 [cited by applicant]
US 20210141798A1 · Henderson · 2021 [cited by applicant]
US 20210141799A1 · Steedman Henderson · 2021 [cited by applicant]
US 20210142164A1 · Liu · 2021 [cited by examiner]
US 20210142789A1 · Gurbani et al. · 2021 [cited by applicant]
US 20210149993A1 · Torres · 2021 [cited by examiner]
US 20210151039A1 · Wu et al. · 2021 [cited by applicant]
US 20210209304A1 · Yang et al. · 2021 [cited by applicant]
US 20210264115A1 · Wang et al. · 2021 [cited by applicant]
US 20210397610A1 · Singh et al. · 2021 [cited by applicant]
US 20220004717A1 · Park et al. · 2022 [cited by applicant]
US 20220059095A1 · Faria et al. · 2022 [cited by applicant]
US 20220093088A1 · Rangarajan Sridhar · 2022 [cited by examiner]
US 20220165257A1 · Singh · 2022 [cited by examiner]
US 20220215159A1 · Qian · 2022 [cited by examiner]
US 20230044079A1 · Bala · 2023 [cited by applicant]
US 20230143110A1 · Han et al. · 2023 [cited by applicant]
US 20230186898A1 · Weisz et al. · 2023 [cited by applicant]
US 20230245649A1 · Singh et al. · 2023 [cited by applicant]
US 20230401251A1 · Majkowska et al. · 2023 [cited by applicant]
EP 3486842A1 · 2019 [cited by applicant]
WO 2018200979A1 · 2018 [cited by applicant]
Li, Zichao, et al., “Decomposable neural paraphrase generation.” arXiv preprint arXiv:1906.09741 (Year: 2019). [cited by examiner]
U.S. Appl. No. 17/455,727, filed Nov. 19, 2021, Ranav Singh. [cited by applicant]
Agnihotri, Souparni. “Hyperparameter Optimization on Neural Machine Translation.” (2019). [cited by applicant]
Cer, Daniel, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant et al. “Universal Sentence Encoder for English.” EMNLP 2018 (2018): 169. [cited by applicant]
Hashemi, Homa B., Amir Asiaee, and Reiner Kraft. “Query intent detection using convolutional neural networks.” In International Conference on Web Search and Data Mining, Workshop on Query Understanding. 2016. [cited by applicant]
Karagkiozis, Nikolaos. “Clustering Semantically Related Questions.” (2019). [cited by applicant]
Klein, Guillaume, François Hernandez, Vincent Nguyen, and Jean Senellart. “The OpenNMT neural machine translation toolkit: 2020 edition.” In Proceedings of the 14th Conference of the Association for Machine Translation … [cited by applicant]
Kriz, Reno, Joao Sedoc, Marianna Apidianaki, Carolina Zheng, Gaurav Kumar, Eleni Miltsakaki, and Chris Callison-Burch. “Complexity-weighted loss and diverse reranking for sentence simplification.” arXiv preprint arXiv:1… [cited by applicant]
Lin, Ting-En, Hua Xu, and Hanlei Zhang. “Discovering new intents via constrained deep adaptive clustering with cluster refinement.” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 05, pp. … [cited by applicant]
OpenNMT, Models—OpenNMT, https://opennmt.net/OpenNMT/training/models/. [cited by applicant]
Reimers, Nils, and Iryna Gurevych. “Sentence-bert: Sentence embeddings using siamese bert-networks.” arXiv preprint arXiv:1908.10084 (2019). [cited by applicant]
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. “Attention is all you need.” arXiv preprint arXiv:1706.03762 (2017). [cited by applicant]
Wang, Peng, Bo Xu, Jiaming Xu, Guanhua Tian, Cheng-Lin Liu, and Hongwei Hao. “Semantic expansion using word embedding clustering and convolutional neural network for improving short text classification.” Neurocomputing … [cited by applicant]
Xu, Jiaming, Bo Xu, Peng Wang, Suncong Zheng, Guanhua Tian, and Jun Zhao. “Self-taught convolutional neural networks for short text clustering.” Neural Networks 88 (2017): 22-31. [cited by applicant]
Lewis, Mike, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. “Bart: Denoising sequence-to-sequence pre-training for natural language generation, transla… [cited by applicant]
Shaw, Andrew. “A multitask music model with bert, transformer-xl and seq2seq.” [cited by applicant]
Geitgey, A. “Faking the News with Natural Language Processing and GPT-2.” Medium. Sep. 27, 2019. [cited by applicant]
Rajapakse, T. “To Distil or Not to Distil: BERT, ROBERTa, and XLNet” towardsdatascience.com. Feb. 7, 2020. [cited by applicant]
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. “Language models are unsupervised multitask learners.” OpenAI blog 1, No. 8 (2019): 9. [cited by applicant]
Zheng, Renjie, Mingbo Ma, and Liang Huang. “Multi-reference training with pseudo-references for neural translation and text generation.” arXiv preprint arXiv:1808.09564 (2018). [cited by applicant]
Juraska, Juraj, and Marilyn Walker. “Characterizing variation in crowd-sourced data for training neural language generators to produce stylistically varied outputs.” arXiv preprint arXiv:1809.05288 (2018). [cited by applicant]
Zhang, Yaoyuan, Zhenxu Ye, Yansong Feng, Dongyan Zhao, and Rui Yan. “A constrained sequence-to-sequence neural model for sentence simplification.” arXiv preprint arXiv:1704.02312 (2017). [cited by applicant]
Wen, Tsung-Hsien, Milica Gasic, Nikola Mrksic, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, and Steve Young. “Multi-domain neural network language generation for spoken dialogue systems.” arXiv preprint arXiv:1603… [cited by applicant]
Tran, Van-Khanh, and Le-Minh Nguyen. “Semantic refinement gru-based neural language generation for spoken dialogue systems.” In International Conference of the Pacific Association for Computational Linguistics, pp. 63-7… [cited by applicant]
Dyer, Chris, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. “Recurrent neural network grammars.” arXiv preprint arXiv:1602.07776 (2016). [cited by applicant]
Li, Zichao, Xin Jiang, Lifeng Shang, and Qun Liu. “Decomposable neural paraphrase generation.” arXiv preprint arXiv:1906.09741 (2019). [cited by applicant]
Guu, Kelvin, Tatsunori B. Hashimoto, Yonatan Oren, and Percy Liang. “Generating sentences by editing prototypes.” Transactions of the Association for Computational Linguistics 6 (2018): 437-450. [cited by applicant]
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention is all you need.” In Advances in neural information processing systems, pp. 5998-… [cited by applicant]
Wen, Tsung-Hsien, Milica Gašic, Nikola Mrkšic, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, and Steve Young. “Toward multi-domain language generation using recurrent neural networks.” In NIPS Workshop on Machine L… [cited by applicant]
Lebret, Rémi, Pedro O. Pinheiro, and Ronan Collobert. “Simple image description generator via a linear phrase-based approach.” arXiv preprint arXiv:1412.8419 (2014). [cited by applicant]
Vinyals, Oriol, Alexander Toshev, Samy Bengio, and Dumitru Erhan. “Show and tell: A neural image caption generator.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3156-3164. 2015. [cited by applicant]
Extended European Search Report of EP21180858.9 by EPO dated Apr. 5, 2022. [cited by applicant]
Examination Report by EPO of European counterpart patent application No. 21180858.9, dated Sep. 10, 2024. [cited by applicant]
JPO office action of counterpart JP patent application No. 2021-103263, dated Dec. 24, 2024. [cited by applicant]
KIPO office action of counterpart Korean patent application No. 2021-0080973, dated Jan. 24, 2025. [cited by applicant]