IP Library Granted Patent US 12,197,872
Granted Patent B2
US 12,197,872 · App. 18/167,946 · Granted Jan 14, 2025

Guided text generation for task-oriented dialogue

Inventors: Abhinav Rastogi (Sunnyvale, CA); Mihir Sanjay Kale (Mountain View, CA)
Assignee: Google LLC
G06F40/30G06F3/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,872
App. No.
18/167,946
Granted
Jan 14, 2025
Kind
B2
Abstract

Systems and methods for guided text generation in task-based dialogue. In some aspects of the technology, an automated assistant system is configured to receive a user request, call multiple APIs, generate dialogue acts based on data received from each API, replace any slot names in the dialogue acts with natural language descriptions of the slots, concatenate the modified dialogue acts, and pass the concatenated result to an NLG model for generation of a natural language response. In some aspects of the technology, the automated assistant may be configured to generate simple templated responses based on the data received from each API, concatenate the simple templated responses, and pass the concatenated sequence to an NLG model trained as a sequence-to-sequence transformer for generation of a final natural language response.

Claims (42)

1. A virtual assistant system, comprising:

memory; and

one or more processors coupled to the memory and configured to:

generate, for each given application of a plurality of applications, a templated response including data obtained in response to an application call;

concatenate each templated response generated for each given application of the plurality of applications to create a concatenated sequence; and

generate a natural language response based on the concatenated sequence.

2. The system of claim 1 , wherein the one or more processors are configured to generate the natural language response based on the concatenated sequence using a learned sequence-to-sequence transformer to transform the concatenated sequence into the natural language response.

3. The system of claim 1 , wherein, for each given application of the plurality of applications, the one or more processors are configured to receive the data from the application call to each given application, the data including the templated response.

4. The system of claim 1 , wherein:

for each given application of the plurality of applications, the one or more processors are configured to receive the data from the application call to each given application, the data including at least one template and information, and

the one or more processors are further configured, for each given application of the plurality of applications, to generate the templated response by combining at least the information and the at least one template.

5. The system of claim 4 , wherein, for at least one application of the plurality of applications, the one or more processors are further configured to generate the templated response by combining at least the information, the at least one template, and data based on input received from a user.

6. The system of claim 1 , wherein the one or more processors are further configured to:

select, for each given application of the plurality of applications, at least one template based on the data associated with the application call to that given application, and

wherein generation, for each given application of the plurality of applications, of the templated response includes combining at least the data and the at least one template.

7. The system of claim 6 , wherein, for at least one application of the plurality of applications, the one or more processors are further configured to generate the templated response by combining at least the data, the at least one template, and data based on input received from a user.

8. The system of claim 1 , wherein the one or more processors are further configured to receive input from a user as a text entry, and to provide the natural language response in response to the received input.

9. The system of claim 1 , wherein the one or more processors are further configured to receive input from a user as a verbal command, and to provide the natural language response in response to the received input.

10. The system of claim 1 , wherein the one or more processors are further configured to receive input from a user as a result user interaction with a user interface, and to provide the natural language response in response to the received input.

11. A computer-implemented method, comprising:

generating, by one or more processors of a processing system, for each given application of a plurality of applications, a templated response including data in response to an application call;

concatenating, by the one or more processors, each templated response generated for each given application of the plurality of applications to create a concatenated sequence; and

generating, by the one or more processors, a natural language response based on the concatenated sequence.

12. The method of claim 11 , wherein generating, by the one or more processors, the natural language response based on the concatenated sequence comprises using a learned sequence-to-sequence transformer to transform the concatenated sequence into the natural language response.

13. The method of claim 11 , wherein, for each given application of the plurality of applications, receiving the data from the application call to each given application, the data including the templated response.

14. The method of claim 11 , wherein:

for each given application of the plurality of applications, the data includes at least one template and information, and

generating, for each given application of the plurality of applications, the templated response further comprises combining at least the information and the at least one template.

15. The method of claim 14 , wherein, for at least one application of the plurality of applications, generating the templated response further comprises combining at least the information, the at least one template, and data based on input received from a user.

16. The method of claim 11 , further comprising:

selecting, by the one or more processors, for each given application of the plurality of applications, at least one template based on the data associated with the application call to that given application, and

wherein generating, for each given application of the plurality of applications, the templated response includes combining at least the data and the at least one template.

17. The method of claim 16 , wherein, for at least one given application of the plurality of applications, generating one templated response further comprises combining at least the data, the at least one template, and data based on input received from a user.

18. The method of claim 11 , further comprising:

receiving input from a user via text entry; and

providing the natural language response in response to the received input.

19. The method of claim 11 , further comprising:

receiving input from a user via a verbal command; and

providing the natural language response in response to the received input.

20. The method of claim 11 , further comprising:

receiving input from a user via user interaction with a user interface; and

providing the natural language response in response to the received input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2023
From: RASTOGI, ABHINAV; KALE, MIHIR SANJAY
To: GOOGLE LLC
Reel/Frame 062671/0237 →
Continuity (2)
Continuation 17007270 · Aug 31, 2020
Related Publication 20230186033A1 · Jun 15, 2023
References Cited (54)
US 10558426B2 · Kothari · 2020 [cited by examiner]
US 11200282B1 · Devenny et al. · 2021 [cited by applicant]
US 11604929B2 · Rastogi · 2023 [cited by examiner]
US 20190286480A1 · Park · 2019 [cited by examiner]
US 20200152181A1 · Woo et al. · 2020 [cited by applicant]
US 20210067460A1 · Lillie et al. · 2021 [cited by applicant]
US 20220147702A1 · Li · 2022 [cited by examiner]
Bapna, Ankur , et al., Towards Zero-Shot Frame Semantic Parsing for Domain Scaling, arXiv:1707.02363v1 [cs.Al] Jul. 7, 2017, pp. 1-5. [cited by applicant]
Barzilay, Regina , et al., Sentence Fusion for Multidocument News Summarization, © 2005 Association for Computational Linguistics, vol. 31, No. 3. [cited by applicant]
Budzianowski, Pawel , et al., Hello, It's GPT-2—How Can I Help You? Towards the Use of Pretrained Language Models for Task-Oriented Dialogue Systems, Proceedings of the 3rd Workshop on Neural Generation and Translation … [cited by applicant]
Budzianowski, Pawel , et al., MultiWOZ—A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 50… [cited by applicant]
Chen, Zhlyu , Few-shot NLG with Pre-trained Language Model, Few-shot NLG with Pre-trained Language Model, arXiv:1904.09521v1 [cs.CL] Apr. 21, 2019, pp. 1-10. [cited by applicant]
Chen, Zhiyu , et al., Few-Shot NLG with Pre-Trained Language Model, arXiv:1904.09521v2 [cs.CL] Sep. 6, 2019, pp. 1-8. [cited by applicant]
Chen, Zhiyu , et al., Few-Shot NLG with Pre-Trained Language Model, arXiv:1904.09521v3 [cs.CL] Apr. 19, 2020, pp. 1-8. [cited by applicant]
Chen, Wenhu , et al., Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 3696-3… [cited by applicant]
Devlin, Jacob , et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Proceedings of NAACL-HLT 2019, pp. 4171-4186. [cited by applicant]
Doddington, George , Automatic Evaluation of Machine Translation Quality Using N-gram Co-Occurrence Statistics, pp. 138-145, 2002. [cited by applicant]
Du, Yuheng , et al., Schema-Guided Natural Language Generation, arXiv:2005.05480v1 [cs.CL] May 11, 2020, pp. 1-13. [cited by applicant]
Dusek, Ondrej , Neural Generation for Czech: Data and Baselines, arXiv:1910.05298v1 [cs.CL] Oct. 11, 2019, pp. 1-12. [cited by applicant]
Dusek, Ondrej , et al., A Context-aware Natural Language Generator for Dialogue Systems, Proceedings of the SIGDIAL 2016 Conference, pp. 185-190. [cited by applicant]
Dusek, Ondrej , et al., Semantic Noise Matters for Neural Natural Language Generation, arXiv:1911.03905v1 [cs. CL] Nov. 10, 2019, pp. 1-8. [cited by applicant]
Dusek, Ondrej , et al., Sequence-to-Sequence Generation for Spoken Dialogue via Deep Syntax Trees and Strings, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pp. 45-51, 2016. [cited by applicant]
Gardent, Claire , et al., The WebNLG Challenge: Generating Text from RDF Data, Proceedings of The 10th International Natural Language Generation conference, pp. 124-133, 2017. [cited by applicant]
Gatt, Albert , et al., Survey of the State of the Art in Natural Language Generation: Core Tasks, applications and evaluation, arXiv:1703.09902v4 [cs.CL] Jan. 29, 2018, pp. 1-118. [cited by applicant]
Kale, Mihir , Text-to-Text Pre-Training for Data-to-Text Tasks, arXiv:2005.10433v2 [cs.CL] May 22, 2020, pp. 1-7. [cited by applicant]
Kale, Mihir , et al., Machine Translation Pre-training for Data-to-Text Generation—A Case Study in Czech, arXiv:2004.02077v1 [cs.CL] Apr. 5, 2020, pp. 1-10. [cited by applicant]
Kale, Mihir , et al., Text-to-Text Pre-Training for Data-to-Text Tasks, arXiv:2005.10433v1 [cs.CL] May 21, 2020, pp. 1-7. [cited by applicant]
Kale, Mihir , et al., “Few-Shot Natural Language Generation by Rewriting Templates,”, arXiv:2004.15006v1 [cs.CL] Apr. 30, 2020, pp. 1-12, 2004. [cited by applicant]
Keskar, Nitish Shirish, et al., CTRL: a Conditional Transformer Language Model for Controllable Generation, arXiv:1909.05858v1 [cs.CL] Sep. 11, 2019, pp. 1-18. [cited by applicant]
Keskar, Nitish Shirish, et al., CTRL: a Conditional Transformer Language Model for Controllable Generation, arXiv:1909.05858v2 [cs.CL] Sep. 20, 2019, pp. 1-18. [cited by applicant]
Lavie, Alon , et al., Meteor: An Automatic Metric for MT Evaluation with High Levels of Correlation with Human Judgments, Proceedings of the Second Workshop on Statistical Machine Translation, pp. 228-231,2007. [cited by applicant]
Lebret, Remi , et al., Neural Text Generation from Structured Data with Application to the Biography Domain, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1203-1213, 2016. [cited by applicant]
Lin, Chin-Yew , Rouge: A Package for Automatic Evaluation of Summaries, Information Sciences Institute, University of Southern California, pp. 1-8, 2004. [cited by applicant]
Liu, Yinhan , et al., RoBERTa: A Robustly Optimized BERT Pretraining Approach, arXiv:1907.11692v1 [cs.CL] Jul. 26, 2019, pp. 1-13. [cited by applicant]
Mi, Fei , et al., Meta-Learning for Low-resource Natural Language Generation in Task-oriented Dialogue Systems, Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI-19), pp. … [cited by applicant]
Nayak, Neha , et al., To Plan or not to Plan? Discourse planning in slot-value informed sequence to sequence models for language generation, INTERSPEECH 2017, pp. 3339-3343. [cited by applicant]
Novikova, Jekaterina , et al., The E2E Dataset: New Challenges For End-to-End Generation, Proceedings of the SIGDIAL 2017 Conference, pp. 201-206. [cited by applicant]
Papineni, Kishore , et al., BLEU: a Method for Automatic Evaluation of Machine Translation, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Philadelphia, Jul. 2002, pp. 311… [cited by applicant]
Peng, Baolin , et al., Few-shot Natural Language Generation for Task-Oriented Dialog, arXiv:2002.12328v1 [cs.CL] Feb. 27, 2020, pp. 1-12. [cited by applicant]
Piwek, Paul , A Flexible Pragmatics-driven Language Generator for Animated Agents, 2003, pp. 1-4. [cited by applicant]
Radford, Alec , et al., Language Models are Unsupervised Multitask Learners, pp. 1-24, 2019. [cited by applicant]
Raffel, Colin , et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, arXiv:1910.10683v1 [cs.LG] Oct. 23, 2019, pp. 1-51. [cited by applicant]
Raffel, Colin , et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, arXiv:1910.10683v2 [cs.LG] Oct. 24, 2019, pp. 1-53. [cited by applicant]
Raffel, Colin , et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, arXiv:1910.10683v3 [cs.LG] Jul. 28, 2020, pp. 1-67. [cited by applicant]
Rastogi, Abhinav , et al., Towards Scalable Multi-domain Conversational Agents: The Schema-Guided Dialogue Dataset, arXiv:1909.05855v1 [cs.CL] Sep. 12, 2019, pp. 1-11. [cited by applicant]
Rastogi, Abhinav , et al., Towards Scalable Multi-Domain Conversational Agents: The schema-Guided Dialogue Dataset, arXiv:1909.05855v2 [cs.CL] Jan. 29, 2020, pp. 1-11. [cited by applicant]
Tran, Van-Khanh , et al., Adversarial Domain Adaptation for Variational Neural Language Generation in Dialogue Systems, Proceedings of the 27th International Conference on Computational Linguistics, pp. 1205-1217, 2018. [cited by applicant]
Van Deemter, Kees , et al., Real vs. template-based natural language generation: a false opposition?, 2003 Association for Computational Linguistics, pp. 1-8. [cited by applicant]
Vedantam, Ramakrishna , et al., CIDEr: Consensus-based Image Description Evaluation, CVPR 2015 paper, pp. 4566-4575. [cited by applicant]
Wen, Tsung-Hsien , et al., A Network-based End-to-End Trainable Task-oriented Dialogue System, Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: vol. 1, Long Pa… [cited by applicant]
Wen, Tsung-Hsien , et al., Multi-domain Neural Network Language Generation for Spoken Dialogue Systems, arXiv:1603.01232v1 [cs.CL] Mar. 3, 2016, pp. 1-11. [cited by applicant]
Wen, Tsung-Hsien , et al., Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 1711-17… [cited by applicant]
Yang, Zhillin , et al., XLNet: Generalized Autoregressive Pretraining for Language Understanding, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), pp. 1-11. [cited by applicant]
Zhu, Chenguang , et al., “Multi-task Learning for Natural Language Generation in Task-Oriented Dialogue”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International … [cited by applicant]