IP Library Granted Patent US 12,374,326
Granted Patent B1
US 12,374,326 · App. 18/141,051 · Granted Jul 29, 2025

Natural language generation

Inventors: Alexandros Potamianos (Santa Monica, CA); Arijit Biswas (Dublin, CA); Bonan Zheng (Torrance, CA); Anushree Venkatesh (San Mateo, CA); Yohan Jo (Sunnyvale, CA); Vincent Auvray (Scotts Valley, CA); Nikolaos Malandrakis (San Jose, CA); Aaron Challenner (Melrose, MA); Xinyan Zhao (Seattle, WA); Angeliki Metallinou (Mountain View, CA); David A Jara (Normandy Park, WA); Jiahui Li (Sunnyvale, CA); Ying Shi (Bellevue, WA); Nikko Strom (Kirkland, WA); Veerdhawal Pande (Walpole, MA)
Assignee: Amazon Technologies, Inc.
G10L15/1815G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,326
App. No.
18/141,051
Granted
Jul 29, 2025
Kind
B1
Abstract

Techniques for determining when speech is directed at another individual of a dialog, and storing a representation of such user-directed speech for use as context when processing subsequently-received system-directed speech are described. A system receives audio data and/or video data and determines therefrom that speech in the audio data is user-directed. Based on this, the system determine whether the speech is able to be used to perform an action by the system. If the speech is able to be used to perform an action, the system stores a natural language representation of the speech. Thereafter, when the system receives system-directed speech, the system generates a rewrite of a natural language representation of the system-directed speech based on the previously-received user-directed speech. The system then determines output data responsive to the system-directed speech using the rewritten natural language representation.

Claims (94)

1. A computer-implemented method comprising:

receiving, from a device, first input audio data corresponding to first speech of a first user;

performing automatic speech recognition (ASR) processing using the first input audio data to determine first ASR output data comprising a first transcript of the first speech;

based on the first input audio data and the first ASR output data, determining the first speech is directed at a second user instead of the device;

determining information included in the first speech is usable to respond to speech directed to the device;

based on determining the information included in the first speech is usable to respond to the speech directed to the device, sending the first ASR output data to a storage;

after sending the first ASR output data to the storage, receiving, from the device, second input audio data corresponding to second speech;

performing ASR processing using the second input audio data to determine second ASR output data comprising a second transcript of the second speech;

based on the second ASR output data, determining the second speech is directed to the device;

based on determining the second speech is directed at the device, receiving the first ASR output data from the storage;

after receiving the first ASR output data from the storage, determining rewritten ASR output data corresponding to the second ASR output data updated to include at least one word from the first ASR output data;

performing natural language understanding (NLU) processing using the rewritten ASR output data to determine NLU output data comprising an intent corresponding to the second speech;

based on the NLU output data, determining second output data responsive to the second speech; and

causing presentation of the second output data.

2. The computer-implemented method of claim 1 , further comprising:

receiving first input image data corresponding to at least a first image representing the first user, wherein determining the first speech is directed at the second user instead of the device is further based on the first input image data; and

receiving second input image data corresponding to at least a second image representing the first user, wherein determining the second speech is directed at the device is further based on the second input image data.

3. The computer-implemented method of claim 1 , further comprising:

based on determining the second speech is directed at the device, receiving, from the storage, third ASR output data comprising a third transcript of third speech, wherein the third speech was received prior to the first speech, and the third ASR output data was stored based on the third speech being directed at a third user instead of the device; and

determining the second ASR output data is semantically similar to the first ASR output data instead of the third ASR output data, wherein the rewritten ASR output data includes the at least one word from the first ASR output data instead of the third ASR output data based on the second ASR output data being semantically similar to the first ASR output data instead of the third ASR output data.

4. The computer-implemented method of claim 1 , further comprising:

determining usage data including third ASR output data corresponding to third speech of a dialog;

determining the first speech is semantically similar to the third speech; and

sending the first ASR output data to the storage is further based on the first speech being semantically similar to the third speech.

5. A computer-implemented method comprising:

receiving, from a device, first input audio data corresponding to first speech of a first user;

determining the first speech is directed at the device;

based on determining the first speech is directed at the device, receiving, from a storage, a transcript of second speech, wherein the second speech was received prior to the first speech, and the transcript was stored based on the second speech being directed at a second user instead of the device;

after receiving the transcript from the storage, determining first natural language data corresponding to the first speech updated to include at least one word from the second speech;

performing natural language understanding (NLU) processing using the first natural language data to determine NLU output data comprising a first intent corresponding to at least one of the first speech and the second speech;

based on the NLU output data, determining first output data responsive to at least one of the first speech and the second speech; and

causing presentation of the first output data.

6. The computer-implemented method of claim 5 , further comprising:

prior to receiving the first input audio data, receiving second input audio data corresponding to the second speech;

determining the second speech is directed at the second user instead of the device;

determining information included in the second speech is usable to respond to speech directed to the device; and

based on determining the information included in the second speech is usable to respond to the speech directed to the device, sending the transcript of the second speech to the storage.

7. The computer-implemented method of claim 6 , further comprising:

determining an entity referenced in the second speech; and

determining the entity corresponds to an entity type capable of being processed using NLU processing, wherein determining the information included in the second speech is usable to respond to the speech directed to the device is based on the entity corresponding to the entity type.

8. The computer-implemented method of claim 6 , further comprising:

determining usage data including a second transcript of third speech; and

determining the second speech is semantically similar to the third speech, wherein sending the second speech to the storage is further based on the second speech being semantically similar to the third speech.

9. The computer-implemented method of claim 5 , further comprising:

receiving input image data corresponding to at least a first image representing the first user, wherein:

determining the first speech is directed at the device is further based on the input image data.

10. The computer-implemented method of claim 5 , further comprising:

based on determining the first speech is directed at the device, receiving, from the storage, a second transcript of third speech, wherein the third speech was received prior to the first speech, and the second transcript was stored based on the third speech being directed at a third user instead of the device; and

determining the first speech is semantically similar to the second speech, instead of the third speech, wherein the first natural language data includes the at least one word from the second speech, instead of the third speech based on the first speech being semantically similar to the second speech, instead of the third speech.

11. The computer-implemented method of claim 5 , further comprising:

receiving third input audio data corresponding to third speech;

determining a first portion of the third speech was spoken by the first user;

determining a second portion of the third speech was spoken by the second user;

determining the first portion of the third speech is directed at the device; and

determining the second portion of the third speech is directed at the first user.

12. The computer-implemented method of claim 5 , further comprising:

determining the second speech was spoken by the first user; and

based on the second speech being spoken by the first user, determining the first output data to include a name of the first user.

13. A computing system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the computing system to:

receive, from a device, first input audio data corresponding to first speech of a first user;

determine the first speech is directed at the device;

based on determining the first speech is directed at the device, receive, from a storage, a transcript of second speech, wherein the second speech was received prior to the first speech, and the transcript was stored based on the second speech being directed at a second user instead of the device;

after receiving the transcript from the storage, determine first natural language data corresponding to the first speech updated to include at least one word from the second speech;

perform natural language understanding (NLU) processing using the first natural language data to determine NLU output data comprising a first intent corresponding to at least one of the first speech and the second speech;

based on the NLU output data, determine first output data responsive to at least one of the first speech and the second speech; and

cause presentation of the first output data.

14. The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

prior to receiving the first input audio data, receive second input audio data corresponding to the second speech;

determine the second speech is directed at the second user instead of the device;

determining information included in the second speech is usable to respond to speech directed to the device; and

based on determining the information included in the second speech is usable to respond to the speech directed to the device, send the transcript of the second speech to the storage.

15. The computing system of claim 14 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine an entity referenced in the second speech; and

determine the entity corresponds to an entity type capable of being processed using NLU processing, wherein determining the information included in the second speech is usable to respond to the speech directed to the device is based on the entity corresponding to the entity type.

16. The computing system of claim 14 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine usage data including a second transcript of third speech; and

determine the second speech is semantically similar to the third speech, wherein sending the second speech to the storage is further based on the second speech being semantically similar to the third speech.

17. The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

receive input image data corresponding to at least a first image representing the first user, wherein:

determine the first speech is directed at the device is further based on the input image data.

18. The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

based on determining the first speech is directed at the device, receive, from the storage, a second transcript of third speech, wherein the third speech was received prior to the first speech, and the second transcript was stored based on the third speech being directed at a third user instead of the device; and

determine the first speech is semantically similar to the second speech, instead of the third speech, wherein the first natural language data includes the at least one word from the second speech, instead of the third speech based on the first speech being semantically similar to the second speech, instead of the third speech.

19. The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

receive third input audio data corresponding to third speech;

determine a first portion of the third speech was spoken by the first user;

determine a second portion of the third speech was spoken by the second user;

determine the first portion of the third speech is directed at the device; and

determine the second portion of the third speech is directed at the first user.

20. The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine the second speech was spoken by the first user; and

based on the second speech being spoken by the first user, determine the first output data to include a name of the first user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2024
From: POTAMIANOS, ALEXANDROS; BISWAS, ARIJIT; ZHENG, BONAN; VENKATESH, ANUSHREE; JO, YOHAN; AUVRAY, VINCENT; MALANDRAKIS, NIKOLAOS; CHALLENNER, AARON; ZHAO, XINYAN; METALLINOU, ANGELIKI; JARA, DAVID A; LI, JIAHUI; SHI, YING; STROM, NIKKO; PANDE, VEERDHAWAL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 068050/0909 →
References Cited (34)
US 11250855B1 · Vozila · 2022 [cited by examiner]
US 12112751B2 · Kim · 2024 [cited by examiner]
US 20120035931A1 · LeBeau · 2012 [cited by examiner]
US 20130144616A1 · Bangalore · 2013 [cited by examiner]
US 20230410801A1 · Mishra · 2023 [cited by examiner]
Andreas, et al., “Task-Oriented Dialogue as Dataflow Synthesis”, Transactions of the Association for Computational Linguistics, 2020, vol. 8, p. 556-571. [cited by applicant]
Chen, et al., “DialogSum: A Real-Life Scenario Dialogue Summarization Dataset”, Findings of the Association for Computational Linguistics: ACL-IJCNLP, 2021, p. 5062-5074. [cited by applicant]
Feigenblat, et al., “TWEETSUMM—A Dialog Summarization Dataset for Customer Service”, Findings of the Association for Computational Linguistics: EMNLP, 2021, p. 245-260. [cited by applicant]
Gliwa, et al., “SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization”, Proceedings of the 2nd Workshop on New Frontiers in Summarization, 2019, p. 70-79. [cited by applicant]
Krishna, et al., “Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization Techniques”, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Int… [cited by applicant]
Lewis, et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”, 2019, https://arxiv.org/abs/1910.13461. [cited by applicant]
Li, et al., “DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset”, Proceedings of the Eighth International Joint Conference on Natural Language Processing (vol. 1: Long Papers), Asian Federation of Natural Lang… [cited by applicant]
Lin, “Rouge: A Package for Automatic Evaluation of Summaries”, Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004, p. 74-81. [cited by applicant]
Lin, et al., “Csds: A Fine-Grained Chinese Dataset for Customer Service Dialogue Summarization”, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, p. 4436-4451. [cited by applicant]
Lin, et al., “Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role Interactions”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), 202… [cited by applicant]
Lui, et al., “Automatic Dialogue Summary Generation for Customer Service”, Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019, p. 1957-1965. [cited by applicant]
Porcheron, et al., “Voice Interfaces in Everyday Life”, Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, 2018, p. 1-12. [cited by applicant]
Radford, et al., “Language Models are Unsupervised Multitask Learners”, Technical report, 2019. [cited by applicant]
Raffel, et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, 2019, https://arxiv.org/abs/1910.10683. [cited by applicant]
Rastogi, et al., “Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue Dataset”, Proceedings of the AAAI Conference on Artificial Intelligence, 2020, vol. 34, No. 5, p. 8689-8696. [cited by applicant]
Wolf, et al., “Transformer: State-of-the-art Natural Language Processing”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Association for Computational Lin… [cited by applicant]
Young, et al., “Fusing task-oriented and open-domain dialogues in conversational agents”, 2021, https://arxiv.org/abs/2109.04137. [cited by applicant]
Zang, et al., “MultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselines”, Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, 2020, p. 109-1… [cited by applicant]
Zhang, et al., “EmailSum: Abstractive Email Thread Summarization”, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language … [cited by applicant]
Zhao, et al., “TODSum: Task-Oriented Dialogue Summarization with State Tracking”, 2021, https://arxiv.org/abs/2110.12680. [cited by applicant]
Zhong, et al., “QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langua… [cited by applicant]
Feng, et al., “A Survey on Dialogue Summarization: Recent Advances and New Frontiers”, Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, 2022, p. 5453-5460. [cited by applicant]
Feng, et al., “Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization”, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Con… [cited by applicant]
Lin, et al., “Dialog Simulation with Realistic Variations for Training Goal-Oriented Conversational Systems”, Human in the Loop Dialogue Systems Workshop (HLDS), 2020. [cited by applicant]
Song, et al., “Summarizing Medical Conversations via Identifying Important Utterances”, Proceedings of the 28th International Conference on Computational Linguistics, 2020, p. 717-729. [cited by applicant]
Vakulenko, et al., “Question Rewriting for Conversational Question Answering”, Proceedings of the 14th ACM International Conference on Web Search and Data Mining, 2021, p. 355-363. [cited by applicant]
Yu, et al., “Few-Shot Generative Conversational Query Rewriting”, Proceedings of the 43rd International Acm Sigir Conference on Research and Development in Information Retrieval, 2020, p. 1933-1936. [cited by applicant]
Zamani, et al., “Conversational Information Seeking”, 2022, arXiv:2201.08808. [cited by applicant]
Zhao, et al., “Domain-Oriented Prefix-Tuning: Towards Efficient and Generalizable Fine-tuning for Zero-Shot Dialogue Summarization”, Proceedings of the 2022 Conference of the North American Chapter of the Association fo… [cited by applicant]
Cited By (1)
US 12,499,883