IP Library › Granted Patent US 12,566,925
Granted Patent B2
US 12,566,925 · App. 18/343,389 · Granted Mar 3, 2026

Dialogue state aware dialogue summarization

Inventors: Haoliang Wang (Sunnyvale, CA); Kaige Xie (Atlanta, GA); Tong Yu (Fremont, CA); Junda Wu (Long Island City, NY); Handong Zhao (Cupertino, CA); Ruiyi Zhang (San Jose, CA); Kanak Vivek Mahadik (San Jose, CA); Ani Nenkova (Philadelphia, PA)
Assignee: Adobe Inc.
G06F40/35G06F16/345G06F16/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,925
App. No.
18/343,389
Granted
Mar 3, 2026
Kind
B2
Abstract

Dialogue state aware dialogue summarization techniques are described that enable generation of dialogue summaries from target domains with limited training data. A content processing system, for instance, generates one or more clusters based on training dialogues from one or more source domains. The clusters represent domain-specific features of the training dialogues and are further based on dialogue states of the training dialogues. The content processing system trains a machine learning model to generate summaries of dialogues by using the one or more clusters as prefixes in a prefix-tuning approach. The content processing system receives an input that includes a dialogue from a target domain. The content processing system generates an input prompt based on the dialogue and the one or more clusters, and the model generates a summary of the dialogue based on the input prompt.

Claims (35)

1 . A method comprising:

receiving, by a processing device, a machine learning model configured for dialogue summary generation, the machine learning model trained using a prefix-tuning approach in which one or more cluster prefixes corresponding to one or more clusters are optimized during training, the one or more clusters based on dialogue states of training dialogs from one or more source domains;

receiving, by the processing device, an input including a dialogue from a target domain different than the one or more source domains;

generating, by the processing device, an input prompt for the machine learning model that includes an embedding of the dialogue and a cluster prefix from the one or more cluster prefixes; and

processing, by the machine learning model during inferencing, the input prompt to generate a summary of the dialogue as guided by the cluster prefix.

2 . The method as described in claim 1 , wherein each training dialogue includes a plurality of dialogue turns, and the one or more clusters are based on dialogue states of the plurality of dialogue turns.

3 . The method as described in claim 2 , wherein the dialogue states of the plurality of dialogue turns are determined using a pretrained dialogue state tracker and include structured data that represent semantic attributes of the plurality of dialogue turns.

4 . The method as described in claim 1 , wherein the one or more clusters are generated using an unsupervised k-means clustering and represent sentence level topical information of the training dialogues.

5 . The method as described in claim 1 , wherein the generating the input prompt includes generating the embedding of the dialogue based on tokens from the dialogue, generating the cluster prefix based on the one or more clusters, and combining the embedding of the dialogue with the cluster prefix.

6 . The method as described in claim 5 , wherein the machine learning model is trained using the one or more clusters as prefixes in a prefix-tuning approach to learn tunable parameters of a default prefix, and the input prompt includes the embedding of the dialogue, the cluster prefix, and the default prefix generated by the prefix-tuning approach to guide the machine learning model during the inferencing.

7 . The method as described in claim 1 , wherein the dialogue is a document that includes a transcript of a conversation particular to the target domain, and the training dialogues each include a transcript of a conversation particular to the one or more source domains, a training summary, and one or more dialogue state annotations.

8 . The method as described in claim 1 , wherein the one or more clusters are based on a first set of hidden representations that includes embeddings of each sentence of the training dialogues and a second set of hidden representations that includes embeddings of dialogue states of dialogue turns of the training dialogues.

9 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

receiving a machine learning model trained using a prefix-tuning approach in which one or more cluster prefixes corresponding to one or more clusters are optimized during training the one or more clusters based on dialogue states of training dialogs from one or more source domains;

receiving an input including a dialogue from a target domain different than the one or more source domains; and

generating, as part of inferencing by the machine learning model, a summary based on an embedding of the dialogue as guided by a cluster prefix selected from the one or more cluster prefixes.

10 . The system as described in claim 9 , wherein the generating includes one or more default prefixes generated using a prefix-tuning approach to train the machine learning model.

11 . The system as described in claim 9 , wherein the one or more clusters are generated using an unsupervised k-means clustering, the one or more clusters representative of sentence level topical information of the dialogues.

12 . The system as described in claim 9 , wherein each annotated dialogue includes a plurality of dialogue turns, and the one or more clusters are based on a first set of hidden representations that includes embeddings of each dialogue turn.

13 . The system as described in claim 12 , wherein the one or more clusters are based on a second set of hidden representations that includes embeddings of dialogue states of each dialogue turn.

14 . The system as described in claim 13 , wherein the dialogue states are determined as slot-value pairs using a pretrained dialogue state tracker.

15 . The system as described in claim 9 , wherein the dialogue includes a transcript of an ongoing conversation, and generating the summary of the dialogue includes updating the summary based on the transcript of the ongoing conversation.

16 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a training dataset that includes a plurality of training dialogues from one or more source domains;

generating one or more clusters that are based in part on dialogue states of the plurality of training dialogues; and

training a machine learning model to generate summaries of input dialogues from unseen domains using a prefix-tuning approach in which one or more cluster prefixes corresponding to one or more clusters are optimized during training to guide the machine learning model during inferencing.

17 . The non-transitory computer-readable storage medium as described in claim 16 , wherein the generating the one or more clusters includes:

generating a first set of hidden representations for a particular training dialogue by embedding each dialogue turn of the particular training dialogue;

determining dialogue states for each dialogue turn of the particular training dialogue; and

generating a second set of hidden representations for the determined dialogue states, and wherein the one or more clusters are based on the first set of hidden representations and the second set of hidden representations.

18 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the generating the one or more clusters includes combining the first set of hidden representations and the second set of hidden representations to generate a combined set of hidden representations; and performing an unsupervised k-means clustering on the combined set of hidden representations to generate the one or more clusters.

19 . The non-transitory computer-readable storage medium as described in claim 16 , further comprising receiving an input including a dialogue from a target domain different than the one or more source domains and generating a summary of the dialogue using the machine learning model.

20 . The system as described in claim 9 , wherein the cluster prefix includes a sequence of continuous task-specific vectors with one or more parameters optimized during training to guide the machine learning model to perform dialogue summarization on dialogues from unseen target domains during inferencing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2023
From: WANG, HAOLIANG; XIE, KAIGE; YU, TONG; WU, JUNDA; ZHAO, HANDONG; ZHANG, RUIYI; MAHADIK, KANAK VIVEK; NENKOVA, ANI
To: ADOBE INC.
Reel/Frame 064100/0188 →
Continuity (1)
Related Publication 20250005289A1 · Jan 2, 2025
References Cited (35)
US 20090157709A1 · Kruger · 2009 [cited by examiner]
US 20200236068A1 · TeNyenhuis · 2020 [cited by examiner]
US 20220108086A1 · Wu · 2022 [cited by examiner]
US 20240249113A1 · Liu et al. · 2024 [cited by applicant]
US 20240427998A1 · Wang et al. · 2024 [cited by applicant]
US 20250028751A1 · Yu et al. · 2025 [cited by applicant]
Razavi, Unsupervised learning: K-means clustering, https://towardsdatascience.com/unsupervised-learning-k-means-clustering-27416b95af27/, Jun. 27, 2022 (Year: 2022). [cited by examiner]
Artley, Understanding Prefix Tuning: A Novel Approach to Fine-Tuning Language Models, https://medium.com/@razavipour6/understanding-prefix-tuning-a-novel-approach-to-fine-tuning-language-models-dc7dafeb32e4, May 16, 202… [cited by examiner]
Lin et al, Topic-Oriented Dialogue Summarization, May 4, 2023, IEEE ACM Transactions on Audio, Speech, and Language Processing, vol. 31, 2023, pp. 1797-1810 (Year: 2023). [cited by examiner]
Brown, Tom , et al., “Language Models are Few-Shot Learners”, Advances in Neural Information Processing Systems [retrieved Mar. 27, 2023]. Retrieved from the Internet <https://armatech.us/OpenLanding/Language%20Models%2… [cited by applicant]
Chen, Jiaao , et al., “Multi-view sequenceto-sequence models with conversational structure for abstractive dialogue summarization”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Inter… [cited by applicant]
Fabbri, Alexander R, et al., “Improving Zero and Few-Shot Abstractive Summarization with Intermediate Fine-tuning and Data Augmentation”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the… [cited by applicant]
Feng, Xiachong , et al., “A Survey on Dialogue Summarization: Recent Advances and New Frontiers”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2107.03… [cited by applicant]
Goo, Chih-Wen , et al., “Abstractive dialogue summarization with sentence-gated modeling optimized by dialogue acts”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://a… [cited by applicant]
Herskowitz, Nicole , “Microsoft Teams Premium: Cut costs and add AI-powered productivity”, Microsoft Teams [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://www.microsoft.com/en-us/microsoft-365/blog/2023/… [cited by applicant]
Kingma, Diederik P, et al., “Adam: A Method for Stochastic Optimization”, Cornell University, arXiv Preprint, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1412.6980.pdf>., Jan. … [cited by applicant]
Lewis, Mike , et al., “Bart: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from th… [cited by applicant]
Li, Xiang Lisa , “Prefix-Tuning: Optimizing Continuous Prompts for Generation”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2101.00190.pdf>., Jan. 1,… [cited by applicant]
Lin, Chin-Yew , “Rouge: A Package for Automatic Evaluation of Summaries”, Information Sciences Institute University of Southern California [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://aclanthology.org… [cited by applicant]
Liu, Zhengyuan , et al., “Topic-aware pointer-generator networks for summarizing spoken conversations”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1… [cited by applicant]
Rajpurkar, Pranav , et al., “Know What You Don't Know: Unanswerable Questions for SQuAD”, Cornell University arXiv, arXiv.org [retrieved Mar. 24, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1806.03822.pdf>… [cited by applicant]
Reimers, Nils , et al., “Sentence-Bert: Sentence Embeddings using Siamese Bert-Networks”, Cornell University arXiv, arXiv.org [retrieved Jul. 12, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1908.10084.pdf>… [cited by applicant]
Saha, Amrita , et al., “Mining Root Cause Knowledge from Cloud Service Incident Investigations for AIOps”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pd… [cited by applicant]
Samanta, Suranjana , et al., “Carbon to Diamond: An Incident Remediation Assistant System From Site Reliability Engineers' Conversations in Hybrid Cloud Operations”, Proceedings of the AAAI Conference on Artificial Inte… [cited by applicant]
Shetty, Manish , et al., “Neural Knowledge Extraction From Cloud Service Incidents”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2007.05505.pdf>., Ja… [cited by applicant]
Shetty, Manish , et al., “SoftNER: Mining Knowledge Graphs From Cloud Incidents”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2101.05961.pdf>., Jun. … [cited by applicant]
Wolf, Thomas , et al., “Transformers: State-of-the-Art Natural Language Processing”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations [retrieved Mar. 23, 2023… [cited by applicant]
Yu, Tiezheng , et al., “AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive Summarization”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/… [cited by applicant]
Zhao, Lulu , et al., “Domain-Oriented Prefix-Tuning: Towards Efficient and Generalizable Fine-tuning for Zero-Shot Dialogue Summarization”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from t… [cited by applicant]
Zhao, Lulu , et al., “TODSum: Task-Oriented Dialogue Summarization with State Tracking”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2110.12680.pdf>.… [cited by applicant]
Zou, Yicheng , et al., “Low-Resource Dialogue Summarization with Domain-Agnostic Multi-Source Pretraining”, Cornell University arXiv, arXiv.org [retrieved Mar. 23, 2023]. Retrieved from the Internet <https://arxiv.org/p… [cited by applicant]
“Non-Final Office Action”, U.S. Appl. No. 18/339,694, Oct. 1, 2025, 15 pages. [cited by applicant]
“Restriction Requirement”, U.S. Appl. No. 18/355,901, Sep. 16, 2025, 6 pages. [cited by applicant]
Liu, et al., “What Makes Good In-Context Examples for GPT-3?”, arXiv Preprint, Cornell University, Jan. 17, 2021, 12 pages. [cited by applicant]
Shwartz, et al., “Unsupervised Commonsense Question Answering with Self-Talk”, arXiv Preprint, Cornell University, Nov. 2020, 15 pages. [cited by applicant]