IP Library › Granted Patent US 12,400,645
Granted Patent B1
US 12,400,645 · App. 18/082,838 · Granted Aug 26, 2025

Virtual assistant humor management

Inventors: Anjali Yuan Narayan-Chen (Santa Clara, CA); Jiao Sun (Los Angeles, CA); Shereen Oraby (Santa Clara, CA); Shuyang Gao (Sunnyvale, CA); Tagyoung Chung (Palo Alto, CA); Jing Huang (Mountain View, CA); Yang Liu (Los Altos, CA); Nanyung Peng (Los Angeles, CA); Arindam Mandal (Redwood City, CA); Premkumar Natarajan (Rolling Hills Estates, CA)
Assignee: Amazon Technologies, Inc.
G10L15/183G10L13/02G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,645
App. No.
18/082,838
Granted
Aug 26, 2025
Kind
B1
Abstract

In accordance with one disclosed method, first data representing one or more first keywords may be processed by at least one first component to determine a first pair of words suitable for generating a pun. The first pair of words may then be processed using a first machine learning model configured to generate an output pun based at least in part on one or more input keywords and an input word pair such that the output pun is contextually related to the one or more input keywords. A first pun that is contextually related to the one or more first keywords may be received from the first machine learning model, and a first device may be caused to output a representation of the first pun.

Claims (65)

1. A computer-implemented method comprising:

receiving, from a first device, first audio data representing an utterance corresponding to a dialog between a user of the first device and a system;

performing speech processing using the first audio data to determine first data representing a plurality of words in the dialog;

processing the first data using a bidirectional encoder representations from transformers (BERT) model configured to use input data to classify a candidate word pair as suitable for generating a pun that is contextually related to the input data;

using an output of the BERT model to select a first word pair suitable for generating a pun;

processing the first word pair and the first data using a text-to-text transfer transformer (T5) model to generate second data representing a first pun that is contextually related to the first data;

sending the second data to a dialog management component;

sending, from the dialog management component to a speech synthesis component, the second data;

processing the second data by the speech synthesis component to generate second audio data representing the first pun; and

sending the second audio data to the first device.

2. The computer-implemented method of claim 1 , wherein the T5 model generates the first pun based at least in part on a first meaning of a first word of the first word pair and a second meaning of a second word of the first word pair.

3. The computer-implemented method of claim 1 , further comprising, prior to receiving the first audio data:

training the BERT model using training data including third data representing one or more keywords, fourth data representing a second word pair, and an indication whether the second word pair is suitable for generating a pun contextually related to the one or more keywords, with the indication being used to supervise the training.

4. The computer-implemented method of claim 1 , further comprising:

training the T5 model using training data including at least third data representing a text prefix corresponding to pun generation, fourth data representing a second word pair, fifth data representing one or more keywords, and sixth data representing a second pun, with the sixth data being used to supervise the training.

5. A computer-implemented method, comprising:

determining first data representing one or more first keywords;

processing the first data by at least one first component to determine a first word pair suitable for generating a pun;

processing the first word pair and the one or more first keywords using a first machine learning model to generate a first pun that is contextually related to the one or more first keywords, wherein the first machine learning model is trained using training data including at least:

second data representing a text prefix corresponding to pun generation,

third data representing a second word pair,

fourth data representing one or more second keywords, and

fifth data representing a second pun; and

causing a first device to output a representation of the first pun.

6. The computer-implemented method of claim 5 , wherein the first machine learning model generates the first pun based at least in part on a first meaning of a first word of the first word pair and a second meaning of a second word of the first word pair.

7. The computer-implemented method of claim 5 , wherein the first machine learning model comprises a text-to-text transfer transformer model.

8. The computer-implemented method of claim 7 , further comprising:

training the text-to-text transfer transformer model using the training data wherein the fifth data is second pun, wherein the fifth data is used to supervise the training.

9. The computer-implemented method of claim 5 , wherein the at least one first component comprises a second machine learning model that uses the first data to classify respective candidate word pairs as suitable for generating a pun that is contextually related to the first data.

10. The computer-implemented method of claim 9 , wherein the at least one first component further comprises a pun word selector configured to select the first word pair based on an output of the second machine learning model.

11. The computer-implemented method of claim 5 , further comprising:

receiving, by a dialog management component, a first communication;

determining the first data based on the first communication;

generating, by the dialog management component, a response representing the first pun; and

sending, to the first device, a second communication corresponding to the response to cause the first device to output the representation of the first pun.

12. The computer-implemented method of claim 5 , wherein:

the first machine learning model is configured to generate an output pun based at least in part on one or more input keywords and an input word pair such that the output pun is contextually related to the one or more input keywords.

13. A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

determine first data representing one or more first keywords;

process the first data by at least one first component to determine a first word pair suitable for generating a pun;

process the first word pair and the one or more first keywords using a first machine learning model to generate a first pun that is contextually related to the one or more first keywords, wherein the first machine learning model is trained using training data including at least:

second data representing a text prefix corresponding to pun generation,

third data representing a second word pair,

fourth data representing one or more second keywords, and

fifth data representing a second pun;

receive, from the first machine learning model, sixth data representing the first pun; and

cause, based on the sixth data, a first device to output a representation of the first pun.

14. The system of claim 13 , wherein the at least one memory further comprises additional instructions that, when executed by the at least one processor, cause the system to:

process a first meaning of a first word of the first word pair and a second meaning of a second word of the first word pair with the first machine learning model to generate the first pun.

15. The system of claim 13 , wherein the first machine learning model comprises a text-to-text transfer transformer model.

16. The system of claim 15 , wherein the at least one memory further comprises additional instructions that, when executed by the at least one processor, cause the system to:

train the text-to-text transfer transformer model using the training data, wherein the fifth data is used to supervise the training.

17. The system of claim 13 , wherein the at least one first component comprises a second machine learning model, and the at least one memory further comprises additional instructions that, when executed by the at least one processor, cause the system to:

process the first data with the second machine learning model to classify respective candidate word pairs as suitable for generating a pun that is contextually related to the first data.

18. The system of claim 17 , wherein the at least one first component further comprises a pun word selector, and the at least one memory further comprises additional instructions that, when executed by the at least one processor, cause the system to:

use the pun word selector to select the first word pair based on an output of the second machine learning model.

19. The system of claim 13 , wherein the at least one memory further comprises additional instructions that, when executed by the at least one processor, cause the system to:

receive, by a dialog management component, a first communication;

determine the first data based on the first communication;

generate, by the dialog management component and based on the sixth data, a response representing the first pun; and

send, to the first device, a second communication corresponding to the response to cause the first device to output the representation of the first pun.

20. The system of claim 19 , wherein:

the first machine learning model is configured to generate an output pun based at least in part on one or more input keywords and an input word pair such that the output pun is contextually related to the one or more input keywords.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2025
From: SUN, JIAO; ORABY, SHEREEN; GAO, SHUYANG; MANDAL, ARINDAM; NATARAJAN, PREMKUMAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 070039/0908 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: NARAYAN-CHEN, ANJALI YUAN; CHUNG, TAGYOUNG; HUANG, JING; LIU, YANG; PENG, NANYUN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 062126/0721 →
Continuity (1)
Provisional Application 63417882 · Oct 20, 2022
References Cited (82)
US 9858925B2 · Gruber · 2018 [cited by applicant]
US 10037768B1 · Akkiraju · 2018 [cited by applicant]
US 10297273B2 · Akkiraju · 2019 [cited by applicant]
US 10311895B2 · Akkiraju · 2019 [cited by applicant]
US 10424319B2 · Akkiraju · 2019 [cited by applicant]
US 10642939B2 · Toplyn · 2020 [cited by applicant]
US 10878817B2 · Toplyn · 2020 [cited by applicant]
US 11080012B2 · Lemay · 2021 [cited by applicant]
US 11080485B2 · Toplyn · 2021 [cited by applicant]
US 11562142B2 · Nijkamp · 2023 [cited by applicant]
US 12211495B2 · Natarajan · 2025 [cited by applicant]
US 20040088327A1 · Shimomura · 2004 [cited by examiner]
US 20190096425A1 · Akkiraju · 2019 [cited by applicant]
US 20190096426A1 · Akkiraju · 2019 [cited by applicant]
US 20190096427A1 · Akkiraju · 2019 [cited by applicant]
US 20190266250A1 · Toplyn · 2019 [cited by applicant]
US 20190355381A1 · Akkiraju · 2019 [cited by applicant]
US 20200227032A1 · Toplyn · 2020 [cited by applicant]
US 20210103700A1 · Toplyn · 2021 [cited by applicant]
US 20220067077A1 · Tupakula · 2022 [cited by examiner]
US 20220277141A1 · Nijkamp · 2022 [cited by applicant]
Karyn Buxman. 2008. Humor in the OR: A Stitch in Time? AORN Journal, Jul. 2008, vol. 88, No. 1, pp. 67-77. [cited by applicant]
Anne Cutler, et al. 1977. On the role of sentence stress in sentence processing*, Language and Speech, pp. 1-10. [cited by applicant]
Joseph L. Fleiss. 1973. The Equivalence of Weighted Kappa and the Intraclass Correlation Coefficient as Measures of Reliability. Educational and Psychological Measurement, 1973, vol. 33, pp. 613-619. [cited by applicant]
Mikyong Kim, et al. 2000. Patterns of Comprehension and Production of Nouns and Verbs in Agrammatism: Implications for Lexical Organization. Brain and Language vol. 74, pp. 1-25. Retrieved from http://www.idealibrary.co… [cited by applicant]
Jacob L. Moreno. 1955. Theory of Spontaneity-Creativity. Sociometry, vol. 18, No. 4, pp. 105-118. Retrieved from JSTOR, https://www.jstor.org/stable/2785848; or https://doi.org/10.2307/2785848. [cited by applicant]
Stuart Rose, et al. 2010. Automatic keyword extraction from individual documents. Text Mining: Applications and Theory, vol. 1, pp. 1-20. [cited by applicant]
Yufeng Diao, et al. 2019. Heterographic Pun Recognition via Pronunciation and Spelling Understanding Gated Attention Network. The World Wide Web Conference, pp. 363-371.A. [cited by applicant]
Abolfazl Horri. 2011. Linguistic mechanisms of humor: Pun and/or ambiguity. Language Related Research, vol. 2, Issue 2, pp. 19-40. Abstract only, retrieved from https://scholar.google.com/citations?view_op=view_citation… [cited by applicant]
U.S. Appl. No. 18/082,790, filed Dec. 16, 2022. [cited by applicant]
Sura Dhiaa Ibraheem, et al. 2016. Pun and (Un)Intentional Humor. Journal of American Academic Research. Retrieved from https://www.researchgate.net/publication/299524825, pp. 1-18. [cited by applicant]
Jacob Devlin, et al. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of North American Association for Computational Linguistics-Human Language Technologies 2019, p… [cited by applicant]
Aparna Garimella, et al. 2020. Judge me by my size (noun), do you? YodaLib: A Demographic-Aware Humor Generation Framework. In Proceedings of the 28th International Conference on Computational Linguistics (Online), pp. … [cited by applicant]
Md Kamrul Hasan, et al. 2019. UR-FUNNY: A Multimodal Language Dataset for Understanding Humor. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Co… [cited by applicant]
Tatsunori B. Hashimoto, et al. 2019. A Retrieve-and-Edit Framework for Predicting Structured Outputs. 32nd Conference on Neural Information Processing Systems (NIPS 2018). Retrieved from https://dl.acm.org/doi/10.5555/3… [cited by applicant]
He He, et al. 2019. Pun Generaration with Surprise. In Proceedings of NAACL HLT 2019, pp. 1734-1744. Association for Computational Linguistics. [cited by applicant]
Pengcheng He, et al. 2021. DeBERTa: Decoding-Enhanced BERT with Dis-Entangled Attention. Published as a conference paper at ICLR 2021. arXiv preprint, arXiv:2006.03654v6, pp. 1-23. [cited by applicant]
Bryan Anthony Hong, et al. 2009. Automatically Extracting Word Relationships as Templates for Pun Generation. In Proceedings of the NAACL HLT Workshop on Computational Approaches to Linguistic Creativity, pp. 24-31. Ass… [cited by applicant]
I-Hung Hsu, et al. 2022. Degree: A Data-Efficient Generation-Based Event Extraction Model. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang… [cited by applicant]
Kuan-Hao Huang, et al. 2021. Generating Syntactically Controlled Paraphrases without Using Annotated Paralell Pairs. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Lin… [cited by applicant]
Mike Lewis, et al. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational … [cited by applicant]
Yinhan Liu, et al. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXIV:1907.11692v1, 13 pages. [cited by applicant]
Ilya Loshchilov, et al. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019. Retrieved from https://openreview.net/pdf?id=Bkg6RiCqY7, 18 pages. [cited by applicant]
Fuli Luo, et al. 2019. Pun-GAN: Generative Adversarial Network for Pun Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on … [cited by applicant]
Tristan Miller, et al. 2015. Automatic disambiguation of English puns. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Lan… [cited by applicant]
Tristan Miller, et al. 2017. SemEval-2017 Task 7: Detection and Interpretation of English puns. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 58-58. Association for Computa… [cited by applicant]
Anirudh Mittal, et al. 2021. “So You Think You're Funny?”: Rating the Humour Quotient in Standup Comedy. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 10073-10079. Associ… [cited by applicant]
Anirudh Mittal, et al. 2022. AmbiPun: Generating humorous puns with ambiguous context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language… [cited by applicant]
Jeffrey Pennington, et al. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532-1543. Association for Computati… [cited by applicant]
Saša Petrovic, et al. 2013. Unsupervised joke generation from big data. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (vol. 2: Short Papers), pp. 228-232. Association for Com… [cited by applicant]
Colin Raffel, et al. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21 (2020). arXiv preprint arXiv:1910.10683v3, 67 pages. [cited by applicant]
Graeme Ritchie. 2005. Computational Mechanisms for Pun Generation. In Proceedings of the Tenth European Workshop, 8 pages. Association for Computational Linguistics. [cited by applicant]
Jiao Sun, et al. 2021. AESOP: Paraphrase generation with adaptive syntactic control. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 5176-5189. Association for Computationa… [cited by applicant]
Jiao Sun, et al., 2022. ExPUNations: Augmenting Puns with Keywords and Explanations. arXiv preprint arXiv:2210/13513v1, 16 pages. [cited by applicant]
Yufei Tian, et al. 2022. Zero-shot Sonnet Generation with Discourse-level Planning and Aesthetics Features. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Lingui… [cited by applicant]
Alessandro Valitutti, et al. 2013. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, pp. 243-248. Association for Computational Linguistics. [cited by applicant]
Thomas Wolf, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38-45, Online. … [cited by applicant]
Ziqing Yang, et al. 2020. TextBrewer: An Open-Source Knowledge Distillation Toolkit for Natural Language Processing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Dem… [cited by applicant]
Zixiaofan Yang, et al. Julia Hirschberg. 2021. CHoRaL: Collecting Humor Reaction Labels from Millions of Social Media Users. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp.… [cited by applicant]
Zhiwei Yu, et al. 2018. A Neural Approach to Pun Generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), pp. 1650-1660. Association for Computational… [cited by applicant]
Zhiwei Yu, et al. 2020. Homophonic Pun Generation with Lexically Constrained Rewriting. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, pp. 2870-2876. Association for C… [cited by applicant]
Yukun Zhu, et al. 2015. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books, pp. 19-27. arXiv preprint arXiv:1506.06724v1, 23 pages. [cited by applicant]
Satanjeev Banerjee, et al. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machin… [cited by applicant]
Dana-Maria Camburu, et al. 2018. e-SNLI: Natural Language Inference with Natural Language Explanations. In Proceedings of Neural Information Processing Systems. Advances in Neural Information Processing Systems 31 (Neur… [cited by applicant]
Santiago Castro, et al. 2018. A Crowd-Annotated Spanish Corpus for Humor Analysis In Proceedings of the Sixth International Workshop on Natural Language Processing for Social Media, pp. 7-11. Association for Computation… [cited by applicant]
Miruna-Adriana Clinciu, et al. 2021. A Study of Automatic Metrics for the Evaluation of Natural Language Explanations. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational L… [cited by applicant]
Alon Jacovi, et al. 2020. Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 419… [cited by applicant]
Peter Jansen, et al. 2018. WorldTree: A Corpus of Explanation Graphs for Elementary Science Questions supporting Multi-Hop Inference. In Proceedings of the Eleventh International Conference on Language Resources and Eva… [cited by applicant]
Justine T. Kao, et al. 2016. A Computational Model of Linguistic Humor in Puns. Cognitive Science, 40(5), pp. 1270-1285. Retrieved from https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5042108. [cited by applicant]
Maxime Kayser, et al. 2021. e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks. e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks, pp. 1-24.… [cited by applicant]
Sawan Kumar, et al. 2020. NILE : Natural Language Inference with Faithful Natural Language Explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8730-8742. Associa… [cited by applicant]
Wang Ling, et al. 2017. Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, pp. 1… [cited by applicant]
Binny Mathew, et al. 2022. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection*. arXiv preprint arXiv 2012.10289v.2, 12 pages. [cited by applicant]
Kishore Papineni, et al. 2002. BLEU: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311-318. Association for Com… [cited by applicant]
Nazneen Fatema Rajani, et al. 2019. Explain yourself! Leveraging Language Models for Commonsense Reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4932-4942. Asso… [cited by applicant]
Jiao Sun, et al. 2022. Context-Situated Pun Generation. arXiv preprint arXiv:2210.13522v1, 14 pages. [cited by applicant]
Orion Weller, et al. 2019. Humor Detection: A Transformer Gets the Last Laugh. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natu… [cited by applicant]
Sarah Wiegreffe, et al. 2021. Measuring Association Between Labels and Free-Text Rationales. In the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 19 pages. Retrieved from https:… [cited by applicant]
Wangchunshu Zhou, et al. 2022. Towards Interpretable Natural Language Understanding with Explanations as Latent Variables. 34th Conference on Neural Information Processing Systems (NeurIPS 2020), pp. 1-16. arXiv preprin… [cited by applicant]
Yanyan Zou, et al. 2019. Joint Detection and Location of English Puns. In Proceedings of NAACL-HLT 2019, pp. 2117-2123. Association for Computational Linguistics. [cited by applicant]
Office Action mailed Feb. 26, 2025 for U.S. Appl. No. 18/082,790. [cited by applicant]
Final Office Action issued Jun. 6, 2025 for U.S. Appl. No. 18/082,790. [cited by applicant]