IP Library › Granted Patent US 12,572,747
Granted Patent B2
US 12,572,747 · App. 18/377,570 · Granted Mar 10, 2026

Multi-turn dialogue response generation with autoregressive transformer models

Inventors: Oluwatobi Olabiyi (Arlington, VA); Erik T. Mueller (Chevy Chase, MD); Rui Zhang (McLean, VA)
Assignee: Capital One Services, LLC
G06F40/30G06F18/2148G06F18/217G06F40/284G06F40/35G06F40/56G06N3/049G06N20/00G10L15/063G10L15/16G10L15/22G10L2015/0631G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,747
App. No.
18/377,570
Filed
Oct 6, 2023
Granted
Mar 10, 2026
Kind
B2
Art Unit
2656
USPC
704/9
Abstract

Machine classifiers in accordance with embodiments of the invention capture long-term temporal dependencies in the dialogue data better than the existing RNN-based architectures. Additionally, machine classifiers may model the joint distribution of the context and response as opposed to the conditional distribution of the response given the context as employed in sequence-to-sequence frameworks. Machine classifiers in accordance with embodiments further append random paddings before and/or after the input data to reduce the syntactic redundancy in the input data, thereby improving the performance of the machine classifiers for a variety of dialogue-related tasks. The random padding of the input data may further provide regularization during the training of the machine classifier and/or reduce exposure bias. In a variety of embodiments, the input data may be encoded based on subword tokenization.

Claims (66)

1 . A computer-implemented method, comprising:

receiving, by a computing device and from a user device, input data comprising a multi-turn dialog corresponding to a plurality of subsequences;

generating an input encoding of the input data, wherein the input encoding comprises token embeddings and position embeddings, wherein the token embeddings comprise separator tokens and a start of sequence token for each subsequence, and the position embeddings indicate an order in which each token appears in the input data;

generating, based on the input encoding, an output sequence comprising:

the start of sequence token; and

one or more output sequence tokens generated by:

providing the input encoding to a machine classifier, wherein the machine classifier comprises a sequence to sequence network comprising a long short-term memory architecture trained based on encoder sequences comprising randomly inserted padding within a portion of each encoder sequence, wherein the randomly inserted padding correspond to a random sampling of encoded tokens in a plurality of training sequences, wherein the machine classifier models a joint distribution of a context of the input data and a response to the input data;

receiving a next output sequence token from the machine classifier; and

appending the next output sequence token to the output sequence until the next output sequence token comprises an end of sequence token; and

outputting, by the computing device, to the user device and based on the output sequence, an automated response to the input data.

2 . The computer-implemented method of claim 1 , further comprising:

training the machine classifier based on the plurality of training sequences comprising an encoder sequence and a decoder sequence corresponding to an encoding of a training sequence.

3 . The computer-implemented method of claim 1 , further comprising:

training the machine classifier based on the plurality of training sequences comprising an encoder sequence and a decoder sequence corresponding to an encoding of a training sequence;

prepending the start of sequence token to the decoder sequence of the encoding; and

appending the end of sequence token to the decoder sequence of the encoding.

4 . The computer-implemented method of claim 1 , wherein the machine classifier comprises the sequence to sequence network architecture comprising an encoder and a decoder.

5 . The computer-implemented method of claim 4 , further comprising:

training the encoder using an encoder sequence corresponding to an encoding of a training sequence; and

training the decoder using a decoder sequence corresponding to the encoding of the training sequence.

6 . The computer-implemented method of claim 1 , wherein the input encoding comprises an attention weight for each token in the input encoding, and

wherein the attention weights are based on a feed-forward analysis using a position of each token in the input encoding.

7 . The computer-implemented method of claim 1 , further comprising:

training an encoder of the machine classifier by updating an attention weight for at least one token in an encoder sequence corresponding to an encoding of a training sequence.

8 . The computer-implemented method of claim 1 , further comprising:

training a decoder of the machine classifier by updating an attention weight for at least one token in a decoder sequence, corresponding to an encoding of a training sequence, to an attention weight associated with at least one token in an encoder sequence, corresponding to the encoding of the training sequence.

9 . The computer-implemented method of claim 1 , wherein an encoding of a sequence comprises a vector representation the sequence.

10 . The computer-implemented method of claim 1 , further comprising:

training the machine classifier based on the plurality of training sequences comprising encoder sequences and decoder sequences, wherein the encoder sequences comprise a set of dialog prompts.

11 . The computer-implemented method of claim 10 , wherein the decoder sequences comprise a set of dialog responses corresponding to the set of dialog prompts.

12 . The computer-implemented method of claim 1 , further comprising:

training the machine classifier based on the plurality of training sequences comprising an encoder sequence and a decoder sequence, wherein the plurality of training sequences comprises a vocabulary, and wherein an encoding for a sequence comprises one hundred percent coverage for the vocabulary.

13 . A device, comprising:

a processor; and

a memory in communication with the processor and storing instructions that, when read by the processor, cause the device to:

receive, from a user device, input data comprising a multi-turn dialog corresponding to a plurality of subsequences;

generate an input encoding of the input data, wherein the input encoding comprises token embeddings and position embeddings, wherein the token embeddings comprise separator tokens and a start of sequence token for each subsequence, and the position embeddings indicate an order in which each token appears in the input data;

generate, based on the input encoding, an output sequence comprising:

the start of sequence token; and

one or more output sequence tokens generated by:

providing the input encoding to a machine classifier, wherein the machine classifier comprises a sequence to sequence network comprising a long short-term memory architecture, wherein the machine classifier models a joint probability distribution of a context of a subsequence in the input data and a response to the subsequence, wherein the machine classifier is trained based on a plurality of training sequences comprising an encoder sequence and a decoder sequence for an encoding of a training sequence, and wherein the encoder sequence comprises randomly inserted padding within a portion of each encoder sequence, wherein the randomly inserted padding corresponds to a random sampling of encoded tokens from the plurality of training sequences;

receiving a next output sequence token from the machine classifier; and

appending the next output sequence token to the output sequence until the next output sequence token comprises an end of sequence token; and

output, to the user device and based on the output sequence, an automated response to the input data.

14 . The device of claim 13 , wherein the machine classifier comprises the sequence to sequence network architecture comprising an encoder and a decoder.

15 . The device of claim 14 , wherein the instructions, when executed by the processor, further cause the device to:

training the encoder using the encoder sequence; and

training the decoder using the decoder sequence.

16 . The device of claim 13 , wherein the instructions, when executed by the processor, further cause the device to:

train an encoder of the machine classifier by updating an attention weight for at least one token in the encoder sequence.

17 . The device of claim 13 , wherein the instructions, when executed by the processor, further cause the device to:

train a decoder of the machine classifier by updating an attention weight for at least one token in the decoder sequence to an attention weight associated with at least one token in the encoder sequence.

18 . The device of claim 13 , wherein an encoding of a sequence comprises a vector representation the sequence.

19 . A non-transitory machine-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform steps comprising:

initializing a machine classifier having a sequence to sequence network architecture comprising an encoder and a decoder;

receiving, from a user device, input data comprising a multi-turn dialog corresponding to a plurality of subsequences;

generating an input encoding of the input data, wherein the input encoding comprises token embeddings and position embeddings, wherein the token embeddings comprise separator tokens and a start of sequence token for each subsequence, and the position embeddings indicate an order in which each token appears in the input data;

generating an output sequence comprising:

the start of sequence token; and

one or more output sequence tokens generated by:

providing the encoding to the machine classifier, wherein the machine classifier comprises a sequence to sequence network comprising a long short-term memory architecture trained based on encoder sequences comprising randomly inserted padding within a portion of each encoder sequence, wherein the randomly inserted padding corresponds to a random sampling of encoded tokens in a plurality of training sequences, and wherein the machine classifier models a joint probability distribution of a context of a subsequence in the input data and a response to the subsequence;

receiving a next output sequence token from the machine classifier; and

appending the next output sequence token to the output sequence until the next output sequence token comprises an end of sequence token; and

outputting, to the user device and based on the output sequence, an automated response to the input data.

20 . The non-transitory machine-readable medium of claim 19 , wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform steps comprising:

training the machine classifier based on the plurality of training sequences comprising an encoder sequence and a decoder sequence corresponding to an encoding of a training sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2025
From: OLABIYI, OLUWATOBI; MUELLER, ERIK T.; ZHANG, RUI
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 073178/0455 →
Continuity (4)
Continuation 18115864 · Mar 1, 2023
Continuation 16935584 · Jul 22, 2020
Provisional Application 62877076 · Jul 22, 2019
Related Publication 20240119233A1 · Apr 11, 2024
References Cited (106)
US 6721706B1 · Strubbe et al. · 2004 [cited by applicant]
US 7853557B2 · Schneider et al. · 2010 [cited by applicant]
US 9473637B1 · Venkatapathy et al. · 2016 [cited by applicant]
US 10019491B1 · Levy · 2018 [cited by applicant]
US 10332513B1 · D'Souza et al. · 2019 [cited by applicant]
US 10664527B1 · Henderson et al. · 2020 [cited by applicant]
US 10860629B1 · Gangadharaiah · 2020 [cited by examiner]
US 10872299B2 · Wayne et al. · 2020 [cited by applicant]
US 10978056B1 · Challa et al. · 2021 [cited by applicant]
US 11200885B1 · Mandal et al. · 2021 [cited by applicant]
US 20090150156A1 · Kennewick et al. · 2009 [cited by applicant]
US 20100312560A1 · Ljolje et al. · 2010 [cited by applicant]
US 20110060587A1 · Phillips et al. · 2011 [cited by applicant]
US 20120221860A1 · Hoornaert et al. · 2012 [cited by applicant]
US 20140222436A1 · Binder et al. · 2014 [cited by applicant]
US 20140358890A1 · Chen et al. · 2014 [cited by applicant]
US 20150100524A1 · Pantel · 2015 [cited by examiner]
US 20160283463A1 · M R et al. · 2016 [cited by applicant]
US 20170300831A1 · Gelfenbeyn et al. · 2017 [cited by applicant]
US 20180040020A1 · Kurian et al. · 2018 [cited by applicant]
US 20180181673A1 · Liu · 2018 [cited by applicant]
US 20180190273A1 · Karimli · 2018 [cited by examiner]
US 20180203852A1 · Goyal et al. · 2018 [cited by applicant]
US 20180293462A1 · Ambati et al. · 2018 [cited by applicant]
US 20180357415A1 · Dhondse et al. · 2018 [cited by applicant]
US 20190057306A1 · Xue et al. · 2019 [cited by applicant]
US 20190188590A1 · Wu et al. · 2019 [cited by applicant]
US 20190266236A1 · Battach et al. · 2019 [cited by applicant]
US 20190287194A1 · Kim · 2019 [cited by examiner]
US 20190341036A1 · Zhang et al. · 2019 [cited by applicant]
US 20190385595A1 · Wabgaonkar et al. · 2019 [cited by applicant]
US 20200027444A1 · Prabhavalkar · 2020 [cited by examiner]
US 20200097814A1 · Devesa · 2020 [cited by applicant]
US 20200099633A1 · D'Agostino et al. · 2020 [cited by applicant]
US 20200125992A1 · Agarwal et al. · 2020 [cited by applicant]
US 20200175374A1 · Hestness et al. · 2020 [cited by applicant]
US 20200202887A1 · Modi et al. · 2020 [cited by applicant]
US 20200218970A1 · Yang · 2020 [cited by examiner]
US 20200401661A1 · Kota · 2020 [cited by examiner]
US 20200410012A1 · Moon et al. · 2020 [cited by applicant]
US 20210027379A1 · Zhu · 2021 [cited by examiner]
US 20210056270A1 · Farhan et al. · 2021 [cited by applicant]
US 20210089877A1 · Shechtman et al. · 2021 [cited by applicant]
US 20210232773A1 · Wang et al. · 2021 [cited by applicant]
US 20210286950A1 · Quamar et al. · 2021 [cited by applicant]
US 20210365636A1 · Li · 2021 [cited by examiner]
US 20230186024A1 · Chen et al. · 2023 [cited by applicant]
CN 110046248A · 2019 [cited by applicant]
CN 110083826A · 2019 [cited by applicant]
Bengio, S., Vinyals, O., Jaitly, N., & Shazeer, N. (2015). Scheduled sampling for sequence prediction with recurrent neural networks Advances in neural information processing systems, 28. (Year: 2015). [cited by examiner]
Mancisidor, R. A., Kampffmeyer, M., Aas, K., & Jenssen, R. (2019). Learning Latent Representations of Bank Customers With The Variational Autoencoder. arXiv preprint arXiv:1903.06580. (Year: 2019). [cited by examiner]
Wu et al. “Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning” Year 2018. [cited by applicant]
Cheng et al, “Long short-term memory-networks for machine reading” arXiv preprint arXiv: 1601.06733, 2016, 10 pages. [cited by applicant]
Chen et al. “Sequential matching model for end-to-end multi-turn response selection” In ICASSP 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) pp. 7350-7354. [cited by applicant]
A.J. Pinheiro et al., “Packet Padding for Improving Privacy in Consumer IoT” 2018 IEEE Symposium on Computers and Communications (ISCC), 2018, pp. 00925-00929, doi: 10.1109/ISCC.2019.8538744. [cited by applicant]
Such R. Ptucha et al. “Intelligent Character Recognition Using Fully Convolution Neural Networks” Pattern Recognition, 88, 604-613 (2019). [cited by applicant]
Iulian V. Serban et al, A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues, arXiv:1605.06069v3 [cs.CL] Jun. 14, 2016. [cited by applicant]
Alec Radford et al, Language Models are Unsupervised Multitask Learners, OpenAI Blog 1.8, 2019. [cited by applicant]
Alec Radford et al, Improving Language Understanding by Generative Pre-Training, 2018. [cited by applicant]
Kishore Papineni et al, BLEU: a Method for Automatic Evaluation of Machine Translation, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Philadelphia, Jul. 2002, pp. 311-318. [cited by applicant]
Oluwatobi Olabiyi et al, Multi-turn Dialogue Response Generation in an Adversarial Learning Framework, arXiv:1805.11752v5 [cs.CL] Jun. 26, 2019. [cited by applicant]
Oluwatobi Olabiyi et al, An Adversarial Learning Framework for a Persona-Based Multi-Turn Dialogue Model, arXiv:1905.01992v2 [cs.CL] Jun. 26, 2019. [cited by applicant]
Oluwatobi O. Olabiyi et al, Adversarial Bootstrapping for Dialogue Model Training, arXiv:1909.00925v2 [cs.CL] Sep. 4, 2019. [cited by applicant]
Oluwatobi O. Olabiyi et al, A Persona-based Multi-turn Conversation Model in an Adversarial Learning Framework, arXiv:1905.01998v1 [cs.CL] Apr. 29, 2019. [cited by applicant]
Ryan Lowe et al, The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems, arXiv:1506.08909v3 [cs.CL] Feb. 4, 2016. [cited by applicant]
Chin-Yew Lin, Rouge: A Package for Automatic Evaluation of Summaries, Association for Computational Linguistics, Jul. 2004. [cited by applicant]
Jiwei Li et al, Deep Reinforcement Learning for Dialogue Generation, arXiv:1606.01541v4 [cs.CL] Sep. 29, 2016. [cited by applicant]
Jiwei Li et al, Adversarial Learning for Neural Dialogue Generation, arXiv:1701.06547v5 [cs.CL] Sep. 24, 2017. [cited by applicant]
Jiwei Li et al, A Diversity-Promoting Objective Function for Neural Conversation Models, arXiv:1510.03055v3 [cs.CL] Jun. 10, 2016. [cited by applicant]
Alex Lamb et al, Professor Forcing: A New Algorithm for Training Recurrent Networks, arXiv:1610.09038v1 [stat.ML] Oct. 27, 2016. [cited by applicant]
Ilya Sutskever et al, Sequence to Sequence Learning with Neural Networks, arXiv:1409.3215v3 [cs.CL] Dec. 14, 2014. [cited by applicant]
Ari Holtzman et al, The Curious Case of Neural Text Degeneration, arXiv:1904.09751v1 [cs.CL] Apr. 22, 2019. [cited by applicant]
Jacob Devlin et al, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv:1810.04805v2 [cs.CL] May 24, 2019. [cited by applicant]
Zihang Dai et al, Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context, arXiv:1901.02860v3 [cs.LG] Jun. 2, 2019. [cited by applicant]
Saizheng Zhang et al, Personalizing Dialogue Agents: I have a dog, do you have pets too?, arXiv:1801.07243v5 [cs.AI] Sep. 25, 2018. [cited by applicant]
Yizhe Zhang et al, Generating Informative and Diverse Conversational Responses via Adversarial Information Maximization, arXiv:1809.05972v5 [cs.CL] Nov. 6, 2018. [cited by applicant]
Rowan Zellers et al, Defending Against Neural Fake News, arXiv:1905.12616v1 [cs.CL] May 29, 2019. [cited by applicant]
Zhilin Yang et al, XLNet: Generalized Autoregressive Pretraining for Language Understanding, arXiv:1906.08237v1 [cs.CL] Jun. 19, 2019. [cited by applicant]
Chen Xing et al, Hierarchical Recurrent Attention Network for Response Generation, arXiv:1701.07149v1 [cs.CL] Jan. 25, 2017. [cited by applicant]
Ronald J. Williams et al, A Learning Algorithm for Continually Running Fully Recurrent Neural Networks, Neural Computation, 1, pp. 270-280, 1989. [cited by applicant]
Oriol Vinyals, A Neural Conversational Model, arXiv:1506.05869v3 [cs.CL] Jul. 22, 2015. [cited by applicant]
Ashish Vaswani et al, Attention Is All You Need, arXiv:1706.03762v5 [cs.CL] Dec. 6, 2017. [cited by applicant]
Yu Sun et al, ERNIE: Enhanced Representation through Knowledge Integration, arXiv:1904.09223v1 [cs.CL] Apr. 19, 2019. [cited by applicant]
Yu Sun et al, ERNIE 2.0: A Continual Pre-Training Framework for Language Understanding, arXiv:1907.12412v1 [cs.CL] Jul. 29, 2019. [cited by applicant]
Iulian Vlad Serban et al, Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation, arXiv:1606.00776v2 [cs.CL] Jun. 14, 2016. [cited by applicant]
Iulian V. Serban et al, Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models, arXiv:1507.04808v3 [cs.CL] Apr. 6, 2016. [cited by applicant]
Pawel Budzianowski et al, Hello, It's GPT-2—How Can I Help You? Towards the Use of Pretrained Language Models for Task-Oriented Dialogue Systems, arXiv:1907.05774v2 [cs.CL] Aug. 4, 2019. [cited by applicant]
Pawel Budzianowski et al, MultiWOZ—A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling, arXiv:1810.00278v3 [cs.CL] Apr. 20, 2020. [cited by applicant]
Whenhu Chen et al, Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention, arXiv:1905.12866v3 [cs.CL] Jun. 9, 2019. [cited by applicant]
Donghoon Ham et al, End-to-End Neural Pipeline for Goal-Oriented Dialogue Systems Using gpt-2, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 583-592, Jul. 5-10, 2020, © 202… [cited by applicant]
Ehsan Hosseini-Asl et al, A Simple Language Model for Task-Oriented Dialogue, arXiv:2005.00796v3 [cs.CL] Jul. 7, 2020. [cited by applicant]
Young-Bum Kim et al, OneNet: Joint Domain, Intent, Slot Prediction for Spoken Language Understanding, arXiv:1801.05149v1 [cs.CL] Jan. 16, 2018. [cited by applicant]
Sungjin Lee et al, ConvLab: Multi-Domain End-to-End Dialog System Platform, arXiv:1904.08637v1 [cs.CL] Apr. 18, 2019. [cited by applicant]
Hwaran Lee et al, SUMBT: Slot-Utterance Matching for Universal and Scalable Belief Tracking, arXiv:1907.07421v1 [cs.CL] Jul. 17, 2019. [cited by applicant]
Shikib Mehri et al, Structured Fusion Networks for Dialog, arXiv:1907.10016v1 [cs.CL] Jul. 23, 2019. [cited by applicant]
Oluwatobi O. Olabiyi et al, DLGNet: A Transformer-based Model for Dialogue Response Generation, arXiv:1908.01841v2 [cs.CL] Sep. 4, 2019. [cited by applicant]
Jiahuan Pei et al, A Modular Task-oriented Dialogue System Using a Neural Mixture-of-Experts, arXiv:1907.05346v1 [cs.CL] Jul. 10, 2019. [cited by applicant]
Baolin Peng et al, Soloist: Few-shot task-oriented dialog with a Single Pre-trained Auto-regressive Model, arXiv:2005.05298v3 [cs.CL] Jun. 22, 2020. [cited by applicant]
Alec Radford et al, Language Models are Unsupervised Multitask Learners, https://d4mucfpksywv.cloudfront.net/better-languagemodels, 2019. [cited by applicant]
Osman Ramadan et al, Large-Scale Multi-Domain Belief Tracking with Knowledge Sharing, arXiv:1807.06517v1 [cs.CL] Jul. 17, 2018. [cited by applicant]
Tsung-Hsien Wen et al, Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems, arXiv:1508.01745v2 [cs.CL] Aug. 26, 2015. [cited by applicant]
Jason Williams et al, The Dialog State Tracking Challenge, Proceedings of the SIGDIAL 2013 Conference, pp. 404-413, Metz, France, Aug. 22-24, 2013, © 2013 Association for Computational Linguistics. [cited by applicant]
Yizhe Zhang et al, DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation, arXiv:1911.00536v3 [cs.CL] May 2, 2020. [cited by applicant]
Tiancheng Zhao et al, Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models, acarXiv:1902.08858v2 [cs.CL] Apr. 15, 2019. [cited by applicant]
Wang, F., “Building high-performance distributed systems with synchronized clocks”, 2019, Available from ProQuest Dissertations and Theses Professional. Retrieved from https://dialog.proquest.com/professional/docview/24… [cited by applicant]
Sordoni, Alessandro, et al. “A hierarchical recurrent encoder-decoder for generative context-aware query suggestion” proceedings of the 24th ACM international on conference on information and knowledge management, 2015,… [cited by applicant]