IP Library › Granted Patent US 12,388,773
Granted Patent B2
US 12,388,773 · App. 18/643,843 · Granted Aug 12, 2025

Systems and methods for generating conversational responses using machine learning models

Inventors: Minh Le (Bentonville, AR); Sara Mikulic (Herndon, VA)
Assignee: Capital One Services, LLC
H04L51/02G06F18/22G06F18/23G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,388,773
App. No.
18/643,843
Granted
Aug 12, 2025
Kind
B2
Abstract

Methods and systems are described for generating dynamic conversational responses using two-tier machine learning models. The dynamic conversational responses may be generated in real time and reflect the likely goals and/or intents of a user. The two-tier machine learning model may include a first tier that determines an intent cluster based on a feature input, and a second tier that determines a specific intent from the cluster.

Claims (65)

1. A system for generating dynamic conversational responses using intent clusters, the system comprising:

cloud-based storage for:

storing a first artificial intelligence model, wherein the first artificial intelligence model is trained to cluster a plurality of specific intents into a plurality of intent clusters through unsupervised hierarchical clustering; and

storing a second artificial intelligence model, wherein the second artificial intelligence model is trained to select a subset of the plurality of intent clusters from the plurality of intent clusters based on a first feature input and a first user action, and wherein each intent cluster of the plurality of intent clusters corresponds to a respective intent of a user following the first user action; and

cloud-based control circuitry for:

generating, at a user interface, an initial response in response to the user initiating, via the first user action, a conversational interaction;

receiving, at the user interface, a second user action during the conversational interaction;

in response to receiving the second user action, generating an output, wherein the output is generated by the second artificial intelligence model that is trained to select the subset of the plurality of intent clusters from the plurality of intent clusters, and wherein each intent cluster of the plurality of intent clusters corresponds to the respective intent, and wherein the plurality of intent clusters is generated by the first artificial intelligence model that is trained to cluster the plurality of specific intents into the plurality of intent clusters through the unsupervised hierarchical clustering;

selecting, based on the output, a dynamic conversational response from a plurality of dynamic conversational responses that include a respective option for each intent cluster of the subset of the plurality of intent clusters; and

generating, at the user interface, the dynamic conversational response during the conversational interaction.

2. A method for generating dynamic conversational responses using intent clusters, the method comprising:

generating, at a user interface, an initial response in response to a user initiating, via a first user action, a conversational interaction;

receiving, at the user interface, a second user action during the conversational interaction;

in response to receiving the second user action, generating an output, wherein the output is generated by a second artificial intelligence model that is trained to select a subset of a plurality of intent clusters from a plurality of intent clusters, and wherein each intent cluster of the plurality of intent clusters corresponds to a respective intent, and wherein the plurality of intent clusters is generated by a first artificial intelligence model that is trained to cluster a plurality of specific intents into the plurality of intent clusters through unsupervised hierarchical clustering;

selecting, based on the output, a dynamic conversational response from a plurality of dynamic conversational responses that include a respective option for each intent cluster of the subset of the plurality of intent clusters; and

generating, at the user interface, the dynamic conversational response during the conversational interaction.

3. The method of claim 2 , further comprising selecting the second artificial intelligence model, from a plurality of artificial intelligence models, based on the plurality of intent clusters that are retrieved.

4. The method of claim 2 , further comprising:

determining a first feature input based on the second user action; and

inputting the first feature input into the second artificial intelligence model.

5. The method of claim 2 , further comprising:

receiving a third user action during the conversational interaction with the user interface; and

in response to receiving the second user action, determining a third feature input for the second artificial intelligence model based on the third user action.

6. The method of claim 5 , further comprising:

inputting the third feature input into the second artificial intelligence model;

receiving a different output from the second artificial intelligence model; and

selecting, based on the different output, a different dynamic conversational response from the plurality of dynamic conversational responses that corresponds to a different subset of the plurality of intent clusters.

7. The method of claim 2 , wherein the subset of the plurality of intent clusters is further selected based on a screen size of a device generating the user interface.

8. The method of claim 2 , wherein the second artificial intelligence model is a factorization machine model, and wherein the first artificial intelligence model is an artificial neural network model.

9. The method of claim 2 , wherein training the first artificial intelligence model comprises:

generating a matrix of pairwise correlations corresponding to the plurality of specific intents; and

clustering the plurality of specific intents based on pairwise distances.

10. The method of claim 2 , further comprising:

receiving a first labeled feature input, wherein the first labeled feature input is labeled with a known intent cluster for the first labeled feature input; and

training the second artificial intelligence model to classify the first labeled feature input with the known intent cluster.

11. The method of claim 2 , further comprising:

determining a first feature input based on the second user action, wherein the first feature input is a conversational detail; and

processing the conversational detail using the second artificial intelligence model.

12. The method of claim 2 , further comprising:

determining a first feature input based on the second user action, wherein the first feature input is information from a user account of the user; and

processing the information using the second artificial intelligence model.

13. The method of claim 2 , further comprising:

determining a first feature input based on the second user action, wherein the first feature input indicates a time at which the user interface was launched; and

processing the time using the second artificial intelligence model.

14. The method of claim 2 , further comprising:

determining a first feature input based on the second user action, wherein the first feature input indicates a webpage from which the user interface was launched; and

processing the webpage using the second artificial intelligence model.

15. One or more non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause operations comprising:

generating, at a user interface, an initial response in response to a user initiating, via a first user action, a conversational interaction;

receiving, at the user interface, a second user action during the conversational interaction;

in response to receiving the second user action, generating an output, wherein the output is generated by a second artificial intelligence model that is trained to select a subset of a plurality of intent clusters from a plurality of intent clusters, and wherein each intent cluster of the plurality of intent clusters corresponds to a respective intent, and wherein the plurality of intent clusters is generated by a first artificial intelligence model that is trained to cluster a plurality of specific intents into the plurality of intent clusters through unsupervised hierarchical clustering;

selecting, based on the output, a dynamic conversational response from a plurality of dynamic conversational responses that include a respective option for each intent cluster of the subset of the plurality of intent clusters; and

generating, at the user interface, the dynamic conversational response during the conversational interaction.

16. The non-transitory computer-readable media of claim 15 , wherein the instructions that, when executed by the one or more processors, further cause operations comprising selecting the second artificial intelligence model, from a plurality of artificial intelligence models, based on the plurality of intent clusters that are retrieved.

17. The non-transitory computer-readable media of claim 15 , wherein the instructions that, when executed by the one or more processors, further cause operations comprising:

determining a first feature input based on the second user action; and

inputting the first feature input into the second artificial intelligence model.

18. The non-transitory computer-readable media of claim 15 , wherein the instructions that, when executed by the one or more processors, further cause operations comprising:

receiving the second user action during the conversational interaction with the user interface;

in response to receiving the second user action, determining a second feature input for the second artificial intelligence model based on the second user action;

inputting the second feature input into the second artificial intelligence model;

receiving a different output from the second artificial intelligence model; and

selecting, based on the different output, a different dynamic conversational response from the plurality of dynamic conversational responses that corresponds to a different subset of the plurality of intent clusters.

19. The non-transitory computer-readable media of claim 15 , wherein the subset of the plurality of intent clusters is further selected based on a screen size of a device generating the user interface.

20. The non-transitory computer-readable media of claim 15 , wherein the second artificial intelligence model is a factorization machine model, and wherein the first artificial intelligence model is an artificial neural network model.

Continuity (3)
Continuation 17823362 · Aug 30, 2022
Continuation 17029925 · Sep 23, 2020
Related Publication 20240275743A1 · Aug 15, 2024
References Cited (7)
US 20190087691A1 · Jelveh · 2019 [cited by applicant]
US 20190179903A1 · Terry et al. · 2019 [cited by applicant]
US 20210232920A1 · Parangi et al. · 2021 [cited by applicant]
US 20210248457A1 · Odibat et al. · 2021 [cited by applicant]
U.S. Notice of Allowance on U.S. Appl. No. 17/029,861 Dated Nov. 22, 2024 (8 pages). [cited by applicant]
EP Extended European Search Report on EP Appl. Ser. No. 21873310.3 Dated Oct. 2, 2024 (10 pages). [cited by applicant]
Yang Zhiming et al: “Multi-Intent Text Classification Using Dual Channel Convolutional Neural Network”, 2019 34rd Youth Academic Annual Conference of Chinese Association of Automation (YAC), IEEE, Jun. 6, 2019 (Jun. 6, … [cited by applicant]
Cited By (1)
US 12,664,822