IP Library Granted Patent US 12682179
Granted Patent B2
US 12682179 · App. 18/620,509 · Granted Jul 14, 2026

Response determination based on contextual attributes and previous conversation content

Inventor: Dino Paul D'Agostino (Richmond Hill, CA)
Assignee: The Toronto-Dominion Bank
G06F40/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682179
App. No.
18/620,509
Granted
Jul 14, 2026
Kind
B2
Abstract

An example operation may include at least one of storing first interaction content with a service provider; receiving second interaction content from a communication session between a source device and a service provider device of the service provider; identifying at least one contextual attribute associated with the source device; determining a response based on execution of at least one large language models (LLMs) on the second interaction content, the at least one contextual attribute associated with the source device, and the first interaction content with the service provider; and outputting the response to at least one of the source device and the service provider device during the communication session.

Claims (37)

1 . An apparatus comprising:

a memory; and

a processor coupled to the memory, the processor configured to:

store first interaction content with a service provider,

receive second interaction content from a communication session between a source device and a service provider device,

identify a plurality of contextual attributes from the second interaction content using a plurality of parallel attention heads of a large language model (LLM) which mask different parts of the second interaction content, during a single execution of the LLM,

determine a response based on execution of a second LLM on the second interaction content, the plurality of contextual attributes, and the first interaction content, and

output the response to at least one of the source device and the service provider device during the communication session.

2 . The apparatus of claim 1 , wherein the processor is configured to record audio from at least one previous call with the service provider, convert the audio from the at least one previous call into a vector, and execute the second LLM on the vector.

3 . The apparatus of claim 1 , wherein the processor is configured to receive real-time second interaction content from the communication session between the source device and the service provider device, convert the real-time second interaction content into a vector, and execute the second LLM on the vector.

4 . The apparatus of claim 1 , wherein the processor is configured to identify an item of interest discussed during the communication session and a sentiment toward the item of interest using the plurality of parallel attention heads.

5 . The apparatus of claim 1 , wherein the processor is configured to receive device data from the source device, and determine the response based on the device data, wherein the device data comprises at least one of a geographical location of the source device, an Internet Protocol (IP) address of the source device, and a type of network connection of the source device.

6 . The apparatus of claim 1 , wherein the processor is configured to determine a first response to display on the source device and determine a second response to display on the service provider device, and simultaneously output the first response to the source device and the second response to the service provider device.

7 . The apparatus of claim 1 , wherein the processor is configured to execute the second LLM at a same time as the communication session occurs between the source device and the service provider device.

8 . The apparatus of claim 1 , wherein the processor is configured to retrieve conversation content from a previous communication session stored in a database based on the plurality of contextual attributes, and determine the response based on the conversation content.

9 . The apparatus of claim 1 , wherein the processor is configured to identify the plurality of contextual attributes from the second interaction content in parallel using a plurality of different subsets of the second interaction content during the single execution of the LLM.

10 . The apparatus of claim 1 , wherein the processor is further configured to retrieve a first portion of the first interaction content while leaving a second portion of the first interaction content, and determine the response based on the retrieved first portion of the interaction content.

11 . A method comprising:

storing first interaction content with a service provider;

receiving second interaction content from a communication session between a source device and a service provider device;

identifying a plurality of contextual attributes from the second interaction content using a plurality of parallel attention heads of a large language model (LLM) which mask different parts of the second interaction content, during a single execution of the LLM;

determining a response based on execution of a second LLM on the second interaction content, the plurality of contextual attributes, and the first interaction content; and

outputting the response to at least one of the source device and the service provider device during the communication session.

12 . The method of claim 11 , wherein the storing comprises recording audio from at least one previous call with the service provider and converting the audio from the at least one previous call into a vector, and the determining comprises executing the second LLM on the vector.

13 . The method of claim 11 , wherein the receiving comprises receiving real-time second interaction content from the communication session between the source device and the service provider device and converting the real-time second interaction content into a vector, and the determining comprises executing the second LLM on the vector.

14 . The method of claim 11 , wherein the identifying comprises identifying an item of interest discussed during the communication session and a sentiment toward the item of interest using the plurality of parallel attention heads.

15 . The method of claim 11 , wherein the method comprises receiving device data from the source device, and the determining the response comprises determining the response based on the device data, wherein the device data comprises at least one of a geographical location of the source device, an Internet Protocol (IP) address of the source device, and a type of network connection of the source device.

16 . The method of claim 11 , wherein the determining comprises determining a first response to display on the source device and determining a second response to display on the service provider device, and the outputting comprises simultaneously outputting the first response to the source device and the second response to the service provider device.

17 . The method of claim 11 , wherein the determining comprises executing the second LLM at a same time as the communication session occurs between the source device and the service provider device.

18 . A non-transitory computer-readable storage medium comprising instructions which when executed by a processor cause the processor to perform:

storing first interaction content with a service provider;

receiving second interaction content from a communication session between a source device and a service provider device of the service provider;

identifying a plurality of contextual attributes from the second interaction content using a plurality of parallel attention heads of a large language model (LLM) which mask different parts of the second interaction content, during a single execution of the LLM;

determining a response based on execution of a second LLM on the second interaction content, the plurality of contextual attributes, and the first interaction content with the service provider; and

outputting the response to at least one of the source device and the service provider device during the communication session.

19 . The non-transitory computer-readable storage medium of claim 18 , wherein the storing comprises recording audio from at least one previous call with the service provider and converting the audio from the at least one previous call into a vector, and the determining comprises executing the second LLM on the vector.

20 . The non-transitory computer-readable storage medium of claim 18 , wherein the receiving comprises receiving real-time second interaction content from the communication session between the source device and the service provider device and converting the real-time second interaction content into a vector, and the determining comprises executing the second LLM on the vector.