IP Library › Granted Patent US 12,581,016
Granted Patent B2
US 12,581,016 · App. 18/589,594 · Granted Mar 17, 2026

Confirming alignment between contemporaneous visual and audio feedback of the agent during customer calls

Inventors: Nipun Mahajan (Lawrenceville, NJ); Ayela Chughtai (New York, NY); Sushama Shelke (Mumbai, IN); John T. Blackmon (Jacksonville, FL); Yogesh Raghuvanshi (Pennington, NJ); Amit Mishra (Chennai, IN); Saravana Prakash Kumaresan (Plainsboro, NJ)
Assignee: Bank of America Corporation
H04M3/5175G06Q10/063114G06Q10/063116G06V20/46G06V40/176G06V40/28G10L15/02G10L15/1815G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,016
App. No.
18/589,594
Granted
Mar 17, 2026
Kind
B2
Abstract

A non-transitory computer-readable medium may store instructions readable by a processor for destressing an agent in a contract center. The instructions may cause the processor to run a monitoring and triggering application (“MTA”) to identify a state of stress of the agent. The MTA may run an artificial intelligence machine learning (“MTA AI/ML”) algorithm to identify the state of stress in which a first data stream feature (“first feature”) may have a first emotion of the agent that registers above a first feature threshold that corresponds to the first feature, and a second data stream feature (“second feature”) that is contemporaneous with the first feature may have a second emotion of the agent that is aligned with the first emotion of the agent when the first emotion registers above the first feature threshold and the second emotion, contemporaneous with the first emotion, registers above a second feature threshold.

Claims (115)

1 . A system for destressing an agent in a contact center of an enterprise, the system comprising:

a processor that is configured to run an application to identify a state of the agent, said state comprising a state of stress;

wherein:

the application is a monitoring and triggering application (“MTA”);

the MTA is configured to run an artificial intelligence machine learning (“MTA AI/ML”) algorithm to identify the state of stress, said MTA AI/ML algorithm comprising:

multidimensional transformers that extract numerical features:

from a video clip to obtain a first data stream feature (“first feature”) comprising a facial expression feature, a gesture recognition feature, or a combination thereof; and

from a speech sample to obtain a second data stream feature (“second feature”) comprising an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof;

transformer units trained separately for the agent;

weighted decision functions that give differential weight to the first feature and the second feature; and

a contemporaneous output function that validates alignment of the first feature and the second feature using timestamp tracking;

wherein:

the first feature has a first emotion of the agent that registers above a first feature threshold, said first feature threshold corresponding to the first feature;

the second feature that is contemporaneous with the first feature has a second feature emotion of the agent that is aligned with the first emotion of the agent;

the second feature is defined as being contemporaneous with the first feature when the first feature and the second feature have timestamps that are separated by sixty seconds or less from each other;

the MTA subscribes to a statistics server for agent performance data;

upon identifying the state of stress, the MTA contacts an agent scheduling workforce management (WFM) tool to coordinate provision of a destress break to avoid efficiency metric penalties for the agent; and

the MTA communicates with a telephony interaction server to route incoming calls away from the agent and to provide the destress break on a desktop computer of the agent.

2 . The system of claim 1 wherein the processor is a graphics processing unit (“GPU”).

3 . The system of claim 1 wherein the MTA AI/ML algorithm is configured to:

measure the first feature and the second feature during a call; and

identify the state of stress of the agent during the call.

4 . The system of claim 1 wherein the second emotion is defined as being aligned with the first emotion if the first emotion registers above the first feature threshold and the second emotion:

is contemporaneous with the first emotion; and

registers above a second feature threshold.

5 . The system of claim 4 wherein:

the first feature threshold is determined before a call; and

the second feature threshold is determined before the call.

6 . The system of claim 1 wherein the processor is configured to, after identifying that the agent is in the state of stress:

provide notification to a supervisor of the agent that the agent is eligible for a destress break;

receive an approval from the supervisor for the agent to receive the destress break;

after the agent completes a call, route incoming calls away from the agent;

provide on a desktop computer of the agent the destress break, said destress break comprises providing the agent with a video recording, an audio recording, a game, or combination thereof; and,

upon completion of the destress break, route incoming calls to the agent.

7 . The system of claim 1 wherein the processor is configured to transform:

a video clip to obtain the first feature, said first feature comprising a facial expression feature, a gesture recognition feature, or a combination thereof; and

a speech sample to obtain the second feature, said second feature comprises an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof.

8 . The system of claim 1 wherein the processor is configured to transform:

a speech sample to obtain the second feature, said second feature comprises an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof; and

a video clip to obtain the first feature, said first feature comprising a facial expression feature, a gesture recognition feature, or a combination thereof.

9 . The system of claim 1 wherein the second feature is defined as being contemporaneous with the first feature when the first feature and the second feature have timestamps that are separated by sixty seconds or less from each other.

10 . A method for destressing an agent in a contact center of an enterprise, the method comprising:

running, using a processor, an application to identify a state of the agent, said state comprising a state of stress;

wherein:

the application is a monitoring and triggering application (“MTA”);

the MTA runs an artificial intelligence machine learning (“MTA AI/ML”) algorithm to identify the state of stress, said MTA AI/ML algorithm comprising:

multidimensional transformers that extract numerical features:

from a video clip to obtain a first data stream feature (“first feature”) comprising a facial expression feature, a gesture recognition feature, or a combination thereof, and

from a speech sample to obtain a second data stream feature (“second feature”) comprising an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof;

transformer units trained separately for the agent;

weighted decision functions that give differential weight to the first feature and the second feature; and

a contemporaneous output function that validates alignment of the first feature and the second feature using timestamp tracking;

wherein:

the first feature has a first emotion of the agent that registers above a first feature threshold, said first feature threshold corresponding to the first feature;

the second feature that is contemporaneous with the first feature has a second emotion of the agent that is aligned with the first emotion of the agent;

the second feature is defined as being contemporaneous with the first feature when the first feature and the second feature have timestamps that are separated by sixty seconds or less from each other;

the MTA subscribes to a statistics server for agent performance data;

upon identifying the state of stress, the MTA contacts an agent scheduling workforce management (WFM) tool to coordinate provision of a destress break to avoid efficiency metric penalties for the agent; and

the MTA communicates with a telephony interaction server to route incoming calls away from the agent and to provide the destress break on a desktop computer of the agent.

11 . The method of claim 10 wherein the processor is a graphics processing unit (“GPU”).

12 . The method of claim 10 wherein:

the MTA AI/ML algorithm measures the first feature and the second feature during a call; and

the MTA AI/ML algorithm identifies the state of stress of the agent during the call.

13 . The method of claim 10 wherein:

the second emotion is defined as being aligned with the first emotion when the first emotion registers above the first feature threshold and the second emotion, contemporaneous with the first emotion, registers above a second feature threshold;

the first feature threshold is determined before a call;

the second feature threshold is determined before the call; and

the second feature is defined as being contemporaneous with the first feature when the first feature and the second feature have timestamps that are separated by sixty seconds or less from each other.

14 . The method of claim 10 further comprising, after identifying that the agent is in the state of stress:

providing, using the processor, notification to a supervisor of the agent that the agent is eligible for a destress break;

receiving, using the processor, an approval from the supervisor for the agent to receive the destress break;

after the agent completes a call, routing, using the processor, incoming calls away from the agent;

providing, using the processor, on a desktop computer of the agent the destress break, said destress break comprises providing the agent with a video recording, an audio recording, a game, or combination thereof; and

upon completion of the destress break, routing, using the processor, incoming calls to the agent.

15 . The method of claim 10 wherein the processor transforms:

a video clip to obtain the first feature, said first feature comprising a facial expression feature, a gesture recognition feature, or a combination thereof; and

a speech sample to obtain the second feature, said second feature comprises an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof.

16 . The method of claim 10 wherein the processor transforms:

a speech sample to obtain the second feature, said second feature comprises an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof; and

a video clip to obtain the first feature, said first feature comprising a facial expression feature, a gesture recognition feature, or a combination thereof.

17 . A non-transitory computer-readable medium storing instructions readable by a processor for destressing an agent in a contract center of an enterprise, the instructions causing the processor to perform processes comprising:

running, using the processor, an application to identify a state of the agent, said state comprising a state of stress;

wherein:

the application is a monitoring and triggering application (“MTA”);

the MTA runs an artificial intelligence machine learning (“MTA AI/ML”) algorithm to identify the state of stress, said MTA AI/ML algorithm comprising:

multidimensional transformers that extract numerical features:

from a video clip to obtain a first data stream feature (“first feature”) comprising a facial expression feature, a gesture recognition feature, or a combination thereof, and;

from a speech sample to obtain a second data stream feature (“second feature”) comprising an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof;

transformer units trained separately for the agent; weighted decision functions that give differential weight to the first feature and the second feature; and a contemporaneous output function that validates alignment of the first feature and the second feature using timestamp tracking;

wherein:

the first feature has a first emotion of the agent that registers above a first feature threshold, said first feature threshold corresponding to the first feature;

the second feature that is contemporaneous with the first feature has a second emotion of the agent that is aligned with the first emotion of the agent:

the second emotion is aligned with the first emotion when the first emotion registers above the first feature threshold and the second emotion, contemporaneous with the first emotion, registers above a second feature threshold;

the first feature threshold is determined before a call;

the second feature threshold is determined before the call;

the second feature is defined as being contemporaneous with the first feature when the first feature and the second feature have timestamps that are separated by sixty seconds or less from each other;

the MTA AI/ML algorithm measures the first feature and the second feature during the call;

the MTA AI/ML algorithm identifies the state of stress of the agent during the call;

the MTA subscribes to a statistics server for agent performance data; when identifying that the agent is in the state of stress:

providing, using the processor, notification to a supervisor of the agent that the agent is eligible for a destress break;

receiving, using the processor, an approval from the supervisor for the agent to receive the destress break;

contacting, using the processor, an agent scheduling workforce management (WFM) tool to coordinate provision of the destress break to avoid efficiency metric penalties for the agent;

after the agent completes the call, communicating, using the processor, with a telephony interaction server to route incoming calls away from the agent;

providing, using the processor, on a desktop computer of the agent the destress break; and

upon completion of the destress break, communicating, using the processor, with the telephony interaction server to route incoming calls to the agent.

18 . The non-transitory computer-readable medium of claim 17 wherein:

the processor is a graphics processing unit (“GPU”);

the destress break comprises providing the agent with a video recording, an audio recording, a game, or combination thereof; and

the second feature is defined as being contemporaneous with the first feature when the first feature and the second feature have timestamps that are separated by sixty seconds or less from each other.

19 . The non-transitory computer-readable medium of claim 17 wherein the processor is configured to transform:

a video clip to obtain the first feature, said first feature comprising a facial expression feature, a gesture recognition feature, or a combination thereof; and

a speech sample to obtain the second feature, said second feature comprises an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof.

20 . The non-transitory computer-readable medium of claim 17 wherein the processor is configured to transform:

a speech sample to obtain the second feature, said second feature comprises an audio sentiment score feature, a transcript tone score feature, a voice tone score feature, or a combination thereof; and

a video clip to obtain the first feature, said first feature comprising a facial expression feature, a gesture recognition feature, or a combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2024
From: CHUGHTAI, AYELA; MAHAJAN, NIPUN; SHELKE, SUSHAMA; BLACKMON, JOHN T.; RAGHUVANSHI, YOGESH; MISHRA, AMIT; KUMARESAN, SARAVANA PRAKASH
To: BANK OF AMERICA CORPORATION
Reel/Frame 066699/0698 →
Continuity (1)
Related Publication 20250272630A1 · Aug 28, 2025
References Cited (14)
US 11076047B1 · Clodore · 2021 [cited by examiner]
US 20120296642A1 · Shammass · 2012 [cited by examiner]
US 20160180277A1 · Skiba · 2016 [cited by examiner]
US 20170103360A1 · Ristock · 2017 [cited by examiner]
US 20170104872A1 · Ristock · 2017 [cited by examiner]
US 20190158671A1 · Feast · 2019 [cited by examiner]
US 20210397824A1 · Mishra · 2021 [cited by examiner]
US 20220292518A1 · Kilicoglu · 2022 [cited by examiner]
US 20220335752A1 · Alghamdi · 2022 [cited by examiner]
US 20250106322A1 · Sakalkar · 2025 [cited by examiner]
Wegge, Jürgen, Joachim Vogt, and Christiane Wecking. “Customer-induced stress in call centre work: A comparison of audio-and videoconference.” Journal of Occupational & Organizational Psychology 80.4 (2007). (Year: 2007… [cited by examiner]
Płaza, Mirosław, et al. “Emotion recognition method for call/contact centre systems.” Applied Sciences 12.21 (2022): 10951. (Year: 2022). [cited by examiner]
Bromuri, Stefano, et al. “Using AI to predict service agent stress from emotion patterns in service interactions.” Journal of Service Management 32.4 (2021): 581-611. (Year: 2021). [cited by examiner]
Lefter, Iulia, Gertjan J. Burghouts, and Léon JM Rothkrantz. “Recognizing stress using semantics and modulation of speech and gestures.” IEEE Transactions on Affective Computing 7.2 (2015): 162-175. (Year: 2015). [cited by examiner]