IP Library Granted Patent US 12701185
Granted Patent B2
US 12701185 · App. 18/521,580 · Granted Aug 4, 2026

Automated conference item selection

Inventor: Nick Swerdlow (Santa Clara, CA)
Assignee: Zoom Communications, Inc.
H04M3/42059G10L15/26H04M3/5175H04M3/5183H04M2201/40H04M2201/42H04M2203/6045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12701185
App. No.
18/521,580
Granted
Aug 4, 2026
Kind
B2
Abstract

Conference item selection is automated by a server that retrieves one or more items from a data store based on a determination that the text includes one or more keywords or a change in the subject of the call. The keywords can include phrases. The retrieved items include one or more of scripts, articles, manuals, daily bulletins regarding a system state, or any resource that can be used to assist with a customer call or interaction. The software running on the server generates a user interface (UI) output based on the retrieve items, and transmits the UI output to an agent device. Software running on the agent device receives the UI output and displays the retrieved items on a display of the agent device.

Claims (54)

1 . A method, comprising:

determining, by a server, that text associated with speech in a call includes one or more keywords;

determining, by the server, a tone of the speech based on a speed of the speech;

determining, by the server, an emotional state of a caller based on the one or more keywords and the tone of the speech, wherein determining the emotional state comprises determining a context of a keyword using a machine learning model that analyzes a neighboring word range of the keyword and selecting a model based on location data of the caller;

retrieving, by the server, one or more items from a data store based on the emotional state of the caller and on at least one of a context of the one or more keywords, a weight assigned to the one or more keywords, a caller attribute, social media data associated with the caller, or spending data associated with the caller;

generating a user interface (UI) output based on an experience level of an agent and the one or more items, wherein the UI output automatically displays a selected script for an agent having a low experience level and the UI output displays a list of the one or more items in a priority order for an agent having a high experience level, and wherein the UI output is automatically adapted in real-time based on a detected change in a subject of the call;

transmitting, by the server, the UI output to an agent device for display; and

monitoring the call for additional keywords to dynamically update the UI output.

2 . The method of claim 1 , wherein determining the tone comprises analyzing a volume of the speech and a pitch of the speech.

3 . The method of claim 1 , wherein the one or more keywords includes a respective weight based on location data of the caller.

4 . The method of claim 1 , wherein retrieving the one or more items from the data store further comprises filtering the items based on a relevance score determined by a probabilistic match between the one or more keywords and data associated with each item.

5 . The method of claim 1 , comprising:

identifying, by the server, a caller identity based on a telephone caller identification (ID).

6 . The method of claim 1 , wherein the server is further configured to retrieve, from the data store, a daily bulletin regarding a system state in a geographic region corresponding to location data of the caller.

7 . The method of claim 1 , comprising:

detecting, by the server, a change in the subject of the call based on a sequence of detected keywords; and

automatically updating the UI output to display a different set of items relevant to a new subject.

8 . A system, comprising:

an agent device; and

a server configured to:

determine that text associated with speech in a call includes one or more keywords;

determine a tone of the speech based on a speed of the speech;

determine an emotional state of a caller based on the one or more keywords and the tone of the speech, wherein the server is configured to determine a context of a keyword using a machine learning model that analyzes a neighboring word range of the keyword and select a model based on location data of the caller;

retrieve one or more items from a data store based on the emotional state of the caller and on at least one of a context of the one or more keywords, a weight assigned to the one or more keywords, a caller attribute, social media data associated with the caller, or spending data associated with the caller;

generate a user interface (UI) output based on an experience level of an agent and the one or more items, wherein the UI output automatically displays a selected script for an agent having a low experience level and the UI output displays a list of the one or more items in a priority order for an agent having a high experience level, and wherein the UI output is automatically adapted in real-time based on a detected change in a subject of the call;

transmit the UI output to the agent device for display; and

monitor the call for additional keywords to dynamically update the UI output.

9 . The system of claim 8 , wherein the server is configured to:

detect an accent of the caller and refine the emotional state using location data.

10 . The system of claim 8 , wherein the one or more keywords includes a respective weight.

11 . The system of claim 8 , wherein the one or more keywords includes a respective weight based on location data of a caller.

12 . The system of claim 8 , wherein a number of social media followers of the caller is correlated with a social media presence strength.

13 . The system of claim 8 , wherein the server is configured to:

obtain location data of the caller; and

retrieve one or more items from the data store based on the location data of the caller.

14 . The system of claim 8 , wherein the server is configured to:

determine a context of the one or more keywords; and

retrieve one or more items from the data store based on the context of the one or more keywords.

15 . A non-transitory computer-readable medium comprising instructions stored on a memory, that when executed by a processor, cause the processor to perform operations comprising:

determining that text associated with speech in a call includes one or more keywords;

determining a tone of the speech based on a speed of the speech;

determining an emotional state of a caller based on the one or more keywords and the tone of the speech, wherein determining the emotional state comprises determining a context of a keyword using a machine learning model that analyzes a neighboring word range of the keyword and selecting a model based on location data of the caller;

retrieving one or more items from a data store based on the emotional state of the caller and on at least one of a context of the one or more keywords, a weight assigned to the one or more keywords, a caller attribute, social media data associated with the caller, or spending data associated with the caller;

generating a user interface (UI) output based on an experience level of an agent and the one or more items, wherein the UI output automatically displays a selected script for an agent having a low experience level and the UI output displays a list of the one or more items in a priority order for an agent having a high experience level, and wherein the UI output is automatically adapted in real-time based on a detected change in a subject of the call;

transmitting the UI output to an agent device for display; and

monitoring the call for additional keywords to dynamically update the UI output.

16 . The non-transitory computer-readable medium of claim 15 , wherein determining the tone comprises analyzing a volume of the speech and a pitch of the speech.

17 . The non-transitory computer-readable medium of claim 15 , wherein retrieving the one or more items from the data store further comprises filtering the items based on a relevance score determined by a probabilistic match between the one or more keywords and data associated with each item.

18 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise retrieving, from the data store, a daily bulletin regarding a system state in a geographic region corresponding to location data of the caller.

19 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

detecting a change in the subject of the call based on a sequence of detected keywords; and

automatically updating the UI output to display a different set of items relevant to a new subject.

20 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

identifying a caller identity based on a telephone caller identification (ID).