IP Library Granted Patent US 9,697,246
Granted Patent B1
US 9,697,246 · App. 14/501,519 · Granted Jul 4, 2017

Themes surfacing for communication data analysis

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,697,246
App. No.
14/501,519
Granted
Jul 4, 2017
Kind
B1
Abstract

An embodiment of the method of processing communication data to identify one or more themes within the communication data includes identifying terms in a set of communication data, wherein a term is a word or short phrase, and defining relations in the set of communication data based on the terms, wherein the relation is a pair of terms that appear in proximity to one another. The method further includes identifying themes in the set of communication data based on the relations, wherein a theme is a group of one or more relations that have similar meanings, and storing the terms, the relations, and the themes in the database.

Claims (50)

1. A method of processing e-communication data by a computer system to identify one or more themes within the communication data, the method comprising:

accessing, by a processing system of a computer system, a set of communication data stored in a storage system of the computer system;

identifying, by the processing system, terms in the set of communication data, wherein a term is a word or short phrase;

defining, by the processing system, relations in the set of communication data based on the terms, wherein a relation is a pair of terms that appear in proximity to one another;

calculating, by the processing system, a relation score for each relation based on a frequency that the terms of the relation appear together in the set of communication data, the number of letters in the terms of the relation, and/or the proximity of the terms to one another, wherein relations that appear relatively frequently in the set of communication data and/or have terms with more letters are given a higher score, wherein the score is lowered for those relations whose terms appear relatively far apart in the set of communication data, wherein scoring each relation includes multiplying a number of times that the terms of the relation appear in the set of communication data by a number of total characters in the terms of the relation, and dividing by 1+an average distance between the terms of the relation as it appears in the set of communication data;

identifying, by the processing system, themes in the set of communication data based on the relations, wherein a theme is a group of one or more relations that have similar meaning; and

storing, by the processing system, the terms, the relations, and the themes in a database.

2. The method of claim 1 further comprising identifying a context vector for each term based on the words appearing before and after that term;

grouping the terms into nodes based on the context vectors, wherein the terms with similar context vectors are put into the same node; and

grouping the relations with the same nodes.

3. The method of claim 2 further comprising assigning a unique node number to each node; and

associating the unique node number with each term grouped into that node;

wherein the step of grouping the relations with the same term nodes includes grouping the relations containing terms with the same node numbers.

4. The method of claim 2 further comprising:

identifying at least one ungrouped term, wherein the ungrouped term is one that is not grouped into the nodes;

determining a similarity score of the ungrouped term to one or more of the terms grouped into the nodes by performing character trigram similarity; and

grouping the ungrouped term into one of the nodes if the similarity score of the ungrouped term to one or more of the terms grouped into that node is at least a threshold similarity score.

5. The method of claim 2 further comprising displaying at least a portion of the grouped relations to a user.

6. The method of claim 2 wherein the step of identifying themes is further based on the grouped relations.

7. The method of claim 1 further comprising eliminating low-scoring relations prior to the step of identifying the themes in the set of communication data.

8. The method of claim 1 further comprising calculating a theme score for each theme by averaging the scores of each relation grouped into that theme, and eliminating low-scoring themes.

9. The method of claim 1 further comprising displaying at least a portion of the themes to a user, prompting the user to eliminate themes.

10. The method of claim 1 further comprising:

calculating a theme score for each theme based on the number of relations associated therewith, the number of unique terms associated therewith, and/or the number of classes associated therewith; and

eliminating themes having a theme score below a threshold.

11. A method of processing communication data by a computer system to identify one or more themes within the communication data, the method comprising:

accessing, by a processing system of a computer system, a set of communication data stored in a storage system of the computer system;

identifying, by the processing system, terms in the set of communication data, wherein a term is a word or short phrase;

defining, by the processing system, relations in the set of communication data based on the terms, wherein a relation is a pair of terms that appear in proximity to one another;

calculating, by the processing system, a relation score for each relation based on a frequency that the terms of the relation appear together in the set of communication data, die number of letters in the terms of the relation, and/or the proximity of the terms to one another, wherein relations that appear relatively frequently in the set of communication data and/or have terms with more letters are given a higher score, wherein the score is lowered for those relations whose terms appear relatively far apart in the set of communication data, wherein scoring each relation includes multiplying a number of times that the terms of the relation appear in the set of communication data by a number of total characters in the terms of the relation, and dividing by 1+an average distance between the terms of the relation as it appears in the set of communication data;

identifying, by the processing system, themes in the set of communication data based on the relations and the relation scores; and

storing, by the processing system, the terms, the relations, scores and the themes in a database.

12. The method of claim 11 further comprising:

calculating a theme score for each theme, wherein the theme score is based on the relation scores of the relations associated with the theme; and

eliminating themes with theme scores below a threshold.

13. The method of claim 12 further comprising:

comparing the themes to a list of important themes, raising the theme score for those themes appearing on the list of important themes.

14. A non-transient computer readable medium programmed with computer readable code that upon execution by a processor causes the processor to execute a method of processing a set of communication data, the method comprising:

accessing a set of communication data;

identifying terms in the set of communication data, wherein a term is a word or short phrase;

defining relations in the set of communication data based on the terms, wherein a relation is a pair of terms that appear in proximity to one another;

identifying a context vector for each term based on the words appearing before and after that term;

grouping the terms with similar context vectors into a node;

grouping relations with the same nodes into groups of relations;

calculating a relation score for each relation based on a frequency that the terms of the relation appear together in the set of communication data, the number of letters in the terms of the relation, and/or die proximity of the terms to one another, wherein relations that appear relatively frequently in the set of communication data and/or have terms with more letters are given a higher score, wherein the score is lowered for those relations whose terms appear relatively far apart in the set of communication data, wherein scoring each relation includes multiplying a number of times that the terms of the relation appear in the set of communication data by a number of total characters in the terms of the relation, and dividing by 1+an average distance between the terms of the relation as it appears in the set of communication data;

identifying theme candidates in the set of communication data based on the groups of relations;

calculating a theme score for each theme candidate;

identifying themes as those theme candidates having at least a threshold theme score; and

storing the terms, the relations, and the themes in a database.

15. The non-transient computer readable medium of claim 14 wherein the theme score is based on the number of relations associated therewith, the number of unique terms associated therewith, and/or the number of classes associated therewith.

Assignments (3)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT SYSTEMS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0919 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2021
From: VERINT SYSTEMS LTD.
To: VERINT SYSTEMS INC.
Reel/Frame 057568/0183 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2014
From: ROMANO, RONI; HORESH, YAIR
To: VERINT SYSTEMS LTD.
Reel/Frame 034572/0687 →