IP Library Granted Patent US 11,336,609
Granted Patent B2
US 11,336,609 · App. 17/159,544 · Granted May 17, 2022

System and method for identifying pairs of related application users

Inventors: Offri Gil (Hod-Hasharon, IL); Pinchas Birenbaum (Rehovot, IL); Yitshak Yishay (Revava, IL)
Assignee: COGNYTE TECHNOLOGIES ISRAEL LTD.
H04L51/28G06K9/6257G06N20/00H04L51/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,336,609
App. No.
17/159,544
Granted
May 17, 2022
Kind
B2
Abstract

Systems and methods for passive monitoring of computer communication that does not require performing any decryption. A monitoring system receives the traffic exchanged with each relevant application server, and identifies, in the traffic, sequences of messages—or “n-grams”—that appear to belong to a communication session between a pair of users. Subsequently, based on the numbers and types of identified n-grams, the system identifies each pair of users that are likely to be related to one another via the application, in that these users used the application to communicate (actively and/or passively) with one another. The system may identify those sequences of messages that, by virtue of the sizes of the messages in the sequence, and/or other properties of the messages that are readily discernable, indicate a possible user-pair relationship.

Claims (56)

1. Apparatus, comprising:

a network interface; and

a processor, configured to:

receive a volume of communication traffic that includes a plurality of messages each of which is exchanged between pairs of users of a plurality of users;

identify pairs of users of the plurality of users with collision counts that exceed a first threshold;

identify, in the received volume, at least one sequence of messages that is exchanged between a particular pair of users of the identified pairs of users and follows one of a predetermined message-sequence pattern;

in response to the identifying, calculate a likelihood that the particular pair of the users communicated with one another, and

in response to the likelihood exceeding a second threshold, generate an output that indicates the particular pair of the users.

2. The apparatus according to claim 1 , wherein the messages are encrypted, and wherein the processor is configured to scan the received volume without decrypting any of the messages.

3. The apparatus according to claim 1 , wherein the processor is configured to scan the received volume for any message sequence that follows any one of the predetermined message-sequence patterns by virtue of a property of the message sequence selected from the group of properties consisting of: respective sizes of messages in the message sequence, respective directionalities of the messages in the message sequence, and respective user-endpoints of the messages in the message sequence.

4. The apparatus according to claim 1 , wherein the processor is configured to identify the sequence in response to the sequence spanning a time interval that is less than a third threshold.

5. The apparatus according to claim 4 , wherein the third threshold is a function of a number of round trips, between a server and the particular pair of users, that is implied by the sequence.

6. The apparatus according to claim 1 ,

wherein the processor is configured to identify a plurality of sequences that collectively follow a plurality of different ones of the predetermined message-sequence patterns, and

wherein the processor is configured to calculate the likelihood, using a machine-learned model, based at least on respective numbers of the identified sequences following the different ones of the predetermined message-sequence patterns.

7. The apparatus according to claim 6 , wherein each message is exchanged between a server for an application and one of the plurality of users, and further wherein the volume is a first volume, and wherein the processor is further configured to:

identify a plurality of true message sequences, each of which follows any one of the predetermined message-sequence patterns and is assumed to belong to a communication session between any two users,

generate a second volume of communication traffic, by intermixing a first sequential series of messages exchanged with the server with a second sequential series of messages exchanged with the server,

identify, in the second volume, a plurality of spurious message sequences, each of which follows any one of the predetermined message-sequence patterns and includes at least one message from the first sequential series and at least one message from the second sequential series, and

train the model, using both the true message sequences and the spurious message sequences.

8. The apparatus according to claim 1 , wherein the processor is further configured to, prior to scanning the volume, learn the message-sequence patterns, by identifying a plurality of ground-truth message sequences, each of which follows any one the message-sequence patterns and is assumed to belong to any one of a plurality of communication sessions between one or more other pairs of users.

9. The apparatus according to claim 1 , wherein the processor is further configured to ascertain that each one of the ground-truth message sequences is assumed to belong to one of the communication sessions, by identifying, for each pair of the other pairs of users, a plurality of instances in which a first message destined to a first member of the pair was received within a given time interval of a second message destined to a second member of the pair.

10. The apparatus according to claim 9 ,

wherein each message is exchanged between a server for an application and one of the plurality of users,

wherein the volume is a first volume,

wherein the processor is further configured to:

generate a second volume of communication traffic, by intermixing a first sequential series of messages exchanged with the server with a second sequential series of messages exchanged with the server, and

identify, in the second volume, a plurality of spurious message sequences, each of which includes at least one message from the first sequential series and at least one message from the second sequential series, and

wherein the processor is configured to, in learning the predetermined message-sequence patterns, exclude at least some patterns followed by the spurious message sequences from the predetermined message-sequence patterns, in response to identifying the spurious message sequences.

11. A method, comprising:

receive a volume of communication traffic that includes a plurality of messages each of which is exchanged between pairs of users of a plurality of users,

identify pairs of users of the plurality of users with collision counts that exceed a first threshold,

identify, in the received volume, at least one sequence of messages that is exchanged between a particular pair of users of the identified pairs of users and follows one of a predetermined message-sequence pattern,

in response to the identifying, calculate a likelihood that the particular pair of the users communicated with one another, and

in response to the likelihood exceeding a second threshold, generate an output that indicates the particular pair of the users.

12. The method according to claim 11 , wherein the messages are encrypted, and further comprising scanning the received volume without decrypting any of the messages.

13. The method according to claim 11 , further comprising scanning the received volume for any message sequence that follows any one of the predetermined message-sequence patterns by virtue of a property of the message sequence selected from the group of properties consisting of: respective sizes of messages in the message sequence, respective directionalities of the messages in the message sequence, and respective user-endpoints of the messages in the message sequence.

14. The method according to claim 11 , further comprising identifying the sequence in response to the sequence spanning a time interval that is less than a third threshold.

15. The method according to claim 14 , wherein the third threshold is a function of a number of round trips, between a server and the particular pair of users, that is implied by the sequence.

16. The method according to claim 11 , further comprising:

identifying a plurality of sequences that collectively follow a plurality of different ones of the predetermined message-sequence patterns, and

calculating the likelihood, using a machine-learned model, based at least on respective numbers of the identified sequences following the different ones of the predetermined message-sequence patterns.

17. The method according to claim 16 , wherein each message is exchanged between a server for an application and one of the plurality of users, and further wherein the volume is a first volume, and further comprising:

identifying a plurality of true message sequences, each of which follows any one of the predetermined message-sequence patterns and is assumed to belong to a communication session between any two users,

generating a second volume of communication traffic, by intermixing a first sequential series of messages exchanged with the server with a second sequential series of messages exchanged with the server,

identifying, in the second volume, a plurality of spurious message sequences, each of which follows any one of the predetermined message-sequence patterns and includes at least one message from the first sequential series and at least one message from the second sequential series, and

training the model, using both the true message sequences and the spurious message sequences.

18. The method according to claim 11 , further comprising, prior to scanning the volume, learning the message-sequence patterns, by identifying a plurality of ground-truth message sequences, each of which follows any one the message-sequence patterns and is assumed to belong to any one of a plurality of communication sessions between one or more other pairs of users.

19. The method according to claim 11 , further comprising identifying, for each pair of the other pairs of users, a plurality of instances in which a first message destined to a first member of the pair was received within a given time interval of a second message destined to a second member of the pair.

20. The method according to claim 19 ,

wherein each message is exchanged between a server for an application and one of the plurality of users,

wherein the volume is a first volume,

and further comprising:

generating a second volume of communication traffic, by intermixing a first sequential series of messages exchanged with the server with a second sequential series of messages exchanged with the server, and

identifying, in the second volume, a plurality of spurious message sequences, each of which includes at least one message from the first sequential series and at least one message from the second sequential series, and

in learning the predetermined message-sequence patterns, excluding at least some patterns followed by the spurious message sequences from the predetermined message-sequence patterns, in response to identifying the spurious message sequences.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2022
From: GIL, OFFRI; BIRENBAUM, PINCHAS; YISHAY, YITSHAK
To: VERINT SYSTEMS LTD.
Reel/Frame 059623/0198 →
CHANGE OF NAME Recorded Apr 18, 2022
From: VERINT SYSTEMS LTD.
To: COGNYTE TECHNOLOGIES ISRAEL LTD
Reel/Frame 059725/0287 →
CHANGE OF NAME Recorded Dec 23, 2021
From: VERINT SYSTEMS LTD.
To: COGNYTE TECHNOLOGIES ISRAEL LTD
Reel/Frame 060751/0532 →
Priority Claims (1)
IL 256690 · Jan 1, 2018 · national
Continuity (2)
Continuation 16228929 · Dec 21, 2018
Related Publication 20210152512A1 · May 20, 2021