IP Library › Granted Patent US 11,888,707
Granted Patent B2
US 11,888,707 · App. 18/088,279 · Granted Jan 30, 2024

Method for detecting anomalies in communications, and corresponding device and computer program product

Inventors: Daniele Ucci (Turin, IT); Filippo Sobrero (Turin, IT); Federica Bisio (Turin, IT)
Assignee: AIZOON S.R.L.
H04L41/16G06F18/29G06N7/01G06N20/10H04L43/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,888,707
App. No.
18/088,279
Granted
Jan 30, 2024
Kind
B2
Abstract

Techniques for detecting anomalies in communication networks are provided. Bayesian networks (first and second) are trained for each feature in first and second lists of features. Third and fourth lists of features are generated and then the first and second Bayesian networks are used to classify each value of the third list of features and of the fourth list of features, respectively, as normal or anomalous. In some examples, a Support Vector Machine can be used for the classification.

Claims (148)

1. A method of detecting anomalies in communications exchanged via a communication network between a respective source and a respective destination, comprising steps of:

obtaining metadata for a plurality of communications in a monitoring interval, wherein said metadata includes for each communication an identifier of said source, an identifier of said destination, and data extracted from an application protocol of the respective communication, wherein said communications comprises Hypertext Transfer Protocol (HTTP) communications;

processing said extracted data to obtain preprocessed data comprising one or more tokens for the respective communication, wherein each token comprises a string, wherein said one or more tokens are extracted from user agent field or referrer field;

dividing said monitoring interval into a training interval and a verification interval;

obtaining the identifier of a given source and generating a first list of a plurality of features (F SRC,TI ) for connections of said given source in said training interval via following steps:

selecting the connections of said given source in said training interval,

determining for said connections of said given source in the said training interval the univocal destination identifiers and for each token the respective univocal values,

determining a first set of enumeration rules by enumerating said univocal destination identifiers and for each token the respective univocal values, and

associating by means of said first set of enumeration rules with each connection of said source in said training interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said first list of features comprises for each connection of said given source in said training interval the respective enumerated destination identifier and the respective one or more enumerated tokens;

obtaining the identifier of a group of devices to which said given source belongs and generating a second list of a plurality of features for the connections of the devices belonging to said group of devices in said training interval via following steps:

selecting the connections of said group of devices in said training interval,

determining for said connections of said group of devices in said training range the univocal destination identifiers and for each token the respective univocal values,

determining a second set of enumeration rules by enumerating said univocal destination identifiers and for each token the respective univocal values, and

associating by means of said second set of enumeration rules with each connection of said group of devices in said training interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said second list of features comprises for each connection of said group of devices in said training interval the respective enumerated destination identifier and the respective one or more enumerated tokens;

generating a first set of Bayesian networks by training for each feature of said first list of features a respective Bayesian network using the data of other features of said first list of features (F SRC,TI ), and generating a second set of Bayesian networks by training for each feature of said second list of features a respective Bayesian network using the data of the other features of said second list of features,

generating a third list of a plurality of features for the connections of said given source in said verification interval via following steps:

selecting the connections of said given source in said verification interval, and

associating by means of said first set of enumeration rules with each connection of said given source in said verification interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said third list of features comprises for each connection of said given source in said verification interval the respective enumerated destination identifier and the respective one or more respective enumerated tokens;

generating a fourth list of a plurality of features for connections of said given source in said verification interval via following steps:

selecting the connections of said given source in said verification interval,

associating by means of said second set of enumeration rules with each connection of said given source in said verification interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said fourth list of features comprises for each connection of said given source in said verification interval the respective enumerated destination identifier and the respective one or more respective enumerated tokens, wherein said first list of a plurality of features, said second list of a plurality of features, said third list of a plurality of features and said fourth list of a plurality of features further comprise at least one of: an enumerated value generated for a destination port of Transmission Control Protocol (TCP) or the User Datagram Protocol (UDP) of the respective communication;

repeating following steps for each connection of said given source in said verification interval:

determining based on the values of the features of said third list of features associated with the respective connection of said given source for each feature of said third list of features the respective most probable value by using said first set of Bayesian networks,

classifying each value of the features of said third list of features associated with the respective connection of said given source via following steps:

in response to determining the value of a feature of said third list of features corresponds to the respective most probable value, classifying the value of the feature of said third list of features as normal, and

in response to determining the value of a feature of said third list of features does not correspond to the respective most probable value:

a) determining for the value of said feature of said third list of features the respective probability of occurrence by using said first set of Bayesian networks, and

b) classifying the value of said feature of said third list of features as normal in response to determining the respective probability of occurrence is greater than a first threshold, and

c) classifying the value of said feature of said third list of features as anomalous in response to determining the respective probability of occurrence is smaller than said first threshold; and

determining based on the values of the feature values of said fourth list of features associated with the respective connection of said given source for each feature of said fourth list of features the respective most probable value by using said second set of Bayesian networks, wherein discretizing one or more of the features of said first list of a plurality of features, said second list of features of a plurality of features, said third list of a plurality of features, and said fourth list of a plurality of features by means of a clustering algorithm, a k-means clustering algorithm, and

classifying each value of the features of said fourth list of features associated with the respective connection of said given source via following steps:

in response to determining the value of a feature of said fourth list of features corresponds to the respective most probable value, classifying the value of the feature of said fourth list of features as normal, and

in response to determining the value of a feature of said fourth list of features does not correspond to the respective most probable value:

a) determining for the value of said feature of said fourth list of features the respective probability of occurrence by using said second set of Bayesian networks, and

b) classifying the value of said feature of said fourth list of features as normal in response to determining the respective probability of occurrence is greater than a second threshold, and

c) classifying the value of said feature of said fourth list of features as anomalous in response to determining the respective probability of occurrence is smaller than said second threshold.

2. The method according to claim 1 , comprising:

repeating following steps for each connection of said given source in said verification interval:

determining a first number of values of the features of said third list of features associated with the respective connection of said given source that are classified as anomalous,

determining a second number of values of the features of said fourth list of features associated with the respective connection of said given source that are classified as anomalous, and

classifying the connection of said given source as anomalous if the first number or the second number is greater than a third threshold.

3. The method according to claim 2 , comprising:

training a first single-class Support Vector Machine, SVM, by using said first list of features,

repeating following steps for each connection of said given source in said verification interval:

classifying the values of the features of said third list of features associated with the respective connection of said given source as normal or anomalous by using said first SVM, and

classifying the connection of said given source as suspicious if the connection of said given source is classified as anomalous and the values of the features of said third list of features associated with the respective connection of said given source are classified as anomalous by said first SVM.

4. The method according to claim 2 , comprising:

training a second single-class SVM by using said second list of features,

repeating following steps for each connection of said given source in said verification interval:

classifying the values of the feature of said fourth list of features associated with the respective connection of said given source as normal or anomalous by using said second SVM, and

classifying the connection of said given source as suspicious if the connection of said given source is classified as anomalous and the feature values of said fourth list of features associated with the respective connection of said given source are classified as anomalous by said second SVM.

5. The method according to claim 1 , comprising:

repeating following steps for each connection of said given source in said verification interval:

determining a first average value of the probabilities of occurrence of the values of the features in said third feature list associated with the respective connection of said given source that are classified as anomalous,

determining a second average value of the probability of occurrence of the values of the features of said fourth feature list associated with the respective connection of said given source that are classified as anomalous, and

classifying the connection of said given source as anomalous if the first average value and/or the second average value is smaller than a fourth threshold.

6. The method according to claim 1 , comprising:

repeating following steps for each source of a plurality of sources:

obtaining a respective identifier of the respective source and generating a respective first list of a plurality of features for the connections of the respective source in said training interval,

calculating for each feature of the respective first list of a plurality of features a respective average value, thereby generating a fifth list of features comprising for each source the respective average values of the features of the respective first list of a plurality of features, and

generating groups of devices by applying a clustering algorithm, a k-means clustering algorithm, to said fifth list of features.

7. The method according to claim 1 , wherein said communications further comprise HTTP communications, and wherein said one or more tokens are selected from: the HTTP method, the host, the mime type; or

wherein said communications comprise Server Message Block (SMB) communications, and wherein said one or more tokens are chosen from: the relative or absolute path to the file or one or more tokens extracted from the path to the file.

8. The method according to claim 1 , wherein said first list of a plurality of features, said second list of a plurality of features, said third list of a plurality of features and said fourth list of a plurality of features further comprise at least one of:

a numerical value identifying the duration of the connection, and

a numeric value identifying the amount of data exchanged.

9. The method according to claim 1 , comprising:

managing a database comprising for each source of a plurality of sources a respective list, wherein each list comprises metadata or preprocessed data of a subset of the connections of the respective source in said training interval, wherein said managing a database comprises:

deleting data that are older than said training interval,

receiving for a given source a list of metadata or preprocessed data of the connections of the respective source in said verification interval,

selecting the list associated with said source and determining a first number of connections saved in said selected list,

determining the number of connections of the respective source in said verification interval,

determining a second number of connections as a function of a maximum number of connections, said first number of connections saved in said selected list and said number of connections of the respective source in said verification interval,

randomly selecting said second connection number from said connections of the respective source in said verification interval and inserting the metadata and/or preprocessed data of said selected connections into said selected list, and

possibly randomly deleting connections from said selected list if the number of connections saved in said selected list exceeds said maximum number of connections.

10. A device comprising:

at least one hardware processor that is configured to perform operations comprising:

obtaining metadata for a plurality of communications in a monitoring interval, wherein said metadata includes for each communication an identifier of said source, an identifier of said destination, and data extracted from an application protocol of the respective communication, wherein said communications comprises Hypertext Transfer Protocol (HTTP) communications;

processing said extracted data to obtain preprocessed data comprising one or more tokens for the respective communication, wherein each token comprises a string, wherein said one or more tokens are extracted from user agent field or referrer field;

dividing said monitoring interval into a training interval and a verification interval;

obtaining the identifier of a given source and generating a first list of a plurality of features (F SRC,TI ) for connections of said given source in said training interval via following steps:

selecting the connections of said given source in said training interval,

determining for said connections of said given source in the said training interval the univocal destination identifiers and for each token the respective univocal values,

determining a first set of enumeration rules by enumerating said univocal destination identifiers and for each token the respective univocal values, and

associating by means of said first set of enumeration rules with each connection of said source in said training interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said first list of features comprises for each connection of said given source in said training interval the respective enumerated destination identifier and the respective one or more enumerated tokens;

obtaining the identifier of a group of devices to which said given source belongs and generating a second list of a plurality of features for the connections of the devices belonging to said group of devices in said training interval via following steps:

selecting the connections of said group of devices in said training interval,

determining for said connections of said group of devices in said training range the univocal destination identifiers and for each token the respective univocal values,

determining a second set of enumeration rules by enumerating said univocal destination identifiers and for each token the respective univocal values, and

associating by means of said second set of enumeration rules with each connection of said group of devices in said training interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said second list of features comprises for each connection of said group of devices in said training interval the respective enumerated destination identifier and the respective one or more enumerated tokens;

generating a first set of Bayesian networks by training for each feature of said first list of features a respective Bayesian network using the data of other features of said first list of features (F SRC,TI ), and generating a second set of Bayesian networks by training for each feature of said second list of features a respective Bayesian network using the data of the other features of said second list of features,

generating a third list of a plurality of features for the connections of said given source in said verification interval via following steps:

selecting the connections of said given source in said verification interval, and

associating by means of said first set of enumeration rules with each connection of said given source in said verification interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said third list of features comprises for each connection of said given source in said verification interval the respective enumerated destination identifier and the respective one or more respective enumerated tokens;

generating a fourth list of a plurality of features for connections of said given source in said verification interval via following steps:

selecting the connections of said given source in said verification interval,

associating by means of said second set of enumeration rules with each connection of said given source in said verification interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said fourth list of features comprises for each connection of said given source in said verification interval the respective enumerated destination identifier and the respective one or more respective enumerated tokens, wherein said first list of a plurality of features, said second list of a plurality of features, said third list of a plurality of features and said fourth list of a plurality of features further comprise at least one of: an enumerated value generated for a destination port of Transmission Control Protocol (TCP) or the User Datagram Protocol (UDP) of the respective communication;

repeating following steps for each connection of said given source in said verification interval:

determining based on the values of the features of said third list of features associated with the respective connection of said given source for each feature of said third list of features the respective most probable value by using said first set of Bayesian networks,

classifying each value of the features of said third list of features associated with the respective connection of said given source via following steps:

in response to determining the value of a feature of said third list of features corresponds to the respective most probable value, classifying the value of the feature of said third list of features as normal, and

in response to determining the value of a feature of said third list of features does not correspond to the respective most probable value:

 a) determining for the value of said feature of said third list of features the respective probability of occurrence by using said first set of Bayesian networks, and

 b) classifying the value of said feature of said third list of features as normal in response to determining the respective probability of occurrence is greater than a first threshold, and

 c) classifying the value of said feature of said third list of features as anomalous in response to determining the respective probability of occurrence is smaller than said first threshold; and

determining based on the values of the feature values of said fourth list of features associated with the respective connection of said given source for each feature of said fourth list of features the respective most probable value by using said second set of Bayesian networks, wherein discretizing one or more of the features of said first list of a plurality of features, said second list of features of a plurality of features, said third list of a plurality of features, and said fourth list of a plurality of features by means of a clustering algorithm, a k-means clustering algorithm, and

classifying each value of the features of said fourth list of features associated with the respective connection of said given source via following steps:

in response to determining the value of a feature of said fourth list of features corresponds to the respective most probable value, classifying the value of the feature of said fourth list of features as normal, and

in response to determining the value of a feature of said fourth list of features does not correspond to the respective most probable value:

a) determining for the value of said feature of said fourth list of features the respective probability of occurrence by using said second set of Bayesian networks, and

b) classifying the value of said feature of said fourth list of features as normal in response to determining the respective probability of occurrence is greater than a second threshold, and

c) classifying the value of said feature of said fourth list of features as anomalous in response to determining the respective probability of occurrence is smaller than said second threshold.

11. A computer-program product stored on a non-transitory computer readable memory that is coupled to at least one processor and comprises portions of software code that are configured to cause the at least one processor to perform operations comprising:

obtaining metadata for a plurality of communications in a monitoring interval, wherein said metadata includes for each communication an identifier of said source, an identifier of said destination, and data extracted from an application protocol of the respective communication, wherein said communications comprises Hypertext Transfer Protocol (HTTP) communications;

processing said extracted data to obtain preprocessed data comprising one or more tokens for the respective communication, wherein each token comprises a string, wherein said one or more tokens are extracted from user agent field or referrer field;

dividing said monitoring interval into a training interval and a verification interval;

obtaining the identifier of a given source and generating a first list of a plurality of features (F SRC,TI ) for connections of said given source in said training interval via following steps:

selecting the connections of said given source in said training interval,

determining for said connections of said given source in the said training interval the univocal destination identifiers and for each token the respective univocal values,

determining a first set of enumeration rules by enumerating said univocal destination identifiers and for each token the respective univocal values, and

associating by means of said first set of enumeration rules with each connection of said source in said training interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said first list of features comprises for each connection of said given source in said training interval the respective enumerated destination identifier and the respective one or more enumerated tokens:

obtaining the identifier of a group of devices to which said given source belongs and generating a second list of a plurality of features for the connections of the devices belonging to said group of devices in said training interval via following steps:

selecting the connections of said group of devices in said training interval,

determining for said connections of said group of devices in said training range the univocal destination identifiers and for each token the respective univocal values,

determining a second set of enumeration rules by enumerating said univocal destination identifiers and for each token the respective univocal values, and

associating by means of said second set of enumeration rules with each connection of said group of devices in said training interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said second list of features comprises for each connection of said group of devices in said training interval the respective enumerated destination identifier and the respective one or more enumerated tokens;

generating a first set of Bayesian networks by training for each feature of said first list of features a respective Bayesian network using the data of other features of said first list of features (F SRC,TI ), and generating a second set of Bayesian networks by training for each feature of said second list of features a respective Bayesian network using the data of the other features of said second list of features,

generating a third list of a plurality of features for the connections of said given source in said verification interval via following steps:

selecting the connections of said given source in said verification interval, and

associating by means of said first set of enumeration rules with each connection of said given source in said verification interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said third list of features comprises for each connection of said given source in said verification interval the respective enumerated destination identifier and the respective one or more respective enumerated tokens;

generating a fourth list of a plurality of features for connections of said given source in said verification interval via following steps:

selecting the connections of said given source in said verification interval,

associating by means of said second set of enumeration rules with each connection of said given source in said verification interval a respective enumerated destination identifier and one or more respective enumerated tokens, wherein said fourth list of features comprises for each connection of said given source in said verification interval the respective enumerated destination identifier and the respective one or more respective enumerated tokens, wherein said first list of a plurality of features, said second list of a plurality of features, said third list of a plurality of features and said fourth list of a plurality of features further comprise at least one of: an enumerated value generated for a destination port of Transmission Control Protocol (TCP) or the User Datagram Protocol (UDP) of the respective communication;

repeating following steps for each connection of said given source in said verification interval:

determining based on the values of the features of said third list of features associated with the respective connection of said given source for each feature of said third list of features the respective most probable value by using said first set of Bayesian networks,

classifying each value of the features of said third list of features associated with the respective connection of said given source via following steps:

in response to determining the value of a feature of said third list of features corresponds to the respective most probable value, classifying the value of the feature of said third list of features as normal, and

in response to determining the value of a feature of said third list of features does not correspond to the respective most probable value:

 a) determining for the value of said feature of said third list of features the respective probability of occurrence by using said first set of Bayesian networks, and

 b) classifying the value of said feature of said third list of features as normal in response to determining the respective probability of occurrence is greater than a first threshold, and

 c) classifying the value of said feature of said third list of features as anomalous in response to determining the respective probability of occurrence is smaller than said first threshold; and

determining based on the values of the feature values of said fourth list of features associated with the respective connection of said given source for each feature of said fourth list of features the respective most probable value by using said second set of Bayesian networks, wherein discretizing one or more of the features of said first list of a plurality of features, said second list of features of a plurality of features, said third list of a plurality of features, and said fourth list of a plurality of features by means of a clustering algorithm, a k-means clustering algorithm, and

classifying each value of the features of said fourth list of features associated with the respective connection of said given source via following steps:

in response to determining the value of a feature of said fourth list of features corresponds to the respective most probable value, classifying the value of the feature of said fourth list of features as normal, and

in response to determining the value of a feature of said fourth list of features does not correspond to the respective most probable value:

a) determining for the value of said feature of said fourth list of features the respective probability of occurrence by using said second set of Bayesian networks, and

b) classifying the value of said feature of said fourth list of features as normal in response to determining the respective probability of occurrence is greater than a second threshold, and

c) classifying the value of said feature of said fourth list of features as anomalous in response to determining the respective probability of occurrence is smaller than said second threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2022
From: UCCI, DANIELE; SOBRERO, FILIPPO; BISIO, FEDERICA
To: AIZOON S.R.L.
Reel/Frame 062197/0212 →
Priority Claims (1)
IT 102021000033203 · Dec 31, 2021 · national
Continuity (1)
Related Publication 20230216746A1 · Jul 6, 2023