Security policy generation and enforcement for device clusters
Techniques for generating and enforcing security policies for device clusters are disclosed. A security manager generates a plurality of clusters of devices for applying security policies. For each cluster of devices, the security manager trains a machine learning model to indicate whether a particular data flow associated with a device in the particular cluster of devices is allowed or denied. The security manager detects a data flow corresponding to a device. If the security manager determines that the device corresponds to a first cluster of devices, the security manager identifies a first trained machine learning model corresponding to the first cluster of devices. The security manager applies the first trained machine learning model to the first data flow to determine whether the first data flow is to be allowed or denied. The security manager allows or denies the first data flow based on the applying operation.
1 . One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause performance of operations comprising:
analyzing behavior for each of a set of devices to generate a plurality of clusters of devices for applying policies;
for each particular cluster of devices in the plurality of clusters of devices:
training a machine learning model to generate a policy indicating whether a particular data flow associated with a device in the particular cluster of devices is allowed or denied at least by:
obtaining training data sets of historical device data, each training data set comprising:
a data flow corresponding to at least one device in the particular cluster of devices, the data flow being associated with a set of data flow attributes;
an indication of whether the data flow was allowed or denied;
training the machine learning model based on the training data sets;
detecting a first data flow corresponding to a first device of the set of devices;
determining the first device corresponds to a first cluster of devices of the plurality of clusters of devices;
identifying a first trained machine learning model corresponding to the first cluster of devices;
applying the first trained machine learning model to the first data flow to generate a first policy indicating whether the first data flow is to be allowed or denied;
allowing or denying the first data flow based on the first policy.
2 . The one or more machine-readable media of claim 1 ,
wherein each training data set further identifies an enforcement point that allows or denies the data flow;
wherein the operations further comprise applying the first trained machine learning model to the first data flow to identify a first enforcement point for the first data flow;
wherein allowing or denying the first data flow is performed by the first enforcement point.
3 . The one or more machine-readable media of claim 2 , wherein the instructions further cause:
detecting a second data flow corresponding to a second device of the set of devices;
wherein the operations further comprise applying the first trained machine learning model to the second data flow to identify a second enforcement point for the second data flow;
wherein allowing or denying the second data flow is performed by the second enforcement point.
4 . The one or more machine-readable media of claim 2 , wherein the instructions further cause:
detecting a second data flow corresponding to the first device of the set of devices;
wherein the operations further comprise applying the first trained machine learning model to the second data flow to identify a second enforcement point for the second data flow;
wherein allowing or denying the second data flow is performed by the second enforcement point.
5 . The one or more machine-readable media of claim 1 , wherein each training data set further comprises:
an enforcement point for allowing or denying the data flow; and
an indication whether enforcement of the allowing or the denying of the data flow was successful;
wherein the operations further comprise applying the first trained machine learning model to the first data flow to identify a first enforcement point for the first data flow;
wherein allowing or denying the first data flow is performed by the first enforcement point.
6 . The one or more machine-readable media of claim 1 , wherein the set of data flow attributes comprises at least one attribute, comprising:
a security risk of the data flow;
a priority level of the data flow;
a destination of the data flow;
a source of the data flow; and
a number of devices, among the set of devices, accessed by the data flow;
wherein training the machine learning model comprises identifying a pattern associating the at least one attribute with the indication of whether the data flow is to be allowed or denied.
7 . The one or more machine-readable media of claim 1 , wherein each training data set further comprises at least one device attribute comprising:
a function of the at least one device;
a location of the at least one device;
an application running on the at least one device;
data stored in the at least one device; and
a level of importance of the at least one device relative to other devices in the particular cluster of devices;
wherein training the machine learning model comprises identifying a pattern associating the at least one device attribute with the indication of whether the data flow is to be allowed or denied.
8 . One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause performance of operations comprising:
analyzing behavior for each of a set of devices to generate a plurality of clusters of devices for applying policies;
for each particular cluster of devices in the plurality of clusters of devices:
training a machine learning model to select enforcement points for applying policies for allowing or denying data flows at least by:
obtaining training data sets of historical device data, each training data set comprising:
a data flow corresponding to at least one device in the particular cluster of devices, the data flow being associated with a set of data flow attributes;
an identification of an enforcement point for enforcing policies of allowing or denying the data flow;
training the machine learning model based on the training data sets;
detecting a first data flow corresponding to a first device of the set of devices;
determining the first device corresponds to a first cluster of devices of the plurality of clusters of devices;
identifying a first trained machine learning model corresponding to the first cluster of devices;
applying the first trained machine learning model to the first data flow to select, by the first trained machine learning model, a first enforcement point for a first policy indicating whether to allow or deny the data flow;
allowing or denying, by the first enforcement point, the first data flow.
9 . The one or more machine-readable media of claim 8 , wherein the instructions further cause:
detecting a second data flow corresponding to a second device of the set of devices;
wherein the operations further comprise applying the first trained machine learning model to the second data flow to select a second enforcement point for the second data flow;
wherein allowing or denying the second data flow is performed by the second enforcement point.
10 . The one or more machine-readable media of claim 8 , wherein the instructions further cause:
detecting a second data flow corresponding to the first device of the set of devices;
wherein the operations further comprise applying the first trained machine learning model to the second data flow to select a second enforcement point for the second data flow;
wherein allowing or denying the second data flow is performed by the second enforcement point.
11 . The one or more machine-readable media of claim 8 , wherein each training data set further comprises:
an enforcement point for allowing or denying the data flow; and
an indication whether enforcement of the allowing or the denying of the data flow was successful;
wherein training the machine learning model comprises identifying a pattern associating the enforcement point with the indication whether enforcement of the allowing or the denying of the data flow was successful.
12 . The one or more machine-readable media of claim 8 , wherein the set of data flow attributes comprises at least one attribute, comprising:
a security risk of the data flow;
a priority level of the data flow;
a destination of the data flow;
a source of the data flow; and
a number of devices, among the set of devices, accessed by the data flow;
wherein training the machine learning model comprises identifying a pattern associating the at least one attribute with the indication of whether the data flow is to be allowed or denied.
13 . The one or more machine-readable media of claim 8 , wherein each training data set further comprises at least one device attribute comprising:
a function of the at least one device;
a location of the at least one device;
an application running on the at least one device;
data stored in the at least one device; and
a level of importance of the at least one device relative to other devices in the particular cluster of devices;
wherein training the machine learning model comprises identifying a pattern associating the at least one device attribute with the indication of whether the data flow is to be allowed or denied.
14 . A method, comprising:
analyzing behavior for each of a set of devices to generate a plurality of clusters of devices for applying policies;
for each particular cluster of devices in the plurality of clusters of devices:
training a machine learning model to generate a policy indicating whether a particular data flow associated with a device in the particular cluster of devices is allowed or denied at least by:
obtaining training data sets of historical device data, each training data set comprising:
a data flow corresponding to at least one device in the particular cluster of devices, the data flow being associated with a set of data flow attributes;
an indication of whether the data flow was allowed or denied;
training the machine learning model based on the training data sets;
detecting a first data flow corresponding to a first device of the set of devices;
determining the first device corresponds to a first cluster of devices of the plurality of clusters of devices;
identifying a first trained machine learning model corresponding to the first cluster of devices;
applying the first trained machine learning model to the first data flow to generate a first policy indicating whether the first data flow is to be allowed or denied;
allowing or denying the first data flow based on the first policy.
15 . The method of claim 14 , wherein each training data set further identifies an enforcement point that allows or denies the data flow;
wherein the method further comprises applying the first trained machine learning model to the first data flow to identify a first enforcement point for the first data flow;
wherein allowing or denying the first data flow is performed by the first enforcement point.
16 . The method of claim 15 , further comprising:
detecting a second data flow corresponding to a second device of the set of devices;
applying the first trained machine learning model to the second data flow to identify a second enforcement point for the second data flow;
wherein allowing or denying the second data flow is performed by the second enforcement point.
17 . The method of claim 15 , further comprising:
detecting a second data flow corresponding to the first device of the set of devices;
applying the first trained machine learning model to the second data flow to identify a second enforcement point for the second data flow;
wherein allowing or denying the second data flow is performed by the second enforcement point.
18 . The method of claim 14 , wherein each training data set further comprises:
an enforcement point for allowing or denying the data flow; and
an indication whether enforcement of the allowing or the denying of the data flow was successful;
wherein the method further comprises applying the first trained machine learning model to the first data flow to identify a first enforcement point for the first data flow;
wherein allowing or denying the first data flow is performed by the first enforcement point.
19 . The method of claim 14 , wherein the set of data flow attributes comprises at least one attribute, comprising:
a security risk of the data flow;
a priority level of the data flow;
a destination of the data flow;
a source of the data flow; and
a number of devices, among the set of devices, accessed by the data flow;
wherein training the machine learning model comprises identifying a pattern associating the at least one attribute with the indication of whether the data flow is to be allowed or denied.
20 . A system, comprising:
one or more processors; and
memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:
analyzing behavior for each of a set of devices to generate a plurality of clusters of devices for applying policies;
for each particular cluster of devices in the plurality of clusters of devices:
training a machine learning model to generate a policy indicating whether a particular data flow associated with a device in the particular cluster of devices is allowed or denied at least by:
obtaining training data sets of historical device data, each training data set comprising:
a data flow corresponding to at least one device in the particular cluster of devices, the data flow being associated with a set of data flow attributes;
an indication of whether the data flow is to be was allowed or denied;
training the machine learning model based on the training data sets;
detecting a first data flow corresponding to a first device of the set of devices;
determining the first device corresponds to a first cluster of devices of the plurality of clusters of devices;
identifying a first trained machine learning model corresponding to the first cluster of devices;
applying the first trained machine learning model to the first data flow to generate a first policy indicating whether the first data flow is to be allowed or denied;
allowing or denying the first data flow based on the first policy.
21 . One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause performance of operations comprising:
generating, for a set of devices, a plurality of clusters of devices for applying policies, wherein the plurality of clusters of devices is determined based on one or more of: user selection or shared attributes;
for each particular cluster of devices in the plurality of clusters of devices:
training a machine learning model to generate a policy indicating whether a particular data flow associated with a device in the particular cluster of devices is allowed or denied at least by:
obtaining training data sets of historical device data, each training data set comprising:
a data flow corresponding to at least one device in the particular cluster of devices, the data flow being associated with a set of data flow attributes;
an indication of whether the data flow is to be was allowed or denied;
training the machine learning model based on the training data sets;
detecting a first data flow corresponding to a first device of the set of devices;
determining the first device corresponds to a first cluster of devices of the plurality of clusters of devices;
identifying a first trained machine learning model corresponding to the first cluster of devices;
applying the first trained machine learning model to the first data flow to generate a first policy indicating whether the first data flow is to be allowed or denied;
allowing or denying the first data flow based on the first policy.
22 . The media of claim 21 , wherein the plurality of clusters of devices is generated based on at least one of:
a geographic location of the devices;
an owner of the devices; and
a device type of the devices.
23 . The media of claim 1 , wherein the operations further comprise:
detecting a second data flow corresponding to a second device of the set of devices;
determining the second device corresponds to a second cluster of devices of the plurality of clusters of devices;
identifying a second trained machine learning model corresponding to the second cluster of devices;
applying the second trained machine learning model to the second data flow to generate a second policy indicating whether the second data flow is to be allowed or denied;
allowing or denying the second data flow based on the second policy.