IP Library › Granted Patent US 12,619,733
Granted Patent B2
US 12,619,733 · App. 17/808,481 · Granted May 5, 2026

Optimizing accuracy of security alerts based on data classification

Inventors: Andrey Karpovsky (Kiryat Motzkin, IL); Sagi Lowenhardt (Hertzeliya, IL); Shimon Ezra (Petach Tikva, IL)
G06F21/577G06F21/6245H04L63/1425G06F2221/033G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,733
App. No.
17/808,481
Granted
May 5, 2026
Kind
B2
Abstract

A computing system and method for training one or more machine-learning models to perform anomaly detection. A training dataset is accessed. An overall sensitivity score is determined that indicates an amount of sensitive data in the training dataset. Machine-learning models are trained based on the training dataset and the overall sensitivity score. The machine-learning models use the overall sensitivity score to determine a threshold. The threshold is relatively low for datasets having a large amount of sensitive data and is relatively high for dataset having a small among of sensitive data. When executed, the machine-learning models determine if a probability score of features extracted from a received dataset are above the determined threshold when a second overall sensitivity score of the received dataset is substantially similar to the overall sensitivity score. When the probability score is above the determined threshold, the machine-learning models cause an alert to be generated.

Claims (49)

1 . A method for a computing system to train one or more machine-learning models to perform anomaly detection, the method comprising:

accessing a group of training datasets, the training datasets including a plurality of data items;

determining overall sensitivity scores for the training datasets in the group, the overall sensitivity scores indicating an amount of sensitive data included in the training datasets and being determined based on individual sensitivity scores of the plurality of data items of the training datasets; and

training one or more machine-learning models to perform anomaly detection based on the group of the training datasets and the overall sensitivity scores of the training datasets in the group, the one or more machine-learning models using the overall sensitivity scores to determine a respective threshold for a training dataset based on an inverse sliding scale relationship, wherein based on the inverse sliding scale relationship, the lower an overall sensitivity score is, the higher the respective threshold is, and the higher an overall sensitivity score is, the lower the respective threshold is, wherein:

when an overall sensitivity score of a received dataset, which is distinct from the group of the training datasets, is closer to an overall sensitivity score of one training dataset in the group than overall sensitivity scores of the other training datasets in the group based on the inverse sliding scale relationship, the one or more machine-learning models are configured to determine if a probability score of one or more features extracted from the received dataset is above a threshold of the one training dataset, and

in response to determining that the probability score is above the threshold of the one training dataset, causing an alert to be generated, the alert indicating that an anomaly has been detected.

2 . The method of claim 1 , further comprising:

determining whether each of the plurality of data items is sensitive or is non-sensitive; and

determining the overall sensitivity score based on the determination of whether each of the plurality of data items is sensitive or is non-sensitive.

3 . The method of claim 2 , wherein the overall sensitivity score is determined based on a total number of the data items that are determined to be sensitive in relation to the total number of data items in the training dataset.

4 . The method of claim 2 , wherein the dataset is scanned to determine whether each of the plurality of data items is sensitive or is non-sensitive.

5 . The method of claim 1 , further comprising:

receiving feedback from a user agent in response to the user agent receiving the generated alert; and

in response to the feedback, adjusting the threshold.

6 . The method of claim 5 , wherein adjusting the threshold comprises adjusting the overall sensitivity score.

7 . The method of claim 1 , further comprising:

receiving feedback from a user agent in response to user interaction with the dataset; and

in response to the feedback, adjusting the threshold.

8 . A method for a computing system to perform anomaly detection, the method comprising:

executing one or more machine-learning models trained based on a group of training datasets, the training datasets including a plurality of data items, wherein training the one or more machine-learning models comprises:

determining overall sensitivity scores for the training datasets in the group, the overall sensitivity scores indicating an amount of sensitive data included in the datasets and being determined based on individual sensitivity scores of the plurality of data items of the training datasets; and

using the overall sensitivity scores to determine a respective threshold for a training dataset based on an inverse sliding scale relationship, wherein based on the inverse sliding scale relationship, the lower an overall sensitivity score is, the higher the respective threshold is, and the higher the overall sensitivity score is, the lower the respective threshold is;

receiving a dataset, which is distinct from the group of the training datasets, at the computing system;

when an overall sensitivity score of the received dataset is closer to an overall sensitivity score of one training dataset than overall sensitivity scores of the other training datasets in the group based on the inverse sliding scale relationship, determining if a probability score of one or more features extracted from the received dataset is above a threshold of the one training dataset; and

in response to determining that the probability score is above the threshold of the one training dataset, generating an alert, the alert indicating that an anomaly has been detected.

9 . The method of claim 8 , further comprising:

determining whether each of the plurality of data items is sensitive or is non-sensitive; and

determining the overall sensitivity score based on the determination of whether each of the plurality of data items is sensitive or is non-sensitive.

10 . A computing system for training one or more machine-learning models to perform anomaly detection, comprising:

one or more processors; and

one or more computer-readable hardware storage devices having stored thereon computer-executable instructions that are structured such that, when executed by the one or more processors, the computer-executable instructions cause the computing system to perform at least:

access a group of training datasets, the training datasets including a plurality of data items;

determine overall sensitivity scores for the training datasets in the group, the overall sensitivity scores indicating an amount of sensitive data included in the training datasets and being determined based on individual sensitivity scores of the plurality of data items of the training datasets; and

train one or more machine-learning models to perform anomaly detection based on the group of the training datasets and the overall sensitivity scores of the training datasets in the group, the one or more machine-learning models using the overall sensitivity scores to determine a respective threshold for a training dataset based on an inverse sliding scale relationship, wherein based on the inverse sliding scale relationship, the lower an overall sensitivity score is, the higher the respective threshold is, and the higher the overall sensitivity score is, the lower the respective threshold is, wherein:

when an overall sensitivity score of a received dataset, which is distinct from the group of the training datasets, is closer to an overall sensitivity score of one training dataset in the group than overall sensitivity scores of the other training datasets in the group based on the inverse sliding scale relationship, the one or more machine-learning models are configured to determine if a probability score of one or more features extracted from the received dataset is above a threshold of the one training dataset, and

in response to determining that the probability score is above the threshold of the one training dataset, the one or more machine-learning models cause an alert to be generated, the alert indicating that an anomaly has been detected.

11 . The computing system of claim 10 , wherein the computing system is further configured to:

determine whether each of the plurality of data items is sensitive or is non-sensitive; and

determine the overall sensitivity score based on the determination of whether each of the plurality of data items is sensitive or is non-sensitive.

12 . The computing system of claim 11 , wherein the overall sensitivity score is determined based on a total number of the data items that are determined to be sensitive in relation to the total number of data items in the training dataset.

13 . The computing system of claim 11 , wherein the dataset is scanned to determine whether each of the plurality of data items is sensitive or is non-sensitive.

14 . The computing system of claim 10 , the computing system is further configured to:

receive feedback from a user agent in response to the user agent receiving the generated alert; and

in response to the feedback, adjust the threshold.

15 . The computing system of claim 14 , wherein adjusting the threshold comprises adjusting the overall sensitivity score.

16 . The computing system of claim 10 , wherein the computing system is further configured to:

receive feedback from a user agent in response to user interaction with the dataset; and

in response to the feedback, adjust the threshold.

17 . The computing system of claim 10 , wherein the one or more machine-learning models are one of a supervised model, a semi-supervised model, or an unsupervised model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2022
From: KARPOVSKY, ANDREY; LOWENHARDT, SAGI; EZRA, SHIMON
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 060323/0648 →
Continuity (1)
Related Publication 20230418948A1 · Dec 28, 2023
References Cited (6)
US 10496842B1 · Ren · 2019 [cited by examiner]
US 20200159947A1 · Shenefiel · 2020 [cited by examiner]
US 20200274894A1 · Argoeti et al. · 2020 [cited by applicant]
US 20230153427A1 · Croteau · 2023 [cited by examiner]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/022633”, Mailed Date: Sep. 14, 2023, 11 Pages. (Ms# 411600-WO-PCT). [cited by applicant]
Xiao, et al., “Secure Mobile Crowdsensing Based On Deep Learning”, In Journal of China Communications, vol. 15, Issue 10, Oct. 2018, pp. 1-11. [cited by applicant]