IP Library › Granted Patent US 12,519,805
Granted Patent B2
US 12,519,805 · App. 17/568,590 · Granted Jan 6, 2026

Bias mitigation in threat disposition systems

Inventors: Aankur Bhatia (Bethpage, NY); Gary I. Givental (Bloomfield Hills, MI); Namrata Tolani (Bangalore, IN); Ajmeera Balaji Naik (Reddigudem Mandal, IN); Oleksandr Shmaliy (Wroclaw, PL)
Assignee: International Business Machines Corporation
H04L63/1416G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,519,805
App. No.
17/568,590
Granted
Jan 6, 2026
Kind
B2
Abstract

Mitigating bias in a machine learning-augmented threat disposition platform can include generating a group of alerts in response to determining a similarity among the alerts. The alerts are generated in real time by a threat monitoring tool in response to one or more potential threats to a networked computing system. One or more alert spikes can be determined by partitioning the group into one or more alert spike subgroups. Each alert spike subgroup corresponds to an alert spike and contains two or more similar alerts that were generated within a predetermined time interval of one another. Duplicate alerts in each alert spike can be eliminated and each non-discarded alert labeled. The labeled alerts are used for training a reduced-bias machine learning model.

Claims (59)

1 . A computer-implemented method of mitigating bias in a machine learning-augmented threat disposition platform, the computer-implemented method comprising:

generating a group of alerts in response to determining a similarity among the alerts, wherein the alerts are generated in real time by a threat monitoring tool in response to one or more potential threats to a networked computing system;

determining one or more alert spikes by partitioning the group into one or more alert spike subgroups, wherein each alert spike subgroup, of the one or more alert spike subgroups, corresponds to an alert spike and contains two or more similar alerts that were generated within a predetermined time interval of one another;

discarding duplicate alerts in each alert spike subgroup and labeling each non-discarded alert;

training a machine learning model using each non-discarded alert as a labeled training example;

relabeling an alert of the non-discarded alerts based at least in part on an age of the alert exceeding a predetermined threshold; and

retraining the machine learning model based at least in part on the relabeled alert, wherein the retraining decreases bias in the machine learning model and increases predictive accuracy of the machine learning model.

2 . The method of claim 1 , comprising:

responsive to determining, in real time, a similarity between a first newly generated alert and a second newly generated alert generated within a predetermined time interval of generation of the first newly generated alert, creating an incipient alert spike subgroup containing the first newly generated alert and the second newly generated alert.

3 . The method of claim 1 , comprising:

closing one of the one or more alert spike subgroups in response to determining that a predetermined time interval has lapsed without adding a newly generated alert to the one of the one or more alert spike subgroups, wherein the closing removes the one of the one or more alert spikes from a list of active alert spike subgroups and discards all but one of the alerts contained in the one or more alert spike subgroups.

4 . The method of claim 1 , comprising:

generating a similar-alerts dataset and assigning weights to each alert contained the similar-alerts dataset, wherein each of the weights corresponds to a time of generation of each alert contained the similar-alerts dataset; and

labeling each alert contained in the similar-alerts dataset according to a ground truth determined based on summing the weights, wherein each label comprises one of escalate or closure.

5 . The method of claim 4 , wherein

changing a current label of one of the alerts contained in the similar-alerts dataset is precluded unless an age of the one of the alerts contained in the similar-alerts dataset is greater than a predetermined threshold.

6 . The method of claim 4 , wherein

the assigning weights assigns to each one of the alerts contained in the similar-alerts dataset a weight computed as an exponentially decreasing function of time.

7 . A system, comprising:

one or more processors; and

one or more memory devices coupled to the one or more processors, wherein the one or more processors are configured to:

generate a group of alerts in response to determining a similarity among the alerts, wherein the alerts are generated in real time by a threat monitoring tool in response to one or more potential threats to a networked computing system;

determine one or more alert spikes by partitioning the group into one or more alert spike subgroups, wherein each alert spike subgroup, of the one or more alert spike groups, corresponds to an alert spike, of the one or more alert spikes, and contains two or more similar alerts that were generated within a predetermined time interval of one another;

discard duplicate alerts in each alert spike, of the one or more alert spikes, and label each non-discarded alert;

train a machine learning model using each non-discarded alert as a labeled training example;

relabel an alert of the non-discarded alerts based at least in part on an age of the alert exceeding a predetermined threshold; and

retrain the machine learning model based at least in part on the relabeled alert, wherein the retraining decreases bias in the machine learning model and increases predictive accuracy of the machine learning model.

8 . The system of claim 7 , wherein the one or more processors are further configured to:

create an incipient alert spike subgroup containing a first newly generated alert and a second newly generated alert responsive to determining, in real time, a similarity between the first newly generated alert and the second newly generated alert generated within a predetermined time interval of generation of the first newly generated alert.

9 . The system of claim 7 , wherein the one or more processors are further configured to:

close one of the one or more alert spike subgroups in response to determining that a predetermined time interval has lapsed without adding a newly generated alert to the one of the one or more alert spike subgroups, wherein the closing removes the one of the one or more alert spikes from a list of active alert spike subgroups and discards all but one of the alerts contained in the one or more alert spike subgroups.

10 . The system of claim 7 , wherein the one or more processors are further configured to:

generate a similar-alerts dataset and assigning weights to each alert contained in the similar-alerts dataset, wherein each of the weights corresponds to a time of generation of each alert contained in the similar-alerts dataset; and

label each alert contained in the similar-alerts dataset according to a ground truth determined based on summing the weights, wherein each label comprises one of escalate or closure.

11 . The system of claim 10 , wherein changing a current label of one of the alerts contained in the similar-alerts dataset is precluded unless an age of the one of the alerts contained in the similar-alerts dataset is greater than a predetermined threshold.

12 . A non-transitory computer-readable medium storing a set of instructions for wireless communication, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

generate a group of alerts in response to determining a similarity

among the alerts, wherein the alerts are generated in real time by a threat monitoring tool in response to one or more potential threats to a networked computing system;

determine one or more alert spikes by partitioning the group into one or more alert spike subgroups, wherein each alert spike subgroup, of the one or more alert spike subgroups, corresponds to an alert spike and contains two or more similar alerts that were generated within a predetermined time interval of one another;

discard duplicate alerts in each alert spike, of the one or more alert spikes, and label each non-discarded alert;

train a machine learning model using each non-discarded alert as a labeled training example;

relabel an alert of the non-discarded alerts based at least in part on an age of the alert exceeding a predetermined threshold; and

retrain the machine learning model based at least in part on the relabeled alert, wherein the retraining decreases bias in the machine learning model and increases predictive accuracy of the machine learning model.

13 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions cause the device to:

create an incipient alert spike subgroup containing a first newly generated alert and a second newly generated alert responsive to determining, in real time, a similarity between the first newly generated alert and the second newly generated alert generated within a predetermined time interval of generation of the first newly generated alert.

14 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions cause the device to:

close one of the one or more alert spike subgroups in response to determining that a predetermined time interval has lapsed without adding a newly generated alert to the one of the one or more alert spike subgroups, wherein the closing removes the one of the one or more alert spikes from a list of active alert spike subgroups and discards all but one of the alerts contained in the one or more alert spike subgroups.

15 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions cause the device to:

generate a similar-alerts dataset and assigning weights to each alert contained in the similar-alerts dataset, wherein each of the weights corresponds to a time of generation of each alert contained in the similar-alerts dataset; and

label each alert contained in the similar-alerts dataset according to a ground truth determined based on summing the weights, wherein each label comprises one of escalate or closure.

16 . The non-transitory computer-readable medium of claim 15 , wherein changing a current label of one of the alerts contained in the similar-alerts dataset is precluded unless an age of the one of the alerts contained in the similar-alerts dataset is greater than a predetermined threshold.

17 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instruction, to assign weights, cause the device to assign to each one of the alerts contained in the similar-alerts dataset a weight computed as an exponentially decreasing function of time.

18 . The method of claim 1 , further comprising:

generating a hash signature corresponding to a vector representation of each of the non-discarded alerts.

19 . The system of claim 7 , wherein the one or more processors are further configured to:

generate a hash signature corresponding to a vector representation of each of the non-discarded alerts.

20 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions cause the device to:

generate a hash signature corresponding to a vector representation of each of the non-discarded alerts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2022
From: BHATIA, AANKUR; GIVENTAL, GARY I.; TOLANI, NAMRATA; NAIK, AJMEERA BALAJI; SHMALIY, OLEKSANDR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058546/0405 →
Continuity (1)
Related Publication 20230216865A1 · Jul 6, 2023
References Cited (17)
US 7949662B2 · Farber et al. · 2011 [cited by applicant]
US 20070094725A1 · Borders · 2007 [cited by examiner]
US 20130046558A1 · Landi et al. · 2013 [cited by applicant]
US 20150163242A1 · Laidlaw · 2015 [cited by examiner]
US 20160357790A1 · Elkington et al. · 2016 [cited by applicant]
US 20160359759A1 · Singh et al. · 2016 [cited by applicant]
US 20200168231A1 · Alagianambi · 2020 [cited by applicant]
US 20200380309A1 · Weider et al. · 2020 [cited by applicant]
US 20230224311A1 · Meshi · 2023 [cited by examiner]
US 20230283629A1 · Boyer · 2023 [cited by examiner]
US 20250039210A1 · Rahmes · 2025 [cited by examiner]
Bontcheva, K. et al., “Balancing Act: Countering Digital Disinformation While Respecting Freedom of Expression,” geneva, Switzerland: United Nations Educational, Scientific and Cultural Organization, Sep. 2020, 348 Pg. [cited by applicant]
Dahir, H. et al., “Dynamic Trust and Risk Scoring Using Last-Known Profile Learning,” [online] IP.com Prior Art Database Technical Disclosure, No. IPCOM000247388D, Copyright 2016 Cisco Systems, Inc., Aug. 31, 2016, retr… [cited by applicant]
Somaraju, A., “Protection Against Adversarial Attacks on Machine Learning and Artificial Intelligence,” [online] IP.com Prior Art Database Technical Disclosure, No. IPCOM000252595D, Copyright 2018 Cisco Systems, Inc., J… [cited by applicant]
“System and Method to Counter Adversaries via Data and Model Biasing,” [online] IP.com Prior Art Database Technical Disclosure, No. IPCOM000259028D, Jul. 4, 2019, retrieved from the Internet: <https://priorart.ip.com/IP… [cited by applicant]
La Diega, G.N., “Against the Dehumanisation of Decision-Making,” J. Intell. Prop. Info. Tech. & Elec. Com. L., 2018, 9, p. 3. [cited by applicant]
Mell, P. et al., The NIST Definition of Cloud Computing, National Institute of Standards and Technology, U.S. Dept. of Commerce, Special Publication 800-145, Sep. 2011, 7 pg. [cited by applicant]
Cited By (1)
US 12,748,417