IP Library Granted Patent US 12,572,657
Granted Patent B2
US 12,572,657 · App. 18/087,290 · Granted Mar 10, 2026

Generating high-quality threat intelligence from aggregated threat reports

Inventors: Mohamed Nabeel (Doha, QA); Saravanan Thirumuruganathan (Doha, QA); Euijin Choo (Doha, QA); Issa M. Khalil (Doha, QA); Ting Yu (Doha, QA)
Assignee: HAMAD BIN KHALIFA UNIVERSITY
G06F21/566G06N5/00G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,657
App. No.
18/087,290
Granted
Mar 10, 2026
Kind
B2
Abstract

Generating high-quality threat intelligence from aggregated threat reports is provided via developing a generative model that identifies relationships between a plurality of threat assessment scanners; pre-training a plurality of individual encoders based on a corresponding plurality of pretext tasks and the generative model; combining the individual encoders into a pre-trained encoder; fine-tuning the pre-trained encoder using threat data; and marking a candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious.

Claims (66)

1 . A method, comprising:

developing a generative model that identifies relationships between a plurality of threat assessment scanners;

pre-training a plurality of individual encoders based on a corresponding plurality of pretext tasks and the generative model;

combining the individual encoders into a pre-trained encoder;

fine-tuning the pre-trained encoder using threat data; and

marking a candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious; and

in response to marking the candidate threat as malicious, quarantining a computing system identified as having downloaded or accessed the candidate threat.

2 . The method of claim 1 , wherein the plurality of pretext tasks comprise:

learning dependencies between the plurality of threat assessment scanners;

modeling a temporal dynamic of threat reports from the plurality of threat assessment scanners; and

learning a temporal consistency of a plurality of embeddings for threats evaluated by the plurality of threat assessment scanners.

3 . The method of claim 1 , wherein fine-tuning is performed as a semi-supervised learning task using a plurality embeddings from the plurality of pretext tasks and a set of ground truth reports.

4 . The method of claim 1 , wherein fine-tuning is performed as an unsupervised learning task that classifies the candidate threat into one of two clusters with other entities, wherein a first cluster of the two clusters is designated as containing benign entities and a second cluster of the two clusters is designated as containing malicious entities.

5 . The method of claim 4 , wherein the first cluster of the two clusters is designated as containing the benign entities and the second cluster of the two clusters is designated as containing the malicious entities based on external verification of a subset of entities in the two clusters as being malicious or benign.

6 . The method of claim 4 , wherein the first cluster of the two clusters is designated as containing the benign entities and the second cluster of the two clusters is designated as containing the malicious entities based on the first cluster containing more entities than the second cluster.

7 . The method of claim 1 , wherein the candidate threat includes at least one of:

a phishing uniform resource locator (URL);

a malware URL;

a malware file; and

a blacklisted internet protocol address, and further comprising:

preventing a second computing device from accessing the candidate threat.

8 . A system, comprising:

a processor; and

a memory storage device including instructions that when executed by the processor perform operations comprising:

developing a generative model that identifies relationships between a plurality of threat assessment scanners;

pre-training a plurality of individual encoders based on a corresponding plurality of pretext tasks and the generative model;

combining the individual encoders into a pre-trained encoder;

fine-tuning the pre-trained encoder using threat data; and

marking a candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious; and

in response to marking the candidate threat as malicious, quarantining a computing system identified as having downloaded or accessed the candidate threat.

9 . The system of claim 8 , wherein the plurality of pretext tasks comprise:

learning dependencies between the plurality of threat assessment scanners;

modeling a temporal dynamic of threat reports from the plurality of threat assessment scanners; and

learning a temporal consistency of a plurality of embeddings for threats evaluated by the plurality of threat assessment scanners.

10 . The system of claim 8 , wherein fine-tuning is performed as a semi-supervised learning task using a plurality embeddings from the plurality of pretext tasks and a set of ground truth reports.

11 . The system of claim 8 , wherein fine-tuning is performed as an unsupervised learning task that classifies the candidate threat into one of two clusters with other entities, wherein a first cluster of the two clusters is designated as containing benign entities and a second cluster of the two clusters is designated as containing malicious entities.

12 . The system of claim 11 , wherein the first cluster of the two clusters is designated as containing the benign entities and the second cluster of the two clusters is designated as containing the malicious entities based on external verification of a subset of entities in the two clusters as being malicious or benign.

13 . The system of claim 11 , wherein the first cluster of the two clusters is designated as containing the benign entities and the second cluster of the two clusters is designated as containing the malicious entities based on the first cluster containing more entities than the second cluster.

14 . The system of claim 8 , wherein the candidate threat includes at least one of:

a phishing uniform resource locator (URL);

a malware URL;

a malware file; and

a blacklisted internet protocol address, and further comprising:

preventing a second computing device from accessing the candidate threat.

15 . A non-transitory memory including instructions that when executed by a processor perform operations, the operations comprising:

developing a generative model that identifies relationships between a plurality of threat assessment scanners;

pre-training a plurality of individual encoders based on a corresponding plurality of pretext tasks and the generative model;

combining the individual encoders into a pre-trained encoder;

fine-tuning the pre-trained encoder using threat data; and

marking a candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious; and

in response to marking the candidate threat as malicious, quarantining a computing system identified as having downloaded or accessed the candidate threat.

16 . The memory of claim 15 , wherein the candidate threat includes at least one of:

a phishing uniform resource locator (URL);

a malware URL;

a malware file; and

a blacklisted internet protocol address, and the operations further comprising:

preventing a second computing device from accessing the candidate threat.

17 . The memory of claim 15 , the operations further comprising:

marking a second candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious; and

in response to marking the second candidate threat as benign, releasing from quarantine a second computing system identified as having downloaded or accessed the candidate threat that was quarantined in response to accessing the second candidate threat.

18 . The method of claim 1 , further comprising:

marking a second candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious; and

in response to marking the second candidate threat as benign, releasing from quarantine a second computing system identified as having downloaded or accessed the candidate threat that was quarantined in response to accessing the second candidate threat.

19 . The method of claim 1 , further comprising:

marking a second candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious; and

permitting a second computing system requesting access to the second candidate threat to access the second candidate threat in response to marking the second candidate threat as benign.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: QATAR FOUNDATION FOR EDUCATION, SCIENCE & COMMUNITY DEVELOPMENT
To: HAMAD BIN KHALIFA UNIVERSITY
Reel/Frame 069936/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2023
From: NABEEL, MOHAMED; THIRUMURUGANATHAN, SARAVANAN; CHOO, EUIJIN; KHALIL, ISSA M.; YU, TING
To: QATAR FOUNDATION FOR EDUCATION, SCIENCE AND COMMUNITY DEVELOPMENT
Reel/Frame 062326/0256 →
Continuity (2)
Provisional Application 63294163 · Dec 28, 2021
Related Publication 20230205884A1 · Jun 29, 2023
References Cited (12)
US 9978067B1 · Sadaghiani · 2018 [cited by examiner]
US 10997608B1 · Veeraraghavan · 2021 [cited by examiner]
US 20210273959A1 · Salji · 2021 [cited by examiner]
Huang, et al., “ITDBERT: Temporal-semantic representation for insider threat detection”, 2021 ISCC, Sep. 5-8, 2021 (Year: 2021). [cited by examiner]
Liu, et al., “Anomaly-based insider threat detection using dep autoencoders”, 2018 IEEE International Conference on Data Mining Workshops (ICDMW), 2018 (Year: 2018). [cited by examiner]
Yu, et al., “Securing critical infrastructures: deep-learning-based threat detection in IIIT”, IEEE Communications Magazine, Oct. 2021 (Year: 2021). [cited by examiner]
Zhu, et al. “Measuring and Modeling the Label Dynamics of Online Anti-Malware Engines”; usenix; Aug. 2020; (19 pages). [cited by applicant]
Sharif, et al.; “Predicting Impending Exposure to Malicious Content from User Behavior”; CCS; 2018; (15 pages). [cited by applicant]
Zhang, et al.; “A Survey on Multi-Task Learning”; IEEE; 2021; (20 pages). [cited by applicant]
Kantchelian, et al.; “Better Malware Ground Truth: Techniques for Weighting Anti-Virus Vendor Labels”; ACM; 2015; (12 pages). [cited by applicant]
Yoon, et al.; “VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain”; 34th Conference on Neural Information Processing Systems; 2020; (11 pages). [cited by applicant]
Thirumuruganathan, et al.; “SIRAJ: A Unified Multi-Source Aggregation Framework for High-Quality Threat Intelligence”; IEEE; 2022; (4 pages). [cited by applicant]