IP Library › Granted Patent US 12,641,112
Granted Patent B2
US 12,641,112 · App. 18/617,133 · Granted May 26, 2026

Proactively detecting malicious domains using graph representation learning

Inventors: Mohamed Nabeel (Doha, QA); Issa Khalil (Doha, QA); Ting Yu (Doha, QA); Fatih Deniz (Doha, QA)
Assignee: Qatar Foundation for Education, Science and Community Development
H04L63/1433H04L41/16H04L63/145
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,641,112
App. No.
18/617,133
Granted
May 26, 2026
Kind
B2
Abstract

Proactively detecting malicious domains using graph representation learning may be provided by extracting seed domains from a uniform resource locator (URL) feed of observed requests for access to domains; expanding the seed domains to a via a passive domain name service (PDNS) crawl to include additional domains with the seed domains; collecting a ground truth, including labeling a first set of the seed domains as benign and a second set of the seed domains as malicious; constructing a graph neural network (GNN) of the additional domains and the seed domains, wherein each domain of the additional domains and the seed domains are represented as a node in the GNN that includes feature values associated that domain; training the GNN to classify unseen domains not associated with a node as either benign or malicious; and classifying, via the GNN, a queried domain as either benign or malicious.

Claims (86)

1 . A method, comprising:

extracting seed domains from a uniform resource locator (URL) feed of observed requests for access to domains;

expanding the seed domains via a passive domain name service crawl to include additional domains with the seed domains;

collecting a ground truth, including labeling a first set of the seed domains as benign and a second set of the seed domains as malicious;

constructing a graph neural network (GNN) of the additional domains and the seed domains, wherein each domain of the additional domains and the seed domains is represented as a node in the GNN that includes feature values associated with that domain;

training the GNN to classify unseen domains not associated with a node as either benign or malicious; and

classifying, via the GNN, a queried domain as either benign or malicious.

2 . The method of claim 1 , wherein constructing the GNN includes:

assembling an ensemble of GNN encoders in a model stack; and

combining outputs of the ensemble of GNN encoders via a metalearner.

3 . The method of claim 1 , wherein labeling the first set of the seed domains as benign includes:

selecting the first set from the URL feed as domains that have not been seen before;

excluding domains that resolve to sinkhole Internet Protocol addresses;

excluding domains that do not have a valid certificate;

excluding domains that have URLs identified as being created via a domain generation algorithm;

excluding domains identified as impersonating a brand name;

excluding domains associated with a top level domain associated with hosting malicious domains by a third party analysis;

excluding domains that have a consensus score assigned by a plurality of consensus sensors in a consensus feed above a consensus threshold;

excluding domains that have been registered from less time than a registration threshold;

adding domains having .gov and .edu TLDs; and

adding domains belonging to a popularity feed.

4 . The method of claim 1 , wherein labeling the second set of the seed domains as malicious includes:

selecting the second set from the URL feed as domains that have not been seen before;

excluding domains that have a consensus score assigned by a plurality of consensus sensors in a consensus feed below a consensus threshold;

excluding domains that have been registered from more time than a registration threshold; and

adding the seed domains.

5 . The method of claim 1 , wherein classifying the queried domain creates a blocklist of several domains classified as malicious from the URL feed and one or more domains not seen in the URL feed that are proactively identified as malicious based on a relationship with domains identified as malicious by the GNN.

6 . The method of claim 1 , wherein classifying the queried domain returns a real-time response that identifies the queried domain as either benign or malicious.

7 . The method of claim 1 , wherein the queried domain is classified as benign or malicious at hosting infrastructure upstream of content delivery without analyzing content hosted by the queried domain.

8 . A system, comprising a processor and a memory including instructions that when executed by the processor, perform operations including:

extracting seed domains from a uniform resource locator (URL) feed of observed requests for access to domains;

expanding the seed domains via a passive domain name service crawl to include additional domains with the seed domains;

collecting a ground truth, including labeling a first set of the seed domains as benign and a second set of the seed domains as malicious;

constructing a graph neural network (GNN) of the additional domains and the seed domains, wherein each domain of the additional domains and the seed domains is represented as a node in the GNN that includes feature values associated with that domain;

training the GNN to classify unseen domains not associated with a node as either benign or malicious; and

classifying, via the GNN, a queried domain as either benign or malicious.

9 . The system of claim 8 , wherein constructing the GNN includes:

assembling an ensemble of GNN encoders in a model stack; and

combining outputs of the ensemble of GNN encoders via a metalearner.

10 . The system of claim 8 , wherein labeling the first set of the seed domains as benign includes:

selecting the first set from the URL feed as domains that have not been seen before;

excluding domains that resolve to sinkhole Internet Protocol addresses;

excluding domains that do not have a valid certificate;

excluding domains that have URLs identified as being created via a domain generation algorithm;

excluding domains identified as impersonating a brand name;

excluding domains associated with a top level domain associated with hosting malicious domains by a third party analysis;

excluding domains that have a consensus score assigned by a plurality of consensus sensors in a consensus feed above a consensus threshold;

excluding domains that have been registered from less time than a registration threshold;

adding domains having .gov and .edu TLDs; and

adding domains belonging to a popularity feed.

11 . The system of claim 8 , wherein labeling the second set of the seed domains as malicious includes:

selecting the second set from the URL feed as domains that have not been seen before;

excluding domains that have a consensus score assigned by a plurality of consensus sensors in a consensus feed below a consensus threshold;

excluding domains that have been registered from more time than a registration threshold; and

adding the seed domains.

12 . The system of claim 8 , wherein classifying the queried domain creates a blocklist of several domains classified as malicious from the URL feed and one or more domains not seen in the URL feed that are proactively identified as malicious based on a relationship with domains identified as malicious by the GNN.

13 . The system of claim 8 , wherein classifying the queried domain returns a real-time response that identifies the queried domain as either benign or malicious.

14 . The system of claim 8 , wherein the queried domain is classified as benign or malicious at hosting infrastructure upstream of content delivery without analyzing content hosted by the queried domain.

15 . A memory including instructions, that when executed by a processor, perform operations including:

extracting seed domains from a uniform resource locator (URL) feed of observed requests for access to domains;

expanding the seed domains via a passive domain name service crawl to include additional domains with the seed domains;

collecting a ground truth, including labeling a first set of the seed domains as benign and a second set of the seed domains as malicious;

constructing a graph neural network (GNN) of the additional domains and the seed domains, wherein each domain of the additional domains and the seed domains is represented as a node in the GNN that includes feature values associated with that domain;

training the GNN to classify unseen domains not associated with a node as either benign or malicious; and

classifying, via the GNN, a queried domain as either benign or malicious.

16 . The memory of claim 15 , wherein constructing the GNN includes:

assembling an ensemble of GNN encoders in a model stack; and

combining outputs of the ensemble of GNN encoders via a metalearner.

17 . The memory of claim 15 , wherein labeling the first set of the seed domains as benign includes:

selecting the first set from the URL feed as domains that have not been seen before;

excluding domains that resolve to sinkhole Internet Protocol addresses;

excluding domains that do not have a valid certificate;

excluding domains that have URLs identified as being created via a domain generation algorithm;

excluding domains identified as impersonating a brand name;

excluding domains associated with a top level domain associated with hosting malicious domains by a third party analysis;

excluding domains that have a consensus score assigned by a plurality of consensus sensors in a consensus feed above a consensus threshold;

excluding domains that have been registered from less time than a registration threshold;

adding domains having .gov and .edu TLDs; and

adding domains belonging to a popularity feed.

18 . The memory of claim 15 , wherein labeling the second set of the seed domains as malicious includes:

selecting the second set from the URL feed as domains that have not been seen before;

excluding domains that have a consensus score assigned by a plurality of consensus sensors in a consensus feed below a consensus threshold;

excluding domains that have been registered from more time than a registration threshold; and

adding the seed domains.

19 . The memory of claim 15 , wherein classifying the queried domain creates a blocklist of several domains classified as malicious from the URL feed and one or more domains not seen in the URL feed that are proactively identified as malicious based on a relationship with domains identified as malicious by the GNN.

20 . The memory of claim 15 , wherein classifying the queried domain returns a real-time response that identifies the queried domain as either benign or malicious.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2024
From: NABEEL, MOHAMED; KHALIL, ISSA; YU, TING; DENIZ, FATIH
To: QATAR FOUNDATION FOR EDUCATION, SCIENCE AND COMMUNITY DEVELOPMENT
Reel/Frame 067118/0621 →
Continuity (2)
Provisional Application 63492395 · Mar 27, 2023
Related Publication 20240333749A1 · Oct 3, 2024
References Cited (17)
US 9762612B1 · Schiffman · 2017 [cited by examiner]
US 10075417B2 · Baughman · 2018 [cited by examiner]
US 11843622B1 · Tellez · 2023 [cited by examiner]
US 20180069883A1 · Meshi · 2018 [cited by examiner]
US 20180293381A1 · Tseng · 2018 [cited by examiner]
US 20220103592A1 · Semel · 2022 [cited by examiner]
US 20230112092A1 · Tymchenko · 2023 [cited by examiner]
US 20230254338A1 · Melicher · 2023 [cited by examiner]
US 20230362176A1 · Jiang · 2023 [cited by examiner]
US 20240039890A1 · Szurdi · 2024 [cited by examiner]
US 20240046107A1 · Chi · 2024 [cited by examiner]
CN 112910929A · 2021 [cited by applicant]
WO WO2017031505A2 · 2017 [cited by examiner]
WO WO2020237613A1 · 2020 [cited by examiner]
Eshete, et al.; “Malicious Website Detection: Effectiveness and Efficiency Issues”; 2011; IEEE; (4 pages). [cited by applicant]
Zhou, et al.; “Graph neural networks: A review of methods and applications”; ScienceDirect; 2020; (25 pages). [cited by applicant]
Hao, et al.; “PREDATOR: Proactive Recognition and Elimination of Domain Abuse at Time-Of-Registration”; Oct. 2016; ACM Digital Library; (12 pages). [cited by applicant]