IP Library › Granted Patent US 12,671,716
Granted Patent B2
US 12,671,716 · App. 17/341,535 · Granted Jun 30, 2026

Machine learning to determine command and control sites

Inventors: Changsha Ma (Palo Alto, CA); Loc Bui (San Jose, CA); Dianhuan Lin (Sunnyvale, CA); Rex Shang (Los Altos, CA); Bryan Lee (San Jose, CA); Shudong Zhou (Fremont, CA); Howie Xu (Palo Alto, CA); Naveen Selvan (Mohali, IN); Nirmal Singh (Chandigarh, IN); Deepen Desai (San Ramon, CA); Parnit Sainion (Morgan Hill, CA); Narinder Paul (Sunnyvale, CA)
Assignee: Zscaler, Inc.
H04L63/1483G06N5/04G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,671,716
App. No.
17/341,535
Filed
Jun 8, 2021
Granted
Jun 30, 2026
Kind
B2
Art Unit
2438
USPC
726/23
Abstract

Systems and methods include receiving a domain for a determination of a likelihood the domain is a command and control site; analyzing the domain with an ensemble of a plurality of trained machine learning models including a Uniform Resource Locator (URL) model that analyzes lexical features of a hostname of the domain and an artifact model that analyzes content features of a webpage associated with the domain; and combining results of the ensemble to predict the likelihood the domain is a command and control site.

Claims (40)

1 . A method comprising the steps of:

receiving a domain for a determination of a likelihood the domain is a command and control site;

analyzing the domain with an ensemble of a plurality of trained machine learning models including a Uniform Resource Locator (URL) model that analyzes lexical features of a hostname of the domain and an artifact model that analyzes content features of a webpage associated with the domain;

combining, by a command-and-control (C2) model of the ensemble, a URL score generated by the URL model and an artifact score generated by the artifact model of the ensemble, the C2 model being a trained gradient-boosted decision tree model configured to receive the URL score and the artifact score as inputs and to output a final score to predict the likelihood the domain is a command and control site; and

performing, by an enforcement node, an action responsive to the likelihood the domain is the command and control site, wherein the action is one or more of adding the domain to blocked list and causing a block of the domain.

2 . The method of claim 1 , wherein the steps include

performing the receiving responsive to a determination by a domain reputation process of a likelihood the domain is malicious.

3 . The method of claim 1 , wherein the steps include

prior to the analyzing, training the URL model and the artifact model.

4 . The method of claim 3 , wherein the training includes

using labeled log data from a cloud-based system that performs monitoring of a plurality of users.

5 . The method of claim 4 , wherein the labeled log data is based on a content classification process.

6 . The method of claim 1 , wherein the ensemble further includes transaction patterns to the domain.

7 . The method of claim 1 , wherein the ensemble further includes an analysis of a reputation of the domain.

8 . The method of claim 1 , wherein the ensemble further includes a malware relation of the domain.

9 . A non-transitory computer-readable medium comprising instructions that, when executed, cause one or more processors to perform steps of:

receiving a domain for a determination of a likelihood the domain is a command and control site;

analyzing the domain with an ensemble of a plurality of trained machine learning models including a Uniform Resource Locator (URL) model that analyzes lexical features of a hostname of the domain and an artifact model that analyzes content features of a webpage associated with the domain;

combining, by a command-and-control (C2) model of the ensemble, a URL score generated by the URL model and an artifact score generated by the artifact model of the ensemble, the C2 model being a trained gradient-boosted decision tree model configured to receive the URL score and the artifact score as inputs and to output a final score to predict the likelihood the domain is a command and control site; and

perform, by a node, an action responsive to the likelihood the domain is the command and control site, wherein the action is one or more adding the domain to a blocked list and causing a block of the domain.

10 . The non-transitory computer-readable medium of claim 9 , wherein the steps include

performing the receiving responsive to a determination by a domain reputation process of a likelihood the domain is malicious.

11 . The non-transitory computer-readable medium of claim 9 , wherein the steps include

prior to the analyzing, training the URL model and the artifact model.

12 . The non-transitory computer-readable medium of claim 11 , wherein the training includes

using labeled log data from a cloud-based system that performs monitoring of a plurality of users.

13 . The non-transitory computer-readable medium of claim 12 , wherein the labeled log data is based on a content classification process.

14 . The non-transitory computer-readable medium of claim 9 , wherein the ensemble further includes transaction patterns to the domain.

15 . The non-transitory computer-readable medium of claim 9 , wherein the ensemble further includes an analysis of a reputation of the domain.

16 . The non-transitory computer-readable medium of claim 9 , wherein the ensemble further includes a malware relation of the domain.

17 . A cloud-based system comprising a plurality of interconnected nodes, each node comprises:

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, cause the node to:

receive a domain for a determination of a likelihood the domain is a command and control site;

analyze the domain with an ensemble of a plurality of trained machine learning models including a Uniform Resource Locator (URL) model that analyzes lexical features of a hostname of the domain and an artifact model that analyzes content features of a webpage associated with the domain;

combine, by a command-and-control (C2) model of the ensemble, a URL score generated by the URL model and an artifact score generated by the artifact model of the ensemble, the C2 model being a trained gradient-boosted decision tree model configured to receive the URL score and the artifact score as inputs and to output a final score to predict the likelihood the domain is a command and control site; and

perform an action responsive to the likelihood the domain is the command and control site, wherein the action is one or more of adding the domain to a blocked list and causing a block of the domain.

18 . The method of claim 1 , wherein the combining includes aggregating predictions of the plurality of trained machine learning models over a predetermined time period to increase a confidence level of the final score.

19 . The method of claim 1 , wherein performing the action includes automatically updating a blocklist database stored at the node at an hourly or daily frequency with domains predicted as command and control sites based on the final score.

20 . The method of claim 1 , wherein analyzing the domain with the ensemble of the plurality of trained machine learning models further includes aggregating transaction data by company identifier, user identifier, hostname, request, response, and user agent over a predefined time for generating input features for the plurality of trained machine learning models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2021
From: MA, CHANGSHA; BUI, LOC; LIN, DIANHUAN; SHANG, REX; LEE, BRYAN; ZHOU, SHUDONG; XU, HOWIE; SELVAN, NAVEEN; SINGH, NIRMAL; DESAI, DEEPEN; SAINION, PARNIT; PAUL, NARINDER
To: ZSCALER, INC.
Reel/Frame 056464/0647 →
Priority Claims (1)
IN 202111018567 · Apr 22, 2021 · national
Continuity (3)
Continuation In Part 17075991 · Oct 21, 2020
Continuation In Part 16889885 · Jun 2, 2020
Related Publication 20210377304A1 · Dec 2, 2021
References Cited (31)
US 8959626B2 · Niemela · 2015 [cited by examiner]
US 9065850B1 · Sobrier · 2015 [cited by applicant]
US 9152789B2 · Natarajan et al. · 2015 [cited by applicant]
US 9838407B1 · Oprea · 2017 [cited by examiner]
US 10142362B2 · Weith et al. · 2018 [cited by applicant]
US 10419477B2 · Desai et al. · 2019 [cited by applicant]
US 10498605B2 · Weith et al. · 2019 [cited by applicant]
US 20070233477A1 · Halowani et al. · 2007 [cited by applicant]
US 20100115621A1 · Staniford et al. · 2010 [cited by applicant]
US 20160021127A1 · Yan · 2016 [cited by examiner]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20160352772A1 · O'Connor · 2016 [cited by examiner]
US 20170063886A1 · Muddu et al. · 2017 [cited by applicant]
US 20170250968A1 · Licht · 2017 [cited by examiner]
US 20170372071A1 · Saxe · 2017 [cited by examiner]
US 20180139235A1 · Desai et al. · 2018 [cited by applicant]
US 20180150758A1 · Niininen et al. · 2018 [cited by applicant]
US 20180150783A1 · Xu · 2018 [cited by examiner]
US 20180293381A1 · Tseng et al. · 2018 [cited by applicant]
US 20190281073A1 · Weith et al. · 2019 [cited by applicant]
US 20190319972A1 · Desai · 2019 [cited by applicant]
US 20200059451A1 · Huang · 2020 [cited by examiner]
US 20200142810A1 · Zingade · 2020 [cited by examiner]
GB 2575052A · 2020 [cited by applicant]
Zargar, “A Survey of Defense Mechanisms Against Distributed Denial of Service (DDoS) Flooding Attacks”, 2013, IEEE, vol. 15, pp. 2046-2065 (Year: 2013). [cited by examiner]
Sadaf, Intrusion Detection Based on Autoencoder and Isolation Forest in Fog Computing, 2020, IEEE, 167059-167067 (Year: 2020). [cited by examiner]
Lalouani et al., “Multi-observable reputation scoring system for flagging suspicious user sessions,” Computer Networks, vol. 182, Aug. 8, 2022, pp. 1-13. [cited by applicant]
Jan. 10, 2022, Extended European Search Report issued for European Patent Application No. EP 21 19 1871. [cited by applicant]
Jordaney, Roberto, et al., “Transcend: Detecting concept drift in malware classification models,” 26th {Usenix} Security Symposium ({Usenix} Security 17), 2017. [cited by applicant]
Kantchelian, Alex, J. D. Tygar, and Anthony Joseph, “Evasion and hardening of tree ensemble classifiers,” International Conference on Machine Learning, 2016. [cited by applicant]
Tolomei, Gabriele, et al., “Interpretable predictions of tree-based ensembles via actionable feature tweaking,” Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 20… [cited by applicant]