IP Library › Granted Patent US 12,346,445
Granted Patent B2
US 12,346,445 · App. 18/488,498 · Granted Jul 1, 2025

Systems and methods for intelligent machine learning-based malware detection

Inventors: Huihsin Tseng (Cupertino, CA); Hao Xu (Palo Alto, CA); Jian L. Zhen (Palo Alto, CA)
Assignee: Zscaler, Inc.
G06F21/566G06N5/01G06N20/00G06N20/20G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,445
App. No.
18/488,498
Granted
Jul 1, 2025
Kind
B2
Abstract

The methods described herein include receiving a plurality of packets associated with a file, each of the plurality of packets comprising content, and a source domain; extracting one or more features from content of a first packet of the plurality of packets; applying a trained machine learning model to the extracted one or more features to determine a probability of maliciousness associated with the first packet; responsive to determining that the probability maliciousness of the first packet is between a first threshold value and a second threshold value, labeling the first packet as having an uncertain maliciousness; extracting one or more features from content of a second packet of the plurality of packets; and applying the trained machine learning model to the extracted one or more features of the first packet and the second packet to determine a probability of maliciousness associated with the second packet.

Claims (38)

1. A non-transitory computer-readable storage medium having computer-readable code stored thereon for programming at least one processor to perform steps of:

receiving a plurality of packets associated with a file, each of the plurality of packets comprising content, and a source domain;

extracting one or more features from content of a first packet of the plurality of packets;

applying a trained machine learning model to the extracted one or more features to determine a probability of maliciousness associated with the first packet;

responsive to determining that the probability maliciousness of the first packet is between a first threshold value and a second threshold value, labeling the first packet as having an uncertain maliciousness;

extracting one or more features from content of a second packet of the plurality of packets; and

applying the trained machine learning model to the extracted one or more features of the first packet and the second packet to determine a probability of maliciousness associated with the second packet.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the steps further comprise:

responsive to labeling the first packet as having an uncertain maliciousness, storing the first packet and its one or more features.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the steps further comprise:

converting the content of the first packet and second packet of the plurality of packets into a digital representation.

4. The non-transitory computer-readable storage medium of claim 3 , wherein the digital representation is any of a decimal representation, a binary representation, a hexadecimal representation, a tokenized script, and a tokenized domain.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the steps further comprise:

labeling the file based on the probability of maliciousness.

6. The non-transitory computer-readable storage medium of claim 1 , wherein the file type is one of a portable executable (PE) file, a portable document format (PDF) file, a Dynamic Loaded Library (DLL), a JavaScript (JS) file, a Hypertext Markup Language (HTML) file, and a Microsoft Office File.

7. The non-transitory computer-readable storage medium of claim 1 , wherein the trained machine learning model comprises one or more decision trees.

8. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more features include n-gram features.

9. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more features include an entropy feature.

10. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more features include a domain feature.

11. A method comprising steps of:

receiving a plurality of packets associated with a file, each of the plurality of packets comprising content, and a source domain;

extracting one or more features from content of a first packet of the plurality of packets;

applying a trained machine learning model to the extracted one or more features to determine a probability of maliciousness associated with the first packet;

responsive to determining that the probability maliciousness of the first packet is between a first threshold value and a second threshold value, labeling the first packet as having an uncertain maliciousness;

extracting one or more features from content of a second packet of the plurality of packets; and

applying the trained machine learning model to the extracted one or more features of the first packet and the second packet to determine a probability of maliciousness associated with the second packet.

12. The method of claim 11 , wherein the steps further comprise:

responsive to labeling the first packet as having an uncertain maliciousness, storing the first packet and its one or more features.

13. The method of claim 11 , wherein the steps further comprise:

converting the content of the first packet and second packet of the plurality of packets into a digital representation.

14. The method of claim 13 , wherein the digital representation is any of a decimal representation, a binary representation, a hexadecimal representation, a tokenized script, and a tokenized domain.

15. The method of claim 11 , wherein the steps further comprise:

labeling the file based on the probability of maliciousness.

16. The method of claim 11 , wherein the file type is one of a portable executable (PE) file, a portable document format (PDF) file, a Dynamic Loaded Library (DLL), a JavaScript (JS) file, a Hypertext Markup Language (HTML) file, and a Microsoft Office File.

17. The method of claim 11 , wherein the trained machine learning model comprises one or more decision trees.

18. The method of claim 11 , wherein the one or more features include n-gram features.

19. The method of claim 11 , wherein the one or more features include an entropy feature.

20. The method of claim 11 , wherein the one or more features include a domain feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2023
From: TRUSTPATH, INC.
To: ZSCALER, INC.
Reel/Frame 065253/0224 →
Continuity (5)
Continuation 17724744 · Apr 20, 2022
Continuation 17067854 · Oct 12, 2020
Continuation 15946706 · Apr 5, 2018
Provisional Application 62483102 · Apr 7, 2017
Related Publication 20240045963A1 · Feb 8, 2024
References Cited (17)
US 9152789B2 · Natarajan et al. · 2015 [cited by applicant]
US 20070233477A1 · Halowani et al. · 2007 [cited by applicant]
US 20100115621A1 · Staniford et al. · 2010 [cited by applicant]
US 20130185230A1 · Zhu · 2013 [cited by examiner]
US 20130198119A1 · Eberhardt, III · 2013 [cited by examiner]
US 20170063886A1 · Sudhakar et al. · 2017 [cited by applicant]
US 20180150758A1 · Niininen et al. · 2018 [cited by applicant]
US 20180191629A1 · Biederman et al. · 2018 [cited by applicant]
US 20180293381A1 · Tseng et al. · 2018 [cited by applicant]
US 20200076835A1 · Ladnai · 2020 [cited by examiner]
US 20200302058A1 · Kenyon · 2020 [cited by examiner]
US 20210019339A1 · Ghulati · 2021 [cited by examiner]
US 20220150275A1 · McNee · 2022 [cited by examiner]
US 20220309360A1 · Zohrevand · 2022 [cited by examiner]
Jordaney, Roberto, et al. “Transcend: Detecting concept drift in malware classification models.” 26th {USENIX} Security Symposium ({USENIX} Security 17). 2017. [cited by applicant]
Kantchelian, Alex, J. D. Tygar, and Anthony Joseph. “Evasion and hardening of tree ensemble classifiers.” International Conference on Machine Learning. 2016. [cited by applicant]
Tolomei, Gabriele, et al. “Interpretable predictions of tree-based ensembles via actionable feature tweaking.” Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 201… [cited by applicant]