IP Library › Granted Patent US 12,609,962
Granted Patent B2
US 12,609,962 · App. 18/755,402 · Granted Apr 21, 2026

Deep learning for malicious URL classification (URLC) with the innocent until proven guilty (IUPG) learning framework

Inventors: Brody James Kutt (Santa Clara, CA); Peng Peng (Santa Clara, CA); Fang Liu (College Station, TX); William Redington Hewlett, II (Mountain View, CA)
Assignee: Palo Alto Networks, Inc.
H04L63/1483G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,962
App. No.
18/755,402
Granted
Apr 21, 2026
Kind
B2
Abstract

Techniques for providing deep learning for malicious URL classification (URLC) using the innocent until proven guilty (IUPG) learning framework are disclosed. In some embodiments, a system, process, and/or computer program product includes storing a set comprising one or more innocent until proven guilty (IUPG) models for static analysis of a sample; performing a static analysis of one or more URLs associated with the sample, wherein performing the static analysis includes using at least one stored IUPG model; and determining that the sample is malicious based at least in part on the static analysis of the one or more URLs associated with the sample, and in response to determining that the sample is malicious, performing an action based on a security policy.

Claims (34)

1 . A system, comprising:

a processor configured to:

receive training data for training an innocent until proven guilty (IUPG) model used to classify malicious content and benign content based on a static analysis, wherein the training data includes a set of input files;

extract a set of tokens from the set of input files to generate a character encoding and a token encoding;

generate an IUPG convolutional neural network (CNN) feature extractor based on the character encoding and the token encoding; and

combine the IUPG CNN feature extractor with another CNN-based feature extractor to generate the IUPG model; and

a memory coupled to the processor and configured to provide the processor with instructions.

2 . The system of claim 1 , wherein the training data includes uniform resource locators (URLs).

3 . The system of claim 1 , wherein the training data includes JavaScript (JS) files.

4 . The system of claim 1 , wherein the training data includes uniform resource locators (URLs), and wherein the processor is configured to preprocess the URLs.

5 . The system of claim 1 , wherein the training data includes uniform resource locators (URLs), wherein the processor is configured to preprocess the URLs, and wherein the preprocessing of the URLs includes to discard scheme and user information of the URLs.

6 . The system of claim 1 , wherein the token encoding includes at least three channels.

7 . The system of claim 1 , wherein the other CNN-based feature extractor includes a Categorical Cross-Entropy (CCE) CNN feature extractor.

8 . A method, comprising:

receiving training data for training an innocent until proven guilty (IUPG) model used to classify malicious content and benign content based on a static analysis, wherein the training data includes a set of input files;

extracting a set of tokens from the set of input files to generate a character encoding and a token encoding;

generating an IUPG convolutional neural network (CNN) feature extractor based on the character encoding and the token encoding; and

combining the IUPG CNN feature extractor with another CNN-based feature extractor to generate the IUPG model.

9 . The method of claim 8 , wherein the training data includes uniform resource locators (URLs).

10 . The method of claim 8 , wherein the training data includes JavaScript (JS) files.

11 . The method of claim 8 , wherein the training data includes uniform resource locators (URLs), and wherein the method further comprises preprocessing the URLs.

12 . The method of claim 8 , wherein the training data includes uniform resource locators (URLs), wherein the method further comprises preprocessing the URLs, and wherein the preprocessing of the URLs includes discarding scheme and user information of the URLs.

13 . The method of claim 8 , wherein the token encoding includes at least three channels.

14 . The method of claim 8 , wherein the other CNN-based feature extractor includes a Categorical Cross-Entropy (CCE) CNN feature extractor.

15 . A computer program product embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

receiving training data for training an innocent until proven guilty (IUPG) model used to classify malicious content and benign content based on a static analysis, wherein the training data includes a set of input files;

extracting a set of tokens from the set of input files to generate a character encoding and a token encoding;

generating an IUPG convolutional neural network (CNN) feature extractor based on the character encoding and the token encoding; and

combining the IUPG CNN feature extractor with another CNN-based feature extractor to generate the IUPG model.

16 . The computer program product recited in claim 15 , wherein the training data includes uniform resource locators (URLs).

17 . The computer program product recited in claim 15 , wherein the training data includes JavaScript (JS) files.

18 . The computer program product recited in claim 15 , wherein the token encoding includes at least three channels.

19 . The computer program product recited in claim 15 , wherein the other CNN-based feature extractor includes a Categorical Cross-Entropy (CCE) CNN feature extractor.

20 . The computer program product recited in claim 15 , wherein the training data includes uniform resource locators (URLs), and further comprising computer instructions for preprocessing the URLs, wherein the preprocessing of the URLs includes discarding scheme and user information of the URLs.

Continuity (5)
Continuation 17408269 · Aug 20, 2021
Continuation In Part 17331549 · May 26, 2021
Provisional Application 63193545 · May 26, 2021
Provisional Application 63034843 · Jun 4, 2020
Related Publication 20240364738A1 · Oct 31, 2024
References Cited (104)
US 7966654B2 · Crawford · 2011 [cited by examiner]
US 8521667B2 · Zhu · 2013 [cited by examiner]
US 9178901B2 · Xue · 2015 [cited by examiner]
US 10203968B1 · Lawson · 2019 [cited by applicant]
US 10616274B1 · Chang · 2020 [cited by examiner]
US 11336689B1 · Miramirkhani · 2022 [cited by examiner]
US 11379577B2 · Patel · 2022 [cited by examiner]
US 11550911B2 · Kutt · 2023 [cited by examiner]
US 11586881B2 · Gronát · 2023 [cited by examiner]
US 11611582B2 · Pryce · 2023 [cited by examiner]
US 11799905B2 · Jones · 2023 [cited by examiner]
US 11856003B2 · Kutt · 2023 [cited by examiner]
US 11924245B2 · Grewal · 2024 [cited by examiner]
US 12063248B2 · Kutt · 2024 [cited by examiner]
US 12261853B2 · Kutt · 2025 [cited by examiner]
US 20070118893A1 · Crawford · 2007 [cited by examiner]
US 20100186088A1 · Banerjee · 2010 [cited by examiner]
US 20120158626A1 · Zhu · 2012 [cited by examiner]
US 20170236118A1 · Laracey · 2017 [cited by examiner]
US 20190104154A1 · Kumar · 2019 [cited by examiner]
US 20190213325A1 · Mckerchar · 2019 [cited by applicant]
US 20200162484A1 · Solis Agea · 2020 [cited by examiner]
US 20210014273A1 · Kipp · 2021 [cited by examiner]
US 20210149788A1 · Downie · 2021 [cited by examiner]
US 20210342651A1 · Shibahara · 2021 [cited by examiner]
US 20210385232A1 · Kutt · 2021 [cited by examiner]
US 20220046057A1 · Kutt · 2022 [cited by examiner]
US 20230082481A1 · Azarafrooz · 2023 [cited by examiner]
US 20230124193A1 · Ito · 2023 [cited by applicant]
US 20240364738A1 · Kutt · 2024 [cited by examiner]
CN 109492395 · 2019 [cited by applicant]
CN 109684835 · 2019 [cited by applicant]
CN 109840417 · 2019 [cited by applicant]
CN 110399300 · 2021 [cited by applicant]
JP 2016091549 · 2016 [cited by applicant]
JP 2017162244 · 2017 [cited by applicant]
JP 2018160172 · 2018 [cited by applicant]
Starov et al., Detecting Malicious Campaigns in Obfuscated JavaScript with Scalable Behavioral Analysis, 2019 IEEE Security and Privacy Workshops (SPW), 2019, pp. 218-223. [cited by applicant]
Stokes et al., Neural Classification of Malicious Scripts: A Study with JavaScript and VBScript, pp. 1-20, 2018. [cited by applicant]
Stokes et al., ScriptNet: Neural Static Analysis for Malicious JavaScript Detection, Apr. 1, 2019. [cited by applicant]
Stuart P. Lloyd, Least Squares Quantization in PCM, IEEE Transactions on Information Theory, vol. IT-28, No. 2, Mar. 1982, pp. 129-137. [cited by applicant]
Suciu et al., Exploring Adversarial Examples in Malware Detection, IEEE SPW, 2019. [cited by applicant]
Teuvo Kohonen, The Self-Organizing Map, Proceedings of the IEEE, vol. 78, No. 9, Sep. 1990. [cited by applicant]
Tin Kam Ho, 1995, Random Decision Forests, In Proceedings of 3rd International Conference on Document Analysis and Recognition, vol. 1. IEEE, pp. 278-282. [cited by applicant]
Tufano et al., Deep Learning Similarities from Different Representations of Source Code, MSR'18, May 28-29, 2018. [cited by applicant]
Van Der Maaten, Visualizing Data Using t-SNE, Journal of Machine Learning Research 9, 2008. [cited by applicant]
Walsh et al., GitHUB, Acorn: A Small, Fast, JaveScript-based parser, 2017. [cited by applicant]
Wang et al., A Deep Learning Approach for Detecting Malicious JavaScript code, Security and Communication Networks, Security Comm. Networks 2016, 9:1520-1534, Published online Feb. 11, 2016 in Wiley Online Library (wile… [cited by applicant]
Wang et al., JSDC: A Hybrid Approach for JavaScript Malware Detection and Classification, ACM, pp. 109-120, 2015. [cited by applicant]
Wiyatno et al., Adversarial Examples in Modern Machine Learning: A Review, Nov. 15, 2019. [cited by applicant]
Yang et al., Robust Classification with Convolutional Prototype Learning, pp. 3474-3482, IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Yoon Kim, Convolutional Neural Networks for Sentence Classification, Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1746-1751. [cited by applicant]
Yu et al., Hyper-Parameter Optimization: A Review of Algorithms and Applications, pp. 1-56, 2020. [cited by applicant]
Zhang et al., Character-Level Convolutional Networks for Text Classification, Apr. 4, 2016, pp. 1-9. [cited by applicant]
Abadi et al., TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems, Nov. 9, 2015, pp. 1-19. [cited by applicant]
Abien Fred M. Agarap, Deep Learning Using Rectified Linear Units (ReLU), 2018. [cited by applicant]
Andra Zaharia, The Ultimate Guide to Angler Exploit Kit for Non-Technical People (Updated), May 18, 2016. [cited by applicant]
Arthur et al., k-means++: The Advantages of Careful Seeding, Proceedings of the 18th Annual ACM-SIAM Symposium on Discrete Algorithms, 2007. [cited by applicant]
Baum et al., Statistical Inference for Probabilistic Function of Finite State Markov Chains, Apr. 4, 1996, pp. 1554-1563. [cited by applicant]
Bergstra et al., Algorithms for Hyper-Parameter Optimization, pp. 1-9, 2011. [cited by applicant]
Brad Duncan, Understanding the Angler Exploit Kit—Part 1: Exploit Kit Fundamentals, Jun. 3, 2016. [cited by applicant]
Bromley et al., Signature Verification using a “Siamese” Time Delay Neural Network, 1994, pp. 737-744. [cited by applicant]
Chen et al., Detecting Filter List Evasion with Event-Loop-Turn Granularity JavaScript Signatures, IEEE, 2021. [cited by applicant]
Chen et al., Robust Out-of-distribution Detection for Neural Networks, Association for the Advancement of Artificial Intelligence, 2022. [cited by applicant]
Cordella et al., A Method for Improving Classification Reliability of Multilayer Perceptrons, IEEE Transactions on Neural Networks, vol. 6, No. 5, pp. 1140-1147, 1995. [cited by applicant]
Cormen et al., Probabilistic Analysis and Randomized Algorithms, Introduction to Algorithms: Second Edition, pp. 94-99, 2001. [cited by applicant]
Cortes et al., Learning with Rejection, 2016. [cited by applicant]
David H. Wolpert, Stacked Generalization, 1992. [cited by applicant]
Fass et al., HIDENOSEEK: Camouflaging Malicious JavaScript in Benign AST's, ACM, 2019. [cited by applicant]
Fass et al., JSTAP: A Static Pre-Filter for Malicious JavaScript Detection, ACSAC, 2019. [cited by applicant]
Fass et al., JSTAP: A Static Pre-Filter for Malicious JavaScript Detection, pp. 1-28, ACSAC, Nov. 12, 2019. [cited by applicant]
Fraser Howard, A Closer Look at the Angler Exploit Kit, Sophos News, Jul. 21, 2015. [cited by applicant]
Gaurav Sood, virustotal: R Client for the virustotal API, 2017. [cited by applicant]
Geifman et al., Selective Classification for Deep Neural Networks, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. [cited by applicant]
Goodfellow et al., Explaining and Harnessing Adversarial Examples, pp. 1-11, ICLR, 2015. [cited by applicant]
Hendrycks et al., A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks, Published as a Conference Paper at ICLR 2017, Oct. 3, 2018. [cited by applicant]
I.T. Jolliffe, Principal Component Analysis, Second Edition, Springer Series in Statistics, 1986. [cited by applicant]
Johnson et al., Effective Use of Word Order for Text Categorization with Convolutional Neural Networks, Mar. 26, 2015. [cited by applicant]
Johnson et al., Semi-Supervised Convolutional Neural Networks for Text Categorization via Region Embedding, Nov. 1, 2015. [cited by applicant]
Kevin P. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press, 2012. [cited by applicant]
Kiefer et al., Stochastic Estimation of the Maximum of a Regression Function, Presented to the American Mathematical Society at New York on Apr. 25, 1952, pp. 462-466. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization”, arXiv preprint, Dec. 22, 2014, pp. 1-15, arXiv:1412.6980v9 [cs.LG]. [cited by applicant]
Landwehr et al., A Taxonomy of Computer Program Security Flaws, with Examples, ACM Computing Surverys, 26, 3 (Sep. 1994). [cited by applicant]
Le et al., URLNet: Learning a URL Representation with Deep Learning for Malicious URL Detection, In Proceedings of ACM Conference, Washington, DC, USA, Jul. 2017 (Conference'17), 13 pages. [cited by applicant]
Lecun et al., Gradient-Based Learning Applied to Document Recognition, Proc. of the IEEE, pp. 2278-2324, Nov. 1998. [cited by applicant]
Lecun et al., The MNIST Database of Handwritten Digits, https://web.archive.org/web/20210514192806/http:/yann.lecun.com/exdb/mnist/, pp. 1-7, May 14, 2021. [cited by applicant]
Lee et al., A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada. [cited by applicant]
Li et al., Deep Learning for Case-Based Reasoning Through Prototypes: A Neural Network That Explains Its Predictions, pp. 3530-3537, 32nd AAAI Conference, 2018. [cited by applicant]
Liu et al., Meta-Learning Based Prototype-relation Network for Few-shot Classification, Neurocomputing, 2019. [cited by applicant]
Liu et al., Prototype Propagation Networks (PPN) for Warkly-supervised Few-shot Learning on Category Graph, 2019. [cited by applicant]
Mettes et al., Hyperspherical Prototype Networks, 33rd NeurIPS, 2019. [cited by applicant]
Microsoft, Microsoft Security Intelligence, JS/Nemucod Threat Description, Mar. 29, 2015. [cited by applicant]
Mucherino et al., k-Nearest Neighbor Classification, Data Mining in Agriculture, Springer Optimization and Its Applications 34, pp. 83-106, 2009. [cited by applicant]
Nikolaev et al., Exploit Kit Website Detection Using HTTP Proxy Logs, pp. 120-125, ICNCC, 2016. [cited by applicant]
Nova et al., A Review of Learning Vector Quantization Classifiers, Sep. 23, 2015. [cited by applicant]
Pochat et al., Rigging Research Results by Manipulating Top Websites Rankings, pp. 1-16, 2018. [cited by applicant]
Ponemon Institute, The Cost of Malware Containment, Sponsored by Damballa, Jan. 2015. [cited by applicant]
Sammut et al., Encyclopedia of Machine Learning, Springer, TF-IDF, pp. 986-987, 2011. [cited by applicant]
Santiago Ontanon, An Overview of Distance and Similarity Functions for Structured Data, Feb. 18, 2020. [cited by applicant]
Sato et al., Generalized Learning Vector Quantization, pp. 423-429, NIPS, 1995. [cited by applicant]
Sehwag et al., Analyzing the Robustness of Open-World Machine Learning, AlSec '19, Nov. 15, 2019, London, United Kingdom. [cited by applicant]
Shalev et al., Out-of-Distribution Detection Using Multiple Semantic Label Representations, pp. 1-11, 32nd Conference on Neural Information Processing Systems, 2018. [cited by applicant]
Snell, Prototypical Networks for Few-shot Learning, NIPS, 2017. [cited by applicant]
Springer, SpringerLink, Encyclopedia of Machine Learning, TF-IDF, downloaded May 23, 2021. [cited by applicant]