IP Library Granted Patent US 12,244,637
Granted Patent B1
US 12,244,637 · App. 18/437,521 · Granted Mar 4, 2025

Machine learning powered cloud sandbox for malware detection

Inventors: Xinjun Zhang (San Jose, CA); Ari Azarafrooz (Rancho Santa Margarita, CA); Zhenxin Zhan (Fremont, CA); Ghanashyam Satpathy (Bangalore, IN); Hung-Ming Chen (Kaohsiung, TW)
Assignee: Netskope, Inc.
H04L63/1441G06F21/53H04L41/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,244,637
App. No.
18/437,521
Filed
Feb 9, 2024
Granted
Mar 4, 2025
Kind
B1
Art Unit
2438
USPC
726/22
Abstract

A cloud-based network security system (NSS) is described. The NSS uses a sandbox to safely detonate and extract information about a document and uses machine learning algorithms to analyze the information to predict whether the document contains malicious software. Specifically, during the detonation, static and dynamic information about the document is captured in the sandbox as well as character strings from images in the document. The dynamic information (and sometimes the static information) is input to an AI or machine learning model trained to provide an output indicating a prediction of whether the document contains malware. The character strings are compared with a batch of phishing keywords to generate a heuristic score. A validation engine combines the output from the AI or machine learning model and the heuristic score to classify the document as malicious or clean. Security policies can then be applied based on the classification.

Claims (79)

1. A computer-implemented method, comprising:

intercepting, by a cloud-based network security system, a request to access a document;

obtaining, by the cloud-based network security system, the document;

detonating, by the cloud-based network security system, the document in a sandbox of the cloud-based network security system;

in response to the detonating, extracting, by the cloud-based network security system, dynamic information about the document;

extracting, by the cloud-based network security system, character strings from images in the document during the detonating in the sandbox;

providing, by the cloud-based network security system, the dynamic information as input to an artificial intelligence model trained to provide an output indicating a prediction of whether the document contains malware based on the input;

generating, by the cloud-based network security system, a heuristic score based on comparing the character strings extracted from the document to a batch of phishing keywords;

providing, by the cloud-based network security system, the output of the artificial intelligence model and the heuristic score as input to a verdict engine, wherein the verdict engine combines the output of the artificial intelligence model and the heuristic score to classify the document as one of malicious or clean; and

implementing, by the cloud-based network security system, a security policy based at least in part on the classification of the document.

2. The computer-implemented method of claim 1 , wherein the extracting the dynamic information about the document comprises:

analyzing behavior of the document during detonation in the sandbox; and

extracting data from the behavior, wherein the dynamic information about the document comprises:

a set of behavior features of the document exhibited during the detonating,

a size of a process tree spawned by the detonating, and

a set of signature features exhibited during the detonating.

3. The computer-implemented method of claim 2 , wherein the set of behavior features comprises at least one of: visited files, visited paths, and pathways explored by processes in the sandbox.

4. The computer-implemented method of claim 2 , wherein the set of signature features comprises:

a signature vector having a dimension for each of a plurality of software signatures, wherein a value of each dimension indicates whether the document invoked the respective software signature in the sandbox, and

severity scores for each of the software signatures the document invoked.

5. The computer-implemented method of claim 2 , wherein the providing the dynamic information as input comprises:

generating a feature vector representing at least a portion of the dynamic information; and

providing the feature vector as the input to the artificial intelligence model.

6. The computer-implemented method of claim 1 , wherein the artificial intelligence model comprises a gradient boosting tree algorithm.

7. The computer-implemented method of claim 1 , wherein the extracting the character strings from the images in the document comprises analyzing, by the cloud-based network security system in the sandbox, the document with optical character recognition.

8. The computer-implemented method of claim 1 , wherein the heuristic score comprises a count of all matches identified during the comparison of the character strings to the batch of the phishing keywords.

9. The computer-implemented method of claim 1 , wherein:

a heuristic rule is triggered based on the heuristic score exceeding a threshold heuristic value;

the output of the artificial intelligence model comprises a prediction score indicating the prediction; and

the verdict engine:

classifies document as malicious based on:

determining the heuristic rule is triggered; and

the prediction score exceeds a pre-defined prediction threshold value, and

classifies the document as clean based on:

determining the heuristic rule is not triggered; or

the prediction score does not exceed the pre-defined prediction threshold value.

10. The computer-implemented method of claim 1 , further comprising:

extracting static information about the document, wherein the providing the dynamic information as input to the artificial intelligence model comprises providing both the dynamic information and the static information as the input to the artificial intelligence model.

11. A cloud-based network security system, comprising:

one or more processors; and

one or more memories having stored thereon instructions that, upon execution by the one or more processors, cause the one or more processors to:

intercept a request to access a document;

obtain the document;

detonate the document in a sandbox of the cloud-based network security system;

in response to the detonating, extract dynamic information about the document;

extract character strings from images in the document during the detonating in the sandbox;

provide the dynamic information as input to an artificial intelligence model trained to provide an output indicating a prediction of whether the document contains malware based on the input;

generate a heuristic score based on comparing the character strings extracted from the document to a batch of phishing keywords;

provide the output of the artificial intelligence model and the heuristic score as input to a verdict engine, wherein the verdict engine combines the output of the artificial intelligence model and the heuristic score to classify the document as one of malicious or clean; and

implement a security policy based on the classification of the document.

12. The cloud-based network security system of claim 11 , wherein the instructions to extract the dynamic information about the document comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

analyze behavior of the document during detonation in the sandbox; and

extract data from the behavior, wherein the dynamic information about the document comprises:

a set of behavior features of the document exhibited during the detonating,

a size of a process tree spawned by the detonating, and

a set of signature features of the document exhibited during the detonating.

13. The cloud-based network security system of claim 12 , wherein the set of behavior features comprises at least one of: visited files, visited paths, and pathways explored by processes in the sandbox.

14. The cloud-based network security system of claim 12 , wherein the set of signature features comprises:

a signature vector having a dimension for each of a plurality of software signatures, wherein a value of each dimension indicates whether the document invoked the respective software signature in the sandbox, and

severity scores for each of the software signatures the document invoked.

15. The cloud-based network security system of claim 12 , wherein the instructions to provide the dynamic information as input comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

generate a feature vector representing the dynamic information; and

provide the feature vector as the input to the artificial intelligence model.

16. The cloud-based network security system of claim 11 , wherein the artificial intelligence model comprises a gradient boosting tree algorithm.

17. The cloud-based network security system of claim 11 , wherein the instructions to extract the character strings from the images in the document comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

analyze the document during the detonation in the sandbox with optical character recognition.

18. The cloud-based network security system of claim 11 , wherein the heuristic score comprises a count of all matches identified during the comparison of the character strings to the batch of phishing keywords.

19. The cloud-based network security system of claim 11 , wherein:

a heuristic rule is triggered based on the heuristic score exceeding a threshold heuristic value;

the output of the artificial intelligence model comprises a prediction score indicating the prediction; and

the verdict engine comprises instructions that, upon execution by the one or more processors, cause the one or more processors to:

classify document as malicious based on:

determining the heuristic rule is triggered; and

the prediction score exceeds a pre-defined prediction threshold value, and classify the document as clean based on:

determining the heuristic rule is not triggered; or

the prediction score does not exceed the pre-defined prediction threshold value.

20. The cloud-based network security system of claim 11 , wherein the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

extract static data about the document; and

provide both the static information and the dynamic information as the input to the artificial intelligence model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2024
From: ZHANG, XINJUN; AZARAFROOZ, ARI; ZHAN, ZHENXIN; SATPATHY, GHANASHYAM; CHEN, HUNG-MING
To: NETSKOPE, INC.
Reel/Frame 066861/0656 →
References Cited (128)
US 5440723A · Arnold et al. · 1995 [cited by applicant]
US 6513122B1 · Magdych et al. · 2003 [cited by applicant]
US 6622248B1 · Hirai · 2003 [cited by applicant]
US 7080408B1 · Pak et al. · 2006 [cited by applicant]
US 7298864B2 · Jones · 2007 [cited by applicant]
US 7376719B1 · Shafer et al. · 2008 [cited by applicant]
US 7735116B1 · Gauvin · 2010 [cited by applicant]
US 7966654B2 · Crawford · 2011 [cited by applicant]
US 8000329B2 · Fendick et al. · 2011 [cited by applicant]
US 8296178B2 · Hudis et al. · 2012 [cited by applicant]
US 8793151B2 · DelZoppo et al. · 2014 [cited by applicant]
US 8839417B1 · Jordan · 2014 [cited by applicant]
US 9177142B2 · Montoro · 2015 [cited by applicant]
US 9197601B2 · Pasdar · 2015 [cited by applicant]
US 9225734B1 · Hastings · 2015 [cited by applicant]
US 9231968B2 · Fang et al. · 2016 [cited by applicant]
US 9280678B2 · Redberg · 2016 [cited by applicant]
US 9811662B2 · Sharpe et al. · 2017 [cited by applicant]
US 10084825B1 · Xu · 2018 [cited by applicant]
US 10169579B1 · Xu et al. · 2019 [cited by applicant]
US 10237282B2 · Nelson et al. · 2019 [cited by applicant]
US 10334442B2 · Vaughn et al. · 2019 [cited by applicant]
US 10382468B2 · Dods · 2019 [cited by applicant]
US 10462173B1 · Aziz et al. · 2019 [cited by applicant]
US 10484334B1 · Lee et al. · 2019 [cited by applicant]
US 10762206B2 · Titonis · 2020 [cited by examiner]
US 10826941B2 · Jain et al. · 2020 [cited by applicant]
US 11025666B1 · Han · 2021 [cited by examiner]
US 11032301B2 · Mandrychenko et al. · 2021 [cited by applicant]
US 11036856B2 · Graun et al. · 2021 [cited by applicant]
US 11281775B2 · Burdett et al. · 2022 [cited by applicant]
US 11310282B1 · Zhang · 2022 [cited by examiner]
US 11444951B1 · Patil · 2022 [cited by examiner]
US 11481709B1 · Liao · 2022 [cited by examiner]
US 20020099666A1 · Dryer et al. · 2002 [cited by applicant]
US 20030055994A1 · Herrmann et al. · 2003 [cited by applicant]
US 20030063321A1 · Inoue et al. · 2003 [cited by applicant]
US 20030172292A1 · Judge · 2003 [cited by applicant]
US 20030204632A1 · Willebeek-Lemair et al. · 2003 [cited by applicant]
US 20040015719A1 · Lee et al. · 2004 [cited by applicant]
US 20050010593A1 · Fellenstein et al. · 2005 [cited by applicant]
US 20050271246A1 · Sharma et al. · 2005 [cited by applicant]
US 20060156401A1 · Newstadt et al. · 2006 [cited by applicant]
US 20070204018A1 · Chandra et al. · 2007 [cited by applicant]
US 20070237147A1 · Quinn et al. · 2007 [cited by applicant]
US 20080069480A1 · Aarabi et al. · 2008 [cited by applicant]
US 20080134332A1 · Keohane et al. · 2008 [cited by applicant]
US 20090144818A1 · Kumar et al. · 2009 [cited by applicant]
US 20090249470A1 · Litvin et al. · 2009 [cited by applicant]
US 20090300351A1 · Lei et al. · 2009 [cited by applicant]
US 20100017436A1 · Wolge · 2010 [cited by applicant]
US 20110119481A1 · Auradkar et al. · 2011 [cited by applicant]
US 20110145594A1 · Jho et al. · 2011 [cited by applicant]
US 20120278896A1 · Fang et al. · 2012 [cited by applicant]
US 20130159694A1 · Chiueh et al. · 2013 [cited by applicant]
US 20130298190A1 · Sikka et al. · 2013 [cited by applicant]
US 20130347085A1 · Hawthorn et al. · 2013 [cited by applicant]
US 20140013112A1 · Cidon et al. · 2014 [cited by applicant]
US 20140068030A1 · Chambers et al. · 2014 [cited by applicant]
US 20140068705A1 · Chambers et al. · 2014 [cited by applicant]
US 20140259093A1 · Narayanaswamy et al. · 2014 [cited by applicant]
US 20140282843A1 · Buruganahalli et al. · 2014 [cited by applicant]
US 20140289852A1 · Evans · 2014 [cited by examiner]
US 20140359282A1 · Shikfa et al. · 2014 [cited by applicant]
US 20140366079A1 · Pasdar · 2014 [cited by applicant]
US 20150100357A1 · Seese et al. · 2015 [cited by applicant]
US 20160323318A1 · Terrill et al. · 2016 [cited by applicant]
US 20160350145A1 · Botzer et al. · 2016 [cited by applicant]
US 20170064005A1 · Lee · 2017 [cited by applicant]
US 20170093917A1 · Chandra et al. · 2017 [cited by applicant]
US 20170250951A1 · Wang et al. · 2017 [cited by applicant]
US 20190222591A1 · Kislitsin · 2019 [cited by examiner]
US 20200050686A1 · Kamalapuram et al. · 2020 [cited by applicant]
US 20200322361A1 · Ravindra · 2020 [cited by examiner]
US 20230297685A1 · Patil et al. · 2023 [cited by applicant]
US 20230342461A1 · Du et al. · 2023 [cited by applicant]
CN 111274583A · 2020 [cited by examiner]
EP 1063833A2 · 2000 [cited by applicant]
WO WO2013184653A1 · 2013 [cited by examiner]
WO WO2014012106A2 · 2014 [cited by examiner]
WO WO2022246131A1 · 2022 [cited by examiner]
Djenna, Amir, et al. “Artificial intelligence-based malware detection, analysis, and mitigation.” Symmetry 15.3 (2023): 677. (Year: 2023). [cited by examiner]
Singh, Jagsir, and Jaswinder Singh. “Detection of malicious software by analyzing the behavioral artifacts using machine learning algorithms.” Information and Software Technology 121 (2020): 106273. (Year: 2020). [cited by examiner]
Kishore, Pushkar, Swadhin Kumar Barisal, and Durga Prasad Mohapatra. “JavaScript malware behaviour analysis and detection using sandbox assisted ensemble model.” 2020 IEEE Region 10 Conference (Tencon). IEEE, 2020. (Yea… [cited by examiner]
U.S. Appl. No. 18/656,895 Non-Final Office Action dated Jul. 5, 2024, USPTO Examiner Fatoumata Traore, 18 pages. [cited by applicant]
Martin, Victoria “Cooperative Security Fabric,” The Fortinet Cookbook, Jun. 8, 2016, 6 pgs., archived Jul. 28, 2016 at https://web.archive.org/web/20160728170025/http://cookbook.fortinet.com/cooperative-security-fabric-… [cited by applicant]
Huckaby, Jeff “Ending Clear Text Protocols,” Rackaid.com, Dec. 9, 2008, 3 pgs. [cited by applicant]
Newton, Harry “fabric,” Newton's Telecom Dictionary, 30th Updated, Expanded, Anniversary Edition, 2016, 3 pgs. [cited by applicant]
Fortinet, “Fortinet Security Fabric Earns 100% Detection Scores Across Several Attack Vectors in NSS Labs' Latest Breach Detection Group Test [press release]”, Aug. 2, 2016, 4 pgs, available at https://www.fortinet.com/… [cited by applicant]
Fortinet, “Fortinet Security Fabric Named 2016 CRN Network Security Product of the Year [press release]”, Dec. 5, 2016, 4 pgs, available at https://www.fortinet.com/corporate/about-US/newsroom/press-releases/2016/fortin… [cited by applicant]
McCullagh, Declan, “How safe is instant messaging? A security and privacy survey,” CNET, Jun. 9, 2008, 14 pgs. [cited by applicant]
Beck et al., “IBM and Cisco: Together for a World Class Data Center,” IBM Redbooks, Jul. 2013, 654 pgs. [cited by applicant]
Martin, Victoria “Installing internal FortiGates and enabling a security fabric,” The Fortinet Cookbook, Jun. 8, 2016, 11 pgs, archived Aug. 28, 2016 at https://web.archive.org/web/20160828235831/http://cookbook.fortine… [cited by applicant]
Zetter, Kim, “Revealed: The Internet's Biggest Security Hole,” Wired, Aug. 26, 2008, 13 pgs. [cited by applicant]
Adya et al., “Farsite: Federated, available, and reliable storage for an incompletely trusted environment,” SIGOPS Oper. Syst. Rev. 36, SI, Dec. 2002, pp. 1-14. [cited by applicant]
Agrawal et al., “Order preserving encryption for numeric data,” In Proceedings of the 2004 Acm Sigmod international conference on Management of data, Jun. 2004, pp. 563-574. [cited by applicant]
Balakrishnan et al., “A layered naming architecture for the Internet,” Acm Sigcomm Computer Communication Review, 34(4), 2004, pp. 343-352. [cited by applicant]
Downing et al., Naming Dictionary of Computer and Internet Terms, (11th Ed.) Barron's, 2013, 6 pgs. [cited by applicant]
Downing et al., Dictionary of Computer and Internet Terms, (10th Ed.) Barron's, 2009, 4 pgs. [cited by applicant]
Zoho Mail, “Email Protocols: What they are & their different types,” 2006, 7 pgs. available at https://www.zoho.com/mail/glossary/email-protocols.html# :˜: text=mode of communication.-, What are the different email prot… [cited by applicant]
NIIT, Special Edition Using Storage Area Networks, Que, 2002, 6 pgs. [cited by applicant]
Chapple, Mike, “Firewall redundancy: Deployment scenarios and benefits,” TechTarget, 2005, 5 pgs. available at https://www.techtarget.com/searchsecurity/tip/Firewall-redundancy-Deployment-scenarios-and-benefits? Offer=a… [cited by applicant]
Fortinet, FortiGate—3600 User Manual (vol. 1, Version 2.50 MR2) Sep. 5, 2003, 329 pgs. [cited by applicant]
Fortinet, FortiGate SOHO and SMB Configuration Example, (Version 3.0 MR5), Aug. 24, 2007, 54 pgs. [cited by applicant]
Fortinet, FortiSandbox—Administration Guide, (Version 2.3.2), Nov. 9, 2016, 191 pgs. [cited by applicant]
Fortinet, FortiSandbox Administration Guide, (Version 4.2.4) Jun. 12, 2023, 245 pgs. available at https://fortinetweb.s3.amazonaws.com/docs.fortinet.com/v2/attachments/fba32b46-b7c0-11ed-8e6d-fa163e15d75b/FortiSandbox-4… [cited by applicant]
Fortinet, FortiOS- Administration Guide, (Versions 6.4.0), Jun. 3, 2021, 1638 pgs. [cited by applicant]
Heady et al., “The Architecture of a Network Level Intrusion Detection System,” University of New Mexico, Aug. 15, 1990, 21 pgs. [cited by applicant]
Kephart et al., “Fighting Computer Viruses,” Scientific American (vol. 277, No. 5) Nov. 1997, pp. 88-93. [cited by applicant]
Wang, L., Chapter 5: Cooperative Security in D2D Communications, “Physical Layer Security in Wireless Cooperative Networks,” 41 pgs. first online on Sep. 1, 2017 at https://link.springer.com/chapter/10.1007/978-3-319-61… [cited by applicant]
Lee et al., “A Data Mining Framework for Building Intrusion Detection Models,” Columbia University, n.d. 13 pgs. [cited by applicant]
Merriam-Webster Dictionary, 2004, 5 pgs. [cited by applicant]
Microsoft Computer Dictionary, (5th Ed.), Microsoft Press, 2002, 8 pgs. [cited by applicant]
Microsoft Computer Dictionary, (4th Ed.), Microsoft Press, 1999, 5 pgs. [cited by applicant]
Mika et al., “Metadata Statistics for a Large Web Corpus,” LDOW2012, Apr. 16, 2012, 6 pgs. [cited by applicant]
Oxford Dictionary of Computing (6th Ed.), 2008, 5 pgs. [cited by applicant]
Paxson, Vern, “Bro: a System for Detecting Network Intruders in Real-Time,” Proceedings of the 7th USENIX Security Symposium, Jan. 1998, 22 pgs. [cited by applicant]
Fortinet Inc., U.S. Appl. No. 62/503,252, “Building a Cooperative Security Fabric of Hierarchically Interconnected Network Security Devices.” n.d., 87 pgs. [cited by applicant]
Song et al., “Practical techniques for searches on encrypted data,” In Proceeding 2000 IEEE symposium on security and privacy. S&P. 2000, May 2000, pp. 44-55. [cited by applicant]
Dean, Tamara, Guide to Telecommunications Technology, Course Technology, 2003, 5 pgs. [cited by applicant]
U.S. Appl. No. 60/520,577, “Device, System, and Method for Defending a Computer Network,” Nov. 17, 2003, 21 pgs. [cited by applicant]
U.S. Appl. No. 60/552,457, “Fortinet Security Update Technology,” Mar. 2004, 6 pgs. [cited by applicant]
Tittel, Ed, Unified Threat Management For Dummies, John Wiley & Sons, Inc., 2012, 76 pgs. [cited by applicant]
Fortinet, FortiOS Handbook: UTM Guide (Version 2), Oct. 15, 2010, 188 pgs. [cited by applicant]
Full Definition of Security, Wayback Machine Archive of Merriam-Webster on Nov. 17, 2016, 1 pg. [cited by applicant]
Definition of Cooperative, Wayback Machine Archive of Merriam-Webster on Nov. 26, 2016, 1 pg. [cited by applicant]
Pfaffenberger, Bryan, Webster's New World Computer Dictionary, (10th Ed.), 2003, 5 pgs. [cited by applicant]
U.S. Appl. No. 18/656,895 Non-Final Office Action dated Jul. 5, 2024, USPTO Examiner Fatoumata Traore, 18 pages (745.0082). [cited by applicant]
Cited By (1)
US 12,657,310