IP Library Granted Patent US 12,430,437
Granted Patent B2
US 12,430,437 · App. 17/582,928 · Granted Sep 30, 2025

Specific file detection baked into machine learning pipelines

Inventors: William Redington Hewlett, II (Mountain View, CA); Anirudh Mittal (Fremont, CA); Ashwin Kumar Dewan (Palm Harbor, FL); Tyler Pals Halfpop (Carmel, IN)
Assignee: Palo Alto Networks, Inc.
G06F21/567G06N5/022G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,437
App. No.
17/582,928
Granted
Sep 30, 2025
Kind
B2
Abstract

A set of features including a first feature and a second feature is received at a server. A subset of the set of features is determined for use in generating a model usable by a device to locally make a malware classification decision. The device has reduced computing resources as compared to computing resources of the server. The subset of the set of features is used to generate the model. The generated model includes the first feature and does not include the second feature. A determination is made, at a time subsequent to the generation of the model, that an updated model should be deployed to the device. An updated model is generated.

Claims (41)

1. A system, comprising:

a processor configured to:

receive, at a server, a set of features including a first feature and a second feature;

determine a subset of the set of features to use in generating a model usable by a device to locally make a malware classification decision, wherein the device has reduced computing resources as compared to computing resources of the server;

use the subset of the set of features to generate the model, wherein the generated model includes the first feature and wherein the generated model does not include the second feature, wherein the first feature includes a specific count of a number of times a specific n-gram appears in a given file, and wherein when a custom benign test file that includes the specific n-gram the specific count of times is encountered by the device, the device will classify the custom benign test file as malicious; and

determine, at a time subsequent to the generation of the model, that an updated model should be deployed to the device, and generate the updated model; and

a memory coupled to the processor and configured to provide the processor with instructions.

2. The system of claim 1 , wherein generating the model includes appending a constructed regression tree to the model.

3. The system of claim 2 , wherein the constructed regression tree classifies a given file based on a count of occurrences of a custom set of n-grams.

4. The system of claim 3 , wherein the custom set of n-grams is selected to classify benign test file as being malicious.

5. The system of claim 1 , wherein the processor is further configured to generate the custom benign file.

6. The system of claim 1 , wherein the custom benign test file includes a specific count of each n-gram included in the custom set of n-grams.

7. The system of claim 2 , wherein the processor is further configured to determine that the constructed regression tree does not return a non-zero value for any samples included in a corpus.

8. The system of claim 2 , wherein the appended constructed regression tree does not reduce accuracy of the model in detecting malicious files.

9. The system of claim 1 , wherein the set of features includes features extracted from a set of known malicious files.

10. The system of claim 1 , wherein the set of features includes features extracted from a set of known benign files.

11. The system of claim 1 , wherein the subset of the set of features is determined using mutual information.

12. The system of claim 1 , wherein the subset of the set of features is determined using Chi-squared score.

13. The system of claim 1 , wherein the determination that the updated model should be deployed is made in response to a false positive result reported by a data appliance.

14. A method, comprising:

receiving, at a server, a set of features including a first feature and a second feature;

determining a subset of the set of features to use in generating a model usable by a device to locally make a malware classification decision, wherein the device has reduced computing resources as compared to computing resources of the server;

using the subset of the set of features to generate the model, wherein the generated model includes the first feature and wherein the generated model does not include the second feature, wherein the first feature includes a specific count of a number of times a specific n-gram appears in a given file, and wherein when a custom benign test file that includes the specific n-gram the specific count of times is encountered by the device, the device will classify the custom benign test file as malicious; and

determining, at a time subsequent to the generation of the model, that an updated model should be deployed to the device, and generate the updated model.

15. The method of claim 14 , wherein generating the model includes appending a constructed regression tree to the model.

16. The method of claim 15 , wherein the constructed regression tree classifies a given file based on a count of occurrences of a custom set of n-grams.

17. The method of claim 16 , wherein the custom set of n-grams is selected to classify the custom benign test file as being malicious.

18. The method of claim 14 , wherein the processor is further configured to generate the custom benign file.

19. The method of claim 14 , wherein the custom benign test file includes a specific count of each n-gram included in the custom set of n-grams.

20. The method of claim 15 , further comprising determining that the constructed regression tree does not return a non-zero value for any samples included in a corpus.

21. The method of claim 15 , wherein the appended constructed regression tree does not reduce accuracy of the model in detecting malicious files.

22. The method of claim 14 , wherein the set of features includes features extracted from a set of known malicious files.

23. The method of claim 14 , wherein the set of features includes features extracted from a set of known benign files.

24. The method of claim 14 , wherein the subset of the set of features is determined using mutual information.

25. The method of claim 14 , wherein the subset of the set of features is determined using Chi-squared score.

26. The system of claim 14 , wherein the determination that the updated model should be deployed is made in response to a false positive result reported by a data appliance.

27. A computer program product embodied in a tangible computer readable storage medium and comprising computer instructions for:

receiving, at a server, a set of features including a first feature and a second feature;

determining a subset of the set of features to use in generating a model usable by a device to locally make a malware classification decision, wherein the device has reduced computing resources as compared to computing resources of the server;

using the subset of the set of features to generate the model, wherein the generated model includes the first feature and wherein the generated model does not include the second feature, wherein the first feature includes a specific count of a number of times a specific n-gram appears in a given file, and wherein when a custom benign test file that includes the specific n-gram the specific count of times is encountered by the device, the device will classify the custom benign test file as malicious; and

determining, at a time subsequent to the generation of the model, that an updated model should be deployed to the device, and generate the updated model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2022
From: HEWLETT, WILLIAM REDINGTON, II; MITTAL, ANIRUDH; DEWAN, ASHWIN KUMAR; HALFPOP, TYLER PALS
To: PALO ALTO NETWORKS, INC.
Reel/Frame 059738/0125 →
Continuity (2)
Continuation In Part 16517465 · Jul 19, 2019
Related Publication 20220245249A1 · Aug 4, 2022
References Cited (64)
US 7792850B1 · Raffill · 2010 [cited by applicant]
US 9037967B1 · Al-Jefri · 2015 [cited by applicant]
US 9396334B1 · Ivanov · 2016 [cited by applicant]
US 10681080B1 · Chen · 2020 [cited by applicant]
US 10878124B1 · Sitaraman · 2020 [cited by applicant]
US 10942963B1 · Huang · 2021 [cited by applicant]
US 20050256716A1 · Bangalore · 2005 [cited by applicant]
US 20060037080A1 · Maloof · 2006 [cited by applicant]
US 20090319536A1 · Parker · 2009 [cited by applicant]
US 20100064369A1 · Stolfo · 2010 [cited by applicant]
US 20110126286A1 · Nazarov · 2011 [cited by applicant]
US 20110320498A1 · Flor · 2011 [cited by applicant]
US 20120317644A1 · Kumar · 2012 [cited by applicant]
US 20140033307A1 · Schmidtler · 2014 [cited by applicant]
US 20140223565A1 · Cohen · 2014 [cited by applicant]
US 20150244730A1 · Vu · 2015 [cited by applicant]
US 20170004306A1 · Zhang · 2017 [cited by applicant]
US 20170085585A1 · Morkovský · 2017 [cited by applicant]
US 20170262633A1 · Miserendino · 2017 [cited by applicant]
US 20180013772A1 · Schmidtler · 2018 [cited by applicant]
US 20180048659A1 · Salsamendi · 2018 [cited by applicant]
US 20180089365A1 · Beal · 2018 [cited by applicant]
US 20180203998A1 · Maisel · 2018 [cited by applicant]
US 20180293381A1 · Tseng · 2018 [cited by applicant]
US 20180300482A1 · Li · 2018 [cited by applicant]
US 20190034632A1 · Tsao · 2019 [cited by applicant]
US 20190087574A1 · Schmidtler · 2019 [cited by applicant]
US 20190095820A1 · Pourmohammad · 2019 [cited by applicant]
US 20190096214A1 · Pourmohammad · 2019 [cited by applicant]
US 20190180175A1 · Meteer · 2019 [cited by applicant]
US 20190354682A1 · Finkelshtein · 2019 [cited by applicant]
US 20200067861A1 · Leddy · 2020 [cited by applicant]
US 20200076835A1 · Ladnai · 2020 [cited by applicant]
US 20200097655A1 · Rihn · 2020 [cited by applicant]
US 20200125728A1 · Savir · 2020 [cited by applicant]
US 20200213325A1 · Scherman · 2020 [cited by applicant]
US 20200364334A1 · Pevny · 2020 [cited by applicant]
US 20210019408A1 · Chrysaidos · 2021 [cited by applicant]
US 20210224534A1 · Miller · 2021 [cited by applicant]
CN 102779249 · 2015 [cited by applicant]
CN 103618744 · 2017 [cited by applicant]
EP 2182458 · 2010 [cited by applicant]
JP 2012003463 · 2012 [cited by applicant]
WO 2010011411 · 2010 [cited by applicant]
WO 2020006415 · 2020 [cited by applicant]
Peter A. Chew, Brett W. Bader, Ahmed Abdelali ; Latent Morpho-Semantic Analysis: Multilingual Information Retrieval with Character N-Grams and Mutual Information; Proceedings of the 22nd International Conference on Comp… [cited by examiner]
Beebe et al., “Sceadan: Using Concatenated N-Gram Vectors for Improved File and Data Type Classification”, IEEE Transactions on Information Forensics and Security, IEEE, USA, vol. 8, No. 9, Sep. 1, 2013 (Sep. 1, 2013), … [cited by applicant]
Gil et al., “Mal-ID: Automatic Malware Detection Using Common Segment Analysis and Meta-Features”, Journal of Machine Learning Research, Feb. 28, 2012 (Feb. 28, 2012), pp. 1-33. [cited by applicant]
Lin et al., “Feature Selection and Extraction for Malware Classification”, Journal of Information Science and Engineering, vol. 31, Jan. 1, 2015 (Jan. 1, 2015), pp. 965-992. [cited by applicant]
Mas'ud et al., “A Comparative Study on Feature Selection Method for N-gram Mobile Malware Detention”, International Journal of Network Security, Sep. 30, 2017, pp. 727-733. [cited by applicant]
Adityaram et al. : “HTTP Attack Detection using N-gram Analysis HTTP Attack Detection using N-gram Analysis”, San Jose State University, May 1, 2013 (May 1, 2013), Retrieved from the Internet: URL:https://scholarworks.s… [cited by applicant]
Li et al., “Fileprints: Identifying File Types by N-Gram Analysis”, Systems, Man and Cybernetics (SMC) Information Assurance Workshop, 200 5, Proceedings From the Sixth Annual IEEE, Jun. 15-17, 2005, Jun. 15, 2005 (Jun.… [cited by applicant]
Wressnegger et al. : “A Close Look on N-Grams in Intrusion Detection”, Artificial Intelligence and Security, ACM, Nov. 4, 2013 (Nov. 4, 2013), pp. 67-76. [cited by applicant]
Chen et al., TinyDroid: A Lightweight and Efficient Model for Android Malware Detection and Classification, Mobile Information Systems, 2018, vol. 2018. [cited by applicant]
Hadžiosmanović et al., N-Gram Against the Machine: On the Feasibility of the N-Gram Network Analysis for Binary Protocols, International Workshop on Recent Advances in Intrusion Detection, 2012, pp. 354-373. [cited by applicant]
Kang et al., N-gram Opcode Analysis for Android Malware Detection, Intl. Journal on Cyber Situational Awareness, 2016, vol. 1, No. 1. [cited by applicant]
Li et al., Fileprints: Identifying File Types by n-gram Analysis, Proceedings of the 2005 IEEE Workshop on Information Assurance, Jun. 2005. [cited by applicant]
Oza et al., HTTP Attack Detection Using N-Gram Analysis, Computers & Security 45, 2014, pp. 242-254. [cited by applicant]
Pektas et al., Proposal of N-gram Based Algorithm for Malware Classification, the Fifth International Conference on Emerging Security Information, Systems and Technologies, Aug. 2011, pp. 7-13. [cited by applicant]
Raff et al., Hash-Grams: Faster N-Gram Features for Classification and Malware Detection, Proceedings of the ACM Symposium on Document Engineering, Jul. 2018, ACM. [cited by applicant]
Santos et al., N-Grams-Based File Signatures for Malware Detection, ICEIS, 2009, pp. 317-320. [cited by applicant]
Shafiq et al, Embedded Malware Detection Using Markov n-Grams, Proceedings of the 5th International Conference on Detection of Intrusions and Malware, and Vulnerability Assessment, Jul. 2008, pp. 88-107, Springer-Verlag. [cited by applicant]
Beebe et al., Sceadan: Using Concatenated N-Gram Vectors for Improved File and Data Type Classification, IEEE Transactions on Information Forensics and Security, Dec. 31, 2013, pp. 1519-1530. [cited by applicant]
Li et al., Fileprints: Identifying File Types by n-gram Analysis, Proceedings of the 2005 IEEE Workshop on Information Assurance and Security, pp. 64-71. [cited by applicant]
Cited By (1)
US 12,566,594