IP Library › Granted Patent US 12,197,574
Granted Patent B2
US 12,197,574 · App. 17/550,420 · Granted Jan 14, 2025

Detecting Microsoft Windows installer malware using text classification models

Inventors: Akshata Krishnamoorthy Rao (Mountain View, CA); Yaron Samuel (Tel Aviv, IL); Lauren Che (Dublin, CA); Wenjun Hu (Santa Clara, CA)
Assignee: Palo Alto Networks, Inc.
G06F21/566G06F21/54G06F21/554G06F21/564
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,574
App. No.
17/550,420
Filed
Dec 14, 2021
Granted
Jan 14, 2025
Kind
B2
Art Unit
2438
USPC
726/22
Abstract

The present application discloses a method, system, and computer system for detecting malicious files. The method includes receiving a sample, extracting an embedded script from the sample, applying a malicious script detector in connection with determining whether the sample is malicious, and in response to determining that the sample is malicious sending, to a security entity, an indication that the sample is malicious.

Claims (52)

1. A system, comprising:

one or more processors configured to:

receive a sample, wherein the sample is a Microsoft Windows Portable Executable (PE) file;

extract an embedded script from the sample, wherein the embedded script is an installer script, wherein the installer script is extracted in a sandboxed environment;

determine one or more features based at least in part on the installer script;

apply a malicious script detector in connection with determining whether the sample is malicious, wherein the malicious script detector determines a sample classification based at least in part on querying a machine learning model based at least in part on the one or more features; and

in response to determining that the sample is malicious,

send, to a security entity, an indication that the sample is malicious; and

a memory coupled to the one or more processors and configured to provide the one or more processors with instructions.

2. The system of claim 1 , wherein extracting the embedded script comprises:

extracting the installer script from the sample; and

decompiling the installer script to obtain code corresponding to the installer script.

3. The system of claim 1 , wherein applying malicious script detector in connection with determining whether the embedded script is malicious comprises:

applying the machine learning model to determine whether the embedded script is malicious.

4. The system of claim 3 , wherein the machine learning model is trained based at least in part on one or more code samples that have been deemed to be malicious.

5. The system of claim 2 , wherein the applying the machine learning model to determine whether the embedded script is malicious comprises:

analyzing code corresponding to the embedded script to determine whether the code comprises one or more elements that are indicative of malicious code.

6. The system of claim 5 , wherein the machine learning model is trained to learn the one or more elements that are indicative of malicious code.

7. The system of claim 5 , wherein analyzing code corresponding to the embedded script to determine whether the code comprises one or more elements that are indicative of malicious code comprises:

applying a text classification machine learning model prediction in connection with detecting, in the code, the one or more elements that are indicative of malicious code.

8. The system of claim 1 , wherein the machine learning model is trained based at least in part on one or more attributes associated with code samples previously deemed to be malicious.

9. The system of claim 1 , wherein:

applying the malicious script detector in connection with determining whether the sample is malicious comprises:

determining a likelihood that code corresponding to the embedded script is malicious; and

the sample is deemed to be malicious in response to a determination that the likelihood that the code corresponding to the embedded script is malicious is greater than a likelihood threshold value.

10. The system of claim 9 , wherein the likelihood that the code corresponding to the embedded script is malicious is determined based at least in part on a degree of similarity between the code and one or more other malicious code samples.

11. The system of claim 1 , wherein the security entity enforces one or more security policies with respect to the sample based at least in part on the indication that the sample is malicious.

12. The system of claim 11 , wherein the one or more security policies are configured based at least in part on a customer setting.

13. The system of claim 11 , wherein the security entity blocks traffic comprising the sample in response to receiving the indication that the sample is malicious.

14. The system of claim 1 , wherein the security entity is a firewall.

15. The system of claim 1 , wherein a signature or file hash corresponding to the sample is sent to the security entity in connection with sending the indication that the sample is malicious.

16. The system of claim 1 , wherein a result of the malicious script detector is used as a factor with one or more other factors to determine whether the sample is malicious.

17. The system of claim 1 , wherein one or more factors used in connection with malicious script detector determining that the code corresponding to the embedded script is malicious comprises: a call to an executable file or to a cryptocurrency wallet.

18. A method, comprising:

receiving, by one or more processors, a sample, wherein the sample is a Microsoft Windows Portable Executable (PE) file;

extracting an embedded script from the sample, wherein the embedded script is an installer script, wherein the installer script is extracted in a sandboxed environment;

determining one or more features based at least in part on the installer script;

applying a malicious script detector in connection with determining whether the sample is malicious, wherein the malicious script detector determines a sample classification based at least in part on querying a machine learning model based at least in part on the one or more features; and

in response to determining that the sample is malicious,

sending, to a security entity, an indication that the sample is malicious.

19. A computer program product embodied in a non-transitory computer readable medium and comprising computer instructions for:

receiving, by one or more processors, a sample, wherein the sample is a Microsoft Windows Portable Executable (PE) file;

extracting an embedded script from the sample, wherein the embedded script is an installer script, wherein the installer script is extracted in a sandboxed environment;

determining one or more features based at least in part on the installer script;

applying a malicious script detector in connection with determining whether the sample is malicious, wherein the malicious script detector determines a sample classification based at least in part on querying a machine learning model based at least in part on the one or more features; and

in response to determining that the sample is malicious,

sending, to a security entity, an indication that the sample is malicious.

20. The system of claim 1 , wherein applying malicious script detector in connection with determining whether the embedded script is malicious comprises:

using the machine learning model to perform text classification on the installer script; and

determining whether the sample is malicious based at least in part on the text classification.

21. The system of claim 1 , wherein the one or more processors are further configured to determine that the sample is malicious based at least in part on one or more of (i) an executable called in the extracted installer script, (ii) a cryptocurrency wallet called in the extracted installer script, and (iii) an alphanumeric string comprised in the extracted installer script.

22. The system of claim 1 , wherein the one or more processors are further configured to determine that the sample is malicious based at least in part on the installer script and a PE file structure for the sample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2022
From: RAO, AKSHATA KRISHNAMOORTHY; SAMUEL, YARON; CHE, LAUREN; HU, WENJUN
To: PALO ALTO NETWORKS, INC.
Reel/Frame 058988/0421 →
Continuity (1)
Related Publication 20230185915A1 · Jun 15, 2023
References Cited (25)
US 8838992B1 · Zhu · 2014 [cited by examiner]
US 9444831B1 · Lee · 2016 [cited by examiner]
US 10956477B1 · Fang · 2021 [cited by examiner]
US 11425151B2 · Kaidi · 2022 [cited by examiner]
US 11574053B1 · Chen · 2023 [cited by examiner]
US 11822657B2 · Tseng · 2023 [cited by examiner]
US 20070067843A1 · Williamson · 2007 [cited by examiner]
US 20090133125A1 · Choi · 2009 [cited by examiner]
US 20130185623A1 · Wicker · 2013 [cited by examiner]
US 20170093893A1 · Davydov · 2017 [cited by examiner]
US 20190034823A1 · Thapliyal · 2019 [cited by examiner]
US 20190102552A1 · Pavlyushchik · 2019 [cited by examiner]
US 20190251251A1 · Carson · 2019 [cited by examiner]
US 20200311259A1 · Schmugar · 2020 [cited by examiner]
US 20210117544A1 · Kurtz · 2021 [cited by examiner]
US 20210326035A1 · Jia · 2021 [cited by examiner]
US 20230342462A1 · Ducau · 2023 [cited by examiner]
KR 20080110894 · 2008 [cited by examiner]
Ajay. Applying Deep Learning for Pe-Malware Classification. Quick Heal Blog. Jan. 10, 2019: https://blogs.quickheal.com/applying-dl-pe-malware-classification/. [cited by applicant]
Bokobza et al. Machine Learning for Cyber Security—Static Detection of Malicious PE Files. Jan. 22, 2019: https://www.cyberbit.com/blog/endpoint-security/machine-learning-for-cyber-security-static-detection/. [cited by applicant]
Gilbert et al. The rise of machine learning for detection and classification of malware: Research developments, trends and challenges. Journal of Network and Computer Applications. vol. 153, Mar. 1, 2020, 102526: https:… [cited by applicant]
Karl Ackerman. Performing Suspect Power Shell Detection with a Yara Rule. Sophos. Aug. 22, 2021: https://community.sophos.com/intercept-x-endpoint/f/recommended-reads/129682/performing-suspect-power-shell-detection-with… [cited by applicant]
Maglaque et al. Fake Installers Drop Malware and Open Doors for Opportunistic Attackers. Sep. 27, 2021: https://www.trendmicro.com/en_us/research/21/i/fake-installers-drop-malware-and-open-doors-for-opportunistic-attack… [cited by applicant]
Microsoft Defender ATP Team. Deep learning rises: New methods for detecting malicious PowerShell. Sep. 3, 2019: https://www.microsoft.com/security/blog/2019/09/03/deep-learning-rises-new-methods-for-detecting-malicious-… [cited by applicant]
Wang et al. A deep learning approach for detecting malicious JavaScript code. Security and Communication Networks Security Comm. Networks 2016; 9:1520-1534. Published online Feb. 11, 2016 in Wiley Online Library (wileyo… [cited by applicant]