IP Library › Granted Patent US 12,688,287
Granted Patent B1
US 12,688,287 · App. 18/474,777 · Granted Jul 21, 2026

Static script malware detection with semantic code representations

Inventors: Ecenaz Erdemir (Jersey City, NJ); Michael James Morais (New York, NY); Marion Marschalek (Albany, OR); Kyuhong Park (Peachtree Corners, GA); Yi Fan (Short Hills, NJ); Vianne Ran Gao (Jersey City, NJ)
Assignee: Amazon Technologies, Inc.
G06F21/56G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,287
App. No.
18/474,777
Filed
Sep 26, 2023
Granted
Jul 21, 2026
Kind
B1
Art Unit
2432
USPC
726/23
Abstract

Techniques for static script malware detection with semantic code representations are described. A request to perform a static malware detection analysis is received, the request identifying a script comprising code. A semantic representation of the script is generated, the semantic representation including context about statements in the code of the script. A feature space embedding is generated based at least in part on the semantic representation of the script. generating, based at least in part on the feature space embedding, A result indicating whether the script is malicious is generated based at least in part on the feature space embedding. An identification of the script and the result is stored in an entry in a log.

Claims (48)

1 . A computer-implemented method comprising:

receiving, by a malware detection service of a cloud provider network, a request to perform a static malware detection analysis, the request identifying a script comprising code;

generating, with a bytestring encoder, a bytestring embedding from a bytestring of a code statement in the script;

generating a semantic representation of the script, the semantic representation including context about statements in the code of the script, wherein the context includes at least one of an identification of language keywords or an identification of function structure;

generating, with a semantics encoder, a semantics embedding from a feature of the semantic representation, the feature corresponding to the code statement in the script;

generating a sequence of feature space embeddings based at least in part on the semantic representation of the script by combining the bytestring embedding and the semantics embedding, wherein the sequence of feature space embeddings is a sequence of vectors that encodes script semantic features in a feature space;

evaluating the sequence of feature space embeddings using a recurrent neural network (RNN)-based sequential classifier model to generate a result that indicates whether the script is malicious; and

storing an identification of the script and the result in an entry in a log of the malware detection service.

2 . The computer-implemented method of claim 1 , wherein each feature space embedding in the sequence of feature space embeddings corresponds to a statement in the code, and wherein the result is generated based on the sequence of feature space embeddings.

3 . The computer-implemented method of claim 1 , wherein the sequence of feature space embeddings represents a graph in a graph feature space, and wherein feature space embeddings in the sequence of feature space embeddings from similar graphs are closer together within the graph feature space than feature space embeddings in the sequence of feature space embeddings from dissimilar graphs.

4 . A computer-implemented method comprising:

receiving a request to perform a static malware detection analysis, the request identifying a script comprising code;

generating, with a bytestring encoder, a bytestring embedding from a bytestring of a code statement in the script;

generating a semantic representation of the script, the semantic representation including context about statements in the code of the script;

generating, with a semantics encoder, a semantics embedding from a feature of the semantic representation, the feature corresponding to the code statement in the script;

generating a sequence of feature space embeddings based at least in part on the semantic representation of the script by combining the bytestring embedding and the semantics embedding;

evaluating the sequence of feature space embeddings using a recurrent neural network (RNN)-based sequential classifier model to generate a result that indicates whether the script is malicious; and

storing an identification of the script and the result in an entry in a log.

5 . The computer-implemented method of claim 4 , wherein the semantic representation of the script is generated using a static code analysis parsing component from at least one of a syntax highlighter application or an abstract syntax tree constructor application.

6 . The computer-implemented method of claim 4 , wherein each feature space embedding in the sequence of feature space embeddings corresponds to a statement in the code.

7 . The computer-implemented method of claim 4 , wherein the sequence of feature space embeddings is ordered based on at least one of a sequence of statements in the code or a serialized traversal of nodes in an abstract syntax tree representing the script.

8 . The computer-implemented method of claim 4 , wherein the semantic representation of the script is generated using a static code analysis parsing component from an abstract syntax tree constructor application, and wherein each feature space embedding in the sequence of feature space embeddings corresponds to a node in an abstract syntax tree of the script.

9 . The computer-implemented method of claim 4 :

wherein the sequence of feature space embeddings represents a graph;

wherein generating the sequence of feature space embeddings and generating the result are performed using a machine learning model, the machine learning model including a graph embedding model and a classifier; and

wherein the graph embedding model is trained to generate feature space embeddings in a graph feature space, and wherein feature space embeddings in the sequence of feature space embeddings from similar graphs are closer together within the graph feature space than feature space embeddings in the sequence of feature space embeddings from dissimilar graphs.

10 . The computer-implemented method of claim 9 , wherein the classifier is trained after learning parameters of the graph embedding model.

11 . The computer-implemented method of claim 4 , wherein the context about statements in the code of the script includes at least one of an identification of language keywords or an identification of function structure.

12 . The computer-implemented method of claim 4 , wherein the RNN-based sequential classifier model includes a bidirectional long short-term memory module.

13 . The computer-implemented method of claim 12 , wherein the RNN-based sequential classifier model further includes multiple layers followed by an attention module and a fully-connected output layer.

14 . A system comprising:

a first one or more computing devices to implement a data store in a multi-tenant provider network; and

a second one or more computing devices to implement a static script malware detection service in the multi-tenant provider network, the static script malware detection service including instructions that upon execution cause the static script malware detection service to:

receive a request to perform a static malware detection analysis, the request identifying a script comprising code;

generate, with a bytestring encoder, a bytestring embedding from a bytestring of a code statement in the script;

generate a semantic representation of the script, the semantic representation including context about statements in the code of the script;

generate, with a semantics encoder, a semantics embedding from a feature of the semantic representation, the feature corresponding to the code statement in the script;

generate a sequence of feature space embeddings based at least in part on the semantic representation of the script by combining the bytestring embedding and the semantics embedding;

evaluating the sequence of feature space embeddings using a recurrent neural network (RNN)-based sequential classifier model to generate a result that indicates whether the script is malicious; and

store an identification of the script and the result in an entry in a log in the data store.

15 . The system of claim 14 , wherein the semantic representation of the script is generated using a static code analysis parsing component from at least one of a syntax highlighter application or an abstract syntax tree constructor application.

16 . The system of claim 14 , wherein each feature space embedding in the sequence of feature space embeddings corresponds to a statement in the code.

17 . The system of claim 14 , wherein the semantic representation of the script is generated using a static code analysis parsing component from an abstract syntax tree constructor application, and wherein each feature space embedding in the sequence of feature space embeddings corresponds to a node in an abstract syntax tree of the script.

18 . The system of claim 14 , wherein the sequence of feature space embeddings represents a graph;

wherein generating the sequence of feature space embeddings and generating the result are performed using a machine learning model, the machine learning model including a graph embedding model and a classifier; and

wherein the graph embedding model is trained to generate feature space embeddings in a graph feature space, and wherein feature space embeddings in the sequence of feature space embeddings from similar graphs are closer together within the graph feature space than feature space embeddings in the sequence of feature space embeddings from dissimilar graphs.

19 . The system of claim 14 , wherein the sequence of feature space embeddings is ordered based on at least one of a sequence of statements in the code or a serialized traversal of nodes in an abstract syntax tree representing the script.

20 . The system of claim 14 , wherein the context about statements in the code of the script includes at least one of an identification of language keywords or an identification of function structure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2023
From: ERDEMIR, ECENAZ; MORAIS, MICHAEL JAMES; MARSCHALEK, MARION; PARK, KYUHONG; FAN, YI; GAO, VIANNE RAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 065047/0072 →
References Cited (36)
US 10521587B1 · Agranonik · 2019 [cited by examiner]
US 10922604B2 · Zhao · 2021 [cited by examiner]
US 10956477B1 · Fang · 2021 [cited by examiner]
US 11574053B1 · Chen · 2023 [cited by examiner]
US 11762990B2 · Gururajan · 2023 [cited by examiner]
US 20120260340A1 · Morris · 2012 [cited by examiner]
US 20180300480A1 · Sawhney · 2018 [cited by examiner]
US 20190377877A1 · Johns · 2019 [cited by examiner]
US 20200311266A1 · Jas · 2020 [cited by examiner]
US 20210240825A1 · Kutt · 2021 [cited by examiner]
US 20230185915A1 · Rao · 2023 [cited by examiner]
Liu et al., A unified multi-task learning model for AST-level and token-level code completion, Apr. 18, 2022, Empirical Software Engineering, 38 pages total (Year: 2022). [cited by examiner]
Fang et al., JStrong: Malicious JavaScript detection based on code semantic representation and graph neural network, Apr. 6, 2022, Computers & Security, Elsevier, 11 pages total (Year: 2022). [cited by examiner]
Hendler et al., Detecting Malicious PowerShell Commands using Deep Neural Networks, Association for Computing Machinery, ASIACCS'18, Jun. 4, 2018, 11 pages total (Year: 2018). [cited by examiner]
Rusak et al., Poster: AST-Based Deep Learning for Detecting Malicious PowerShell, Oct. 15, 2018, Proceedings of 2018 ACMSIGSACConference on Computer & Communications Security (CCS '18), 3 pages total (Year: 2018). [cited by examiner]
KDNuggets, Adding an attention mechanism to RNNs, Mar. 10, 2022, 5 total pages (Year: 2022). [cited by examiner]
Alex Delamotte, Dissecting alienfox: The cloud spammer's swiss army knife. Mar. 30, 2023., 1-13. [cited by applicant]
Andreas Moser et al., Limits of static analysis for malware detection. In Twenty-Third Annual Computer Security Applications Conference (ACSAC 2007), pp. 421-430, 2007. [cited by applicant]
Colin Clement et al., Pymt5: multi-mode translation of natural language and python code with transformers. ArXiv, Nov. 2020., 1-14. [cited by applicant]
Edward Raff, Classifying sequences of extreme length with constant memory applied to malware detection. Proceedings of the AAAI Conference on Artificial Intelligence, 35(11):9386-9394, May 2021., 1-9. [cited by applicant]
Edward Ruff et al., Malware detection by eating a whole EXE. In the Workshops of the the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, Louisiana, USA, Feb. 2-7, 2018, 1-9. [cited by applicant]
GitHub Team, Tree-sitter: a new parsing system for programming tools., 2017., 1-9. [cited by applicant]
Guillermo Suarez-Tangil et al., Dendroid: A text mining approach to analyzing and classifying code structures in android malware families. Expert Systems with Applications, 2014., 1-9. [cited by applicant]
Julian Georg Zilly et al., Recurrent highway networks., Aug. 2017, 1-10. [cited by applicant]
Kesu Wang et al., Unified abstract syntax tree representation learning for cross-language program classification. In 2022 IEEE/ACM 30th International Conference on Program Comprehension (ICPC), pp. 390-400, 2022. [cited by applicant]
Li Yujia et al., Graph matching networks for learning the similarity of graph structured objects. In Proceedings of the International conference on machine learning, 2019., 1-18. [cited by applicant]
Matt Muir, Legion: an aws credential harvester and smtp hijacker. Apr. 13, 2023., 1-12, https://www. cadosecurity.com/legion-an-aws-credential-harvester-and-smtp-hijacker/. [cited by applicant]
Milhai Christodorescu et al.,, Static analysis of executables to detect malicious patterns. In 12th USENIX Security Symposium (USENIX Security 03), Washington, D.C., Aug. 2003. USENIX Association. [cited by applicant]
Nwokedi Idika et al., A survey of malware detection techniques. Purdue University, Feb. 2, 2007, 1-48. [cited by applicant]
Paras Jain et al., Contrastive code representation learning. arXiv preprint, 2020., 1-20. [cited by applicant]
Ponemon Institute. The third annual study on the state of endpoint security risk. Jan. 2020., 1-35. [cited by applicant]
Sahil Suneja et al., Learning to map source code to software vulnerability using code-as-a-graph. ArXiv, abs/2006.08614, 2020, 1-8. [cited by applicant]
Tianqi Chen et al., XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '16,, Jun. 10, 2016, 1-13. [cited by applicant]
Tristan Hume and et al. Syntect., 2017, https://github.com/trishume/syntect. [cited by applicant]
Yuanzhi Ke et al., Cnn-encoded radical-level representation for japanese processing. Transactions of the Japanese Society for Artificial Intelligence, 33:D-123, 07 2018., 1-8. [cited by applicant]
Zhangyin Feng et al., CodeBERT: A pre-trained model for programming and natural languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1536-1547, Online, Nov. 2020. [cited by applicant]