MULTI-LEVEL MALWARE CLASSIFICATION MACHINE-LEARNING METHOD AND SYSTEM
A cyber security method and system for detecting malware via an anti-malware application employing a fast locality-sensitive hashing evaluation using a vantage-point tree (VPT) structure for the indication of malicious files and non-malicious files. The locality-sensitive hashing evaluation using the VPT structure can be performed prior to initiating the deeper, more computationally intensive evaluation and is used to identify with high confidence a scanned file or data object being (i) a malicious file, (ii) a non-malicious file, or a low confidence measure of the two.
1 . A method for generating a malware classification output for a target code, the method comprising:
receiving the target code;
identifying, via a similarity-based operation comprising a locality-sensitive hashing operation assessed using one or more vantage-point tree structures, as a first malware classification operation, a malware classification output with respect to the target code, wherein the similarity-based operation is performed entirely using CPU caching;
in an instance in which the first malware classification output fails to satisfy a first confidence threshold associated with a malware classification or a second confidence threshold associated with a non-malware classification, generating, using a trained malware classification machine learning model, a second malware classification output; and
performing one or more malware-based actions, including to reject, pass, and/or quarantine the target code, based on the first malware classification output or the second malware classification output.
2 . The method of claim 1 , wherein the second malware classification output comprises a trained neural network model.
3 . The method of claim 1 , wherein the trained malware classification machine learning model is not executed until after the first malware classification output is generated.
4 . The system of claim 1 , wherein the similarity-based operation is assessed with respect to a library of malware code.
5 . The method of claim 4 , wherein the similarity-based operation is further assessed with respect to a library of non-malware code.
6 . The method of claim 1 , wherein the similarity-based operation calculates a first distance value of fuzzy hashes of the target code to nodes in a first vantage-point tree structure of the one or more vantage-point tree structures, wherein the nodes in the first vantage-point tree structure are generated by a set of malware code, and
wherein the similarity-based operation calculates a second distance value of fuzzy hashes of the target code to nodes in a second vantage-point tree structure or the first vantage-point tree structure, wherein the nodes in the second vantage-point tree structure or the first vantage-point tree structure employed for the second distance value calculation are generated by a set of non-malware code.
7 . The method of claim 1 , wherein the similarity-based operation generates multiple search results.
8 . A system comprising:
a processor; and
a memory having instructions stored thereon for generating a malware classification output for a target code, wherein execution of the instructions by the processor causes the processor to:
receive the target code;
identify, via a similarity-based operation comprising a fast locality-sensitive hashing operation assessed using one or more vantage-point tree structures, as a first malware classification operation, a malware classification output with respect to the target code;
in an instance in which the first malware classification output fails to satisfy a first confidence threshold associated with a malware classification or a second confidence threshold associated with a non-malware classification, generate, using a trained malware classification machine learning model, a second malware classification output; and
perform one or more malware-based actions, including to reject, pass, and/or quarantine the target code, based on the first malware classification output or the second malware classification output.
9 . The system of claim 8 , wherein the second malware classification output comprises a trained neural network model.
10 . The system of claim 8 , wherein the trained malware classification machine learning model is not executed until after the first malware classification output is generated.
11 . The system of claim 8 , wherein the similarity-based operation is assessed with respect to a library of malware code.
12 . The system of claim 11 , wherein the similarity-based operation is further assessed with respect to a library of non-malware code.
13 . The system of claim 8 , wherein the similarity-based operation calculates a first distance value of fuzzy hashes of the target code to nodes in a first vantage-point tree structure of the one or more vantage-point tree structures, wherein the nodes in a first vantage-point tree structure are generated by a set of malware code, and
wherein the similarity-based operation calculates a second distance value of fuzzy hashes of the target code to nodes in a second vantage-point tree structure or the first vantage-point tree structure, wherein the nodes in the second vantage-point tree structure or the first vantage-point tree structure employed for the second distance value calculation are generated by a set of non-malware code.
14 . The system of claim 8 , wherein the fuzzy hashes of the target code are added to the set of non-malware code or the set of malware code to be subsequently used to update at least one of the first vantage-point tree structure or the second vantage-point tree structure.
15 . A non-transitory computer-readable medium having instructions stored thereon for generating a malware classification output for a target code, wherein execution of the instructions by a processor causes the processor to:
receive the target code;
identify, via a similarity-based operation comprising a fast locality-sensitive hashing operation assessed using one or more vantage-point tree structures, as a first malware classification operation, a malware classification output with respect to the target code;
in an instance in which the first malware classification output fails to satisfy a first confidence threshold associated with a malware classification or a second confidence threshold associated with a non-malware classification, generate, using a trained malware classification machine learning model, a second malware classification output; and
perform one or more malware-based actions, including to reject, pass, and/or quarantine the target code, based on the first malware classification output or the second malware classification output.
16 . The computer-readable medium of claim 15 , wherein the second malware classification output comprises a trained neural network model.
17 . The computer-readable medium of claim 15 , wherein the trained malware classification machine learning model is not executed until after the first malware classification output is generated.
18 . The computer-readable medium of claim 15 , wherein the similarity-based operation is assessed with respect to a library of malware code.
19 . The computer-readable medium of claim 18 , wherein the similarity-based operation is further assessed with respect to a library of non-malware code.
20 . The computer-readable medium of claim 15 ,
wherein the similarity-based operation calculates a first distance value of fuzzy hashes of the target code to nodes in a first vantage-point tree structure of the one or more vantage-point tree structure, wherein the nodes in a first vantage-point tree structure are generated by a set of malware code, and
wherein the similarity-based operation calculates a second distance value of fuzzy hashes of the target code to nodes in a second vantage-point tree structure or the first vantage-point tree structure, wherein the nodes in the second vantage-point tree structure or the first vantage-point tree structure employed for the second distance value calculation are generated by a set of non-malware code.