IP Library › Granted Patent US 12,511,389
Granted Patent B2
US 12,511,389 · App. 18/366,875 · Granted Dec 30, 2025

Multi-level malware classification machine- learning method and system

Inventor: Mantas Briliauskas (Vilnius, LT)
Assignee: UAB 360 IT
G06F21/566G06F2221/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,389
App. No.
18/366,875
Granted
Dec 30, 2025
Kind
B2
Abstract

A cyber security method and system for detecting malware via an anti-malware application employing a fast locality-sensitive hashing evaluation using a vantage-point tree (VPT) structure for the indication of malicious files and non-malicious files. The locality-sensitive hashing evaluation using the VPT structure can be performed prior to initiating the deeper, more computationally intensive evaluation and is used to identify with high confidence a scanned file or data object being (i) a malicious file, (ii) a non-malicious file, or a low confidence measure of the two.

Claims (50)

1 . A method for detecting malware in a target file, the method comprising:

receiving the target file;

generating one or more fuzzy hashes from the target file using a locality-sensitive hashing operation;

determining a malware classification with respect to the target file by assessing the one or more fuzzy hashes using one or more vantage-point tree structures, wherein assessing the one or more fuzzy hashes comprises iteratively:

comparing a fuzzy hash of the one or more fuzzy hashes to a node of the one or more vantage-point tree structures to determine a distance that measures a similarity of the node and the fuzzy hash;

comparing the distance to a threshold; and

modifying, based on the comparison, a tree data structure by adding two nodes corresponding to children of the node of the one or more vantage-point tree structures;

wherein determining the malware classification comprises analyzing the tree data structure; and

performing a malware-based responsive action based on the malware classification.

2 . The method of claim 1 , wherein the malware classification is a prediction of whether the target file is malicious.

3 . The method of claim 1 , wherein assessing the one or more fuzzy hashes further comprises comparing the one or more fuzzy hashes to hashes generated using at least one of a library of malicious code or a library of non-malicious code.

4 . The method of claim 1 , wherein the node represents a known malware.

5 . The method of claim 1 , wherein the one or more vantage-point tree structures comprises:

a first vantage-point tree structure, wherein the nodes in the first vantage-point tree structure are generated by a first set of malicious code; and

a second vantage-point tree structure, wherein the nodes in the second vantage-point tree structure are generated by a set of non-malicious code or a second set of malicious code that is different from the first set of malicious code.

6 . The method of claim 1 , wherein assessing the one or more fuzzy hashes comprises searching for neighboring fuzzy hashes using the one or more vantage-point tree structures.

7 . The method of claim 6 , wherein searching for the neighboring fuzzy hashes is limited to a CPU cache level.

8 . The method of claim 7 , wherein the malware-based responsive action comprises at least one of rejecting, passing, or quarantining the target file.

9 . A system for determining whether a target file is malicious, the system comprising:

one or more processors; and

memory having instructions stored thereon that, when executed by the one or more processors, cause the system to:

receive the target file;

generate one or more fuzzy hashes from the target file using a locality-sensitive hashing operation;

determining a malware classification with respect to the target file by assessing the one or more fuzzy hashes using one or more vantage-point tree structures, wherein assessing the one or more fuzzy hashes comprises iteratively:

comparing a fuzzy hash of the one or more fuzzy hashes to a node of the one or more vantage-point tree structures to determine a distance that measures a similarity of the node and the fuzzy hash;

comparing the distance to a threshold; and

modifying, based on the comparison, a tree data structure by adding two nodes corresponding to children of the node of the one or more vantage-point tree structures;

wherein determining the malware classification comprises analyzing the tree data structure; and

perform a malware-based responsive action based on the malware classification.

10 . The system of claim 9 , wherein the malware classification is a prediction of whether the target file is malicious.

11 . The system of claim 9 , wherein assessing the one or more fuzzy hashes further comprises comparing the one or more fuzzy hashes to hashes generated using a library of malicious code.

12 . The system of claim 9 , wherein assessing the one or more fuzzy hashes further comprises comparing the one or more fuzzy hashes to hashes generated using a library of non-malicious code.

13 . The system of claim 9 , wherein the node represents a known malware.

14 . The system of claim 9 , wherein the one or more vantage-point tree structures comprises:

a first vantage-point tree structure, wherein the nodes in a first vantage-point tree structure are generated by a first set of malicious code; and

a second vantage-point tree structure, wherein the nodes in the second vantage-point tree structure are generated by a set of non-malicious code or a second set of malicious code that is different from the first set of malicious code.

15 . The system of claim 14 , wherein a fuzzy hash of the one or more fuzzy hashes are added to the set of non-malicious code or the first set of malicious code or the second set of malicious code to be subsequently used to update at least one of the first vantage-point tree structure or the second vantage-point tree structure.

16 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause a device to:

receive a target file;

generate one or more fuzzy hashes from the target file using a locality-sensitive hashing operation;

determining a malware classification with respect to the target file by assessing the one or more fuzzy hashes using one or more vantage-point tree structures, wherein assessing the one or more fuzzy hashes comprises iteratively:

comparing a fuzzy hash of the one or more fuzzy hashes to a node of the one or more vantage point tree structures to determine a distance that measures a similarity of the node and the fuzzy hash;

comparing the distance to a threshold; and

modifying, based on the comparison, a tree data structure by adding two nodes corresponding to children of the node of the one or more vantage-point tree structures;

wherein determining the malware classification comprises analyzing the tree data structure; and

perform a malware-based responsive action based on the malware classification.

17 . The computer-readable medium of claim 16 , wherein the malware classification is a prediction of whether the target file is malicious.

18 . The computer-readable medium of claim 16 , wherein the one or more vantage-point tree structures comprises:

a first vantage-point tree structure, wherein the nodes in the first vantage-point tree structure are generated by a first set of malicious code; and

a second vantage-point tree structure, wherein the nodes in the second vantage-point tree structure are generated by a set of non-malicious code or a second set of malicious code that is different from the first set of malicious code.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2024
From: BRILIAUSKAS, MANTAS
To: UAB 360 IT
Reel/Frame 067639/0120 →
Continuity (2)
Continuation 18152476 · Jan 10, 2023
Related Publication 20240232355A1 · Jul 11, 2024
References Cited (22)
US 11182481B1 · Oliver · 2021 [cited by examiner]
US 11354409B1 · Kenefick · 2022 [cited by examiner]
US 11636161B1 · Chang · 2023 [cited by examiner]
US 20180097822A1 · Huang · 2018 [cited by applicant]
US 20200082083A1 · Choi · 2020 [cited by examiner]
US 20210073890A1 · Lee · 2021 [cited by examiner]
US 20210232682A1 · Mitra · 2021 [cited by applicant]
US 20220036208A1 · Rao · 2022 [cited by applicant]
US 20230029679A1 · Lancioni · 2023 [cited by examiner]
US 20230098919A1 · Kulaga · 2023 [cited by examiner]
Oliver, Jonathan, et al. “Fast Clustering of High Dimensional Data Clustering the Malware Bazaar Dataset.” 2022, 11 pages. http://tlsh.org/papersDir/n21_opt_cluster.pdf. [cited by applicant]
Choi, Sunoh. “Combined kNN Classification and hierarchical similarity hash for fast malware detection.” Applied Sciences 10.15 (2020): 5173. [cited by applicant]
Unpublished U.S. Appl. No. 17/725,718, filed Apr. 21, 2022. [cited by applicant]
Combing through the fuzz: Using fuzzy hashing and deep learning to counter malware detection evasion techniques. 2021. on-line at: https://www.microsoft.com/security/blog/2021/07/27/combing-through-the-fuzz-using-fuzzy-… [cited by applicant]
Kumar, Neeraj, Li Zhang, and Shree Nayar. “What is a good nearest neighbors algorithm for finding similar patches in images?. ” European conference on computer vision. Springer, Berlin, Heidelberg, 2008. [cited by applicant]
Fuzzy hash: https://www.microsoft.com/security/blog/2021/07/27/combing-through-the-fuzz-using-fuzzy-hashing-and-deep-learning-to-counter-malware-detection-evasion-techniques/. [cited by applicant]
Non-final Office Action in connection to U.S. Appl. No. 18/152,476, dated Oct. 24, 2024. [cited by applicant]
Interview Summary in connection to U.S. Appl. No. 18/152,476, dated Jan. 13, 2025. [cited by applicant]
Advisory Action in connection to U.S. Appl. No. 18/152,476, dated May 29, 2025. [cited by applicant]
Final Office Action in connection to U.S. Appl. No. 18/152,476, dated Mar. 21, 2025. [cited by applicant]
Non-Final Office Action in connection to U.S. Appl. No. 18/448,466, dated Jul. 29, 2025. [cited by applicant]
Hu, Xin et al., Large-Scale Malware Indexing Using Function-Call Graphs, In Proceedings of the 16th ACM conference on Computer and communications security (CCS '009) (2009), 10 pages. [cited by applicant]