IP Library › Granted Patent US 12,645,799
Granted Patent B2
US 12,645,799 · App. 18/897,632 · Granted Jun 2, 2026

Methods and apparatus to augment classification coverage for low prevalence samples through neighborhood labels proximity vectors

Inventors: German Lancioni (San Jose, CA); Jonathan King (Hillsboro, OR)
Assignee: McAfee, LLC
G06F21/566G06F21/56G06F21/567G06N7/01G06F2221/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,799
App. No.
18/897,632
Granted
Jun 2, 2026
Kind
B2
Abstract

Disclosed examples include obtaining malicious neighbor samples based on a non-classified sample; obtaining clean neighbor samples based on the non-classified sample; generating a malicious proximity vector representing first distances between the non-classified sample and a first malicious neighbor sample from the malicious neighbor samples; generating a clean proximity vector representing second distances between the non-classified sample and a first clean neighbor sample from the clean neighbor samples; and classifying the non-classified sample as a clean sample or a malicious sample based on at least one of the malicious proximity vector or the clean proximity vector.

Claims (50)

1 . An apparatus comprising:

interface circuitry;

machine-readable instructions; and

at least one processor circuit to be programmed by the machine-readable instructions to:

obtain malicious neighbor samples based on a non-classified sample;

obtain clean neighbor samples based on the non-classified sample;

generate a malicious proximity vector representing first distances between the non-classified sample and a first malicious neighbor sample from the malicious neighbor samples;

generate a clean proximity vector representing second distances between the non-classified sample and a first clean neighbor sample from the clean neighbor samples; and

classify the non-classified sample as a clean sample or a malicious sample based on at least one of the malicious proximity vector or the clean proximity vector.

2 . The apparatus of claim 1 , wherein the non-classified sample corresponds to a file that is not classified as malicious or clean.

3 . The apparatus of claim 1 , wherein the first distances include at least one of a hamming distance, a Euclidean distance, or a token set ratio distance.

4 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

sort the first distances in the malicious proximity vector based on a first one of the first distances first and a second one of the first distances second; and

sort the second distances in the clean proximity vector based on a first one of the second distances first and a second one of the second distances second.

5 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

generate a hash table based on the non-classified sample;

obtain the malicious neighbor samples from a malicious locality sensitive hashing (LSH) forest based on the hash table; and

obtain the clean neighbor samples from a clean LSH forest based on the hash table.

6 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to output a classification of the non-classified sample as the clean sample or the malicious sample after a determination that a feature-based classification of the non-classified sample does not satisfy a threshold.

7 . The apparatus of claim 1 , wherein a quantity of the first distances is based on a quantity of datatypes in the at least one of the malicious neighbor samples or the clean neighbor samples, a quantity of the second distances is based on the quantity of the datatypes.

8 . At least one non-transitory machine-readable medium comprising machine-readable instructions to cause at least one processor circuit to at least:

request malicious neighbor samples based on a non-classified sample;

request clean neighbor samples based on the non-classified sample;

cause storing of a malicious proximity vector, the malicious proximity vector representing first distances between the non-classified sample and a first malicious neighbor sample from the malicious neighbor samples;

cause storing of a clean proximity vector, the clean proximity vector representing second distances between the non-classified sample and a first clean neighbor sample from the clean neighbor samples; and

classify the non-classified sample as clean or malicious based on at least one of the malicious proximity vector or the clean proximity vector.

9 . The at least one non-transitory machine-readable medium of claim 8 , wherein the non-classified sample corresponds to a file.

10 . The at least one non-transitory machine-readable medium of claim 8 , wherein the first distances include at least one of a hamming distance, a Euclidean distance, or a token set ratio distance.

11 . The at least one non-transitory machine-readable medium of claim 8 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to:

sort the first distances in the malicious proximity vector based on a first one of the first distances first and a second one of the first distances second; and

sort the second distances in the clean proximity vector based on a first one of the second distances first and a second one of the second distances second.

12 . The at least one non-transitory machine-readable medium of claim 8 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to:

generate a hash table based on the non-classified sample;

request the malicious neighbor samples from a malicious locality sensitive hashing (LSH) forest based on the hash table; and

request the clean neighbor samples from a clean LSH forest based on the hash table.

13 . The at least one non-transitory machine-readable medium of claim 8 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to output a malicious classification or a clean classification of the non-classified sample based on the at least one of the malicious proximity vector or the clean proximity vector after a determination that a feature-based classification of the non-classified sample does not satisfy a threshold.

14 . The at least one non-transitory machine-readable medium of claim 8 , wherein a quantity of the first distances is based on a quantity of datatypes in the malicious neighbor samples and the clean neighbor samples, a quantity of the second distances is based on the quantity of the datatypes.

15 . A method comprising:

identifying malicious neighbor samples based on a non-classified sample;

identifying clean neighbor samples based on the non-classified sample;

generating, by at least one processor circuit programmed by at least one instruction, a malicious proximity vector that includes first distance values, the first distance values representing first distances between the non-classified sample and a first malicious neighbor sample from the malicious neighbor samples;

generating, by one or more of the at least one processor circuit, a clean proximity vector that includes second distance values, the second distance values representing second distances between the non-classified sample and a first clean neighbor sample from the clean neighbor samples; and

generating, by one or more of the at least one processor circuit, a classification of the non-classified sample as a clean sample or a malicious sample based on at least one of the malicious proximity vector or the clean proximity vector.

16 . The method of claim 15 , wherein the non-classified sample corresponds to a file that is not classified as malicious or clean.

17 . The method of claim 15 , wherein the first distances include at least one of a hamming distance, a Euclidean distance, or a token set ratio distance.

18 . The method of claim 15 , including:

sorting the first distance values in the malicious proximity vector based on a first one of the first distances first and a second one of the first distances second; and

sorting the second distance values in the clean proximity vector based on a first one of the second distances first and a second one of the second distances second.

19 . The method of claim 15 , including generating a hash table based on the non-classified sample, the identifying of the malicious neighbor samples is from a malicious locality sensitive hashing (LSH) forest based on the hash table, the identifying of the clean neighbor samples is from a clean LSH forest based on the hash table.

20 . The method of claim 15 , including outputting the classification of the non-classified sample as the clean sample or the malicious sample after a determination that a feature-based classification of the non-classified sample does not satisfy a threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: LANCIONI, GERMAN; KING, JONATHAN
To: MCAFEE, LLC
Reel/Frame 075538/0366 →
Continuity (3)
Continuation 17566760 · Dec 31, 2021
Provisional Application 63227305 · Jul 29, 2021
Related Publication 20250013748A1 · Jan 9, 2025
References Cited (36)
US 8621626B2 · Alme · 2013 [cited by applicant]
US 10924503B1 · Pereira et al. · 2021 [cited by applicant]
US 10970650B1 · Abusorrah · 2021 [cited by examiner]
US 11262742B2 · Srinivasamurthy · 2022 [cited by examiner]
US 11676069B2 · Hazard · 2023 [cited by examiner]
US 11880775B1 · Hazard · 2024 [cited by examiner]
US 20130326625A1 · Anderson et al. · 2013 [cited by applicant]
US 20160098561A1 · Keller et al. · 2016 [cited by applicant]
US 20160132521A1 · Reininger et al. · 2016 [cited by applicant]
US 20170126736A1 · Urias et al. · 2017 [cited by applicant]
US 20190026466A1 · Krasser et al. · 2019 [cited by applicant]
US 20190199736A1 · Howard et al. · 2019 [cited by applicant]
US 20200311262A1 · Nguyen · 2020 [cited by examiner]
US 20200410091A1 · Kimon · 2020 [cited by examiner]
US 20210295209A1 · Lancioni · 2021 [cited by examiner]
US 20220172105A1 · Nia · 2022 [cited by examiner]
US 20220318383A1 · Huang · 2022 [cited by examiner]
US 20230029679A1 · Lancioni et al. · 2023 [cited by applicant]
US 20230030136A1 · Lancioni et al. · 2023 [cited by applicant]
US 20230032194A1 · Lancioni et al. · 2023 [cited by applicant]
US 20230171277A1 · Giaconi et al. · 2023 [cited by applicant]
CN 110991538 · 2020 [cited by applicant]
EP 4446916B1 · 2025 [cited by examiner]
United States Patent and Trademark Office, “Non-Final Office Action,” issued Mar. 12, 2024 in connection with U.S. Appl. No. 17/645,921, 22 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued Mar. 14, 2024 in connection with U.S. Appl. No. 17/561,475, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance,” issued Aug. 30, 2024 in connection with U.S. Appl. No. 17/561,475, 6 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued Dec. 7, 2023 in connection with U.S. Appl. No. 17/566,760, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action,” issued Apr. 29, 2024 in connection with U.S. Appl. No. 17/566,760, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance,” issued Jun. 11, 2024 in connection with U.S. Appl. No. 17/566,760, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Corrected Notice of Allowability,” issued in connection with U.S. Appl. No. 17/566,760, dated Sep. 26, 2024, 9 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action,” issued in connection with U.S. Appl. No. 17/645,921, mailed on Oct. 10, 2024, 21 pages. [cited by applicant]
Alaeiyan, et al., “A Multilabel Fuzzy Relevance Clustering System for Malware Attack Attribution in the Edge Layer of Cyber-Physical Networks,” ACM Transactions on Cyber-Physical Systems, vol. 4, No. 3, 22 pages, Mar. 2… [cited by applicant]
D'Elia, et al., “On the Dissection of Evasive Malware,” IEEE Transactions on Information Forensics and Security, 15 pages, 2020. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/645,921, mailed Jun. 17, 2025, 10 pages. [cited by applicant]
United States Patent and Trademark Office, “Advisory Action,” issued in connection with U.S. Appl. No. 17/645,921, dated Jan. 30, 2025, 3 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/645,921, dated Feb. 12, 2025, 22 pages. [cited by applicant]