IP Library › Granted Patent US 12,483,590
Granted Patent B2
US 12,483,590 · App. 17/714,023 · Granted Nov 25, 2025

Methods and apparatus to visualize machine learning based malware classification

Inventors: Yonghong Huang (Hillsboro, OR); Steven Grobman (Plano, TX); Jonathan King (Hillsboro, OR)
Assignee: McAfee, LLC
H04L63/145H04L41/16H04L63/20H04L63/0227H04L63/1408
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,483,590
App. No.
17/714,023
Granted
Nov 25, 2025
Kind
B2
Abstract

Methods, apparatus, systems, and articles of manufacture are disclosed. An example apparatus includes at least one memory, instructions, and processor circuitry to execute the instructions. The processor circuitry executes the instructions to identify a test data distribution, generate a first visualization of the identified test data distribution, select a visualization type for a machine learning model, generate a second visualization including an indication of features extracted from the test data by the machine learning model, and generate a third visualization of results of inference performed by the machine learning model, the inference performed on the test data.

Claims (36)

1 . An apparatus comprising:

at least one memory;

machine-readable instructions; and

processor circuitry to execute the machine-readable instructions to:

identify a distribution associated with test data;

generate a first visualization of the identified distribution;

generate a recommendation to adjust a machine learning model, the recommendation based on a classification error, the recommendation to be displayed with the first visualization, the recommendation to include an instruction to reduce a malware classification error of the machine learning model;

select a visualization type for the machine learning model;

generate a second visualization including an indication of features extracted from the test data by the machine learning model, the second visualization generated using gradient weighted class activation mapping t-distributed stochastic neighbor embedding feature projections; and

generate a third visualization of results of an inference performed by the machine learning model, the inference performed on the test data, wherein the inference is at least one of a classification of a sample as malware or a classification of the sample as benign.

2 . The apparatus of claim 1 , wherein the processor circuitry is to generate a recommendation to improve inference accuracy based on a comparison of at least two of: the first visualization, the second visualization, or the third visualization.

3 . The apparatus of claim 1 , wherein the processor circuitry is to generate a fourth visualization of a machine learning pipeline, the fourth visualization including the first visualization, the second visualization, and the third visualization.

4 . The apparatus of claim 1 , wherein at least one of the first visualization, the second visualization, and the third visualization includes an indication that a data is an outlier data.

5 . The apparatus of claim 1 , wherein the third visualization includes a receiver operating characteristic curve and an indication of a partial area under the receiver operating characteristic curve.

6 . A non-transitory computer readable medium comprising instructions which, when executed, cause processor circuitry to:

identify a distribution associated with test data;

generate a first visualization of the identified distribution;

generate a recommendation to adjust a machine learning model, the recommendation based on a classification error, the recommendation to be displayed with the first visualization, the recommendation to include an instruction to reduce a malware classification error of the machine learning model;

select a visualization type for the machine learning model;

generate a second visualization including an indication of features extracted from the test data by the machine learning model, the second visualization generated using t-distributed stochastic neighbor embedding feature projections; and

generate a third visualization of results of an inference performed by the machine learning model, the inference performed on the test data, wherein the inference is at least one of a classification of a sample as malware or a classification of the sample as benign.

7 . The non-transitory computer readable medium of claim 6 , wherein the instructions, when executed, cause the processor circuitry to generate a recommendation to improve inference accuracy based on a comparison of at least two of: the first visualization, the second visualization, or the third visualization.

8 . The non-transitory computer readable medium of claim 6 , wherein the instructions, when executed, cause the processor circuitry to generate a fourth visualization of a machine learning pipeline, the fourth visualization including the first visualization, the second visualization, and the third visualization.

9 . The non-transitory computer readable medium of claim 6 , wherein at least one of the first visualization, the second visualization, and the third visualization includes an indication that a data is an outlier data.

10 . The non-transitory computer readable medium of claim 6 , wherein the third visualization includes a receiver operating characteristic curve and an indication of a partial area under the receiver operating characteristic curve.

11 . A method for error analysis in a machine learning pipeline, the method comprising:

identifying, by executing an instruction with processor circuitry, a distribution associated with test data;

generating, by executing an instruction with the processor circuitry, a first visualization of the identified distribution;

generating, by executing an instruction with the processor circuitry, a recommendation to adjust a machine learning model, the recommendation based on a classification error, the recommendation to be displayed with the first visualization, the recommendation to include an instruction to reduce a malware classification error of the machine learning model;

selecting, by executing an instruction with processor circuitry, a visualization type for the machine learning model;

generating, by executing an instruction with processor circuitry, a second visualization including an indication of features extracted from the test data by the machine learning model, the second visualization generated using t-distributed stochastic neighbor embedding feature projections; and

generating, by executing an instruction with processor circuitry, a third visualization of results of an inference performed by the machine learning model, the inference performed on the test data, wherein the inference is at least one of a classification of a sample as malware or a classification of the sample as benign.

12 . The method of claim 11 , further including generating a recommendation to improve inference accuracy based on a comparison of at least two of: the first visualization, the second visualization, and the third visualization.

13 . The method of claim 11 , further including generating a fourth visualization of the machine learning pipeline, the fourth visualization including the first visualization, the second visualization, and the third visualization.

14 . The method of claim 11 , wherein at least one of the first visualization, the second visualization, and the third visualization includes an indication that a data is an outlier data.

15 . The method of claim 11 , wherein the third visualization includes a receiver operating characteristic curve and an indication of a partial area under the receiver operating characteristic curve.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: HUANG, YONGHONG; GROBMAN, STEVEN; KING, JONATHAN
To: MCAFEE, LLC
Reel/Frame 060501/0384 →
Continuity (2)
Provisional Application 63170650 · Apr 5, 2021
Related Publication 20220321579A1 · Oct 6, 2022
References Cited (12)
US 11710034B2 · Anderson · 2023 [cited by examiner]
US 20150295945A1 · Canzanese, Jr. · 2015 [cited by examiner]
US 20180083903A1 · El-Alfy · 2018 [cited by examiner]
US 20180122508A1 · Wilde · 2018 [cited by examiner]
US 20200301955A1 · Ludlow · 2020 [cited by examiner]
US 20210110288A1 · Poothiyot · 2021 [cited by examiner]
US 20220067580A1 · Rho · 2022 [cited by examiner]
US 20220094709A1 · Sharma · 2022 [cited by examiner]
Selvaraju, R., et al., Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization, 2017, Proceedings of the IEEE International Conferenceon Computer Vision, p. 618-626. (Year: 2017). [cited by examiner]
Lundberg et al., “Explainable IA for Trees: From Local Explanations to Global Understanding”, University of Washington, May 11, 2019, 72 pages. (Retrieved from: https://arxiv.org/pdf/1905.04610.pdf on Apr. 5, 2022). [cited by applicant]
Selvaraju et al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization”, Georgia Institute of Technology, Dec. 3, 2019, 23 pages. (Retrieved from: https://arxiv.org/pdf/1610.02391.pdf on Apr… [cited by applicant]
Maaten et al., “Visualizing Data using t-SNE”, Journal of Machine Learning Research 9 (2008) 2579-2605, Published Nov. 8, 2008, 27 pages (Retrieved from: https://www.jmlr.org/papers/volume9/vandermaaten08a/vandermaaten0… [cited by applicant]