IP Library Granted Patent US 11,620,471
Granted Patent B2
US 11,620,471 · App. 15/800,603 · Granted Apr 4, 2023

Clustering analysis for deduplication of training set samples for machine learning based computer threat analysis

Inventor: John Brock (Irvine, CA)
Assignee: Cylance Inc.
G06K9/6218G05B13/0265G06F21/563G06F21/564G06K9/622G06K9/6255G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,620,471
App. No.
15/800,603
Granted
Apr 4, 2023
Kind
B2
Abstract

A method, a system, and a computer program product for performing analysis of data to detect presence of malicious code are disclosed. Reduced dimensionality vectors are generated from a plurality of original dimensionality vectors representing features in a plurality of samples. The reduced dimensionality vectors have a lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors. A first plurality of clusters is determined by applying a first clustering algorithm to the reduced dimensionality vectors. A second plurality of clusters is determined by applying a second clustering algorithm to one or more clusters in the first plurality of clusters using the original dimensionality. An exemplar for a cluster in the second plurality of clusters is added to a training set, which is used to train a machine learning model for identifying a file containing malicious code.

Claims (36)

1. A computer-implemented method comprising

generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors;

first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors;

second determining a second plurality of clusters and a plurality of outliers not forming part of one of the second plurality of clusters, the second determining comprising second applying a second clustering algorithm to original dimensionality vectors corresponding to the reduced dimensionality vectors in the first plurality of clusters, the second plurality of clusters being smaller than the first plurality of clusters;

randomly selecting at least one exemplar from each of at least a portion of the clusters of the second plurality of clusters;

adding the randomly selected exemplars to a training set;

adding the plurality of outliers to the training set; and

training a machine learning model for identifying a file containing malicious code, the training comprising use of the training set;

wherein at least one of the generating, the first determining, the second determining, the adding, and the training is performed by at least one processor of at least one computing system.

2. The method according to claim 1 , wherein the first and second clustering algorithms are same.

3. The method according to claim 1 , where the generating of the reduced dimensionality vectors comprises applying a random projection to the original dimensionality vectors.

4. The method according to claim 3 , wherein the random projection approximately preserves all pairwise distances between the original dimensionality vectors.

5. The method according to claim 3 , wherein the random projection has a predetermined size.

6. The method according to claim 1 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between points contained in the cluster of the second plurality of clusters are less than the predetermined radius.

7. The method according to claim 1 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between some points contained in the cluster of the second plurality of clusters are greater than the predetermined radius.

8. The method according to claim 1 , wherein the cluster of the second plurality of clusters has a predetermined minimum number of points.

9. A system comprising:

at least one programmable processor; and

memory storing instructions which, when executed by the at least one programmable processor, execute operations comprising:

generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors;

first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors;

second determining a second plurality of clusters and a plurality of outliers not forming part of one of the second plurality of clusters, the second determining comprising second applying a second clustering algorithm to original dimensionality vectors corresponding to the reduced dimensionality vectors in the first plurality of clusters, the second plurality of clusters being smaller than the first plurality of clusters;

selecting at least one exemplar from each of at least a portion of the clusters of the second plurality of clusters which correspond to a point approximately close to a center of the associated cluster in the second plurality of clusters;

adding the selected exemplars to a training set;

adding the plurality of outliers to the training set; and

training a machine learning model for identifying a file containing malicious code, the training comprising use of the training set.

10. The system according to claim 9 , wherein the first and second clustering algorithms are same.

11. The system according to claim 9 , where the generating of the reduced dimensionality vectors comprises applying a random projection to the original dimensionality vectors.

12. The system according to claim 11 , wherein the random projection approximately preserves all pairwise distances between the original dimensionality vectors.

13. The system according to claim 11 , wherein the random projection has a predetermined size.

14. A computer program product comprising a non-transitory machine readable medium storing instructions that, when executed by one or more programmable processors, cause the one or more programmable processors to perform operations comprising:

generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors;

first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors;

second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to original dimensionality vectors corresponding to the reduced dimensionality vectors in the first plurality of clusters, the second plurality of clusters being smaller than the first plurality of clusters;

adding, from each cluster in the second plurality of clusters, a single exemplar to a training set which is based on a date of creation of the corresponding cluster ; and

training a machine learning model for identifying a file containing malicious code, the training comprising use of the training set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: BROCK, JOHN
To: CYLANCE INC.
Reel/Frame 044011/0849 →
Continuity (2)
Provisional Application 62428402 · Nov 30, 2016
Related Publication 20180150724A1 · May 31, 2018
Cited By (2)
US 12,189,551 US 12,430,899