IP Library Granted Patent US 8,776,231
Granted Patent B2
US 8,776,231 · App. 12/471,529 · Granted Jul 8, 2014

Unknown malcode detection using classifiers with optimal training sets

Inventors: Robert Moskovitch (Ashkelon, IL); Yuval Elovici (Moshav Arugot, IL)
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,776,231
App. No.
12/471,529
Granted
Jul 8, 2014
Kind
B2
Abstract

A method for detecting unknown malicious code is provided. A data set is created, which is a collection of files that includes a first subset with malicious code and a second subset with benign code files, whereas the malicious and benign files are identified by an antivirus program. Subsequently, all files are parsed and a set of top features of all-n grams of the files is selected and reduced by using features selection methods. After determining the optimal number of features, they will be used as training and test sets.

Claims (13)

1. A method for detecting unknown malicious code in a computer readable file, comprising:

a) creating a Data Set being a collection of files that includes a first subset with malicious code files and a second subset with benign code files;

b) identifying malicious code files and benign code files by an antivirus program;

c) parsing all files using n-gram moving windows of several lengths;

d) for each n-gram in each file, computing a term frequency (TF) representation, wherein the TF representation of an n-gram is the frequency of this n-gram in one of the files;

e) selecting an initial set of top features of all n-grams, based on a document frequency (DF) measure, wherein the DF measure indicates those of the files containing a given n-gram;

f) reducing the number of the top features to comply with the computation resources required for classifier training, by using features selection methods, wherein the features selection methods are DF measure, Gain Ratio and Fisher Score;

g) determining an optimal number of features based on the evaluation of the detection accuracy of several sets of reduced top features;

h) preparing different data sets with different distributions of benign code files and malicious code files, based on said optimal number, which will be used as training and test sets, wherein a percentage of malicious code files in each data set is in a range between 10% and 40% of all the files in the said data set;

i) for each classifier:

i.1) iteratively evaluating the detection accuracy for all combinations of training and test sets distributions, while in each iteration training a classifier using a specific distribution and testing the trained classifier on all distributions; and

i.2) selecting the optimal distribution that results with the highest detection accuracy.

2. A method according to claim 1 , wherein the malicious code is a malicious executable.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2009
From: MOSKOVITCH, ROBERT; ELOVICI, YUVAL
To: BEN-GURION UNIVERSITY OF THE NEGEV RESEARCH & DEVELOPMENT AUTHORITY
Reel/Frame 022730/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2009
From: BEN-GURION UNIVERSITY OF THE NEGEV RESEARCH & DEVELOPMENT AUTHORITY
To: DEUTSCHE TELEKOM AG
Reel/Frame 022730/0308 →
Priority Claims (1)
IL 191744 · May 27, 2008 · national
Continuity (1)
Related Publication 20090300765A1 · Dec 3, 2009