IP Library › Granted Patent US 9,471,882
Granted Patent B2
US 9,471,882 · App. 14/234,747 · Granted Oct 18, 2016

Information identification method, program product, and system using relative frequency

Inventors: Shohei Hido (Kanagawa-ken, JP); Michiaki Tatsubori (Tokyo, JP)
Assignee: International Business Machines Corporation
G06N99/005G06F21/552G06Q10/10G06Q40/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,471,882
App. No.
14/234,747
Filed
Jan 24, 2014
Granted
Oct 18, 2016
Kind
B2
Examiner
CHANG, LI WU
Art Unit
2129
USPC
706/12
Abstract

In a case where supervised (learning) data is prepared and the case where test data is prepared, the data is recorded with time information attached to the data. The method includes clustering the learning data in a target class and clustering the test data in the target class. Then, the probability density for each of identified subclasses is calculated for each of time intervals having various time points and widths for the learning data, and is calculated for each of time intervals in the latest time period which have various widths, for the test data. Then, a ratio between a probability density obtained when learning is performed and a probability density obtained when testing is performed is obtained as a relative frequency in each of the time intervals for each of the subclasses. Input having a relative frequency that statistically and markedly increases is detected as an anomaly.

Claims (56)

1. A computer implemented information identification method for detecting an attack carried out using irregular data against a classifier that is configured by means of supervised machine learning, the method comprising of:

preparing a plurality of pieces of training data each including a feature vector, a class label, and a time stamp of each piece of the training data;

generating the classifier by using the feature vector and the class label of the plurality of pieces of training data;

clustering the plurality of pieces of training data based on a distance between the feature vectors of the plurality of pieces of training data into a plurality of subclasses;

preparing a plurality of pieces of test data each including a feature vector, a class label, and a time stamp of each piece of the test data;

classifying the plurality of pieces of test data by using the classifier, the classifier adding the class label to each piece of the test data;

clustering the plurality of pieces of test data, which have been classified by using the classifier, into the plurality of subclasses;

for each subclass, calculating statistical data representing a ratio of a frequency of the plurality of pieces of test data clustered into the corresponding subclass and a frequency of the plurality of pieces of training data clustered into the corresponding subclass; and

warning of a possibility of occurrence of the attack, carried out using the irregular data, in one or more subclasses of the plurality of subclasses, in response to a value of the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding a predetermined threshold.

2. The information identification method according to claim 1 , wherein

the feature vector is obtained by converting an answer to a question item in a financial application document into an electronic form, and the class label represents classes including an acceptance class and a rejection class.

3. The information identification method according to claim 1 , wherein

the classifier is configured with a support vector machine.

4. The information identification method according to claim 1 , wherein clustering the plurality of pieces of training data is performed based on a K-means algorithm.

5. The information identification method according to claim 2 , wherein

the irregular data is falsely accepted data.

6. The information identification method according to claim 1 , wherein

the statistical data is calculated by using a moving average of the ratio and a variance of the moving average.

7. A non-transitory storage medium readable by a processor, the storage medium storing a program of instructions executable by the processor to perform a method of detecting an attack carried out using irregular data against a classifier that is configured by means of supervised machine learning, the method comprising:

preparing a plurality of pieces of training data each including a feature vector, a class label, and a time stamp of each piece of the training data;

generating the classifier by using the feature vector and the class label of the plurality of pieces of training data;

clustering the plurality of pieces of training data based on a distance between the feature vectors of the plurality of pieces of training data into a plurality of subclasses;

preparing a plurality of pieces of test data each including a feature vector, a class label, and a time stamp of each piece of the test data;

classifying the plurality of pieces of test data by using the classifier, the classifier adding the class label to each piece of the test data;

clustering the plurality of pieces of test data, which have been classified by using the classifier, into the plurality of subclasses;

for each subclass, calculating statistical data representing a ratio of a frequency of the plurality of pieces of test data clustered into the corresponding subclass and a frequency of the plurality of pieces of training data clustered into the corresponding subclass; and

warning of a possibility of occurrence of the attack, carried out using the irregular data, in one or more subclasses of the plurality of subclasses, in response to a value of the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding a predetermined threshold.

8. The information identification program product according to claim 7 , wherein the feature vector is obtained by converting an answer to a question item in a financial application document into an electronic form, and the class label represents classes including an acceptance class and a rejection class.

9. The information identification program product according to claim 7 , wherein the classifier is configured with a support vector machine.

10. The information identification program product according to claim 7 , wherein clustering the plurality of pieces of training data is performed based on a K-means algorithm.

11. The information identification program product according to claim 8 , wherein the irregular data is falsely accepted data.

12. The information identification program product according to claim 7 , wherein the statistical data is calculated by using a moving average of the ratio and a variance of the moving average.

13. A computer implemented information identification system for detecting an attack carried out using irregular data against a classifier that is configured by means of supervised machine learning, the information identification system comprising:

a storage device and a processor coupled with the storage device, the processor configured to process or execute data or routines included in the storage device,

wherein the storage device includes:

a plurality of pieces of training data each including a feature vector, a class label, and a time stamp of each piece of the training data, and being stored in the storage device;

a classifier generated by using the feature vector and the class label of plurality of pieces of training data;

a sub-classifier obtained by clustering the plurality of pieces of training data based on a distance between the feature vectors of the plurality of pieces of training data into a plurality of subclasses;

a plurality of pieces of test data each including a feature vector, a class label, and a time stamp of each piece of test data, and the test data being stored in the storage device, wherein the plurality of pieces of test data are classified by using the classifier, wherein the plurality of pieces of test data, which have been classified by using the classifier, are clustered into the plurality of subclasses, and wherein the classifier adds the class label to each piece of the test data;

calculation routine, for each subclass, calculating statistical data representing a ratio of a frequency of the plurality of pieces of test data clustered into the corresponding subclass and a frequency of the plurality of pieces of training data clustered into the corresponding subclass; and

warning routine warning of a possibility of occurrence of the attack, carried out using the irregular data, in one or more subclasses of the plurality of subclasses, in response to a value of the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding a predetermined threshold.

14. The information identification system according to claim 13 , wherein

the feature vector is obtained by converting an answer to a question item in a financial application document into an electronic form, and the class label represents classes including an acceptance class and a rejection class.

15. The information identification system according to claim 13 , wherein

the classifier is configured with a support vector machine.

16. The information identification system according to claim 13 , wherein

the sub-classifier uses a K-means algorithm.

17. The information identification system according to claim 14 , wherein

the irregular data is falsely accepted data.

18. The information identification system according to claim 13 , wherein

the statistical data is calculated by using a moving average of the ratio and a variance of the moving average.

19. The information identification method according to claim 1 , wherein

the warning of a possibility of occurrence of the attack includes displaying a first time window in which the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding the predetermined threshold.

20. The information identification method according to claim 19 , further comprising:

analyzing the pieces of test data in the first time window in which the statistical data exceeds the predetermined threshold; and

modifying the class label of the pieces of test data of the first time window into a rejection class, in response to the pieces of test data in the first time window is determined to be misclassified.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2014
From: HIDO, SHOHEI; TATSUBORI, MICHIAKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 032038/0672 →
Priority Claims (1)
JP 2011-162082 · Jul 25, 2011 · national
Continuity (1)
Related Publication 20140180980A1 · Jun 26, 2014