IP Library Granted Patent US 12,153,604
Granted Patent B2
US 12,153,604 · App. 17/967,999 · Granted Nov 26, 2024

Apparatus and method for generating data set

Inventors: Hyun-Jin Kim (Daejeon, KR); Jong-Hoon Lee (Daejeon, KR); Young-Soo Kim (Sejong-si, KR); Jong-Geun Park (Sejong-si, KR); Cheol-Hee Park (Gongju-si, KR)
Assignee: Electronics and Telecommunications Research Institute
G06F16/285G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,153,604
App. No.
17/967,999
Granted
Nov 26, 2024
Kind
B2
Abstract

Disclosed herein are an apparatus and method for generating a data set. The apparatus includes one or more processors and executable memory for storing at least one program executed by the one or more processors. The at least one program classifies collected data into numerical feature data and categorical feature data using a filter method, performs correlation analysis on the numerical feature data and the categorical feature data using an analysis of variance (ANOVA) method and a Chi-Squared method, and generates a data set for supervised learning and a data set for unsupervised learning using correlation scores calculated through correlation analysis.

Claims (23)

1. An apparatus for generating a data set, comprising:

one or more processors; and

executable memory for storing at least one program executed by the one or more processors,

wherein the at least one program is configured to

classify collected data into numerical feature data and categorical feature data using a filter method, the collected data including network traffic data, system log data, and security event data,

perform correlation analysis on the numerical feature data and the categorical feature data using an analysis of variance (ANOVA) method and a Chi-Squared method,

generate a data set for supervised learning and a data set for unsupervised learning using correlation scores calculated through the correlation analysis, and

generate a neural network model for cyber breach threat detection based on the data set for supervised learning and a data set for unsupervised learning,

wherein the at least one program is configured to:

rank importance of features according to predefined feature criteria using the filter method and measure a correlation between data features based on the ranked importance of the features, thereby classifying the collected data into the numerical feature data and the categorical feature data,

normalize the numerical feature data using a min-max scaling method and convert the categorical feature data into numerical values using a one-hot encoding method, and

determine that data corresponds to the data set for supervised learning as a correlation score calculated using the ANOVA method is higher and a correlation score calculated using the Chi-Squared method is lower.

2. The apparatus of claim 1 , wherein the at least one program determines that data corresponds to the data set for unsupervised learning as the correlation score calculated using the ANOVA method is lower and the correlation score calculated using the Chi-Squared method is higher.

3. A method for generating a data set, performed by an apparatus for generating a data set, comprising:

classifying collected data into numerical feature data and categorical feature data using a filter method, the collected data including network traffic data, system log data, and security event data;

performing correlation analysis on the numerical feature data and the categorical feature data using an analysis of variance (ANOVA) method and a Chi-Squared method;

generating a data set for supervised learning and a data set for unsupervised learning using correlation scores calculated through the correlation analysis; and

generate a neural network model for cyber breach threat detection based on the data set for supervised learning and a data set for unsupervised learning,

wherein classifying the collected data comprises ranking importance of features according to predefined feature criteria using the filter method and measuring a correlation between data features based on the ranked importance of the features, thereby classifying the collected data into the numerical feature data and the categorical feature data,

wherein performing the correlation analysis comprises:

normalizing the numerical feature data using a min-max scaling method and converting the categorical feature data into numerical values using a one-hot encoding method; and

determining that data corresponds to the data set for supervised learning as a correlation score calculated using the ANOVA method is higher and as a correlation score calculated using the Chi-Squared method is lower.

4. The method of claim 3 , wherein performing the correlation analysis comprises determining that data corresponds to the data set for unsupervised learning as the correlation score calculated using the ANOVA method is lower and the correlation score calculated using the Chi-Squared method is higher.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: KIM, HYUN-JIN; LEE, JONG-HOON; KIM, YOUNG-SOO; PARK, JONG-GEUN; PARK, CHEOL-HEE
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 061450/0300 →
Priority Claims (1)
KR 10-2021-0139656 · Oct 19, 2021 · national
Continuity (1)
Related Publication 20230123045A1 · Apr 20, 2023