IP Library Granted Patent US 9,378,266
Granted Patent B2
US 9,378,266 · App. 14/347,778 · Granted Jun 28, 2016

System and method for supporting cluster analysis and apparatus supporting the same

Inventors: Minsoeng Kim (Seoul, KR); Doyung Yoon (Seoul, KR); Chaehyun Lee (Seoul, KR); Junesup Lee (Yongin-si, KR)
Assignee: SK PLANET CO., LTD.
G06F17/30598G06K9/00979G06K9/6223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,378,266
App. No.
14/347,778
Granted
Jun 28, 2016
Kind
B2
Abstract

Disclosed is a cluster analysis supporting system, with respect to providing a cluster analysis function, including a cluster analysis service apparatus configured to request a distributed processing service apparatus to perform a k-means clustering based on k values within a predetermined range and a preset iteration frequency until a predefined converge condition is satisfied, and if center values of the k values are calculated from the distributed processing service apparatus, select an optimum center value among the center values, and control calculation and application of an optimum k value through an index calculation with respect to applying clustered indexes assigned based on the selected optimum center value to data, and the distributed processing service apparatus configured to perform the k-means clustering based on the k values and the preset iteration frequency provided from the cluster analysis service apparatus upon the request by the cluster analysis service apparatus.

Claims (30)

1. A system for supporting a cluster analysis, the system comprising:

a cluster analysis service apparatus configured to select an optimum center values corresponding to each of k values within a predetermined range through a simultaneous cluster analysis of a k-means which corresponds to a preset iteration frequency for each of the k values within the predetermined range provided for the cluster analysis and determine an optimum k value among the k values within the predetermined range through an index calculation with respect to applying clustered indexes assigned based on the selected optimum center values to data; and

at least one distributed processing service apparatus configured to provide the clustering analysis service apparatus with the optimum center values selected by simultaneously performing the cluster analysis of a k-means which corresponds to a preset iteration frequency for each of the k values within the predetermined range upon a request by the cluster analysis service apparatus, wherein the distributed processing service apparatus comprises at least one data node configured to provide the clustered indexes on the data to the cluster analysis service apparatus.

2. A cluster analysis service apparatus for supporting a cluster analysis, the cluster analysis service apparatus comprising:

an apparatus storage unit configured to store data;

an apparatus input unit configured to generate an input signal related to k values within a predetermined range, a convergence condition, and an iteration frequency, which are provided for a cluster analysis of the stored data; and

an apparatus control unit configured to select an optimum center values corresponding to each of the k values by simultaneously performing a cluster analysis of a k-means which corresponds to a preset iteration frequency for each of the k values, and determine an optimum k value among the k values within the predetermined range through an index calculation with respect to applying clustered indexes assigned based on the selected optimum center values to data.

3. The cluster analysis service apparatus of claim 2 , wherein the apparatus storage unit stores a previously calculated previous k value.

4. The cluster analysis service apparatus of claim 3 , wherein the apparatus control unit comprises:

a data distribution unit configured to distribute data such that a cluster analysis of a k-means which corresponds to a preset iteration frequency for each of the k values is simultaneously performed;

an analysis result selection unit configured to select optimum center values;

an analysis index application unit configured to perform an index calculation with respect to k value efficiency of data assigned clustered indexes obtained by applying a result of selection having the optimum center values to the data; and

an optimum value update unit configured to update a previously stored k value based on a k value having an optimum result of the index calculation.

5. The cluster analysis service apparatus of claim 2 , wherein the apparatus control unit is configured to support the cluster analysis to automatically perform result calculation of center values with respect to a plurality of k values simultaneously by a plurality of number of times while having a different initialization value each time of the calculation.

6. A method of supporting a cluster analysis, the method comprising:

by a cluster analysis service apparatus, transmitting data in a distributed manner to data nodes based on a preset iteration frequency for each of k values within a predetermined range, which are input to perform a cluster analysis;

by the data nodes, performing a cluster analysis of a k-means which corresponds to each of the k values;

by the data nodes, selecting optimum center values corresponding to each of the k values based on the results of the cluster analysis, and providing the selected optimum center values to the cluster analysis service apparatus;

by the cluster analysis service apparatus, sharing the selected optimum center values with the data nodes;

by the data nodes, assigning clustered indexes obtained by applying the selected optimum center values to data;

by the cluster analysis service apparatus, performing an index calculation on the data to which the clustered indexes are assigned; and

by the cluster analysis service apparatus, determining an optimum k value among the k values within the predetermined range based on the results of the index calculation.

7. The method of claim 6 , wherein in the step of performing an index calculation, the cluster analysis service apparatus calculates data based on which the index calculation is performed according to a sampling condition, and performs the index calculation based on the calculated data.

8. The method of claim 6 , further comprising by the cluster analysis service apparatus, performing sampling on the data assigned the clustered indexes that are provided by the data nodes.

9. The method of claim 6 , wherein the performing of the index calculation comprises:

performing calculation of a clustered index for each k value; and

selecting a k value having a highest clustered index.

10. The method of claim 9 , wherein the performing of the index calculation further comprises applying a plurality of indexing methods to the calculation of the clustered index for each k value to select a relatively higher k value from the plurality of indexing methods.

11. The method of claim 6 , wherein in the step of performing the k-means clustering, a result of the k-means clustering is automatically calculated with respect to a plurality of k values simultaneously by a plurality of number of times while having a different initialization value each time of the calculation.

12. A computer readable recording medium recording a program executing the method recited in claim 6 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2014
From: KIM, MINSOENG; YOON, DOYUNG; LEE, CHAEHYUN; LEE, JUNESUP
To: SK PLANET CO., LTD.
Reel/Frame 032541/0676 →
Priority Claims (1)
KR 10-2012-0097498 · Sep 4, 2012 · national
Continuity (1)
Related Publication 20140236950A1 · Aug 21, 2014