IP Library Granted Patent US 9,489,627
Granted Patent B2
US 9,489,627 · App. 13/932,299 · Granted Nov 8, 2016

Hybrid clustering for data analytics

Inventor: Jerzy W. Bala (Potomac Falls, VA)
Assignee: Bottomline Technologies (DE), Inc.
G06N5/025G06F19/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,489,627
App. No.
13/932,299
Granted
Nov 8, 2016
Kind
B2
Abstract

The invention includes methods and systems for analyzing data to determine trends in the data and to identify outliers. The methods and systems include a learning algorithm whereby a data space is co-populated with artificial, evenly distributed data, and then the data space is carved into smaller portions whereupon the number of real and artificial data points are compared. Through an iterative process, clusters having less than evenly distributed real data are discarded. Additionally, a final quality control measurement is used to merge clusters that are too similar to be meaningful. The invention is widely applicable to data analytics, generally, including financial transactions, retail sales, elections, and sports.

Claims (47)

1. A computer implemented method for clustering data, the method comprising:

(a) obtaining real data from a computer readable medium;

(b) identifying a rule that clusters the real data, wherein identifying comprises:

i) generating an evenly-distributed set of synthesized data having an identical number of points as the real data;

ii) creating a database on the computer readable medium including the real data and the synthesized data;

iii) executing a supervised machine-learning algorithm on a processor to determine a rule for clustering the real data, such that the real data is differentiated from the synthesized data:

iv) clustering the real data using the determined rule;

v) evaluating the quality of the determined rule for clustering the real data by:

calculating a first quality statistic based on a percentage of the real data clustered by the rule; and

calculating a second quality statistic based on a percentage of the real data and the synthesized data covered by the rule that is synthesized data;

wherein a rule resulting in a larger first quality statistic and a smaller second quality statistic is evaluated more favorably than a rule resulting in a smaller first quality statistic and a larger second quality statistic;

vi) correlating the clustered real data with a data space using a processor; and

vii) validating the rule by comparing the location of the clustered real data with a location of other clustered real data using a processor;

(c) clustering the real data with the rule using a processor; and

(d) storing the clustered real data on the computer readable medium.

2. The method of claim 1 , wherein validating comprises calculating a centroid of the clustered real data using a processor and calculating a distance between the centroid of the clustered real data and the centroid of a nearest neighboring cluster formed by the same rule.

3. The method of claim 2 , wherein validating further comprises calculating an average distance between each real data point in the cluster and the centroid of the clustered real data.

4. The method of claim 3 , further comprising calculating a ratio of (the average distance between each real data point in the cluster and the centroid) to (the distance between the centroid and the centroid of the nearest neighboring cluster formed by the same rule).

5. The method of claim 1 , further comprising calculating a centroid for each cluster of real data on the computer readable medium, and comparing the distribution of centroids from the clusters of real data on the computer readable medium to determine a trend in the real data.

6. The method of claim 1 , further comprising calculating a centroid for each cluster of real data on the computer readable medium, and comparing the distribution of centroids from the clusters of real data on the computer readable medium to determine an outlier in the real data.

7. The method of claim 1 , wherein the real data relates to genetic sequences.

8. The method of claim 1 , wherein the real data relates to stock, commodity, currency, or retail sales.

9. The method of claim 1 , wherein the real data relates to athletic performance.

10. The method of claim 1 , wherein the real data is data reported to a government authority.

11. A system for clustering data comprising a processor, memory, and a computer readable medium, wherein the computer readable medium comprises instructions that when executed cause the processor to:

obtain real data from the computer readable medium;

identify a rule that clusters the real data, wherein identifying comprises:

i) generating an evenly-distributed set of synthesized data having an identical number of points as the real data;

ii) creating a database on the computer readable medium including the real data and the synthesized data;

iii) executing a supervised machine-learning algorithm to determine a rule for clustering the real data, such that the real data is differentiated from the synthesized data;

iv) clustering the real data using the determined rule;

v) evaluating the Quality of the determined rule for clustering the real data by:

calculating a first quality statistic based on a percentage of the real data clustered by the rule; and

calculating a second quality statistic based on a percentage of the real data and the synthesized data covered by the rule that is synthesized data;

vi) correlating the clustered real data with a data space; and

vii) validating the rule by comparing the location of the clustered real data with a location of other clustered real data;

cluster the real data with the rule; and

store the clustered real data on the computer readable medium.

12. The system of claim 11 , wherein validating comprises calculating a centroid of the clustered real data and calculating a distance between the centroid of the clustered real data and the centroid of a nearest neighboring cluster formed by the same rule.

13. The system of claim 12 , wherein validating further comprises calculating an average distance between each real data point in the cluster and the centroid of the clustered real data.

14. The system of claim 13 , further comprising calculating a ratio of (the average distance between each real data point in the cluster and the centroid) to (the distance between the centroid and the centroid of the nearest neighboring cluster formed by the same rule).

15. The system of claim 11 , wherein the computer readable medium further comprises instructions that cause the processor to calculate a centroid for each cluster of real data on the computer readable medium, and compare the distribution of centroids from the clusters of real data on the computer readable medium to determine a trend in the real data.

16. The system of claim 11 , wherein the computer readable medium further comprises instructions that cause the processor to calculate a centroid for each cluster of real data on the computer readable medium, and compare the distribution of centroids from the clusters of real data on the computer readable medium to determine an outlier in the real data.

17. The system of claim 11 , wherein the real data relates to genetic sequences.

18. The system of claim 11 , wherein the real data relates to stock, commodity, currency, or retail sales.

19. The system of claim 11 , wherein the real data relates to athletic performance.

20. The system of claim 11 , wherein the real data is data reported to a government authority.

Assignments (6)
RELEASE OF SECURITY INTEREST IN REEL/FRAME: 040882/0908 Recorded May 13, 2022
From: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
To: BOTTOMLINE TECHNOLOGIES (DE), INC.
Reel/Frame 060063/0701 →
SECURITY INTEREST Recorded May 13, 2022
From: BOTTOMLINE TECHNOLOGIES, INC.
To: ARES CAPITAL CORPORATION
Reel/Frame 060064/0275 →
CHANGE OF NAME Recorded Mar 19, 2021
From: BOTTOMLINE TECHNOLOGIES (DE), INC.
To: BOTTOMLINE TECHNLOGIES, INC.
Reel/Frame 055661/0461 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Dec 12, 2016
From: BOTTOMLINE TECHNOLOGIES (DE), INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 040882/0908 →
MERGER Recorded Feb 26, 2014
From: RATIONALWAVE ANALYTICS, INC.
To: BOTTOMLINE TECHNOLOGIES (DE), INC.
Reel/Frame 032335/0070 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2013
From: BALA, JERZY W.
To: RATIONALWAVE ANALYTICS, INC.
Reel/Frame 031078/0238 →
Continuity (3)
Provisional Application 61727815 · Nov 19, 2012
Provisional Application 61776021 · Mar 11, 2013
Related Publication 20140143186A1 · May 22, 2014