IP Library Granted Patent US 8,918,397
Granted Patent B2
US 8,918,397 · App. 13/561,468 · Granted Dec 23, 2014

Clustering customers

Inventors: Heng Cao (Shanghai, CN); Jin Dong (Beijing, CN); Jacqueline Giang Huong Morris (Brooklyn, NY); Ming Xie (Beijing, CN); Wen Jun Yin (Beijing, CN); Bin Zhang (Beijing, CN)
Assignee: International Business Machines Corporation
G06Q30/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,918,397
App. No.
13/561,468
Granted
Dec 23, 2014
Kind
B2
Abstract

A computer implemented method for clustering customers includes receiving a source set of customer records, wherein each customer record represents one customer, and each customer record includes at least one data attribute, and each data attribute has an attribute value; pre-processing the source set of customer records to generate a pre-processed set of customer records; executing a clustering algorithm on the pre-processed set of customer records to group the pre-processed set of customer records into clusters of a pre-defined number. The pre-processing comprises: determining the type of a customer in the source set of customer records; using a type attribute value to indicate the type of the customer in its customer record; normalizing data attribute values and type attribute values; weighting to the data attribute values and the type attribute values respectively to obtain weighted attribute values of the data attribute and weighted attribute values of the type attribute.

Claims (13)

1. A computer-implemented method for clustering customers, comprising:

receiving a set of customer records, wherein each customer record in the set of customer records represents one customer, each customer record includes at least one data attribute, and each data attribute has a data attribute value;

pre-processing the set of customer records to generate a pre-processed set of customer records;

wherein the pre-processing of the set of customer records comprises:

determining the type of the customer represented by each record in the set of customer records;

using a type attribute to represent the type of the customer in the corresponding customer record, wherein the type attribute indicates whether the corresponding customer is a seed customer or a non-seed customer;

normalizing the data attribute values and the type attribute values; and

weighting the data attribute values and the type attribute values, wherein the weighting comprises multiplying the data attribute values by a dispersion weighting factor and multiplying the type attribute values by a purity weighting factor;

executing, by a computer processor, a clustering algorithm on the pre-processed set of customer records to cluster the pre-processed set of customer records into a pre-defined number of clusters, wherein each of the clusters comprises two or more customer records representing two or more customers, and wherein each pre-processed customer record in the pre-processed set of customer records being clustered comprises at least a normalized and weighted data attribute and a normalized and weighted type attribute;

wherein the dispersion weighting factor used to weight the data attribute values and the purity weighting factor used to weight the type attribute values are adjustable to affect the dispersion and purity of a clustering result of the clustering algorithm applied to the set of customer records.

2. The method of claim 1 , wherein the sum of dispersion weighting factor used to weight the data attribute values and the purity weighting used to weight the type attribute values is 1.

3. The method of claim 1 , wherein dividing the customer records into a current set of clusters comprises use of a K-means clustering algorithm.

4. The method of claim 1 , further comprising selecting clusters with higher purity from the pre-defined number of clusters, and outputting data attribute values of customer records in the clusters, wherein the purity of a cluster is the ratio of the number of customer records with specified type attribute in the cluster to the total number of customer records in the cluster.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: UTOPUS INSIGHTS, INC.
Reel/Frame 042700/0530 →
Priority Claims (1)
CN 2011 1 0080939 · Mar 31, 2011 · national
Continuity (2)
Continuation 13432361 · Mar 28, 2012
Related Publication 20120290580A1 · Nov 15, 2012