IP Library Granted Patent US 7,783,648
Granted Patent B2
US 7,783,648 · App. 11/772,343 · Granted Aug 24, 2010

Methods and systems for partitioning datasets

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,783,648
App. No.
11/772,343
Granted
Aug 24, 2010
Kind
B2
Abstract

A partitioning system that provides a fast, simple and flexible method for partitioning a dataset. The process, executed within a computer system, retrieves product and sales data from a data store. Data items are selected and sorted by a data attribute of interest to a user and a distribution curve is determined for the selected data and data attribute. The total length of the distribution curve is calculated, and then the curve is divided into k equal pieces, where k is the number of the partitions. The selected data is thereafter partitioned into k groups corresponding to the curve divisions.

Claims (28)

1. A computer-implemented method for partitioning a dataset, the method comprising the steps of:

maintaining the dataset in an electronic database, the dataset including at least ten data items, the data items containing a first attribute, the first attribute having a range of attribute values;

sorting the dataset in an ascending or descending sequence by the range of attribute values present in the first attribute;

generating a distribution curve for the dataset by plotting the data items in the dataset in ascending or descending sort sequence;

calculating a total length of said distribution curve, the length of the distribution curve being estimated through use of the equation:

Length= f (range,count)≈√{square root over ((range) 2 +(count) 2 )}{square root over ((range) 2 +(count) 2 )},

where count represents the number of data items within the dataset, and range represents the range of attribute values;

dividing said distribution curve, using the total length of the distribution curve, into a plurality of equal length curve segments; and

grouping said data items into a plurality of partitions, each partition comprising data items corresponding to one of said curve segments.

2. The computer-implemented method for partitioning a dataset in accordance with claim 1 , wherein:

said step of dividing said distribution curve into a plurality of equal curve segments comprises the step of dividing said distribution curve into k equal pieces, where k equals a number of desired partitions; and

said step of grouping said data items into a plurality of partitions comprises grouping said data items into k partitions, each one of said k partitions corresponding to one of said k curve segments.

3. The computer-implemented method for partitioning a dataset in accordance with claim 2 , further comprising the step of:

normalizing the values of said attribute prior to determining said distribution curve.

4. A system for partitioning a dataset, comprising:

an electronic database containing the dataset, the dataset including at least ten data items, the data items containing a first attribute, the first attribute having a range of attribute values;

sorting the dataset in an ascending or descending sequence by the range of attribute values present in the first attribute;

generating a distribution curve for the dataset by plotting the data items in the dataset in ascending or descending sort sequence;

calculating a total length of said distribution curve, the length of the distribution curve being estimated through use of the equation:

Length= f (range,count)≈√{square root over ((range) 2 +(count) 2 )}{square root over ((range) 2 +(count) 2 )},

where count represents the number of data items within the dataset, and range represents the range of attribute values;

dividing said distribution curve, using the total length of the distribution curve, into a plurality of equal length curve segments; and

grouping said data items into a plurality of partitions, each partition comprising data items corresponding to one of said curve segments.

5. The system in accordance with claim 4 , wherein:

dividing said distribution curve into a plurality of equal curve segments comprises the dividing said distribution curve into k equal pieces, where k equals a number of desired partitions; and

grouping said data items into a plurality of partitions comprises grouping said data items into k partitions, each one of said k partitions corresponding to one of said k curve segments.

6. The system in accordance with claim 5 , wherein:

the values of said attribute are normalized prior to determining said distribution curve.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2008
From: NCR CORPORATION
To: TERADATA US, INC.
Reel/Frame 020666/0438 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2007
From: BATENI, ARASH; KIM, EDWARD; BALENDRAN, PRATHAYANA; CHAN, ANDREW
To: NCR CORPORATION
Reel/Frame 019874/0611 →