IP Library Granted Patent US 7,565,369
Granted Patent B2
US 7,565,369 · App. 10/857,030 · Granted Jul 21, 2009

System and method for mining time-changing data streams

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,565,369
App. No.
10/857,030
Granted
Jul 21, 2009
Kind
B2
Abstract

A general framework for mining concept-drifting data streams using weighted ensemble classifiers. An ensemble of classification models, such as C4.5, RIPPER, naive Bayesian, etc., is trained from sequential chunks of the data stream. The classifiers in the ensemble are judiciously weighted based on their expected classification accuracy on the test data under the time-evolving environment. Thus, the ensemble approach improves both the efficiency in learning the model and the accuracy in performing classification. An empirical study shows that the proposed methods have substantial advantage over single-classifier approaches in prediction accuracy, and the ensemble framework is effective for a variety of classification models.

Claims (22)

1. An apparatus for mining concept-drifting data streams, said apparatus comprising:

a processor;

an arrangement for accepting an ensemble of classifiers;

an arrangement for training the ensemble of classifiers from data in a data stream; and

an arrangement for weighting classifiers in the classifier ensembles, wherein said weighting arrangement is adapted to weight classifiers based on an expected prediction accuracy with regard to current data, and to apply a weight to each classifier such that the weight is inversely proportional to an expected prediction error of the classifier with regard to a current group of data;

wherein the weighted classifiers are stored in a computer memory.

2. The apparatus according to claim 1 , wherein said arrangement for training the ensemble of classifiers from data in a data stream is adapted to train the ensemble of classifiers from groups of sequential data in a data stream.

3. The apparatus according to claim 1 , further comprising an arrangement for pruning a classifier ensemble.

4. The apparatus according to claim 3 , wherein the pruning arrangement is adapted to apply an instance-based pruning technique to data streams with conceptual drifts.

5. A computer implemented method for mining concept-drifting data streams, said method comprising the steps of:

accepting an ensemble of classifiers;

training the ensemble of classifiers from data in a data stream; and

weighting classifiers in the classifier ensembles, wherein said weighting step comprises weighting classifiers based on an expected prediction accuracy with regard to current data, and applying a weight to each classifier such that the weight is inversely proportional to an expected prediction error of the classifier with regard to a current group of data;

wherein the weighted classifiers are stored in a computer memory.

6. The method according to claim 5 , wherein said step of training the ensemble of classifiers from data in a data stream comprises training the ensemble of classifiers from groups of sequential data in a data stream.

7. The method according to claim 5 , further comprising the step of pruning a classifier ensemble.

8. The method according to claim 7 , wherein said pruning step comprises applying an instance-based pruning technique to data streams with conceptual drifts.

9. A program storage device readable by machine, tangibly embodying a program of instructions executable by the machine to perform method steps for mining concept-drifting data streams, said method comprising the steps of:

accepting an ensemble of classifiers;

training the ensemble of classifiers from data in a data stream; and

weighting classifiers in the classifier ensembles, wherein said weighting step comprises weighting classifiers based on an expected prediction accuracy with regard to current data, and applying a weight to each classifier such that the weight is inversely proportional to an expected prediction error of the classifier with regard to a current group of data;

wherein the weighted classifiers are stored in memory readable by the machine.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE 1ST ASSIGNEE NAME 50% INTEREST PREVIOUSLY RECORDED AT REEL: 043418 FRAME: 0692. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 1, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SERVICENOW, INC.; INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044348/0451 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SERVICENOW, INC.
Reel/Frame 043418/0692 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2004
From: FAN, WEI; WANG, HAIXUN; YU, PHILIPS S.
To: IBM, CORPORATION
Reel/Frame 015020/0923 →