IP Library Granted Patent US 8,219,383
Granted Patent B2
US 8,219,383 · App. 13/030,295 · Granted Jul 10, 2012

Method and system for automated supervised data analysis

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,219,383
App. No.
13/030,295
Granted
Jul 10, 2012
Kind
B2
Abstract

The invention relates to a method for automatically analyzing data and constructing data classification models based on the data. In an embodiment of the method, the method includes selecting a best combination of methods from a plurality of classification, predictor selection, and data preparatory methods; and determining a best model that corresponds to one or more best parameters of the classification, predictor selection, and data preparatory methods for the data to be analyzed. The method also includes estimating the performance of the best model using new data that was not used in selecting the best combination of methods or in determining the best model; and returning a small set of predictors sufficient for the classification task.

Claims (16)

1. A tangible computer-readable medium having stored thereon a plurality of executable instructions that when executed by a processor perform a method for automatically analyzing data and constructing data classification models based on the data, the method comprising:

selecting a best combination of methods from a plurality of classification, predictor selection, and data preparatory methods;

determining a best model that corresponds to one or more best parameters of the classification, predictor selection, and data preparatory methods for the data to be analyzed;

estimating the performance of the best model using new data that was not used in selecting the best combination of methods or in determining the best model;

returning a small set of predictors sufficient for the classification task; and

performing optimal model selection and estimating performance of the data classification models using the N non-overlapping and balanced subsets and a nested cross-validation procedure.

2. The computer-readable medium of claim 1 wherein the selecting a best combination of methods comprises:

selecting the best combination of methods from a plurality of predetermined classification, predictor selection, and data preparatory methods.

3. The computer-readable medium of claim 1 wherein the method requires minimal or no knowledge of data analysis by a user of the method.

4. The computer-readable medium of claim 1 wherein the method requires minimal or no knowledge about the domain where the data comes from by a user of the method.

5. The computer-readable medium of claim 1 wherein the method is executed by a fully automated system.

6. The computer-readable medium of claim 1 wherein the method performs comparable to or better than human experts in a plurality of applications.

7. The computer-readable medium of claim 1 wherein the method uses a pre-selected set of classification, predictor selection, and data preparatory methods.

8. The computer-readable medium of claim 7 wherein using the pre-selected set of classification, predictor selection, and data preparatory method comprises:

using a limited pre-selected set of classification, predictor selection, and data preparatory methods.

9. The computer-readable medium of claim 8 wherein the pre-selected set of classification, predictor selection, and data preparatory methods are selected on the basis of extensive tests in application domains of interest.

Assignments (1)
CONFIRMATORY LICENSE Recorded Oct 30, 2019
From: VANDERBILT UNIVERSITY
To: NATIONAL INSTITUTES OF HEALTH - DIRECTOR DEITR
Reel/Frame 050864/0316 →
Continuity (3)
Division 11510847 · Aug 28, 2006
Provisional Application 60711402 · Aug 26, 2005
Related Publication 20110246403A1 · Oct 6, 2011