IP Library Patent Application 14935742
Patent Application
App. No. 14/935,742

DATA CLEAN-UP METHOD FOR IMPROVING PREDICTIVE MODEL TRAINING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/935,742
Abstract

A method that improves the training of predictive models. Better trained predictive models make better predictions, and can classify transactions with reduced levels of false positives and false negative. Included is an apparatus for executing a data clean-up algorithm that harmonizes a wide range of real world supervised and unsupervised training data into a single, error-free, uniformly formatted record file that has every field coherent and well populated with information.

Claims (27)

1 . A method that improves the training of predictive models, comprising:

converting and transforming a variety of inconsistent and incoherent supervised and unsupervised training data for predictive models received by a network server as electronic data files, and storing that in a computer data storage mechanism, and then into another single, error-free, uniformly formatted record file stored in the computer data storage mechanism with an apparatus for executing a data integrity analysis algorithm that harmonizes a range of supervised and unsupervised training data into flat-data records in which every field of every record file is modified to be coherent and well-populated with information;

comparing and correcting any data values in each data field in the inconsistent and incoherent supervised and unsupervised training data according to a user-service consumer preference and a predefined data dictionary of valid data values with an apparatus for executing an algorithm that substitutes data values in the data fields of incoming supervised and unsupervised training data with at least one value representing a minimum, a maximum, a null, an average, and a default;

discerning the context of any text included in the inconsistent and incoherent supervised and unsupervised training data with an apparatus for executing a contextual dictionary algorithm that employs a thesaurus of alternative contexts of ambiguous words for find a common context denominator, and to then record the context determined into the computer data storage mechanism for later access by a predictive model;

cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing an algorithm for cleaning up raw data in stored data records, field-by-field, record-by-record in which some types of fields are restricted in what is legal or allowed, and includes fetching raw data from the computer data storage mechanism and testing each field if a data value reported is numeric or symbolic, and if numeric, a data dictionary is used to see if such data value is previously listed as valid, and if symbolic, using another data dictionary to see if such data value is listed there as valid;

cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing a Smith-Waterman algorithm for a local-sequence alignment and to determine if there are any similar regions between two strings or sequences, and in which a consistent, coherent terminology is then enforceable in each data field without data loss, and in which the Smith-Waterman algorithm compares segments of all possible lengths and optimizes a similarity measure without looking at any total sequence;

cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for replacing a numeric value, wherein a numeric value to use as a replacement depends on any flags or preferences that were set to use a default, the average, a minimum, a maximum, or a null;

sampling cleaned, raw-data from the flat-data records in the computer data storage mechanism with an apparatus for executing an algorithm that tests if data are supervised, and if so, that creates a plurality of individual data sets for each class with a stratified selection as needed, and then testing if a selected class is abnormal or uncharacteristic, and if not, down-sampling and producing sampled records of the classes and splitting any remaining data into separate training sets, separate test sets, and separate blind sets all then stored in the computer data storage mechanism for later use in subsequent steps to train a predictive model;

if the test for each record of each class in supervised data is abnormal or uncharacteristic, then skipping a down-sampling for that instance; and

if in a previous step the cleaned, raw-data from the flat-data records in the computer data storage mechanism was determined by the apparatus for executing an algorithm that tests if data are supervised are, in fact, unsupervised, then down-sampling all records and splitting a remaining a sampled record data into a separate a training set, a separate test set, and a separate blind set for later use in subsequent steps to train a predictive model.

2 . A method that improves the training of predictive models, comprising:

converting and transforming a variety of inconsistent and incoherent supervised and unsupervised training data for predictive models received by a network server as electronic data files, and storing that in a computer data storage mechanism, and then into another single, error-free, uniformly formatted record file stored in the computer data storage mechanism with an apparatus for executing a data integrity analysis algorithm that harmonizes a range of supervised and unsupervised training data into flat-data records in which every field of every record file is modified to be coherent and well-populated with information.

3 . The method of claim 2 , further comprising:

comparing and correcting any data values in each data field in the inconsistent and incoherent supervised and unsupervised training data according to a user-service consumer preference and a predefined data dictionary of valid data values with an apparatus for executing an algorithm that substitutes data values in the data fields of incoming supervised and unsupervised training data with at least one value representing a minimum, a maximum, a null, an average, and a default.

4 . The method of claim 2 , further comprising:

discerning the context of any text included in the inconsistent and incoherent supervised and unsupervised training data with an apparatus for executing a contextual dictionary algorithm that employs a thesaurus of alternative contexts of ambiguous words for find a common context denominator, and to then record the context determined into the computer data storage mechanism for later access by a predictive model.

5 . The method of claim 2 , further comprising:

cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing an algorithm for cleaning up raw data in stored data records, field-by-field, record-by-record in which some types of fields are restricted in what is legal or allowed, and includes fetching raw data from the computer data storage mechanism and testing each field if a data value reported is numeric or symbolic, and if numeric, a data dictionary is used to see if such data value is previously listed as valid, and if symbolic, using another data dictionary to see if such data value is listed there as valid.

6 . The method of claim 2 , further comprising:

cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for executing a Smith-Waterman algorithm for a local-sequence alignment and to determine if there are any similar regions between two strings or sequences, and in which a consistent, coherent terminology is then enforceable in each data field without data loss, and in which the Smith-Waterman algorithm compares segments of all possible lengths and optimizes a similarity measure without looking at any total sequence.

7 . The method of claim 2 , further comprising:

cleaning up inconsistent, missing, and illegal data in each data field by removal or reconstitution with an apparatus for replacing a numeric value, wherein a numeric value to use as a replacement depends on any flags or preferences that were set to use a default, the average, a minimum, a maximum, or a null.

8 . The method of claim 2 , further comprising:

sampling cleaned, raw-data from the flat-data records in the computer data storage mechanism with an apparatus for executing an algorithm that tests if data are supervised, and if so, that creates a plurality of individual data sets for each class with a stratified selection as needed, and then testing if a selected class is abnormal or uncharacteristic, and if not, down-sampling and producing sampled records of the classes and splitting any remaining data into separate training sets, separate test sets, and separate blind sets all then stored in the computer data storage mechanism for later use in subsequent steps to train a predictive model; and

if the test for each record of each class in supervised data is abnormal or uncharacteristic, then skipping a down-sampling for that instance.

9 . The method of claim 8 , further comprising:

if in a previous step the cleaned, raw-data from the flat-data records in the computer data storage mechanism was determined by the apparatus for executing an algorithm that tests if data are supervised are, in fact, unsupervised, then down-sampling all records and splitting a remaining a sampled record data into a separate a training set, a separate test set, and a separate blind set for later use in subsequent steps to train a predictive model.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2018
From: ADJAOUTE, AKLI
To: BRIGHTERION, INC.
Reel/Frame 045686/0918 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2017
From: BRIGHTERION, INC
To: ADJAOUTE, AKLI
Reel/Frame 042048/0621 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2017
From: ADJAOUTE, AKLI
To: BRIGHTERION INC
Reel/Frame 041545/0279 →