IP Library Patent Application 16359608
Patent Application
App. No. 16/359,608

Data Mining Technique With Maintenance of Ancestry Counts

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/359,608
Abstract

Roughly described, a computer-implemented evolutionary data mining system includes a memory storing a candidate gene database in which each candidate individual has a respective fitness estimate; a gene pool processor which tests individuals from the candidate gene pool on training data and updates the fitness estimate associated with the individuals in dependence upon the tests; and a gene harvesting module for deploying selected individuals from the gene pool, wherein the gene pool processor includes a competition module which selects individuals for discarding in dependence upon their updated fitness estimate. The system maintains the ancestry count for each of the candidate individuals, and may use this information to adjust the competition among the individuals, to adjust the selection of individuals for further procreation, and/or for other purposes.

Claims (34)

1 . A computer-implemented data mining method, for use with a data mining training database containing training data, comprising the steps of:

providing a system having a candidate gene database identifying a candidate pool of candidate individuals, each candidate individual identifying a plurality of conditions and at least one corresponding proposed output in dependence upon the conditions, each candidate individual further having associated therewith an indication of a respective fitness estimate and an ancestry count, wherein ancestry count corresponds to a number of procreation events that occurred in a procreation history of each candidate individual;

testing on the training data by a gene testing module of the system, each individual in a testing subset of at least one of the candidate individuals, each individual in the testing subset undergoing at least one trial, each trial applying the conditions of the respective individual to the training data to propose a result;

calculating by the gene testing module, an estimated fitness of each of the individuals in the testing subset in dependence upon the training data and the results proposed by the individual in the step of testing; and

populating an elitist pool with elite individuals by a competition module in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals; and

further wherein the competition module adjusts the fitness estimate of each individual in dependence upon the individual's ancestry count and selects an individual with a higher fitness estimate over an individual with a lower fitness estimate for inclusion in the elitist pool.

2 . The computer-implemented data mining method of claim 1 , wherein the competition module applies a handicap of a fixed percentage against individuals whose ancestry count exceeds a predetermined number of generations.

3 . The computer-implemented data mining method of claim 1 , wherein in populating the elitist pool with elite individuals, the competition module further (iii) compares an experience level of each individual to a predetermined threshold experience level, wherein experience level is determined by a total number of trials the individual has undergone.

4 . The computer-implemented data mining method of claim 2 , wherein the handicap applied to each given one of the individuals by the competition module varies non-decreasingly as a function of the ancestry count of the given individual.

5 . A computer-implemented data mining method for use with a data mining training database containing a plurality of data samples, comprising:

providing a system having a candidate gene database identifying a candidate pool of candidate individuals, each candidate individual identifying a plurality of conditions and at least one corresponding proposed output in dependence upon the conditions, each candidate individual further having associated therewith an indication of a respective fitness estimate and an ancestry count, wherein ancestry count corresponds to a number of procreation events that occurred in a procreation history of each candidate individual;

performing a procreation step by a procreation module of a gene pool processor of forming new candidate individuals in the candidate pool of candidate individuals, wherein each new individual is related to one or more previously existing parent candidate individuals;

testing by a gene testing module of the gene pool processor each individual in a testing subset of at least one of the candidate individuals, wherein the testing subset includes at least one new candidate individual, each of the tests applying the conditions of the respective individual to a respective subset of the data samples in a training database to propose a result, each individual in the testing subset being tested on at least one data sample and at least one of the individuals in the testing subset being tested on more than one data sample;

calculating by a competition module of the gene pool processor an overall fitness estimate for each of the individuals in the testing subset, in dependence upon the results proposed by the respective individual when the conditions of the respective individual were applied to the respective subset of the data samples; and

selecting one or more individuals for discarding from the candidate pool by the competition module in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals;

further wherein the competition module adjusts the fitness estimate of each individual in dependence upon the individual's ancestry count and discards an individual with a lower fitness estimate over an individual with a higher fitness estimate; and

harvesting by a gene harvesting module selected ones of non-discarded individuals from the candidate pool of candidate individuals for deployment of selected individuals in a predetermined production environment, wherein candidate individuals are selected in dependence upon comparisons among their respective ancestry counts.

6 . The computer-implemented data mining method of claim 5 , the procreation module selects the one or more previously existing parent candidate individuals for forming a new candidate individual in accordance with an ancestry count for each of the one or more previously existing parent candidate individuals.

7 . A computer-implemented data mining method for use with a client-server architecture, comprising:

providing a server having a candidate gene database identifying a candidate pool of candidate individuals, each candidate individual identifying a plurality of conditions and at least one corresponding proposed output in dependence upon the conditions, each candidate individual further having associated therewith an indication of a respective fitness estimate and an ancestry count, wherein ancestry count corresponds to a number of procreation events that occurred in a procreation history of each candidate individual;

delegating by the server testing subsets containing at least one of the candidate individuals to individual clients, wherein the testing by the individual clients includes:

testing each individual in a delegated testing subset on training data by a gene testing module of the client, each individual in the delegated testing subset undergoing at least one trial, each trial applying the conditions of the respective individual to the training data to propose a result;

calculating by the gene testing module, an estimated fitness of each of the individuals in the delegated testing subset in dependence upon the training data and the results proposed by the individual in the step of testing; and

providing tested individuals, including evaluation data, to the server, wherein the evaluation data includes the calculated estimated fitness and ancestry count;

adjusting the calculated estimated fitness for received tested individuals by a competition module on the server in dependence upon the individual's ancestry count; and

updating the candidate pool by the competition module to include individuals in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals, wherein the competition module selects an individual with a higher fitness estimate over an individual with a lower fitness estimate for inclusion in the candidate pool.

8 . The computer-implemented data mining method of claim 7 , further comprising:

performing a procreation step by a procreation module of the gene pool processor of each client of forming new candidate individuals in the candidate pool of candidate individuals, wherein each new individual is related to one or more previously existing parent candidate individuals;

testing by a gene testing module of the gene pool processor of each client each individual in a testing subset of at least one of the candidate individuals, wherein the testing subset includes at least one new candidate individual, each of the tests applying the conditions of the respective individual to a respective subset of the data samples in a training database to propose a result, each individual in the testing subset being tested on at least one data sample and at least one of the individuals in the testing subset being tested on more than one data sample;

providing tested procreated individuals, including evaluation data, to the server, wherein the evaluation data includes the calculated estimated fitness and ancestry count;

calculating by a competition module on the server an overall fitness estimate for each of the procreated individuals in the testing subset, in dependence upon the results proposed by the respective procreated individual when the conditions of the respective procreated individual were applied to the respective subset of the data samples; and

selecting one or more procreated individuals for discarding from the candidate pool by the competition module in accordance with (i) a predetermined requirement for fitness estimate and (ii) a maximum number of allowed candidate individuals;

further wherein the competition module adjusts the fitness estimate of each procreated individual in dependence upon the individual's ancestry count and discards an individual with a lower fitness estimate over an individual with a higher fitness estimate; and

harvesting by a gene harvesting module selected ones of non-discarded procreated individuals from the candidate pool of candidate individuals for deployment of selected procreated individuals in a predetermined production environment, wherein candidate procreated individuals are selected in dependence upon comparisons among their respective ancestry counts.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2019
From: SENTIENT TECHNOLOGIES (BARBADOS) LIMITED; SENTIENT TECHNOLOGIES HOLDINGS LIMITED; SENTIENT TECHNOLOGIES (USA) LLC
To: COGNIZANT TECHNOLOGY SOLUTIONS U.S. CORPORATION
Reel/Frame 050151/0408 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2019
From: FINK, DANIEL E.; SHAHRZAD, HORMOZ
To: SENTIENT TECHNOLOGIES (BARBADOS) LIMITED
Reel/Frame 048652/0336 →