Generating a decision tree model during query execution via a relational database system
A database system is operable to execute a request to generate a decision tree model. A training set of rows are determined based on accessing a plurality of rows of a relational database table of a relational database. First query data is generated for execution based on the training set of rows. First query output is generated based on executing the first query data. A first portion of the decision tree model data is built based on the first query output. Additional query data is generated for execution based on the first query output. Additional query output is generated based on executing the additional query data. An additional portion of the decision tree model data is built based on the additional query output. Model output for the decision tree model is generated via processing input data in conjunction with processing the decision tree model data.
1 . A database system comprising:
a plurality of computing device clusters, wherein a computing device cluster of the plurality of computing device clusters includes a plurality of computing devices, wherein a computing device of the plurality of computing devices includes a plurality of computing nodes, and wherein a computing node of the plurality of computing nodes includes a plurality of processing core resources;
wherein a set of computing nodes of the pluralities of computing nodes is operable to:
obtain a plurality of queries that include a plurality of training query operations regarding training of a plurality of machine learning models, wherein a first a query of the plurality of queries includes a first training query operation of the plurality of training query operations regarding a first machine learning model of the plurality of machine learning models;
identify a plurality of training data based on the plurality of training query operations;
wherein pluralities of sets of processing core resources of the set of computing nodes is operable to:
receive the plurality of training query operations from the set of computing nodes;
in a distributed manner, receive the plurality of training data from the set of computing nodes;
execute, substantially in parallel, a first training query operation of the plurality of training query operations on at least a portion of a first machine learning model of the plurality of machine learning models based on respective sub-sets of sets of first training data of the plurality of training data to produce a plurality of first partial training results;
execute, substantially in parallel, a second training query operation of the plurality of training query operations on at least a portion of a second machine learning model of the plurality of machine learning models based on respective sub-sets of sets of second training data of the plurality of training data to produce a plurality of second partial training results; and
wherein the set of computing nodes is further operable to:
compile the plurality of first partial training results to produce a first training result;
when the first training result compares favorably to a first training threshold, update the first machine learning model based on the first training result;
compile the plurality of second partial training results to produce a second training result; and
when the second training result compares favorably to a second training threshold, update the second machine learning model based on the second training result.
2 . The database system of claim 1 , wherein a machine learning model of the plurality of machine learning models comprises:
an equation, wherein the equation defines a dependent variable in terms of a number of independent variables and a number of coefficients, and wherein the training of the machine learning model is to determine values for the coefficients that provide an acceptable level of modeling error.
3 . The database system of claim 1 further comprises:
the set of computing nodes providing the data of the plurality of training data to the pluralities of sets of processing core resources in accordance with a random shuffle function; and
the set of computing nodes providing a corresponding set of the plurality of training data to a corresponding set of the pluralities of sets of processing core resources in the distributed manner via a multiplex function.
4 . The database system of claim 3 further comprises:
the training data is organized as a plurality of rows of columnar data;
the set of processing core resources is operable to:
replicate rows of columnar data of the plurality of rows data based on an overwrite factor of the random shuffle function to produce a plurality of replicated rows of columnar data; and
provide, as the data of the plurality of training data, the plurality of rows of data and the plurality of replicated rows of columnar data to the pluralities of sets of processing core resources.
5 . The database system of claim 1 further comprises:
the training query operation including a particle swarm optimization operation that includes a direction value and a gravity value, wherein the direction value causes a particle of the particle swarm optimization operation to move in arbitrary direction and wherein the gravity value causes the particle of the particle swarm operation to be pulled towards a best known position.
6 . The database system of claim 5 further comprises:
the training query operation including a linear search algorithm that is executed after a number of iterations of the particle swarm optimization operation, wherein the linear search algorithm uses current position of particles of the particle swarm optimization operation to further improve position of the particles when possible.
7 . The database system of claim 6 further comprises:
the training query operation including a golden section search to is executed, in a serial manner, on coefficients of an equation of the machine learning model, to further improve position of the particles when possible and return to the particle swarm optimization operation when the position of the particles are not improved.
8 . A computer-readable memory comprises:
a first memory sections that stores operational instructions that, when executed by a set of computing nodes of pluralities of computing nodes of a database system, causes the set of computing nodes to:
obtain a plurality of queries that include a plurality of training query operations regarding training of a plurality of machine learning models, wherein a first a query of the plurality of queries includes a first training query operation of the plurality of training query operations regarding a first machine learning model of the plurality of machine learning models;
identify a plurality of training data based on the plurality of training query operations;
a second memory sections that stores operational instructions that, when executed by pluralities of sets of processing core resources of the set of computing nodes, causes the pluralities of sets of processing core resources to:
receive the plurality of training query operations from the set of computing nodes;
in a distributed manner, receive the plurality of training data from the set of computing nodes;
execute, in substantially in parallel, a first training query operation of the plurality of training query operations on at least a portion of a first machine learning model of the plurality of machine learning models based on respective sub-sets of sets of first training data of the plurality of training data to produce a plurality of first partial training results;
execute, in substantially in parallel, a second training query operation of the plurality of training query operations on at least a portion of a second machine learning model of the plurality of machine learning models based on respective sub-sets of sets of second training data of the plurality of training data to produce a plurality of second partial training results; and
wherein the first memory section further stores operational instructions that, when executed by the set of computing nodes, causes the set of computing nodes to:
compile the plurality of first partial training results to produce a first training result;
when the first training result compares favorably to a desired result, update the first machine learning model based on the first training result;
compile the plurality of second partial training results to produce a second training result; and
when the second training result compares favorably to a desired result, update the second machine learning model based on the second training result.
9 . The computer-readable memory of claim 8 , wherein a machine learning model of the plurality of machine learning models comprises:
an equation, wherein the equation defines a dependent variable in terms of a number of independent variables and a number of coefficients, and wherein the training of the machine learning model is to determine values for the coefficients that provide an acceptable level of modeling error.
10 . The computer-readable memory of claim 8 further comprises:
wherein the first memory section further stores operational instructions that, when executed by the set of computing nodes, causes the set of computing nodes to:
provide the data of the plurality of training data to the pluralities of sets of processing core resources in accordance with a random shuffle function; and
provide a corresponding set of the plurality of training data to a corresponding set of the pluralities of sets of processing core resources in the distributed manner via a multiplex function.
11 . The computer-readable memory of claim 8 further comprises:
the plurality of training data is organized as a plurality of rows of columnar data;
wherein the first memory section further stores operational instructions that, when executed by the set of computing nodes, causes the set of computing nodes to:
replicate rows of columnar data of the plurality of rows data based on an overwrite factor of the random shuffle function to produce a plurality of replicated rows of columnar data; and
provide, as the data of the plurality of training data, the plurality of rows of data and the plurality of replicated rows of columnar data to the pluralities of sets of processing core resources.
12 . The computer-readable memory of claim 8 , wherein the training query operation including a particle swarm optimization operation that includes a direction value and a gravity value, wherein the direction value causes a particle of the particle swarm optimization operation to move in arbitrary direction and wherein the gravity value causes the particle of the particle swarm operation to be pulled towards a best known position.
13 . The computer-readable memory of claim 12 , wherein the training query operation including a linear search algorithm that is executed after a number of iterations of the particle swarm optimization operation, wherein the linear search algorithm uses current position of particles of the particle swarm optimization operation to further improve position of the particles when possible.
14 . The computer-readable memory of claim 13 , wherein the training query operation including a golden section search to is executed, in a serial manner, on coefficients of an equation of the machine learning model, to further improve position of the particles when possible and return to the particle swarm optimization operation when the position of the particles are not improved.