IP Library › Granted Patent US 12,748,761
Granted Patent B2
US 12,748,761 · App. 18/457,568 · Granted Sep 29, 2026

Generating a decision tree model during query execution via a relational database system

Inventor: Jason Arnold (Chicago, IL)
Assignee: Ocient Holdings LLC
G06F16/24566G06F16/242
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,761
App. No.
18/457,568
Granted
Sep 29, 2026
Kind
B2
Abstract

A database system is operable to execute a request to generate a decision tree model. A training set of rows are determined based on accessing a plurality of rows of a relational database table of a relational database. First query data is generated for execution based on the training set of rows. First query output is generated based on executing the first query data. A first portion of the decision tree model data is built based on the first query output. Additional query data is generated for execution based on the first query output. Additional query output is generated based on executing the additional query data. An additional portion of the decision tree model data is built based on the additional query output. Model output for the decision tree model is generated via processing input data in conjunction with processing the decision tree model data.

Claims (59)

1 . A database system comprising:

a plurality of computing device clusters, wherein a computing device cluster of the plurality of computing device clusters includes a plurality of computing devices, wherein a computing device of the plurality of computing devices includes a plurality of computing nodes, and wherein a computing node of the plurality of computing nodes includes a plurality of processing core resources;

wherein a set of computing nodes of the pluralities of computing nodes is operable to:

obtain a plurality of queries that include a plurality of training query operations regarding training of a plurality of machine learning models, wherein a first a query of the plurality of queries includes a first training query operation of the plurality of training query operations regarding a first machine learning model of the plurality of machine learning models;

identify a plurality of training data based on the plurality of training query operations;

wherein pluralities of sets of processing core resources of the set of computing nodes is operable to:

receive the plurality of training query operations from the set of computing nodes;

in a distributed manner, receive the plurality of training data from the set of computing nodes;

execute, substantially in parallel, a first training query operation of the plurality of training query operations on at least a portion of a first machine learning model of the plurality of machine learning models based on respective sub-sets of sets of first training data of the plurality of training data to produce a plurality of first partial training results;

execute, substantially in parallel, a second training query operation of the plurality of training query operations on at least a portion of a second machine learning model of the plurality of machine learning models based on respective sub-sets of sets of second training data of the plurality of training data to produce a plurality of second partial training results; and

wherein the set of computing nodes is further operable to:

compile the plurality of first partial training results to produce a first training result;

when the first training result compares favorably to a first training threshold, update the first machine learning model based on the first training result;

compile the plurality of second partial training results to produce a second training result; and

when the second training result compares favorably to a second training threshold, update the second machine learning model based on the second training result.

2 . The database system of claim 1 , wherein a machine learning model of the plurality of machine learning models comprises:

an equation, wherein the equation defines a dependent variable in terms of a number of independent variables and a number of coefficients, and wherein the training of the machine learning model is to determine values for the coefficients that provide an acceptable level of modeling error.

3 . The database system of claim 1 further comprises:

the set of computing nodes providing the data of the plurality of training data to the pluralities of sets of processing core resources in accordance with a random shuffle function; and

the set of computing nodes providing a corresponding set of the plurality of training data to a corresponding set of the pluralities of sets of processing core resources in the distributed manner via a multiplex function.

4 . The database system of claim 3 further comprises:

the training data is organized as a plurality of rows of columnar data;

the set of processing core resources is operable to:

replicate rows of columnar data of the plurality of rows data based on an overwrite factor of the random shuffle function to produce a plurality of replicated rows of columnar data; and

provide, as the data of the plurality of training data, the plurality of rows of data and the plurality of replicated rows of columnar data to the pluralities of sets of processing core resources.

5 . The database system of claim 1 further comprises:

the training query operation including a particle swarm optimization operation that includes a direction value and a gravity value, wherein the direction value causes a particle of the particle swarm optimization operation to move in arbitrary direction and wherein the gravity value causes the particle of the particle swarm operation to be pulled towards a best known position.

6 . The database system of claim 5 further comprises:

the training query operation including a linear search algorithm that is executed after a number of iterations of the particle swarm optimization operation, wherein the linear search algorithm uses current position of particles of the particle swarm optimization operation to further improve position of the particles when possible.

7 . The database system of claim 6 further comprises:

the training query operation including a golden section search to is executed, in a serial manner, on coefficients of an equation of the machine learning model, to further improve position of the particles when possible and return to the particle swarm optimization operation when the position of the particles are not improved.

8 . A computer-readable memory comprises:

a first memory sections that stores operational instructions that, when executed by a set of computing nodes of pluralities of computing nodes of a database system, causes the set of computing nodes to:

obtain a plurality of queries that include a plurality of training query operations regarding training of a plurality of machine learning models, wherein a first a query of the plurality of queries includes a first training query operation of the plurality of training query operations regarding a first machine learning model of the plurality of machine learning models;

identify a plurality of training data based on the plurality of training query operations;

a second memory sections that stores operational instructions that, when executed by pluralities of sets of processing core resources of the set of computing nodes, causes the pluralities of sets of processing core resources to:

receive the plurality of training query operations from the set of computing nodes;

in a distributed manner, receive the plurality of training data from the set of computing nodes;

execute, in substantially in parallel, a first training query operation of the plurality of training query operations on at least a portion of a first machine learning model of the plurality of machine learning models based on respective sub-sets of sets of first training data of the plurality of training data to produce a plurality of first partial training results;

execute, in substantially in parallel, a second training query operation of the plurality of training query operations on at least a portion of a second machine learning model of the plurality of machine learning models based on respective sub-sets of sets of second training data of the plurality of training data to produce a plurality of second partial training results; and

wherein the first memory section further stores operational instructions that, when executed by the set of computing nodes, causes the set of computing nodes to:

compile the plurality of first partial training results to produce a first training result;

when the first training result compares favorably to a desired result, update the first machine learning model based on the first training result;

compile the plurality of second partial training results to produce a second training result; and

when the second training result compares favorably to a desired result, update the second machine learning model based on the second training result.

9 . The computer-readable memory of claim 8 , wherein a machine learning model of the plurality of machine learning models comprises:

an equation, wherein the equation defines a dependent variable in terms of a number of independent variables and a number of coefficients, and wherein the training of the machine learning model is to determine values for the coefficients that provide an acceptable level of modeling error.

10 . The computer-readable memory of claim 8 further comprises:

wherein the first memory section further stores operational instructions that, when executed by the set of computing nodes, causes the set of computing nodes to:

provide the data of the plurality of training data to the pluralities of sets of processing core resources in accordance with a random shuffle function; and

provide a corresponding set of the plurality of training data to a corresponding set of the pluralities of sets of processing core resources in the distributed manner via a multiplex function.

11 . The computer-readable memory of claim 8 further comprises:

the plurality of training data is organized as a plurality of rows of columnar data;

wherein the first memory section further stores operational instructions that, when executed by the set of computing nodes, causes the set of computing nodes to:

replicate rows of columnar data of the plurality of rows data based on an overwrite factor of the random shuffle function to produce a plurality of replicated rows of columnar data; and

provide, as the data of the plurality of training data, the plurality of rows of data and the plurality of replicated rows of columnar data to the pluralities of sets of processing core resources.

12 . The computer-readable memory of claim 8 , wherein the training query operation including a particle swarm optimization operation that includes a direction value and a gravity value, wherein the direction value causes a particle of the particle swarm optimization operation to move in arbitrary direction and wherein the gravity value causes the particle of the particle swarm operation to be pulled towards a best known position.

13 . The computer-readable memory of claim 12 , wherein the training query operation including a linear search algorithm that is executed after a number of iterations of the particle swarm optimization operation, wherein the linear search algorithm uses current position of particles of the particle swarm optimization operation to further improve position of the particles when possible.

14 . The computer-readable memory of claim 13 , wherein the training query operation including a golden section search to is executed, in a serial manner, on coefficients of an equation of the machine learning model, to further improve position of the particles when possible and return to the particle swarm optimization operation when the position of the particles are not improved.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2023
From: ARNOLD, JASON
To: OCIENT HOLDINGS LLC
Reel/Frame 064746/0884 →
Continuity (5)
Continuation In Part 18449282 · Aug 14, 2023
Continuation 16921226 · Jul 6, 2020
Provisional Application 63374819 · Sep 7, 2022
Provisional Application 63374821 · Sep 7, 2022
Related Publication 20230401217A1 · Dec 14, 2023
References Cited (57)
US 5548770A · Bridges · 1996 [cited by applicant]
US 6230200B1 · Forecast · 2001 [cited by applicant]
US 6317738B1 · Lohman · 2001 [cited by applicant]
US 6633772B2 · Ford · 2003 [cited by applicant]
US 7499907B2 · Brown · 2009 [cited by applicant]
US 7908242B1 · Achanta · 2011 [cited by applicant]
US 20010051949A1 · Carey · 2001 [cited by applicant]
US 20020032676A1 · Reiner · 2002 [cited by applicant]
US 20040162853A1 · Brodersen · 2004 [cited by applicant]
US 20080133456A1 · Richards · 2008 [cited by applicant]
US 20090063893A1 · Bagepalli · 2009 [cited by applicant]
US 20090183167A1 · Kupferschmidt · 2009 [cited by applicant]
US 20100082577A1 · Mirchandani · 2010 [cited by applicant]
US 20100241646A1 · Friedman · 2010 [cited by applicant]
US 20100274983A1 · Murphy · 2010 [cited by applicant]
US 20100312756A1 · Zhang · 2010 [cited by applicant]
US 20110219169A1 · Zhang · 2011 [cited by applicant]
US 20120078951A1 · Hsu et al. · 2012 [cited by applicant]
US 20120109888A1 · Zhang · 2012 [cited by applicant]
US 20120151118A1 · Flynn · 2012 [cited by applicant]
US 20120185866A1 · Couvee · 2012 [cited by applicant]
US 20120226710A1 · Bowers et al. · 2012 [cited by applicant]
US 20120254252A1 · Jin · 2012 [cited by applicant]
US 20120311246A1 · Mcwilliams · 2012 [cited by applicant]
US 20130332484A1 · Gajic · 2013 [cited by applicant]
US 20140047095A1 · Breternitz · 2014 [cited by applicant]
US 20140074771A1 · He et al. · 2014 [cited by applicant]
US 20140136510A1 · Parkkinen · 2014 [cited by applicant]
US 20140188841A1 · Sun · 2014 [cited by applicant]
US 20150205607A1 · Lindholm · 2015 [cited by applicant]
US 20150244804A1 · Warfield · 2015 [cited by applicant]
US 20150248366A1 · Bergsten · 2015 [cited by applicant]
US 20150293966A1 · Cai · 2015 [cited by applicant]
US 20150310045A1 · Konik · 2015 [cited by applicant]
US 20160034547A1 · Lerios · 2016 [cited by applicant]
US 20170116276A1 · Ziauddin · 2017 [cited by applicant]
US 20180357565A1 · Syed · 2018 [cited by examiner]
US 20190138642A1 · Pal et al. · 2019 [cited by applicant]
US 20190332600A1 · Gillespie · 2019 [cited by applicant]
US 20200193332A1 · Zhang · 2020 [cited by examiner]
KR 1020180077830A · 2018 [cited by applicant]
A new high performance fabric for HPC, Michael Feldman, May 2016, Intersect360 Research. [cited by applicant]
Alechina, N. (2006-2007). B-Trees. School of Computer Science, University of Nottingham, http://www.cs.nott.ac.uk/~psznza/G5BADS06/lecture13-print.pdf. 41 pages. [cited by applicant]
Amazon DynamoDB: ten things you really should know, Nov. 13, 2015, Chandan Patra, http://cloudacademy. .com/blog/amazon-dynamodb-ten-thing. [cited by applicant]
An Inside Look at Google BigQuery, by Kazunori Sato, Solutions Architect, Cloud Solutions team, Google Inc., 2012. [cited by applicant]
Big Table, a NoSQL massively parallel table, Paul Krzyzanowski, Nov. 2011, https://www.cs.rutgers.edu/pxk/417/notes/contentlbigtable.html. [cited by applicant]
Distributed Systems, Fall2012, Mohsen Taheriyan, http://www-scf.usc.edu/-csci57212011Spring/presentations/Taheriyan.pptx. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2017/054773; Feb. 13, 2018; 17 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2017/054784; Dec. 28, 2017; 10 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2017/066145; Mar. 5, 2018; 13 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2017/066169; Mar. 6, 2018; 15 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2018/025729; Jun. 27, 2018; 9 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2018/034859; Oct. 30, 2018; 8 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2021/035619; Sep. 17, 2021; 11 pgs. [cited by applicant]
MapReduce: Simplified Data Processing on Large Clusters, OSDI 2004, Jeffrey Dean and Sanjay Ghemawat, Google, Inc., 13 pgs. [cited by applicant]
Rodero-Merino, L.; Storage of Structured Data: Big Table and HBase, New Trends in Distributed Systems, MSc Software and Systems, Distributed Systems Laboratory; Oct. 17, 2012; 24 pages. [cited by applicant]
Step 2: Examine the data model and implementation details, 2016, Amazon Web Services, Inc., http://docs.aws.amazon.com/amazondynamodb/latestldeveloperguide!Ti . . . . [cited by applicant]