IP Library › Granted Patent US 12,657,377
Granted Patent B2
US 12,657,377 · App. 18/241,645 · Granted Jun 16, 2026

Automatic fitting and/or proposal of predictive models to data entries of a spreadsheet file representative of a larger dataset

Inventor: Oscar Castañeda-Villagrán (Guatemala City, GT)
G06F40/18G06F9/5077G06F40/186G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,377
App. No.
18/241,645
Granted
Jun 16, 2026
Kind
B2
Abstract

Disclosed is a system, a method, and/or a device of automatic fitting and/or proposal of prediction models to data entries of a spreadsheet file representative of a larger dataset. In one embodiment, a system for automatic determination of a predictive model for scaled data analysis includes two or more servers that process a spreadsheet file including data from a dataset, each data entry of the spreadsheet file including one or more independent variables in one or more cells and a dependent variable. The system automatically determines the predictive model fits a data entry of the spreadsheet. The system proposes an algorithm in response to a fitting of the predictive model, the algorithm accepting as inputs the one or more independent variables and outputting the dependent variable. The system applies the algorithm against the dataset utilizing parallel processing to generate the dependent variable for each data entry of the dataset.

Claims (154)

1 . A system for automatic proposal of a predictive model for scaled data analysis, the system comprising:

a server comprising:

a processor of the server; and

a memory of the server comprising computer readable instructions that when executed:

process a spreadsheet file comprising data from a dataset, the dataset comprising two or more data entries of the dataset and the spreadsheet file comprising one or more data entries of the spreadsheet file that are a subset of the two or more data entries of the dataset, each data entry of the one or more data entries of the spreadsheet file comprising one or more independent variables in one or more cells and a dependent variable in a cell of the spreadsheet file,

automatically determine the predictive model of one or more predictive models fits at least one of the one or more data entries of the spreadsheet file, and

propose an algorithm in response to a fitting of the one or more predictive models to the one or more data entries of the spreadsheet file, the algorithm accepting as inputs the one or more independent variables and outputting the dependent variable,

an execution server communicatively coupled to the server, the execution server comprising:

a processor of the execution server,

a memory of the execution server comprising computer readable instructions that when executed:

apply the algorithm against the dataset utilizing parallel processing to generate an output data comprising a value for the dependent variable calculated for each of the two or more data entries of the dataset.

2 . The system of claim 1 , further comprising:

a model server comprising:

a processor of the model server; and

a memory of the model server comprising computer readable instructions that when executed:

select a model family comprising the predictive model,

optimize a model parameter of the predictive model,

tune a hyper parameter of the predictive model,

run an automatic machine learning process to automatically apply the one or more predictive models to the at least one of the two or more data entries of the spreadsheet file,

verify calculation independence of a data entry of the dataset when the algorithm to be applied to the two or more data entries of the dataset is applied to the data entry of the dataset,

run the automatic machine learning process to automatically apply the one or more predictive models to the dataset,

determine a different predictive model of the one or more predictive models fits the dataset,

propose a different algorithm in response to an application of one or more predictive models to the dataset, and

re-submit a first computation block comprising a different subset of the two or more data entries of the dataset and the different algorithm for parallel processing.

3 . The system of claim 1 , wherein:

the memory of the execution server further comprising computer readable instructions that when executed:

specify a first computation block comprising a different subset of the two or more data entries of the dataset,

extract from the dataset each of the different subset of the two or more data entries within the first computation block, and

submit the first computation block and the algorithm for parallel processing.

4 . The system of claim 1 , further comprising:

a translation server comprising:

a processor of the translation server, and

a memory of the translation server comprising computer readable instructions that when executed:

import the data entry of the one or more data entries of the spreadsheet file from the dataset as a prototype data representative the dataset, the data entry of the one or more data entries of the spreadsheet file comprising two or more pieces of data, and

map each of the two or more pieces of data of the data entry to cells of the spreadsheet file comprising the one or more cells of the one or more independent variables and the cell of the dependent variable; and

a client device comprising:

a processor of the client device, and

a memory of the client device comprising a spreadsheet application for reading the spreadsheet file and a browser application for accessing the spreadsheet file as a software-as-a-service.

5 . The system of claim 1 , the system further comprising:

a translation server comprising:

a processor of the translation server, and

a memory of the translation server comprising computer readable instructions that when executed:

process the spreadsheet file comprising a formula algorithm to be applied to the dataset comprising the two or more data entries of the dataset,

wherein the formula algorithm outputting the dependent variable and accepting as inputs the one or more independent variables, and

wherein the formula algorithm comprising one or more spreadsheet formulas stored in a different set of one or more cells of the spreadsheet file, the one or more independent variables referenced from at least one of the different set of one or more cells of the spreadsheet file, and the dependent variable is output in the cell of the spreadsheet file.

6 . The system of claim 5 , wherein the memory of the translation server further comprising computer readable instructions that when executed:

generate from the formula algorithm an extrapolated algorithm expressed in a programming language that is at least one of a query language, an interpreted programming language, and a functional programming language,

wherein each of the one or more spreadsheet formulas equivalent to one or more functions of the programming language and each of the one or more independent variables define a declared variable of at least one of the one or more functions of the programming language, and

verify calculation independence of the data entry of the dataset when the extrapolated algorithm to be applied to the two or more data entries of the dataset is applied to the data entry.

7 . The system of claim 1 , wherein the memory of the execution server further comprising computer readable instructions that when executed:

provision at least one of a computing virtual machine and a computing process container;

store the output data in at least one of the computing virtual machine and the computing process container; and

communicatively couple the computing process container to an application programming interface (API) of a visualization application.

8 . The method of claim 6 , wherein:

the memory of the translation server further comprising computer readable instructions that when executed:

determine the data entry of the two or more data entries of the dataset references data of a different data entry of the two or more data entries of the dataset, and

replicate the data of the different data entry of the dataset into the data entry of the dataset to ensure the calculation independence of the data entry of the dataset; and

the memory of the server further comprising computer readable instructions that when executed:

process an extra-spreadsheet instruction stored in a third set of one or more cells of the spreadsheet file to at least one of run a data analysis process, report the output data to a user, parametrize the formula algorithm, and parameterize at least one of the one or more independent variables,

wherein the dependent variable is a prediction metric,

wherein the calculation independence based on a syntax format comprising:

(i) confining data of the data entry to a row of the spreadsheet file, with each instance of a cell of the row comprising at least one of a null, an independent variable of the one or more independent variables, the dependent variable, a spreadsheet formula of the one or more spreadsheet formulas, the extra-spreadsheet instruction, and an analysis instruction, and

(ii) confining to the row of the spreadsheet file the one or more spreadsheet formulas comprising the formula algorithm,

wherein the spreadsheet file is accessed as a software-as-a-service through a browser application,

wherein the output data is re-combined from data comprising a first output block generated through parallel processing resulting in one or more additional output blocks, and

wherein the programming language comprises a structured query language (SQL), and

wherein a format of the spreadsheet file is at least one of: .123, .12M, ._XLS, ._XLSX, .AST, .AWS, .BKS, .CELL, .DEF, .DEX, .DFG, .DIS, .EDX, .EDXZ, .ESS, .FCS, .FM, .FODS, .FP, .GNM, .GNUMERIC, .GSHEET, .HCDT, .IMP, .MAR, .NB, .NCSS, .NMBTEMPLATE, .NUMBERS, .NUMBERS-TEF, .ODS, .OGW, .OGWU, .OTS, .PMD, .PMDX, .PMV, .PMVX, .QPW, .RDF, .SDC, .STC, .SXC, .TMV, .TMVT, .UOS, .WKI, .WKQ, .WKS, .WKU, .WQ1, .WQ2, .WR1, .XAR, .XL, .XLR, .XLS, .XLSB, .XLSHTML, .XLSM, .XLSMHTML, .XLSX, .XLTHTML, .XLTM, and .XLTX.

9 . A method of automatic proposal of a predictive model for scaled data analysis, the method comprising:

processing a spreadsheet file comprising data from a dataset, the dataset comprising two or more data entries of the dataset and the spreadsheet file comprising one or more data entries of the spreadsheet file that are a subset of the two or more data entries of the dataset, each data entry of the one or more data entries of the spreadsheet file comprising one or more independent variables in one or more cells and a dependent variable in a cell of the spreadsheet file;

automatically determining the predictive model of one or more predictive models fits at least one of the one or more data entries of the spreadsheet file;

proposing an algorithm in response to a fitting of the one or more predictive models to the one or more data entries of the spreadsheet file, the algorithm accepting as inputs the one or more independent variables and outputting the dependent variable; and

applying the algorithm against the dataset utilizing parallel processing to generate an output data comprising a value for the dependent variable calculated for each of the two or more data entries of the dataset.

10 . The method of claim 9 , further comprising:

running an automatic machine learning process to automatically apply the one or more predictive models to the at least one of the two or more data entries of the spreadsheet file; and

verifying calculation independence of a data entry of the dataset when the algorithm to be applied to the two or more data entries of the dataset is applied to the data entry of the dataset.

11 . The method of claim 9 , further comprising:

specifying a first computation block comprising a different subset of the two or more data entries of the dataset;

extracting from the dataset each of the different subset of the two or more data entries within the first computation block;

submitting the first computation block and the algorithm for parallel processing;

running an automatic machine learning process to automatically apply the one or more predictive models to the dataset;

determining a different predictive model of the one or more predictive models fits the dataset;

proposing a different algorithm in response to an application of one or more predictive models to the dataset; and

re-submitting the first computation block and the different algorithm for parallel processing.

12 . The method of claim 9 , further comprising:

importing the data entry of the one or more data entries of the spreadsheet file from the dataset as a prototype data representative the dataset, the data entry of the one or more data entries of the spreadsheet file comprising two or more pieces of data; and

mapping each of the two or more pieces of data of the data entry to cells of the spreadsheet file comprising the one or more cells of the one or more independent variables and the cell of the dependent variable.

13 . The method of claim 9 , further comprising:

processing the spreadsheet file comprising a formula algorithm to be applied to the dataset comprising the two or more data entries of the dataset, the formula algorithm outputting the dependent variable and accepting as inputs the one or more independent variables,

wherein the formula algorithm comprising one or more spreadsheet formulas stored in a different set of one or more cells of the spreadsheet file, the one or more independent variables referenced from at least one of the different set of one or more cells of the spreadsheet file, and the dependent variable is output in the cell of the spreadsheet file.

14 . The method of claim 13 , further comprising:

generating from the formula algorithm an extrapolated algorithm expressed in a programming language that is at least one of a query language, an interpreted programming language, and a functional programming language,

wherein each of the one or more spreadsheet formulas equivalent to one or more functions of the programming language and each of the one or more independent variables define a declared variable of at least one of the one or more functions of the programming language; and

verifying calculation independence of the data entry of the dataset when the extrapolated algorithm to be applied to the two or more data entries of the dataset is applied to the data entry.

15 . The method of claim 9 , further comprising:

selecting a model family comprising the predictive model;

optimizing a model parameter of the predictive model;

tuning a hyper parameter of the predictive model;

provisioning at least one of a computing virtual machine and a computing process container;

storing the output data in at least one of the computing virtual machine and the computing process container; and

communicatively coupling the computing process container to an application programming interface (API) of a visualization application.

16 . The method of claim 14 , further comprising:

determining the data entry of the two or more data entries of the dataset references data of a different data entry of the two or more data entries of the dataset;

replicating the data of the different data entry of the dataset into the data entry of the dataset to ensure the calculation independence of the data entry of the dataset; and

processing an extra-spreadsheet instruction stored in a third set of one or more cells of the spreadsheet file to at least one of run a data analysis process, report the output data to a user, parametrize the formula algorithm, and parameterize at least one of the one or more independent variables,

wherein the dependent variable is a prediction metric,

wherein the calculation independence based on a syntax format comprising:

(i) confining data of the data entry to a row of the spreadsheet file, with each instance of a cell of the row comprising at least one of a null, an independent variable of the one or more independent variables, the dependent variable, a spreadsheet formula of the one or more spreadsheet formulas, the extra-spreadsheet instruction, and an analysis instruction, and

(ii) confining to the row of the spreadsheet file the one or more spreadsheet formulas comprising the formula algorithm,

wherein the spreadsheet file is accessed as a software-as-a-service through a browser application,

wherein the output data is re-combined from data comprising a first output block generated through parallel processing resulting in one or more additional output blocks, and

wherein the programming language comprises a structured query language (SQL), and

wherein a format of the spreadsheet file is at least one of: .123, .12M, ._XLS, ._XLSX, .AST, .AWS, .BKS, .CELL, .DEF, .DEX, .DFG, .DIS, .EDX, .EDXZ, .ESS, .FCS, .FM, .FODS, .FP, .GNM, .GNUMERIC, .GSHEET, .HCDT, .IMP, .MAR, .NB, .NCSS, .NMBTEMPLATE, .NUMBERS, .NUMBERS-TEF, .ODS, .OGW, .OGWU, .OTS, .PMD, .PMDX, .PMV, .PMVX, .QPW, .RDF, .SDC, .STC, .SXC, .TMV, .TMVT, .UOS, .WKI, .WKQ, .WKS, .WKU, .WQ1, .WQ2, .WR1, .XAR, .XL, .XLR, .XLS, .XLSB, .XLSHTML, .XLSM, .XLSMHTML, .XLSX, .XLTHTML, .XLTM, and .XLTX.

17 . A system for automatic determination of a predictive model for scaled data analysis, the system comprising:

two or more servers communicatively coupled over a network, the two or more servers comprising two or more processors and two or more memories, the two or more memories comprising computer readable instructions that when executed on any of the two or more servers:

process a spreadsheet file comprising data from a dataset, the dataset comprising two or more data entries of the dataset and the spreadsheet file comprising one or more data entries of the spreadsheet file that are a subset of the two or more data entries of the dataset, each data entry of the one or more data entries of the spreadsheet file comprising one or more independent variables in one or more cells and a dependent variable in a cell of the spreadsheet file;

automatically determine the predictive model of one or more predictive models fits at least one of the one or more data entries of the spreadsheet file;

propose an algorithm in response to a fitting of the one or more predictive models to the one or more data entries of the spreadsheet file, the algorithm accepting as inputs the one or more independent variables and outputting the dependent variable; and

apply the algorithm against the dataset utilizing parallel processing to generate an output data comprising a value for the dependent variable calculated for each of the two or more data entries of the dataset.

18 . The system of claim 17 , wherein the two or more memories further comprising computer readable instructions that when executed on any of the two or more servers:

run an automatic machine learning process to automatically apply the one or more predictive models to the at least one of the two or more data entries of the spreadsheet file; and

verify calculation independence of a data entry of the dataset when the algorithm to be applied to the two or more data entries of the dataset is applied to the data entry of the dataset.

19 . The system of claim 17 , further comprising:

specify a first computation block comprising a different subset of the two or more data entries of the dataset;

extract from the dataset each of the different subset of the two or more data entries within the first computation block;

submit the first computation block and the algorithm for parallel processing;

run an automatic machine learning process to automatically apply the one or more predictive models to the dataset;

determine a different predictive model of the one or more predictive models fits the dataset;

propose a different algorithm in response to an application of one or more predictive models to the dataset; and

re-submit the first computation block and the different algorithm for parallel processing.

20 . The system of claim 17 , further comprising:

import the data entry of the one or more data entries of the spreadsheet file from the dataset as a prototype data representative the dataset, the data entry of the one or more data entries of the spreadsheet file comprising two or more pieces of data;

map each of the two or more pieces of data of the data entry to cells of the spreadsheet file comprising the one or more cells of the one or more independent variables and the cell of the dependent variable;

process the spreadsheet file comprising a formula algorithm to be applied to the dataset comprising the two or more data entries of the dataset, the formula algorithm outputting the dependent variable and accepting as inputs the one or more independent variables,

wherein the formula algorithm comprising one or more spreadsheet formulas stored in a different set of one or more cells of the spreadsheet file, the one or more independent variables referenced from at least one of the different set of one or more cells of the spreadsheet file, and the dependent variable is output in the cell of the spreadsheet file;

generate from the formula algorithm an extrapolated algorithm expressed in a programming language that is at least one of a query language, an interpreted programming language, and a functional programming language,

wherein each of the one or more spreadsheet formulas equivalent to one or more functions of the programming language and each of the one or more independent variables define a declared variable of at least one of the one or more functions of the programming language;

verify calculation independence of the data entry of the dataset when the extrapolated algorithm to be applied to the two or more data entries of the dataset is applied to the data entry;

select a model family comprising the predictive model;

optimize a model parameter of the predictive model;

tune a hyper parameter of the predictive model;

provision at least one of a computing virtual machine and a computing process container;

store the output data in at least one of the computing virtual machine and the computing process container;

communicatively couple the computing process container to an application programming interface (API) of a visualization application;

determine the data entry of the two or more data entries of the dataset references data of a different data entry of the two or more data entries of the dataset;

replicate the data of the different data entry of the dataset into the data entry of the dataset to ensure the calculation independence of the data entry of the dataset; and

process an extra-spreadsheet instruction stored in a third set of one or more cells of the spreadsheet file to at least one of run a data analysis process, report the output data to a user, parametrize the formula algorithm, and parameterize at least one of the one or more independent variables,

wherein the dependent variable is a prediction metric,

wherein the calculation independence based on a syntax format comprising:

(i) confining data of the data entry to a row of the spreadsheet file, with each instance of a cell of the row comprising at least one of a null, an independent variable of the one or more independent variables, the dependent variable, a spreadsheet formula of the one or more spreadsheet formulas, the extra-spreadsheet instruction, and an analysis instruction, and

(ii) confining to the row of the spreadsheet file the one or more spreadsheet formulas comprising the formula algorithm,

wherein the spreadsheet file is accessed as a software-as-a-service through a browser application,

wherein the output data is re-combined from data comprising a first output block generated through parallel processing resulting in one or more additional output blocks, and

wherein the programming language comprises a structured query language (SQL), and

wherein a format of the spreadsheet file is at least one of: .123, .12M, ._XLS, ._XLSX, .AST, .AWS, .BKS, .CELL, .DEF, .DEX, .DFG, .DIS, .EDX, .EDXZ, .ESS, .FCS, .FM, .FODS, .FP, .GNM, .GNUMERIC, .GSHEET, .HCDT, .IMP, .MAR, .NB, .NCSS, .NMBTEMPLATE, .NUMBERS, .NUMBERS-TEF, .ODS, .OGW, .OGWU, .OTS, .PMD, .PMDX, .PMV, .PMVX, .QPW, .RDF, .SDC, .STC, .SXC, .TMV, .TMVT, .UOS, .WKI, .WKQ, .WKS, .WKU, .WQ1, .WQ2, .WR1, .XAR, .XL, .XLR, .XLS, .XLSB, .XLSHTML, .XLSM, .XLSMHTML, .XLSX, .XLTHTML, .XLTM, and .XLTX.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2023
From: CASTAÑEDA-VILLAGRÁN, OSCAR
To: SCIENCESHEET INC.
Reel/Frame 064790/0392 →
Continuity (6)
Continuation 17869410 · Jul 20, 2022
Continuation 17134299 · Dec 26, 2020
Continuation 16867659 · May 6, 2020
Continuation 16150262 · Oct 2, 2018
Provisional Application 62575430 · Oct 21, 2017
Related Publication 20240012987A1 · Jan 11, 2024
References Cited (10)
US 10685175B2 · Castañeda-Villagrán · 2020 [cited by examiner]
US 20060074824A1 · Li · 2006 [cited by examiner]
US 20060112123A1 · Clark · 2006 [cited by examiner]
US 20080243823A1 · Baris · 2008 [cited by examiner]
US 20100169758A1 · Thomsen · 2010 [cited by examiner]
US 20130124957A1 · Oppenheimer · 2013 [cited by examiner]
US 20140164895A1 · Matheson · 2014 [cited by examiner]
US 20140372346A1 · Phillipps · 2014 [cited by examiner]
US 20160070805A1 · Malkin · 2016 [cited by examiner]
Cunha, Jácome, et al. “Model inference for spreadsheets.” Automated Software Engineering 23 (2016): 361-392. (Year: 2016). [cited by examiner]