IP Library › Granted Patent US 12,493,817
Granted Patent B2
US 12,493,817 · App. 16/903,872 · Granted Dec 9, 2025

Performing data pre-processing operations during data preparation of a machine learning lifecycle

Inventors: Petr Novotny (Mount Kisco, NY); Qi Zhang (Elmsford, NY); Lei Yu (Sleepy Hollow, NY); Hong Min (Poughkeepsie, NY)
Assignee: International Business Machines Corporation
G06N20/00G06F16/21G06F16/2433G06F16/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,817
App. No.
16/903,872
Filed
Jun 17, 2020
Granted
Dec 9, 2025
Kind
B2
Art Unit
2165
USPC
706/12
Abstract

In accordance with an embodiment of the invention, a method is provided for performing data pre-processing operations during data preparation of a machine learning lifecycle. The method includes defining one or more data pre-processing functions for applying to data stored in a dataset, executing one or more learn functions for learning the data, and executing one or more transform functions for transforming the data. Each of the one or more learn functions generates a first Structured Query Language (SQL) statement representing a definition of corresponding learn function for corresponding defined data pre-processing function. Each of the one or more transform functions generates a second SQL statement representing a definition of corresponding transform function for corresponding defined data pre-processing function. The dataset is stored in a database.

Claims (36)

1. A method for performing data pre-processing operations during data preparation of a machine learning lifecycle, the method comprising:

defining, by a user on a computing system, one or more data pre-processing functions for applying to data stored in a dataset, the dataset stored in a database of a database management system, wherein each of the one or more data pre-processing functions performs a different operation on a different set of data within the dataset, and wherein each of the one or more data pre-processing functions are defined to apply on data stored in different respective specific columns of the dataset, and wherein the different operation is selected from the group consisting of an operation to scale data, an operation to threshold data, and an operation to normalize labels of data;

executing, at the database management system, one or more learn functions for learning the data, each of the one or more learn functions corresponding to a respective data pre-processing function of the one or more data pre-processing functions, each of the one or more learn functions generating a first Structured Query Language (SQL) statement for respectively corresponding data pre-processing function, each first SQL statement representing a definition of respectively corresponding learn function for corresponding defined data pre-processing function, each first SQL statement executes against a respective specific column, of the dataset, defined in respectively corresponding data pre-processing function to produce learned information of data stored in the respective specific column;

executing, at the database management system, one or more transform functions to transform the data, each of the one or more transform functions corresponding to a respective executed learn function of the one or more learn functions, wherein executing each of the one or more transform functions generates a respective second SQL statement based on learned information produced by respectively corresponding executed learn function, each second SQL statement representing a definition of corresponding transform function for corresponding defined data pre-processing function;

aggregating each respective second SQL statement into a single SQL statement and transforming the data stored in each specified column of the specified table based on the learned information of data; and

storing the transformed data to a location of machine learning models.

2. The method of claim 1 , wherein executing the one or more learn functions further comprises:

executing each first SQL statement for respectively corresponding data pre-processing function in the database;

producing learned information of each first SQL statement for respectively corresponding data pre-processing function; and

storing the learned information produced from each first SQL statement corresponding data pre-processing function in a storage mechanism.

3. The method of claim 2 , wherein executing the one or more transform functions further comprises:

executing the aggregated SQL statement in the database.

4. The method of claim 1 , wherein the user is a data scientist.

5. The method of claim 1 , wherein the dataset is comprised of one or more of financial data, demographic data, and medical data.

6. The method of claim 1 , wherein the dataset comprises at least one of text data type, numeric data type, and image data type.

7. The method of claim 2 , wherein the storage mechanism is a memory, a file, a temporary database storage.

8. The method of claim 7 , wherein the temporary database storage is a temporary table.

9. The method of claim 1 , wherein the database is a relational database.

10. The method of claim 1 , wherein the database is configured in a relational database management system.

11. The method of claim 7 , wherein the learned information is stored in the temporary database storage.

12. The method of claim 11 , wherein the learned information stored in the temporary database storage is a dictionary of labels.

13. The method of claim 1 , wherein the data pre-processing functions are defined with a machine learning software tool and library.

14. The method of claim 13 , wherein the machine learning software tool and library is Python scikit-learn library.

15. A system for performing data pre-processing operations during data preparation of a machine learning lifecycle, the system comprising:

a database management system, the database management system including a database, the database including a dataset for performing the data pre-processing operations; and

a computing system capable of communicating with the database management system, the computing system including a machine learning software tool, the machine learning software tool having a data pre-processing module and SQL generating module,

wherein the computing system comprises one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions of the machine learning software tool stored on at least one of the one or more computer-readable tangible storage devices for execution by at least one of the one or more processors via at least one of the one or more memories,

wherein the data pre-processing module of the computing system includes a first library for defining data pre-processing functions, and wherein each of the data pre-processing functions performs a different operation on a different set of data within the dataset, and wherein each of the one or more data pre-processing functions are defined to apply on data stored in different respective specific columns of the dataset, and wherein the different operation is selected from the group consisting of an operation to scale data, an operation to threshold data, and an operation to normalize labels of data,

wherein the SQL generating module includes a second library for translating logic of each of defined data pre-processing functions to respective SQL instructions and generating a corresponding first SQL statement for each defined data pre-processing function, for executing in the database of the database management system, each first SQL statement executes against data in a respective specific column, of the dataset, defined in respectively corresponding data pre-processing function to produce learned information of data stored in the respective specific column,

wherein the SQL generating module also generates a corresponding second SQL statement, for executing in the database of the database management system, for each first SQL statement, generation of each second SQL statement is based on learned information produced by respectively corresponding executed first SQL statement, each second SQL statement representing a definition of corresponding transform function for corresponding defined data pre-processing function,

wherein each second SQL statement is aggregated into a single SQL statement and transforming the data stored in each specified column of the specified table based on the learned information of data; and

wherein the transformed data is stored to a location of machine learning models.

16. The system of claim 15 , wherein the database is a relational database.

17. The system of claim 15 , wherein the database management system is a relational database management system.

18. The system of claim 15 , wherein the second library is a customized library which includes a plurality of classes.

19. The system of claim 15 , wherein the first library is Python scikit-learn library.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2020
From: NOVOTNY, PETR; ZHANG, QI; YU, LEI; MIN, HONG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052964/0906 →
Continuity (1)
Related Publication 20210398012A1 · Dec 23, 2021
References Cited (36)
US 7318051B2 · Weston et al. · 2008 [cited by applicant]
US 7359913B1 · Ordonez · 2008 [cited by examiner]
US 9652714B2 · Achin et al. · 2017 [cited by applicant]
US 10169433B2 · Lerios et al. · 2019 [cited by applicant]
US 10338968B2 · Bequet · 2019 [cited by examiner]
US 10445657B2 · Qian · 2019 [cited by examiner]
US 10497250B1 · Hayward · 2019 [cited by examiner]
US 10679012B1 · Salimov · 2020 [cited by examiner]
US 20020083067A1 · Tamayo · 2002 [cited by examiner]
US 20050102292A1 · Tamayo · 2005 [cited by examiner]
US 20090006346A1 · C N · 2009 [cited by examiner]
US 20170243140A1 · Achin et al. · 2017 [cited by applicant]
US 20180018602A1 · DiMaggio · 2018 [cited by examiner]
US 20190012403A1 · Bequet · 2019 [cited by examiner]
US 20190034767A1 · Sainani · 2019 [cited by examiner]
US 20190095801A1 · Saillet · 2019 [cited by examiner]
US 20190235484A1 · Ristovski · 2019 [cited by examiner]
US 20200007158A1 · Cooper · 2020 [cited by examiner]
US 20200050612A1 · Bhattacharjee · 2020 [cited by examiner]
US 20200065303A1 · Bhattacharjee · 2020 [cited by examiner]
US 20200134486A1 · Jiang · 2020 [cited by examiner]
US 20200167914A1 · Stamatoyannopoulos · 2020 [cited by examiner]
US 20200210525A1 · Yang · 2020 [cited by examiner]
US 20200279181A1 · O'Reilly · 2020 [cited by examiner]
US 20200279200A1 · Makhija · 2020 [cited by examiner]
US 20200327371A1 · Sharma · 2020 [cited by examiner]
US 20210074269A1 · Duong · 2021 [cited by examiner]
US 20210224585A1 · Schmidt · 2021 [cited by examiner]
US 20240012810A1 · Lee · 2024 [cited by examiner]
Google: “Data Set Preprocessing and Transformation in a Database System”, Ordonez, C.; 2011. [cited by applicant]
Google: “In-Database Machine Learning: Gradient Descent and Tensor Algebra for Main Memory Database Systems”, Schule, M. et al.; 2019. [cited by applicant]
ip.com: “Automatic Denormalization of Databases”, Anonymously; Aug. 17, 2018. [cited by applicant]
ip.com: “A Method for Time Prediction on Database System Statistics Collection and SQL Rebind by Machine Learning”, Anonymously; May 31, 2018. [cited by applicant]
ip.com: “Methodology for Performing Machine Learning on Database Data in a SQL Statement”, Anonymously; Jan. 13, 2018. [cited by applicant]
http://keystone-ml.org/, “KeystoneML”, Accessed on Aug. 5, 2022, 3 pages. [cited by applicant]
https://mldb.ai/, “MLDB is the Machine Learning Database”, Accessed on Aug. 5, 2022, 4 pages. [cited by applicant]