IP Library › Granted Patent US 11,429,878
Granted Patent B2
US 11,429,878 · App. 15/712,691 · Granted Aug 30, 2022

Cognitive recommendations for data preparation

Inventors: Yannick Saillet (Stuttgart, DE); Martin A. Oberhofer (Pondorf, DE); Jens P. Seifert (Gaertringen, DE)
Assignee: International Business Machines Corporation
G06N5/04G06N5/00G06N20/00G06F16/215
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,878
App. No.
15/712,691
Granted
Aug 30, 2022
Kind
B2
Abstract

A method, computer system, and computer program product for providing recommendations about processing datasets. A set of machine learning models are provided for use in respectively determining data processing action performable on a dataset based on a respective set of features of the dataset. A current dataset is received. A set of features of the current dataset are determined. One or more data processing actions are generated to be executed on the current dataset, which are determined by at least two machine learning models of the provided set, based on the determined set of features of the current dataset. One or more of the data processing actions are performed on the current dataset.

Claims (55)

1. A computer-implemented method for providing recommendations about processing datasets, each dataset comprising records containing a respective set of data fields, the method comprising:

receiving, by a processor, a current dataset;

determining, by the processor, a set of features of the current dataset;

creating, by the processor, a set of data processing actions based on the set of features of the current dataset, each data processing action in the set of data processing actions being output from a corresponding machine learning model in a provided set of machine learning models, and wherein each data processing action is associated with a confidence score and is selected from a group consisting of: data standardization, data masking, row deduplication, data filtering, data enrichment, data transformation and replacement of missing or incorrect values;

generating, by the processor, one or more recommended data processing actions based on the set of data processing actions and the associated confidence scores;

performing, by the processor, one or more data processing actions in the set of data processing actions on the current dataset;

determining, by the processor, a prior data processing action that was previously performed on the current dataset;

updating, by the processor, training data used in training the machine learning model based on the prior data processing action that was previously performed on the current dataset; and

retraining, by the processor, the machine learning model using the updated training data.

2. The method of claim 1 , wherein the confidence score indicates a relevance of an associated data processing action.

3. The method of claim 2 , wherein the confidence score indicates one or more of: a level of confidence of a determination of an associated data processing action, and a level of confidence of the machine learning model.

4. The method of claim 2 , wherein the performing the one or more data processing actions in the set of data processing actions on the current dataset further comprises:

selecting, by the processor, the recommended data processing actions, based on the associated confidence score; and

performing, by the processor, the recommended data processing action on the current dataset.

5. The method of claim 1 , wherein the performing the one or more data processing actions in the set of data processing actions on the current dataset further comprises:

displaying, by the processor, the one or more data processing actions in the set of data processing actions to a user for selection; and

performing, by the processor, the one or more data processing actions in the set of data processing actions that are selected by the user on the current dataset.

6. The method of claim 1 , wherein the performing the one or more data processing actions in the set of data processing actions on the current dataset further comprises:

receiving, by the processor, a selection of one or more data processing actions in the set of data processing actions; and

performing, by the processor, the selection of the one or more data processing actions in the set of data processing actions on the current dataset.

7. The method of claim 1 , wherein the set of features of the current dataset are determined using one or more of: data profiling algorithms, data classification algorithms, and data quality algorithms.

8. The method of claim 1 , further comprising:

determining a quality metric of the current dataset after performance of the one or more data processing actions in the set of data processing actions; and

retraining the machine learning model using the determined quality metric.

9. The method of claim 1 , wherein the machine learning model comprises one or more of: an association model for associating terms with classifications, a classification model, and a text classification model.

10. A computer system for providing recommendations about processing datasets, the computer system comprising:

one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions, when executed by at least one of the one or more computer processors, causing the computer system to perform a method comprising:

receiving a current dataset;

determining a set of features of the current dataset;

creating a set of data processing actions based on the set of features of the current dataset, each data processing action in the set of data processing actions being output from a corresponding machine learning model in a provided set of machine learning models, and wherein each data processing action is associated with a confidence score and is selected from a group consisting of: data standardization, data masking, row deduplication, data filtering, data enrichment, data transformation and replacement of missing or incorrect values;

generating, one or more recommended data processing actions based on the set of data processing actions and the associated confidence scores;

performing one or more data processing actions in the set of data processing actions on the current dataset;

determining a prior data processing action that was previously performed on the current dataset;

updating training data used in training the machine learning model based on the prior data processing action that was previously performed on the current dataset; and

retraining the machine learning model using the updated training data.

11. The computer system of claim 10 , wherein the confidence score indicates a relevance of an associated data processing action.

12. The computer system of claim 11 , wherein the confidence score indicates one or more of: a level of confidence of a determination of an associated data processing action, and a level of confidence of the machine learning model.

13. The computer system of claim 11 , wherein the performing the one or more data processing actions in the set of data processing actions on the current dataset further comprises:

selecting the recommended data processing actions, based on the associated confidence score; and

performing, by the processor, the recommended data processing action on the current dataset.

14. A computer program product for providing recommendations about processing datasets, comprising:

one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices for execution by at least one or more computer processors of a computer system, the program instructions, when executed by at least one of the one or more computer processors, causing the computer system to perform a method comprising:

receiving a current dataset;

determining a set of features of the current dataset;

creating a set of data processing actions based on the set of features of the current dataset, each data processing action in the set of data processing actions being output from a corresponding machine learning model in a provided set of machine learning models, and wherein each data processing action is associated with a confidence score and is selected from a group consisting of: data standardization, data masking, row deduplication, data filtering, data enrichment, data transformation and replacement of missing or incorrect values;

generating, one or more recommended data processing actions based on the set of data processing actions and the associated confidence scores;

performing one or more data processing actions in the set of data processing actions on the current dataset;

determining a prior data processing action that was previously performed on the current dataset;

updating training data used in training the machine learning model based on the prior data processing action that was previously performed on the current dataset; and

retraining the machine learning model using the updated training data.

15. The computer program product of claim 14 , wherein the confidence score indicates a relevance of an associated data processing action.

16. The computer program product of claim 15 , wherein the confidence score indicates one or more of: a level of confidence of a determination of an associated data processing action, and a level of confidence of the machine learning model.

17. The computer program product of claim 15 , wherein the performing the one or more data processing actions in the set of data processing actions on the current dataset further comprises:

selecting the recommended data processing actions, based on the associated confidence score; and

performing, by the processor, the recommended data processing action on the current dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2017
From: SAILLET, YANNICK; OBERHOFER, MARTIN A.; SEIFERT, JENS P.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 043975/0822 →
Continuity (1)
Related Publication 20190095801A1 · Mar 28, 2019
Cited By (1)
US 12,505,383