Erroneous cell detection using an artificial intelligence model
Classification of cell data includes obtaining a target dataset and an artificial intelligence (AI) model trained to identify relationship(s) between cells of a row and classify whether a focus cell of the row is erroneous based on the identified relationship(s), and applying the AI model to the target dataset to identify erroneous cell(s) thereof. The applying includes selecting a row of cells of the target dataset, inputting the selected row of cells to the AI model with an identification of a focus cell, the focus cell to be classified by the AI model, classifying the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous, and outputting an indication of the classification of the focus cell.
1 . A computer-implemented method comprising:
performing data modification on selected data cells of an initial dataset, wherein the data modification imparts errors in the selected data cells;
labeling the selected data cells as being erroneous, wherein the data modification and the labeling provide a training dataset with the errors imparted in the selected data cells,
training an artificial intelligence (AI) model using the training dataset to identify one or more relationships between cells of a row that inform a context of a selected focus cell of the row based at least in part on cells neighboring the selected focus cell, and classify whether the selected focus cell of the row is erroneous based at least in part on the identified one or more relationships and the informed context;
obtaining a target dataset, the target dataset comprising rows and columns of data cells,
applying the AI model to the target dataset to identify one or more erroneous cells of the target dataset, the applying comprising:
selecting a row of cells of the target dataset;
inputting the selected row of cells to the AI model with an identification of a focus cell of the selected row of cells, the focus cell to be classified by the AI model, wherein the inputting comprises:
building a string of cell data of a plurality of cells of the selected row of cells, the plurality of cells comprising the focus cell, and the string of cell data of the plurality of cells comprising cell data of the focus cell wherein the building provides a delimiter in the string between cell data of different cells of the selected row of cells; and
inputting to the AI model the string of cell data of the plurality of cells;
classifying, by the AI model, the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous; and
outputting an indication of the classification of the focus cell.
2 . The method of claim 1 , wherein the classifying comprises one selected from the group consisting of:
a binary classification that classifies the focus cell as being either erroneous or not erroneous; and
a multi-classification that identifies whether the focus cell is erroneous and, if so, one or more errors of the focus cell.
3 . The method of claim 1 , wherein the identification of the focus cell comprises an identification of the cell data of the focus cell within the string of cell data of the plurality of cells, and wherein the classifying the focus cell classifies the cell data of the focus cell.
4 . The method of claim 3 , wherein the identification of the cell data of the focus cell comprises insertion of a selected character into the string at one or more positions relative to the cell data of the focus cell, wherein the AI model is configured to identify the focus cell by locating the inserted selected character.
5 . The method of claim 1 , wherein the applying further comprises:
iterating, over one or more other cells of the selected row of cells, the inputting the selected row of cells, the classifying the focus cell, and the outputting an indication of the classification of the focus cell, wherein, at each iteration of the iterating, the applying:
identifies a next focus cell of the selected row of cells, the next focus cell being a different cell than each prior-identified and classified focus cell of the selected row of cells;
inputs the selected row of cells to the AI model with an identification of the next focus cell, the next focus cell to be classified by the AI model;
classifies the next focus cell to obtain a classification of the next focus cell, the classifying identifying whether the focus cell is erroneous; and
outputs an indication of the classification of the next focus cell.
6 . The method of claim 5 , wherein the applying further comprises iterating, over one or more other rows of the target dataset:
the selecting a row of cells;
the inputting the selected row of cells to the AI model with an identification of a focus cell;
the classifying the focus cell;
the outputting an indication of the classification of the focus cell; and
the iterating, over one or more other cells of the selected row of cells, the inputting the selected row of cells, the classifying the focus cell, and the outputting an indication of the classification of the focus cell, wherein at each iteration of the iterating over the one or more other rows of the target dataset, the applying identifies a next row of cells of the target dataset, the next row of cells being a different row of cells than each prior-selected row of cells, and wherein the identified next row of cells is the selected row of cells.
7 . The method of claim 1 , wherein the AI model is trained to identify the one or more relationships by way of a multi-head attention component that comprises a plurality of attention heads, and wherein the AI model is configured to classify the focus cell based at least in part on the informed context.
8 . The method of claim 1 , wherein the AI model comprises:
a plurality of encoder layers, each encoder layer of the plurality of encoder layers comprising a respective at least one attention layer and a feed-forward layer; and
a perceptron layer comprising an activation function configured to provide the output classification of the focus cell.
9 . The method of claim 1 , wherein the training dataset comprises at least the selected data cells labeled as erroneous and at least some cells labeled as correct.
10 . The method of claim 9 , wherein the training dataset is based on a first proper subset of a larger dataset and wherein the target dataset is a second proper subset of that larger dataset, the second proper subset being different from the first proper subset.
11 . The method of claim 1 , wherein the data modification on a selected data cell of the initial dataset comprises at least one selected from the group consisting of: random character replacement, random character insertion, random character deletion, random character swapping, value deletion, column value swapping, and row value swapping.
12 . The method of claim 1 , wherein the data modification imparts in a selected data cell of the initial dataset at least one error selected from the group consisting of: a typographical error, a value swap across columns error, a value violating data constraint error, a formatting error, and a missing value error.
13 . The method of claim 1 , further comprising performing automatic processing, the automatic processing comprising raising an electronic alert to a user that indicates one or more erroneous cells of the target dataset.
14 . A computer system comprising:
a memory; and
a processor in communication with the memory, wherein the computer system is configured to perform a method comprising:
performing data modification on selected data cells of an initial dataset, wherein the data modification imparts errors in the selected data cells;
labeling the selected data cells as being erroneous, wherein the data modification and the labeling provide a training dataset with the errors imparted in the selected data cells;
training an artificial intelligence (AI) model using the training dataset to identify one or more relationships between cells of a row that inform a context of a selected focus cell of the row based at least in part on cells neighboring the selected focus cell, and classify whether the selected focus cell of the row is erroneous based at least in part on the identified one or more relationships and the informed context;
obtaining a target dataset, the target dataset comprising rows and columns of data cells;
applying the AI model to the target dataset to identify one or more erroneous cells of the target dataset, the applying comprising:
selecting a row of cells of the target dataset;
inputting the selected row of cells to the AI model with an identification of a focus cell of the selected row of cells, the focus cell to be classified by the AI model, wherein the inputting comprises:
building a string of cell data of a plurality of cells of the selected row of cells, the plurality of cells comprising the focus cell, and the string of cell data of the plurality of cells comprising cell data of the focus cell, wherein the building provides a delimiter in the string between cell data of different cells of the selected row of cells; and
inputting to the AI model the string of cell data of the plurality of cells;
classifying, by the AI model, the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous; and
outputting an indication of the classification of the focus cell.
15 . The computer system of claim 14 , wherein the identification of the focus cell comprises an identification of the cell data of the focus cell within the string of cell data of the plurality of cells, wherein the identification of the cell data of the focus cell comprises insertion of a selected character into the string at one or more positions relative to the cell data of the focus cell, wherein the AI model is configured to identify the focus cell by locating the inserted selected character, and wherein the classifying the focus cell classifies the cell data of the focus cell.
16 . The computer system of claim 14 , wherein the AI model is trained to identify the one or more relationships by way of a multi-head attention component that comprises a plurality of attention heads, and wherein the AI model is configured to classify the focus cell based at least in part on the informed context.
17 . The computer system of claim 14 , wherein the training dataset comprises at least the selected data cells labeled as erroneous and at least some cells labeled as correct, wherein the training dataset is based on a first proper subset of a larger dataset and wherein the target dataset is a second proper subset of that larger dataset, the second proper subset being different from the first proper subset.
18 . A computer program product comprising:
a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising:
performing data modification on selected data cells of an initial dataset, wherein the data modification imparts errors in the selected data cells;
labeling the selected data cells as being erroneous, wherein the data modification and the labeling provide a training dataset with the errors imparted in the selected data cells;
training an artificial intelligence (AI) model using the training dataset to identify one or more relationships between cells of a row that inform a context of a selected focus cell of the row based at least in part on cells neighboring the selected focus cell, and classify whether the selected focus cell of the row is erroneous based at least in part on the identified one or more relationships and the informed context;
obtaining a target dataset, the target dataset comprising rows and columns of data cells;
applying the AI model to the target dataset to identify one or more erroneous cells of the target dataset, the applying comprising:
selecting a row of cells of the target dataset;
inputting the selected row of cells to the AI model with an identification of a focus cell of the selected row of cells, the focus cell to be classified by the AI model, wherein the inputting comprises:
building a string of cell data of a plurality of cells of the selected row of cells, the plurality of cells comprising the focus cell, and the string of cell data of the plurality of cells comprising cell data of the focus cell, wherein the building provides a delimiter in the string between cell data of different cells of the selected row of cells; and
inputting to the AI model the string of cell data of the plurality of cells;
classifying, by the AI model, the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous; and
outputting an indication of the classification of the focus cell.
19 . The computer program product of claim 18 , wherein the AI model is trained to identify the one or more relationships by way of a multi-head attention component that comprises a plurality of attention heads, and wherein the AI model is configured to classify the focus cell based at least in part on the informed context.
20 . The computer program product of claim 18 , wherein the training dataset comprises at least the selected data cells labeled as erroneous and at least some cells labeled as correct, wherein the training dataset is based on a first proper subset of a larger dataset and wherein the target dataset is a second proper subset of that larger dataset, the second proper subset being different from the first proper subset.