IP Library › Granted Patent US 12,639,550
Granted Patent B2
US 12,639,550 · App. 17/455,461 · Granted May 26, 2026

Erroneous cell detection using an artificial intelligence model

Inventors: Shaikh Shahriar Quader (Scarborough, CA); Omar Al-Shamali (Edmonton, CA); James Miller (Edmonton, CA); Yannick Saillet (Stuttgart, DE); Albert Maier (Tuebingen, DE); Remus Lazar (Morgan Hill, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/04G06F16/215G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,550
App. No.
17/455,461
Granted
May 26, 2026
Kind
B2
Abstract

Classification of cell data includes obtaining a target dataset and an artificial intelligence (AI) model trained to identify relationship(s) between cells of a row and classify whether a focus cell of the row is erroneous based on the identified relationship(s), and applying the AI model to the target dataset to identify erroneous cell(s) thereof. The applying includes selecting a row of cells of the target dataset, inputting the selected row of cells to the AI model with an identification of a focus cell, the focus cell to be classified by the AI model, classifying the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous, and outputting an indication of the classification of the focus cell.

Claims (70)

1 . A computer-implemented method comprising:

performing data modification on selected data cells of an initial dataset, wherein the data modification imparts errors in the selected data cells;

labeling the selected data cells as being erroneous, wherein the data modification and the labeling provide a training dataset with the errors imparted in the selected data cells,

training an artificial intelligence (AI) model using the training dataset to identify one or more relationships between cells of a row that inform a context of a selected focus cell of the row based at least in part on cells neighboring the selected focus cell, and classify whether the selected focus cell of the row is erroneous based at least in part on the identified one or more relationships and the informed context;

obtaining a target dataset, the target dataset comprising rows and columns of data cells,

applying the AI model to the target dataset to identify one or more erroneous cells of the target dataset, the applying comprising:

selecting a row of cells of the target dataset;

inputting the selected row of cells to the AI model with an identification of a focus cell of the selected row of cells, the focus cell to be classified by the AI model, wherein the inputting comprises:

building a string of cell data of a plurality of cells of the selected row of cells, the plurality of cells comprising the focus cell, and the string of cell data of the plurality of cells comprising cell data of the focus cell wherein the building provides a delimiter in the string between cell data of different cells of the selected row of cells; and

inputting to the AI model the string of cell data of the plurality of cells;

classifying, by the AI model, the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous; and

outputting an indication of the classification of the focus cell.

2 . The method of claim 1 , wherein the classifying comprises one selected from the group consisting of:

a binary classification that classifies the focus cell as being either erroneous or not erroneous; and

a multi-classification that identifies whether the focus cell is erroneous and, if so, one or more errors of the focus cell.

3 . The method of claim 1 , wherein the identification of the focus cell comprises an identification of the cell data of the focus cell within the string of cell data of the plurality of cells, and wherein the classifying the focus cell classifies the cell data of the focus cell.

4 . The method of claim 3 , wherein the identification of the cell data of the focus cell comprises insertion of a selected character into the string at one or more positions relative to the cell data of the focus cell, wherein the AI model is configured to identify the focus cell by locating the inserted selected character.

5 . The method of claim 1 , wherein the applying further comprises:

iterating, over one or more other cells of the selected row of cells, the inputting the selected row of cells, the classifying the focus cell, and the outputting an indication of the classification of the focus cell, wherein, at each iteration of the iterating, the applying:

identifies a next focus cell of the selected row of cells, the next focus cell being a different cell than each prior-identified and classified focus cell of the selected row of cells;

inputs the selected row of cells to the AI model with an identification of the next focus cell, the next focus cell to be classified by the AI model;

classifies the next focus cell to obtain a classification of the next focus cell, the classifying identifying whether the focus cell is erroneous; and

outputs an indication of the classification of the next focus cell.

6 . The method of claim 5 , wherein the applying further comprises iterating, over one or more other rows of the target dataset:

the selecting a row of cells;

the inputting the selected row of cells to the AI model with an identification of a focus cell;

the classifying the focus cell;

the outputting an indication of the classification of the focus cell; and

the iterating, over one or more other cells of the selected row of cells, the inputting the selected row of cells, the classifying the focus cell, and the outputting an indication of the classification of the focus cell, wherein at each iteration of the iterating over the one or more other rows of the target dataset, the applying identifies a next row of cells of the target dataset, the next row of cells being a different row of cells than each prior-selected row of cells, and wherein the identified next row of cells is the selected row of cells.

7 . The method of claim 1 , wherein the AI model is trained to identify the one or more relationships by way of a multi-head attention component that comprises a plurality of attention heads, and wherein the AI model is configured to classify the focus cell based at least in part on the informed context.

8 . The method of claim 1 , wherein the AI model comprises:

a plurality of encoder layers, each encoder layer of the plurality of encoder layers comprising a respective at least one attention layer and a feed-forward layer; and

a perceptron layer comprising an activation function configured to provide the output classification of the focus cell.

9 . The method of claim 1 , wherein the training dataset comprises at least the selected data cells labeled as erroneous and at least some cells labeled as correct.

10 . The method of claim 9 , wherein the training dataset is based on a first proper subset of a larger dataset and wherein the target dataset is a second proper subset of that larger dataset, the second proper subset being different from the first proper subset.

11 . The method of claim 1 , wherein the data modification on a selected data cell of the initial dataset comprises at least one selected from the group consisting of: random character replacement, random character insertion, random character deletion, random character swapping, value deletion, column value swapping, and row value swapping.

12 . The method of claim 1 , wherein the data modification imparts in a selected data cell of the initial dataset at least one error selected from the group consisting of: a typographical error, a value swap across columns error, a value violating data constraint error, a formatting error, and a missing value error.

13 . The method of claim 1 , further comprising performing automatic processing, the automatic processing comprising raising an electronic alert to a user that indicates one or more erroneous cells of the target dataset.

14 . A computer system comprising:

a memory; and

a processor in communication with the memory, wherein the computer system is configured to perform a method comprising:

performing data modification on selected data cells of an initial dataset, wherein the data modification imparts errors in the selected data cells;

labeling the selected data cells as being erroneous, wherein the data modification and the labeling provide a training dataset with the errors imparted in the selected data cells;

training an artificial intelligence (AI) model using the training dataset to identify one or more relationships between cells of a row that inform a context of a selected focus cell of the row based at least in part on cells neighboring the selected focus cell, and classify whether the selected focus cell of the row is erroneous based at least in part on the identified one or more relationships and the informed context;

obtaining a target dataset, the target dataset comprising rows and columns of data cells;

applying the AI model to the target dataset to identify one or more erroneous cells of the target dataset, the applying comprising:

selecting a row of cells of the target dataset;

inputting the selected row of cells to the AI model with an identification of a focus cell of the selected row of cells, the focus cell to be classified by the AI model, wherein the inputting comprises:

building a string of cell data of a plurality of cells of the selected row of cells, the plurality of cells comprising the focus cell, and the string of cell data of the plurality of cells comprising cell data of the focus cell, wherein the building provides a delimiter in the string between cell data of different cells of the selected row of cells; and

inputting to the AI model the string of cell data of the plurality of cells;

classifying, by the AI model, the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous; and

outputting an indication of the classification of the focus cell.

15 . The computer system of claim 14 , wherein the identification of the focus cell comprises an identification of the cell data of the focus cell within the string of cell data of the plurality of cells, wherein the identification of the cell data of the focus cell comprises insertion of a selected character into the string at one or more positions relative to the cell data of the focus cell, wherein the AI model is configured to identify the focus cell by locating the inserted selected character, and wherein the classifying the focus cell classifies the cell data of the focus cell.

16 . The computer system of claim 14 , wherein the AI model is trained to identify the one or more relationships by way of a multi-head attention component that comprises a plurality of attention heads, and wherein the AI model is configured to classify the focus cell based at least in part on the informed context.

17 . The computer system of claim 14 , wherein the training dataset comprises at least the selected data cells labeled as erroneous and at least some cells labeled as correct, wherein the training dataset is based on a first proper subset of a larger dataset and wherein the target dataset is a second proper subset of that larger dataset, the second proper subset being different from the first proper subset.

18 . A computer program product comprising:

a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising:

performing data modification on selected data cells of an initial dataset, wherein the data modification imparts errors in the selected data cells;

labeling the selected data cells as being erroneous, wherein the data modification and the labeling provide a training dataset with the errors imparted in the selected data cells;

training an artificial intelligence (AI) model using the training dataset to identify one or more relationships between cells of a row that inform a context of a selected focus cell of the row based at least in part on cells neighboring the selected focus cell, and classify whether the selected focus cell of the row is erroneous based at least in part on the identified one or more relationships and the informed context;

obtaining a target dataset, the target dataset comprising rows and columns of data cells;

applying the AI model to the target dataset to identify one or more erroneous cells of the target dataset, the applying comprising:

selecting a row of cells of the target dataset;

inputting the selected row of cells to the AI model with an identification of a focus cell of the selected row of cells, the focus cell to be classified by the AI model, wherein the inputting comprises:

building a string of cell data of a plurality of cells of the selected row of cells, the plurality of cells comprising the focus cell, and the string of cell data of the plurality of cells comprising cell data of the focus cell, wherein the building provides a delimiter in the string between cell data of different cells of the selected row of cells; and

inputting to the AI model the string of cell data of the plurality of cells;

classifying, by the AI model, the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous; and

outputting an indication of the classification of the focus cell.

19 . The computer program product of claim 18 , wherein the AI model is trained to identify the one or more relationships by way of a multi-head attention component that comprises a plurality of attention heads, and wherein the AI model is configured to classify the focus cell based at least in part on the informed context.

20 . The computer program product of claim 18 , wherein the training dataset comprises at least the selected data cells labeled as erroneous and at least some cells labeled as correct, wherein the training dataset is based on a first proper subset of a larger dataset and wherein the target dataset is a second proper subset of that larger dataset, the second proper subset being different from the first proper subset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: QUADER, SHAIKH SHAHRIAR; SAILLET, YANNICK; MAIER, ALBERT; LAZAR, REMUS; AL-SHAMALI, OMAR; MILLER, JAMES
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058148/0777 →
Continuity (1)
Related Publication 20230153566A1 · May 18, 2023
References Cited (48)
US 5943663A · Mouradian · 1999 [cited by applicant]
US 6766512B1 · Khosrowshahi · 2004 [cited by examiner]
US 10599954B2 · Sun · 2020 [cited by applicant]
US 11734937B1 · Pushkin · 2023 [cited by examiner]
US 20060155797A1 · Jeang et al. · 2006 [cited by applicant]
US 20130346844A1 · Graepel · 2013 [cited by examiner]
US 20150046785A1 · Byron · 2015 [cited by examiner]
US 20160019198A1 · Welton · 2016 [cited by examiner]
US 20180203836A1 · Singh · 2018 [cited by examiner]
US 20180285349A1 · Mineno et al. · 2018 [cited by applicant]
US 20190332966A1 · Gidney et al. · 2019 [cited by applicant]
US 20190354613A1 · Zoldi et al. · 2019 [cited by applicant]
US 20200302009A1 · Stegmaier · 2020 [cited by examiner]
US 20200349169A1 · Venkatesan et al. · 2020 [cited by applicant]
US 20210182476A1 · Cervelli et al. · 2021 [cited by applicant]
US 20210209297A1 · Dong · 2021 [cited by examiner]
US 20210256214A1 · Gretz et al. · 2021 [cited by applicant]
US 20210383191A1 · Ainslie · 2021 [cited by examiner]
US 20220342583A1 · Scott · 2022 [cited by examiner]
US 20220391414A1 · Baur · 2022 [cited by examiner]
CA 2243120C · 1997 [cited by applicant]
CN 1285556A · 2001 [cited by applicant]
CN 103150454A · 2013 [cited by applicant]
CN 107038157A · 2017 [cited by applicant]
CN 112817863A · 2021 [cited by applicant]
CN 113010503A · 2021 [cited by applicant]
CN 113411405A · 2021 [cited by applicant]
WO WO2020013956A1 · 2020 [cited by applicant]
WO 2023088109A1 · 2023 [cited by applicant]
Liu, et al., “Outlier Detection Algorithm Based on SOM Neural Network for Spatial Series Dataset”, 2018 Tenth International Conference on Advanced Computational Intelligence (ICACI), Mar. 2018, pp. 162-168. [cited by applicant]
Abutbul, et al., “DNF-Net: A Neural Architecture for Tabular Data”, arXiv:2006.06465v1 [cs.LG] Jun. 2020, 17 pgs. [cited by applicant]
Mell, et al., “The NIST Definition of Cloud Computing”, NIST Special Publication 800-145, Sep. 2011, Gaithersburg, MD, 7 pgs. [cited by applicant]
Singh, et al., “Melford: Using Neural Networks to Find Spreadsheet Errors”, Microsoft Tech Report No. MSR-TR-2017-5, Jan. 2017, 13 pgs. Retrieved on Nov. 8, 2021 from Internet URL: https://www.microsoft.com/en-us/resear… [cited by applicant]
Rekatsinas, et al., “HoloClean: holistic data repairs with probabilistic inference.” VLDB Endowment, Aug. 2017, 13 pgs. Retrieved on Nov. 9, 2021 from Internet URL: https://arxiv.org > pdf > 1702.00820.pdf. [cited by applicant]
Huang, et al., “Auto-Detect: Data-Driven Error Detection in Tables” in Proceedings of the 2018 International Conference on Management of Data, 2018, 16 pgs. Retrieved on Nov. 9, 2021 from Internet URL: https://www.micro… [cited by applicant]
Dallachiesa et al., “NADEEF: A Commodity Data Cleaning System,” in Proceedings of the 2013 International Conference on Management of Data—SIGMOD '13, 2013, 12 pgs., doi: 10.1145/2463676.2465327. Retrieved on Nov. 9, 202… [cited by applicant]
Schelter, et al., “Automating Large-Scale Data Quality Verification.” VLDB Endowment, Aug. 2018, pp. 1781-1794. Retrieved on Nov. 9, 2021 from Internet URL: https://ssc.io/publication/. [cited by applicant]
Rashid, et al., “Completeness and Consistency Analysis for Evolving Knowledge Bases,” J. Web Semant., vol. 54, Jan. 2019, 13 pgs. doi: 10.1016/j.websem.2018.11.004. Retrieved on Nov. 9, 2021 from Internet URL: https://a… [cited by applicant]
Saxena, H., et al., “Distributed Discovery of Functional Dependencies”, 2019 IEEE 35th International Conference on Data Engineering (ICDE), Apr. 2019, 4 pgs. Retrieved on Nov. 9, 2021 from Internet URL: https://cs.uwate… [cited by applicant]
Rezig, et al., “Pattern-Driven Data Cleaning”, ArXiv171209437 Cs, Dec. 2017, 13 pgs. Retrieved on Nov. 9, 2021 from Internet URL: https://arxiv.org/abs/1712.09437. [cited by applicant]
Qahtan, et al., “Pattern functional dependencies for data cleaning”, VLDB Endowment, Jan. 2020, 14 pgs. Retrieved on Nov. 9, 2021 from Internet URL: http://dspace.library.uu.nl/handle/1874/396267. [cited by applicant]
Reddy, et al., “Using Gaussian Mixture Models to Detect Outliers in Seasonal Univariate Network Traffic,” in 2017 IEEE Security and Privacy Workshops (SPW), May 2017, pp. 229-234. Retrieved on Nov. 9, 2021 from Internet… [cited by applicant]
Pit-Claudel, et al., “Outlier Detection in Heterogeneous Datasets Using Automatic Tuple Expansion”, Feb. 2016, 12 pgs. Retrieved on Nov. 9, 2021 from Internet URL: https://dspace.mit.edu > bitstream > handle > 1721.1 > … [cited by applicant]
Riahi, et al., “Model-based Exception Mining for Object-Relational Data,” Data Min. Knowl. Discov., vol. 34, No. 3, pp. 681-722, May 2020. Abstract retrieved on Nov. 9, 2021 from Internet URL: https://arxiv.org/abs/1807… [cited by applicant]
Liu, et al., “Generative Adversarial Active Learning for Unsupervised Outlier Detection”, IEEE Trans. Knowl. Data Eng., pp. 1-1, 2019, doi: 10.1109/TKDE.2019.2905606. Retrieved on Nov. 9, 2021 from Internet URL: https:/… [cited by applicant]
Vaswani, et al., “Attention is All you Need”, in Advances in Neural Information Processing Systems 30, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, … [cited by applicant]
Wei, et al., “EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks”, arXiv190111196 Cs, Aug. 2019, 9 pgs. Retrieved on Nov. 9, 2021 from Internet URL: https://arxiv.org/abs/1901.1… [cited by applicant]
International Search Report and Written Opinion for PCT/CN2022/129891 completed Jan. 18, 2023, 13 pgs. [cited by applicant]