IP Library Granted Patent US 11,481,667
Granted Patent B2
US 11,481,667 · App. 16/255,885 · Granted Oct 25, 2022

Classifier confidence as a means for identifying data drift

Inventors: Orna Raz (Haifa, IL); Marcel Zalmanovici (Kiriat Motzkin, IL); Aviad Zlotnick (Mitzpeh Netofah, IL)
Assignee: International Business Machines Corporation
G06N20/00G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,667
App. No.
16/255,885
Granted
Oct 25, 2022
Kind
B2
Abstract

Embodiments of the present systems and methods may provide improved machine learning performance even though data drift has occurred. For example, a method may comprise providing a machine learning model in a computer system, operating the machine learning model using a first dataset to obtain results of the first dataset, operating the machine learning model using a second dataset to obtain results of the second dataset, performing statistical testing on a confidence distribution of results of the first dataset and of results of the second dataset to determine a difference in a result confidence distribution between the first dataset and of the second dataset, and determining whether data included in the second dataset has data drift relative to the first dataset based on the difference in a result confidence distribution between the first dataset and of the second dataset.

Claims (25)

1. A method comprising:

providing a non-binary machine learning classification model in a computer system comprising a processor, memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor;

operating, at the computer system, the non-binary machine learning classification model using a first dataset to obtain a plurality of labels of the first dataset;

operating, at the computer system, the non-binary machine learning classification model using a second dataset to obtain a plurality of labels of the second dataset, wherein second data set simulates data drift relative to the first data set by the plurality of labels of the second data set having at least one more label than the plurality of labels of the first data set;

performing, at the computer system, a non-parametric statistical test for identity, over a classifier confidence-per-label distribution for at least one label of the plurality of labels of the first dataset and of the second dataset to determine a difference in a resulting confidence-per-label distribution between the first dataset and of the second dataset;

determining, at the computer system, whether the second dataset has data drift relative to the first dataset based on the difference in the resulting confidence-per-label distribution between the first dataset and of the second dataset; and

training the non-binary machine learning classification model based on the determination of whether data drift has occurred from the first data set to the second data set.

2. The method of claim 1 , wherein the first dataset comprises at least a portion of a training dataset and the second dataset comprises as least a portion of a production dataset.

3. The method of claim 1 , wherein the first dataset comprises a first portion of a production dataset and the second dataset comprises a second portion of the production dataset.

4. A system comprising a processor, memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor to perform:

operating a non-binary machine learning classification model using a first dataset to obtain a plurality of labels of the first dataset;

operating the non-binary machine learning classification model using a second dataset to obtain a plurality of labels of the second dataset, wherein second data set simulates data drift relative to the first data set by the plurality of labels of the second data set having at least one more label than the plurality of labels of the first data set;

performing a non-parametric statistical test for identity, over a classifier confidence-per-label distribution for at least one label of the plurality of labels of the first dataset and of the second dataset to determine a difference in a resulting confidence-per-label distribution between the first dataset and of the second dataset;

determining whether the second dataset has data drift relative to the first dataset based on the difference in the resulting confidence-per-label distribution between the first dataset and of the second dataset; and

training the non-binary machine learning classification model based on the determination of whether data drift has occurred from the first data set to the second data set.

5. The system of claim 4 , wherein the first dataset comprises at least a portion of a training dataset and the second dataset comprises as least a portion of a production dataset.

6. The system of claim 4 , wherein the first dataset comprises a first portion of a production dataset and the second dataset comprises a second portion of the production dataset.

7. A computer program product comprising a non-transitory computer readable storage having program instructions embodied therewith, the program instructions executable by a computer, to cause the computer to perform a method comprising:

operating a non-binary machine learning classification model using a first dataset to obtain a plurality of labels of the first dataset;

operating the non-binary machine learning classification model using a second dataset to obtain a plurality of labels of the second dataset, wherein second data set simulates data drift relative to the first data set by the plurality of labels of the second data set having at least one more label than the plurality of labels of the first data set;

performing a non-parametric statistical test for identity, over a classifier confidence-per-label distribution for at least one label of the plurality of labels of the first dataset and of the second dataset to determine a difference in a resulting confidence-per-label distribution between the first dataset and of the second dataset;

determining whether the second dataset has data drift relative to the first dataset based on the difference in the resulting confidence-per-label distribution between the first dataset and of the second dataset; and

training the non-binary machine learning classification model based on the determination of whether data drift has occurred from the first data set to the second data set.

8. The computer program product of claim 7 , wherein the first dataset comprises at least a portion of a training dataset and the second dataset comprises as least a portion of a production dataset.

9. The computer program product of claim 7 , wherein the first dataset comprises a first portion of a production dataset and the second dataset comprises a second portion of the production dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2019
From: RAZ, ORNA; ZALMANOVICI, MARCEL; ZLOTNICK, AVIAD
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048117/0156 →
Continuity (1)
Related Publication 20200242505A1 · Jul 30, 2020