IP Library › Granted Patent US 11,914,629
Granted Patent B2
US 11,914,629 · App. 17/394,994 · Granted Feb 27, 2024

Practical supervised classification of data sets

Inventors: Arunav Mishra (Ludwigshafen, DE); Henning Schwabe (Ludwigshafen, DE); Lalita Shaki Uribe Ordonez (Ludwigshafen, DE)
Assignee: BASF SE
G06F16/35G06F16/338G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,914,629
App. No.
17/394,994
Granted
Feb 27, 2024
Kind
B2
Abstract

The present invention relates to information retrieval. In order to facilitate a search and identification of documents, there is provided a computer-implemented method for training a classifier model for data classification in response to a search query. The computer-implemented method comprises: a) obtaining a dataset that comprises a seed set of labeled data representing a training dataset; b) training the classifier model by using the training dataset to fit parameters of the classifier model; c) evaluating a quality of the classifier model using a test dataset that comprises unlabeled data from the obtained dataset to generate a classifier confidence score indicative of a probability of correctness of the classifier model working on the test dataset; d) determining a global risk value of misclassification and a reward value based on the classifier confidence score on the test dataset; e) iteratively updating the parameters of the classifier model and performing steps b) to d) until the global risk value falls within a predetermined risk limit value or an expected reward value is reached.

Claims (47)

1. A computer-implemented method for training a classifier model for data classification, in particular in response to a search query, comprising:

a) obtaining a dataset that comprises a seed set of labeled data representing a training dataset;

b) training the classifier model by using the training dataset to fit parameters of the classifier model;

c) evaluating a quality of the classifier model using a test dataset that comprises unlabeled data from the obtained dataset to generate a classifier confidence score indicative of a probability of correctness of the classifier model working on the test dataset;

d) determining a global risk value of misclassification and a reward value based on the classifier confidence score on the test dataset;

e) iteratively updating the parameters of the classifier model and performing steps b) to d) until the global risk value falls within a predetermined risk limit value or an expected reward value is reached to obtain a trained classifier model for data classification,

wherein step d) further comprises:

d1) generating a classifier confidence score indicative of a probability of correctness of the classifier model working on the test dataset;

d2) computing a classifier metric at different thresholds on classifier confidence score, the classifier metrics representing a measure of a test's accuracy;

d3) determining a reference threshold that corresponds to a peak in a distribution of the classifier metric over the threshold on classifier confidence score;

d4) determining a threshold range that defines a recommended window according to a predefined criteria, wherein the reference threshold is located within the threshold range; and

d5) computing the reward value at different thresholds on classifier confidence score, and

wherein the reward value includes at least one of a measure of information gain and a measure of decrease in uncertainty.

2. The computer-implemented method according to claim 1 ,

wherein the classifier metrics comprises F 2 score.

3. The computer-implemented method according to claim 1 ,

wherein step d) further comprises:

d6) receiving a user-defined threshold value that yields a desirable risk-reward pair for a use case based on a comparison of the global risk value and the reward value at the user-defined threshold value.

4. The computer-implemented method according to claim 1 ,

wherein step a) further comprises:

a1) performing a search to find seed documents based on the search query;

a2) selecting and annotating relevant and irrelevant seed documents; and

a3) repeating steps a1) and a2) until the seed set has sufficient labeled seed documents.

5. The computer-implemented method according to claim 1 ,

wherein step a) further comprises:

a4) augmenting the seed set of labeled data by a data augmentation technique; and

wherein in step b), the classifier model is trained using the augmented seed set.

6. The computer-implemented method according to claim 5 ,

wherein the data augmentation technique comprises data set expansion through sampling by similarity and active learning.

7. The computer-implemented method according to claim 1 ,

wherein the global risk value and the reward value are expressed as an algebraic expression in an objective function.

8. The computer-implemented method according to claim 1 ,

wherein the data comprises at least one of: text data; image data; experimental data from chemical, biological, and/or physical experiments; plant operations data, business operations data; and machine-generated data in log files.

9. A non-transitory computer readable medium having stored thereon a program element of, which when being executed by a processor is configured to carry out a computer-implemented method for training a classifier model for data classification, in particular in response to a search query, the computer-implemented method comprising:

a) obtaining a dataset that comprises a seed set of labeled data representing a training dataset;

b) training the classifier model by using the training dataset to fit parameters of the classifier model;

c) evaluating a quality of the classifier model using a test dataset that comprises unlabeled data from the obtained dataset to generate a classifier confidence score indicative of a probability of correctness of the classifier model working on the test dataset;

d) determining a global risk value of misclassification and a reward value based on the classifier confidence score on the test dataset;

e) iteratively updating the parameters of the classifier model and performing steps b) to d) until the global risk value falls within a predetermined risk limit value or

an expected reward value is reached to obtain a trained classifier model for data classification, wherein step d) further comprises:

d1) generating a classifier confidence score indicative of a probability of correctness of the classifier model working on the test dataset;

d2) computing a classifier metric at different thresholds on classifier confidence score, the classifier metrics representing a measure of a test's accuracy;

d3) determining a reference threshold that corresponds to a peak in a distribution of the classifier metric over the threshold on classifier confidence score;

d4) determining a threshold range that defines a recommended window according to a predefined criteria, wherein the reference threshold is located within the threshold range; and

d5) computing the reward value at different thresholds on classifier confidence score, and

wherein the reward value includes at least one of a measure of information gain and a

measure of decrease in uncertainty.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2023
From: MISHRA, ARUNAV; SCHWABE, HENNING; URIBE ORDONEZ, LALITA SHAKI
To: BASF SE
Reel/Frame 065702/0902 →
Priority Claims (1)
EP 20190061 · Aug 7, 2020 · regional
Continuity (1)
Related Publication 20220043850A1 · Feb 10, 2022