IP Library › Granted Patent US 11,663,515
Granted Patent B2
US 11,663,515 · App. 16/059,700 · Granted May 30, 2023

Machine learning classification with model quality prediction

Inventor: Baskar Jayaraman (Fremont, CA)
Assignee: ServiceNow, Inc.
G06F18/2148G06F18/211G06F18/217G06F18/40G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,663,515
App. No.
16/059,700
Granted
May 30, 2023
Kind
B2
Abstract

An embodiment may include a machine learning based classifier that maps input observations into respective categories and a database containing a corpus of training data for the classifier. The training data includes a plurality of entries, each entry having an observation respectively associated with a ground truth category thereof. A computing device may be configured to select, from the training data, a plurality of subsets each containing a different number of entries. The computing device may also be configured to, for each particular subset: (i) divide the particular subset into a training portion and a validation portion, (ii) train the classifier with the training portion, (iii) provide the validation portion as input to the classifier as trained, and (iv) based on how entries of the validation portion are mapped to the categories, determine a respective precision for the particular subset.

Claims (42)

1. A computing system comprising:

a machine learning based classifier that maps input observations into respective categories, wherein the observations include textual descriptions of problems related to information technology usage, and wherein the categories include types of problems related to information technology usage; and

a computing device configured to:

select a plurality of subsets of training data from a corpus of the training data, wherein the corpus of the training data includes a plurality of entries, each entry having an observation respectively associated with a ground truth category of the observation, and wherein each subset of the plurality of subsets of the training data contains a different number of entries;

for each subset of the plurality of subsets of the training data: (i) divide the subset into a training portion and a validation portion, (ii) train the classifier with the training portion, (iii) provide the validation portion as input to the classifier as trained, and (iv) based on how entries of the validation portion are mapped to the categories, determine a respective precision for the subset, wherein a largest subset of the plurality of subsets includes all of the entries in the corpus and has a particular precision;

determine, from the plurality of subsets, one or more subsets that have respective precisions that are no more than a pre-determined amount lower than the particular precision; and

recommend, from the one or more subsets, a particular subset having a smallest number of entries to use in training the classifier for a production environment.

2. The computing system of claim 1 , wherein the respective precision for the subset is calculated as a percentage of all entries of the validation portion that were mapped to their ground truth categories.

3. The computing system of claim 1 , wherein the respective precision for the subset is calculated as a percentage of entries of the validation portion associated with a particular ground truth category that were mapped to the particular ground truth category.

4. The computing system of claim 1 , wherein the computing device is configured to:

train the classifier using the recommended particular subset of the plurality of subsets of the training data; and

deploy the classifier as trained into the production environment.

5. The computing system of claim 1 , wherein the largest subset has a highest precision of any of the plurality of subsets of the training data, and wherein the computing device is configured to:

determine, from the plurality of subsets of the training data, one or more particular subsets that have precisions that are no more than a particular pre-determined amount lower than the highest precision; and

recommend, from the one or more particular subsets, a subset with a smallest number of entries.

6. The computing system of claim 1 , wherein the computing device is configured to generate, for display on a graphical user interface of a client device, a representation of a graph that plots the number of entries in each of the plurality of subsets of the training data versus the respective precision for each of the plurality of subsets of the training data, and wherein the representation of the graph plots the number of entries in each of the plurality of subsets of the training data on an x-axis and plots the respective precision for each of the plurality of subsets of the training data on a y-axis.

7. The computing system of claim 1 , wherein the computing device is configured to generate, for display on a graphical user interface of a client device, a representation of a graph that plots the number of entries in each of the plurality of subsets of the training data versus the respective precision for each of the plurality of subsets of the training data, wherein the graphical user interface allows selection of one or more of the categories, and wherein the computing device is configured to:

in response to receiving a selection of any of the categories, generate, for display on the graphical user interface, a second representation of a second graph that plots the number of entries in each of the plurality of subsets of the training data versus the respective precision of the category for each of the plurality of subsets of the training data.

8. The computing system of claim 1 , wherein the computing system is disposed within a computational instance of a remote network management platform, and wherein the computational instance is configured to remotely manage a particular managed network.

9. A computer-implemented method comprising:

selecting, by a computing device, a plurality of subsets of training data from a corpus of the training data, wherein the corpus of the training data includes a plurality of entries, each entry having an observation respectively associated with a ground truth category of the observation, wherein each of the observations includes a textual description of a problem related to information technology usage, wherein the ground truth category include a type of problems related to information technology usage, and wherein each of the plurality of subsets of the training data contains a different number of entries;

for each subset of the plurality of subsets of the training data, the computing device: (i) dividing the subset into a training portion and a validation portion, (ii) training a machine learning based classifier with the training portion, wherein the classifier maps input observations into respective categories, (iii) providing the validation portion as input to the classifier as trained, and (iv) based on how entries of the validation portion are mapped to the categories, determining a respective precision for the subset, wherein a largest subset of the plurality of subsets includes all of the entries in the corpus and has a particular precision;

determining, by the computing device, from the plurality of subsets, one or more subsets that have respective precisions that are no more than a pre-determined amount lower than the particular precision; and

recommending, by the computing device, from the one or more subsets, a particular subset having a smallest number of entries to use in training the classifier for a production environment.

10. The computer-implemented method of claim 9 , wherein the respective precision for the subset is calculated as a percentage of all entries of the validation portion that were mapped to their ground truth categories.

11. The computer-implemented method of claim 9 , wherein the respective precision for the subset is calculated as a percentage of entries of the validation portion associated with a particular ground truth category that were mapped to the particular ground truth category.

12. The computer-implemented method of claim 9 , wherein the largest subset has a highest precision of any of the plurality of subsets of the training data, and comprising:

determining, from the plurality of subsets, one or more particular subsets that have precisions that are no more than a particular pre-determined amount lower than the highest precision; and

recommending, from the one or more particular subsets, a subset with a smallest number of entries.

13. The computer-implemented method of claim 9 , comprising:

training the classifier using the particular subset of the plurality of subsets of the training data; and

deploying the classifier as trained into the production environment.

14. The computer-implemented method of claim 9 , comprising:

generating, for display on a graphical user interface of a client device, a representation of a graph that plots the number of entries in each of the plurality of subsets of the training data versus the respective precision for each of the plurality of subsets of the training data.

15. The computer-implemented method of claim 14 , wherein the representation of the graph plots the number of entries in each of the plurality of subsets of the training data on an x-axis and plots the respective precision for each of the plurality of subsets of the training data on a y-axis.

16. The computer-implemented method of claim 14 , wherein the graphical user interface allows selection of one or more of the categories, the method comprising:

in response to receiving a selection of any of the categories, generating, for display on the graphical user interface, a second representation of a second graph that plots the number of entries in each of the plurality of subsets of the training data versus the respective precision of the category for each of the plurality of subsets of the training data.

17. An article of manufacture including a non-transitory computer-readable medium, having stored thereon program instructions that, upon execution by a computing system, cause the computing system to perform operations comprising:

selecting a plurality of subsets of training data from a corpus of the training data, wherein the corpus of the training data includes a plurality of entries, each entry having an observation respectively associated with a ground truth category of the observation, wherein each of the observations includes a textual description of a problem related to information technology usage, wherein the ground truth category include a type of problems related to information technology usage, and wherein each of the plurality of subsets of the training data contains a different number of entries;

for each subset of the plurality of subsets of the training data: (i) dividing the subset into a training portion and a validation portion, (ii) training a machine learning based classifier with the training portion, wherein the classifier maps input observations into respective categories, (iii) providing the validation portion as input to the classifier as trained, and (iv) based on how entries of the validation portion are mapped to the categories, determining a respective precision for the subset, wherein a largest subset of the plurality of subsets includes all of the entries in the corpus and has a particular precision;

determining, from the plurality of subsets, one or more subsets that have respective precisions that are no more than a pre-determined amount lower than the particular precision; and

recommending, from the one or more subsets, a particular subset having a smallest number of entries to use in training the classifier for a production environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2018
From: JAYARAMAN, BASKAR
To: SERVICENOW, INC.
Reel/Frame 046608/0023 →
Continuity (1)
Related Publication 20200050896A1 · Feb 13, 2020