IP Library › Granted Patent US 11,586,987
Granted Patent B2
US 11,586,987 · App. 16/293,225 · Granted Feb 21, 2023

Dynamically updated text classifier

Inventors: Aron Szanto (New York, NY); Sireesh Gururaja (New York, NY); Domenic Puzio (Arlington, VA)
Assignee: Kensho Technologies, LLC
G06N20/20G06F40/30G06K9/6267
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,586,987
App. No.
16/293,225
Granted
Feb 21, 2023
Kind
B2
Abstract

Methods and systems for dynamically updating machine learning models such as text classifiers. One of the methods includes: receiving first data; producing a first machine learning model using the first data; releasing the first machine learning model for use; receiving second data after receipt of the first data; determining that the second data has a difference metric relative to the first data that exceeds a difference threshold; retraining the first machine learning model using at least part of the second data, the retraining producing a second machine learning model; and releasing the second machine learning model for use.

Claims (48)

1. A method comprising:

receiving first data;

producing a first machine learning model using the first data;

releasing the first machine learning model for use;

receiving second data after receipt of the first data;

determining that the second data has a difference metric relative to the first data that exceeds a difference threshold, wherein the first machine learning model is a first text classifier and the second machine learning model is a second text classifier, wherein the first data are first documents and the second data are second documents, and wherein determining that the second data has a difference metric relative to the first data that exceeds a difference threshold comprises determining that the second documents have a difference metric relative to the first documents that exceeds a difference threshold based at least in part on the words used in the second documents:

determining, at a concept labeling engine, at least one concept for a document and a concept label confidence score;

determining, at an active learning engine, whether a document labeled by the concept labeling engine should be assigned a ground-truth label based at least in part on whether a concept label confidence score is below a threshold;

retraining the first machine learning model using at least part of the second data, the retraining producing a second machine learning model; and

releasing the second machine learning model for use.

2. The method of claim 1 , wherein releasing the first machine learning model for use comprises:

determining a performance metric for the first machine learning model; and

determining that the performance metric for the first machine learning model exceeds a performance threshold.

3. The method of claim 2 , wherein using the first data to produce a first machine learning model comprises using at least a portion of the first data to produce a plurality of machine learning models and wherein releasing the first machine learning model for use comprises selecting a first machine learning model to release, from the plurality of machine learning models, based at least in part on the performance metric.

4. A method comprising:

receiving first data;

using the first data to produce a first machine learning model;

forwarding the first machine learning model for use prior to receiving post release data;

receiving post release data after the first machine learning model has been released, the first machine learning model having been trained using first data;

determining that the post release data has a difference metric relative to the first data that exceeds a difference threshold, wherein the first machine learning model is a first text classifier and the second machine learning model is a second text classifier, wherein the first data are first documents and the post release data are second documents, and wherein determining that the post release data has a difference metric relative to the first data that exceeds a difference threshold comprises determining that the second documents have a difference metric relative to the first documents that exceeds a difference threshold based at least in part on the words used in the second documents;

determining, at a concept labeling engine, at least one concept for a document and a concept label confidence score;

determining, at an active learning engine, whether a document labeled by the concept labeling engine should be assigned a ground-truth label based at least in part on whether a concept label confidence score is below a threshold;

retraining the first machine learning model to produce a second machine learning model using at least part of the post release data; and

forwarding the second machine learning model for use.

5. The method of claim 4 , wherein receiving first data comprises automatically sampling data to produce sampled first data and wherein using the first data to produce a first machine learning model comprises using the sampled first data to produce the first machine learning model.

6. The method of claim 5 , wherein receiving first data comprises augmenting the sampled first data to produce augmented sampled first data and wherein using the first data to produce a first machine learning model comprises using the augmented sampled first data to produce the first machine learning model.

7. The method of claim 4 , wherein the method further comprises selecting data from the post release data to use at least in part in retraining the first machine learning model.

8. A system comprising:

one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving first data;

using the first data to produce a first machine learning model;

releasing the first machine learning model for use;

receiving second data after receipt of the first data;

determining that the second data has a difference metric relative to the first data that exceeds a difference threshold, wherein the first machine learning model is a first text classifier and the second machine learning model is a second text classifier, wherein the first data are first documents and the second data are second documents, and wherein determining that the post first release data has a difference metric relative to the first data that exceeds a difference threshold comprises determining that the second documents have a difference metric relative to the first documents that exceeds a difference threshold based at least in part on the words used in the second documents;

determining, at a concept labeling engine, at least one concept for a document and a concept label confidence score;

determining, at an active learning engine, whether a document labeled by the concept labeling engine should be assigned a ground-truth label based at least in part on whether a concept label confidence score is below a threshold;

retraining the first machine learning model using at least part of the second data, the retraining producing a second machine learning model; and

releasing the second machine learning model for use.

9. A method comprising:

receiving first documents;

using the first documents to train a first text classifier; and

forwarding the first text classifier for use prior to receiving post release documents;

receiving the post release documents after the first text classifier has been released, the first text classifier having been trained using first documents;

determining that the post release documents have a difference metric relative to the first documents that exceeds a difference threshold, wherein the first text classifier comprises at least one of data describing a concept and data predictive of the concept, and wherein determining that the post release documents have a difference metric relative to the first documents that exceeds a difference threshold comprises determining that the post release documents have a difference metric relative to the first documents that exceeds a difference threshold based at least in part on the words used in the post release documents;

determining, at a concept labeling engine, at least one concept for a document and a concept label confidence score;

determining, at an active learning engine, whether a document labeled by the concept labeling engine should be assigned a ground-truth label based at least in part on whether a concept label confidence score is below a threshold;

retraining the first text classifier to produce a second text classifier using at least part of the post release documents; and

forwarding the second text classifier for use.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2019
From: SZANTO, ARON; GURURAJA, SIREESH; PUZIO, DOMENIC
To: KENSHO TECHNOLOGIES, LLC
Reel/Frame 048664/0363 →
Continuity (1)
Related Publication 20200286002A1 · Sep 10, 2020
Cited By (1)
US 12,412,393