IP Library Granted Patent US 8,856,123
Granted Patent B1
US 8,856,123 · App. 11/780,803 · Granted Oct 7, 2014

Document classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,856,123
App. No.
11/780,803
Granted
Oct 7, 2014
Kind
B1
Abstract

Provided are, among other things, systems, methods and techniques for classifying a collection of documents. A term is identified based on an indication of ability of the term's presence within a given document to predict whether the given document should be classified into an identified category. A document index is then queried using the identified term and, in response, search results that define a candidate set of documents are received. Finally, a classifier is applied to documents within the candidate set to determine which of the documents should be classified into the identified category.

Claims (28)

1. A method comprising:

for a given category, receiving a positive set of training documents within the given category and a negative set of training documents not within the given category, by a processor of a computing device;

using a feature selector on the positive set and the negative set to determine a first set of features that are predictive for the given category, by the processor, each feature comprising a word or a phrase of words;

training a first classifier from the positive set and the negative set to assign a weight to each feature of the first set of features, by the processor;

after training the first classifier, querying a document index of a plurality of production documents for each feature of the first set of features to yield a sub-plurality of the production documents that are likely but not necessarily within the given category, by the processor, by formulating a query that includes the first set of features as weighted by the weights thereof to locate the sub-plurality of the production documents, where each of the sub-plurality of the production documents resulting from querying the document index has a total sum of the weights of the first set of features greater than a threshold; and

applying a second classifier that uses a second set of features greater in number than the first set to determine whether each production document of the sub-plurality is predicted to be within the given category, by the processor,

wherein the document index is queried to decrease a number of the production documents against which the second classifier is applied to just the sub-plurality of the production documents yielded by querying the document index,

and wherein the first classifier is one of a Naïve Bayesian classifier, a support vector machine classifier, and a logistic regression classifier.

2. The method of claim 1 , further comprising training the second classifier from the positive set and the negative set, by the processor.

3. The method of claim 1 , further comprising outputting the production documents of the sub-plurality that are predicted to be within the given category, by the processor.

4. The method of claim 1 , wherein the sub-plurality of the documents resulting from querying the document index includes the production documents that contain any of the first set of features.

5. The method of claim 1 , wherein some of the sub-plurality of production documents are not within the given category.

6. The method of claim 1 , wherein the sub-plurality of production documents is smaller in number than the plurality of production documents.

7. The method of claim 6 , wherein the sub-plurality of production documents that are within the given category is smaller in number than a total number of the sub-plurality of production documents.

8. A non-transitory machine-readable medium storing machine-executable process steps for classifying a collection of documents, said process steps comprising:

for a given category, receiving a positive set of training documents within the given category and a negative set of training documents not within the given category;

using a feature selector on the positive set and the negative set to determine a first set of features that are predictive for the given category, each feature comprising a word or a phrase of words;

training a first classifier from the positive set and the negative set to assign a weight to each feature of the first set of features;

after training the first classifier, querying a document index of a plurality of production documents for each feature of the first set of features to yield a sub-plurality of the production documents that are likely but not necessarily within the given category, by formulating a query that includes the first set of features as weighted by the weights thereof to locate the sub-plurality of the production documents, where each of the sub-plurality of the production documents resulting from querying the document index has a total sum of the weights of the first set of features greater than a threshold; and

applying a second classifier that uses a second set of features greater in number than the first set to determine whether each production document of the sub-plurality is predicted to be within the given category,

wherein the document index is queried to decrease a number of the production documents against which the second classifier is applied to just the sub-plurality of the production documents yielded by querying the document index,

and wherein the first classifier is one of a Naïve Bayesian classifier, a support vector machine classifier, and a logistic regression classifier.

9. The non-transitory machine-readable medium of claim 8 , wherein said process steps further comprise training the second classifier from the positive set and the negative set.

10. The non-transitory machine-readable medium of claim 8 , wherein said process steps further comprise outputting the production documents of the sub-plurality that are predicted to be within the given category.

11. The non-transitory machine-readable medium of claim 8 , wherein the sub-plurality of the documents resulting from querying the document index includes the production documents that contain any of the first set of features.

12. The non-transitory machine-readable medium of claim 8 , wherein some of the sub-plurality of production documents are not within the given category.

13. The non-transitory machine-readable medium of claim 8 , wherein the sub-plurality of production documents is smaller in number than the plurality of production documents.

14. The non-transitory machine-readable medium of claim 13 , wherein the sub-plurality of production documents that are within the given category is smaller in number than a total number of the sub-plurality of production documents.

Assignments (12)
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 063546/0181) Recorded Jun 21, 2024
From: BARCLAYS BANK PLC
To: MICRO FOCUS LLC
Reel/Frame 067807/0076 →
SECURITY INTEREST Recorded Aug 30, 2023
From: MICRO FOCUS LLC
To: THE BANK OF NEW YORK MELLON
Reel/Frame 064760/0862 →
SECURITY INTEREST Recorded May 4, 2023
From: MICRO FOCUS LLC
To: BARCLAYS BANK PLC
Reel/Frame 063546/0181 →
SECURITY INTEREST Recorded May 4, 2023
From: MICRO FOCUS LLC
To: BARCLAYS BANK PLC
Reel/Frame 063546/0190 →
SECURITY INTEREST Recorded May 4, 2023
From: MICRO FOCUS LLC
To: BARCLAYS BANK PLC
Reel/Frame 063546/0230 →
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0577 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC)
Reel/Frame 063560/0001 →
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0718 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC); BORLAND SOFTWARE CORPORATION; MICRO FOCUS (US), INC.; SERENA SOFTWARE, INC; ATTACHMATE CORPORATION; MICRO FOCUS SOFTWARE INC. (F/K/A NOVELL, INC.); NETIQ CORPORATION
Reel/Frame 062746/0399 →
CHANGE OF NAME Recorded Aug 8, 2019
From: ENTIT SOFTWARE LLC
To: MICRO FOCUS LLC
Reel/Frame 050004/0001 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ENTIT SOFTWARE LLC; ARCSIGHT, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0577 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ENTIT SOFTWARE LLC; ATTACHMATE CORPORATION; BORLAND SOFTWARE CORPORATION; NETIQ CORPORATION; MICRO FOCUS (US), INC.; MICRO FOCUS SOFTWARE, INC.; ARCSIGHT, LLC; SERENA SOFTWARE, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0718 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2017
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
To: ENTIT SOFTWARE LLC
Reel/Frame 042746/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2007
From: FORMAN, GEORGE
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 019590/0617 →