IP Library Granted Patent US 9,235,812
Granted Patent B2
US 9,235,812 · App. 13/693,075 · Granted Jan 12, 2016

System and method for automatic document classification in ediscovery, compliance and legacy information clean-up

Inventor: Johannes Cornelis Scholtes (Bussum, NL)
Assignee: MSC INTELLECTUAL PROPERTIES B.V.
G06N99/005G06F17/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,235,812
App. No.
13/693,075
Granted
Jan 12, 2016
Kind
B2
Abstract

A system, method and computer program product for automatic document classification, including an extraction module configured to extract structural, syntactical and/or semantic information from a document and normalize the extracted information; a machine learning module configured to generate a model representation for automatic document classification based on feature vectors built from the normalized and extracted semantic information for supervised and/or unsupervised clustering or machine learning; and a classification module configured to select a non-classified document from a document collection, and via the extraction module extract normalized structural, syntactical and/or semantic information from the selected document, and generate via the machine learning module a model representation of the selected document based on feature vectors, and match the model representation of the selected document against the machine learning model representation to generate a document category, and/or classification for display to a user.

Claims (27)

1. A computer implemented system for automatic document classification, the system comprising:

an extraction module configured to extract structural, syntactical and semantic information from a document and normalize the extracted information;

a machine learning module configured to generate a model representation for automatic document classification based on feature vectors built from the normalized and extracted semantic information for supervised and unsupervised clustering and machine learning; and

a classification module configured to select a non-classified document from a document collection, and via the extraction module extract normalized structural, syntactical and semantic information from the selected document, and generate via the machine learning module a model representation of the selected document based on feature vectors, and match the model representation of the selected document against the machine learning model representation to generate a document category, and classification for display to a user,

wherein the extracted information includes named entities, properties of entities, noun-phrases, facts, events, and concepts.

2. The system of claim 1 , wherein extraction module employs text-mining, language identification, gazetteers, regular expressions, noun-phrase identification with part-of-speech taggers, and statistical models and rules, and is configured to identify patterns, and

the patterns include libraries, and algorithms shared among cases, and which can be tuned for a specific case, to generate case-specific semantic information.

3. The system of claim 1 , wherein the extracted information is normalized by using normalization rules, groupers, thesauri, taxonomies, and string-matching algorithms.

4. The system of claim 1 , wherein the model representation of the document is a term frequency-inverse document frequency (TF-IDF) document representation of the extracted information, and the clustering and machine learning includes a classifier based on decision trees, support vector machines (SVM), naïve-bayes classifiers, k-nearest neighbors, rules-based classification, Linear discriminant analysis (LDA), Maximum Entropy Markov Model (MEMM), scatter-gather clustering, and hierarchical agglomerate clustering (HAC).

5. A computer implemented method for automatic document classification, the method comprising:

extracting with an extraction module structural, syntactical and semantic information from a document and normalizing with the extraction module the extracted information;

generating with a machine learning module a model representation for automatic document classification based on feature vectors built from the normalized and extracted semantic information for supervised and unsupervised clustering and machine learning; and

selecting with a classification module a non-classified document from a document collection, and extracting via the extraction module normalized structural, syntactical and semantic information from the selected document, and generating via the machine learning module a model representation of the selected document based on feature vectors, and matching with the classification module the model representation of the selected document against the machine learning model representation and generating with the classification module a document category, and classification for display to a user,

wherein the extracted information includes named entities, properties of entities, noun-phrases facts, events, and concepts.

6. The method of claim 5 , wherein extraction module employs text-mining, language identification, gazetteers, regular expressions, noun-phrase identification with part-of-speech taggers, and statistical models and rules, and is configured to identify patterns, and

the patterns include libraries, and algorithms shared among cases, and which can be tuned for a specific case, to generate case-specific semantic information.

7. The method of claim 5 , wherein the extracted information is normalized by using normalization rules, groupers, thesauri, taxonomies, and string-matching algorithms.

8. The method of claim 5 , wherein the model representation of the document is a term frequency-inverse document frequency (TF-IDF) document representation of the extracted information, and the clustering and machine learning includes a classifier based on decision trees, support vector machines (SVM), naïve-bayes classifiers, k-nearest neighbors, rules-based classification, Linear discriminant analysis (LDA), Maximum Entropy Markov Model (MEMM), scatter-gather clustering, and hierarchical agglomerate clustering (HAC).

9. A computer program product for automatic document classification and including one or more computer readable instructions embedded on a tangible, non-transitory computer readable medium and configured to cause one or more computer processors to perform the steps of:

extracting with an extraction module structural, syntactical and semantic information from a document and normalizing with the extraction module the extracted information;

generating with a machine learning module a model representation for automatic document classification based on feature vectors built from the normalized and extracted semantic information for supervised and unsupervised clustering and machine learning; and

selecting with a classification module a non-classified document from a document collection, and extracting via the extraction module normalized structural, syntactical and semantic information from the selected document, and generating via the machine learning module a model representation of the selected document based on feature vectors, and matching with the classification module the model representation of the selected document against the machine learning model representation and generating with the classification module a document category, and classification for display to a user,

wherein the extracted information includes named entities, properties of entities, noun-phrases, facts, events, and concepts.

10. The computer program product of claim 9 , wherein extraction module employs text-mining, language identification, gazetteers, regular expressions, noun-phrase identification with part-of-speech taggers, and statistical models and rules, and is configured to identify patterns, and

the patterns include libraries, and algorithms shared among cases, and which can be tuned for a specific case, to generate case-specific semantic information.

11. The computer program product of claim 9 , wherein the extracted information is normalized by using normalization rules, groupers, thesauri, taxonomies, and string-matching algorithms.

12. The computer program product of claim 9 , wherein the model representation of the document is a term frequency-inverse document frequency (TF-IDF) document representation of the extracted information, and the clustering and machine learning includes a classifier based on decision trees, support vector machines (SVM), naïve-bayes classifiers, k-nearest neighbors, rules-based classification, Linear discriminant analysis (LDA), Maximum Entropy Markov Model (MEMM), scatter-gather clustering, and hierarchical agglomerate clustering (HAC).

Assignments (3)
RELEASE OF SECURITY INTERESTS IN PATENTS RECORDED AT R/F 057484/0493 Recorded Sep 10, 2023
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: MSC INTELLECTUAL PROPERTIES B.V.
Reel/Frame 064853/0202 →
SECURITY INTEREST Recorded Sep 15, 2021
From: MSC INFORMATION RETRIEVAL TECHNOLOGIES B.V.; MSC INTELLECTUAL PROPERTIES B.V.; ZYLAB TECHNOLOGIES B.V.; ZYLAB DISTRIBUTION B.V.; ZYLAB BENELUX B.V.; ZYLAB EDISCOVERY & COMPLIANCE SERVICES (DCS) B.V.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 057484/0493 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2014
From: SCHOLTES, JOHANNES CORNELIS
To: MSC INTELLECTUAL PROPERTIES B.V.
Reel/Frame 033314/0634 →
Continuity (1)
Related Publication 20140156567A1 · Jun 5, 2014