IP Library Granted Patent US 9,507,857
Granted Patent B2
US 9,507,857 · App. 13/845,989 · Granted Nov 29, 2016

Apparatus and method for classifying document, and computer program product

Inventors: Masumi Inaba (Tokyo, JP); Toshihiko Manabe (Kanagawa, JP); Tomoharu Kokubu (Kanagawa, JP); Wataru Nakano (Kanagawa, JP)
Assignees: KABUSHIKI KAISHA TOSHIBA; TOSHIBA SOLUTIONS CORPORATION
G06F17/30705G06F3/048G06F17/30707
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,507,857
App. No.
13/845,989
Granted
Nov 29, 2016
Kind
B2
Abstract

According to an embodiment, a document classification apparatus includes an extraction unit, a clustering unit, a classification unit, and a label assignment unit. The extraction unit is configured to extract feature words from documents. The clustering unit is configured to cluster the feature words into clusters so that a difference between the number of documents each including any one of the feature words belonging to one cluster and the number of documents each including any one of the feature words belonging to another cluster is equal to or less than a predetermined reference value. The classification unit is configured to classify the documents into the clusters so that each document belongs to the cluster to which the feature word included in the each document belongs. The label assignment unit is configured to assign a classification label to each cluster as a word representative of the corresponding feature words.

Claims (37)

1. A document classification apparatus comprising:

a storage device that stores at least one or more thesauruses; and

a processor that performs operations, comprising:

extracting feature words from documents included in a document set;

clustering the extracted feature words into a plurality of clusters by using the one or more thesauruses stored in the storage device so that a difference between the number of documents each including any one of the feature words belonging to one cluster and the number of documents each including any one of the feature words belonging to another cluster is equal to or less than a predetermined reference value, the clusters corresponding respectively to subtrees of the thesaurus having a tree structure;

classifying the documents included in the document set into the clusters so that each document belongs to the cluster to which the feature word included in the each document belongs;

assigning a classification label to each cluster, the classification label being a word representative of the feature words belonging to the each cluster; and

presenting a document classification result in association with the classification label assigned to the corresponding cluster, wherein

the clustering of the extracted feature words further comprises clustering a plurality of feature words which do not each correspond to any subtree in the thesaurus, into one cluster, and

the assigning of the classification label comprises assigning a classification label representing that the cluster is a set of the feature words not corresponding to any subtree of the thesaurus, to the cluster to which the feature words that do not each correspond to any subtree in the thesaurus belongs.

2. The apparatus according to claim 1 , wherein the extracting of the feature words comprises extracting words subjected to intention representation by using more than one type of intention representation and selecting words from the extracted words based on a predefined criteria, as the feature words.

3. The apparatus according to claim 2 , wherein the selecting of the words comprises selecting words, each word having a weight calculated based on appearance frequency that is equal to or greater than a predetermined value as the feature words.

4. The apparatus according to claim 2 , wherein the presenting of the document classification result comprises presenting the document classification result in association with the classification label assigned to the corresponding cluster and the feature words belonging to the corresponding cluster.

5. The apparatus according to claim 4 , wherein the presenting of the document classification result comprises presenting each of the feature words that is to be presented in association with the document classification result, in a distinguishable form for each type of intention representation used for the extracting of the feature words.

6. The apparatus according to claim 1 , wherein the extracting of the feature words further comprises extracting one or more feature words from a designated document that is other than the documents included in the document set,

the clustering of the extracted feature words further comprises, when the one or more feature words are extracted from the designated document, clustering the one or more feature words extracted from the designated document into one cluster, and

the classifying of the documents further comprises, when a document included in the document set includes the one or more feature words extracted from the designated document, classifying the document into the cluster to which the one or more feature words extracted from the designated document belong.

7. The apparatus according to claim 2 , wherein the storage device further stores a dictionary for view point subject to intention representation, and

the selecting of words comprises selecting words that are included in the dictionary for viewpoint among words subjected to intention representation, as the feature word.

8. The apparatus according to claim 2 , wherein each of the documents included in the document set is a structured document that is separated into document elements each of which corresponds to one of intention representations, and

the extracting of the feature words comprises extracting the feature word from the document elements.

9. A document classification method comprising:

extracting feature words from documents included in a document set;

clustering the extracted feature words into a plurality of clusters by using a thesaurus stored in a storage device so that a difference between the number of documents each including any one of the feature words belonging to one cluster and the number of documents each including any one of the feature words belonging to another cluster is equal to or less than a predetermined reference value, the clusters corresponding respectively to subtrees of the thesaurus having a tree structure;

classifying the documents included in the document set into the clusters so that each document belongs to the cluster to which the feature word included in the each document belongs;

assigning a classification label to each cluster, the classification label being a word representative of the feature words belonging to the each cluster; and

presenting a document classification result in association with the classification label assigned to the corresponding cluster, wherein

the clustering of the extracted feature words further comprises clustering a plurality of feature words which do not each correspond to any subtree in the thesaurus, into one cluster, and

the assigning of the classification label comprises assigning a classification label representing that the cluster is a set of the feature words not corresponding to any subtree in the thesaurus, to the cluster to which the feature words that do not each correspond to any subtree in the thesaurus belongs.

10. A computer program product comprising a non-transitory computer-readable medium containing a program executed by a computer, the program causing the computer to execute:

extracting feature words from documents included in a document set;

clustering the extracted feature words into a plurality of clusters by using a thesaurus stored in a storage device so that a difference between the number of documents each including any one of the feature words belonging to one cluster and the number of documents each including any one of the feature words belonging to another cluster is equal to or less than a predetermined reference value, the clusters corresponding respectively to subtrees of the thesaurus having a tree structure;

classifying the documents included in the document set into the clusters so that each document belongs to the cluster to which the feature word included in the each document belongs;

assigning a classification label to each cluster, the classification label being a word representative of the feature words belonging to the each cluster; and

presenting a document classification result in association with the classification label assigned to the corresponding cluster, wherein

the clustering of the extracted feature words further comprises clustering a plurality of feature words which do not each correspond to any subtree in the thesaurus, into one cluster, and

the assigning of the classification label comprises assigning a classification label representing that the cluster is a set of the feature words not corresponding to any subtree of the thesaurus, to the cluster to which the feature words that do not each correspond to any subtree in the thesaurus belongs.

Assignments (5)
CHANGE OF CORPORATE NAME AND ADDRESS Recorded Feb 8, 2021
From: TOSHIBA SOLUTIONS CORPORATION
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 055259/0587 →
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0098. ASSIGNOR(S) HEREBY CONFIRMS THE CHANGE OF ADDRESS. Recorded May 28, 2019
From: TOSHIBA SOLUTIONS CORPORATION
To: TOSHIBA SOLUTIONS CORPORATION
Reel/Frame 051297/0742 →
CHANGE OF ADDRESS Recorded Mar 8, 2019
From: TOSHIBA SOLUTIONS CORPORATION
To: TOSHIBA SOLUTIONS CORPORATION
Reel/Frame 048547/0098 →
CHANGE OF NAME Recorded Mar 8, 2019
From: TOSHIBA SOLUTIONS CORPORATION
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0215 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2013
From: INABA, MASUMI; MANABE, TOSHIHIKO; KOKUBU, TOMOHARU; NAKANO, WATARU
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA SOLUTIONS CORPORATION
Reel/Frame 030650/0654 →
Priority Claims (1)
JP 2011-202281 · Sep 15, 2011 · national
Continuity (2)
Continuation PCTJP2012066184 · Jun 25, 2012
Related Publication 20130268535A1 · Oct 10, 2013