IP Library Granted Patent US 7,529,719
Granted Patent B2
US 7,529,719 · App. 11/378,095 · Granted May 5, 2009

Document characterization using a tensor space model

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,529,719
App. No.
11/378,095
Granted
May 5, 2009
Kind
B2
Abstract

Computer-readable media having computer-executable instructions and apparatuses categorize documents or corpus of documents. A Tensor Space Model (TSM), which models the text by a higher-order tensor, represents a document or a corpus of documents. Supported by techniques of multilinear algebra, TSM provides a framework for analyzing the multifactor structures. TSM is further supported by operations and presented tools, such as the High-Order Singular Value Decomposition (HOSVD) for a reduction of the dimensions of the higher-order tensor. The dimensionally reduced tensor is compared with tensors that represent possible categories. Consequently, a category is selected for the document or corpus of documents. Experimental results on the dataset for 20 Newsgroups suggest that TSM is advantageous to a Vector Space Model (VSM) for text classification.

Claims (34)

1. A computer-readable medium having computer-executable instructions for controlling a processor of a computer system to categorize a document by a method comprising:

for each of a plurality of categories, providing documents within that category, each document having words with characters;

for each document,

generating a high-order tensor having an order of at least three, each order represented by a coordinate with characters as dimensions of the coordinate, each element of the high-order tensor representing a sequence of at least three characters and being set to a weight based on number of occurrences of that sequence of at least three characters within the document, the weight being based on term frequency by inverse document frequency; and

generating a core tensor by reducing dimensionality of the generated high-order tensor using high-order singular value decomposition;

training a support vector machine (“SVM”) classifier using the generated core tensors for the documents and the categories of the documents; and

categorizing a document by generating a high-order tensor for the document, generating a core tensor for the generated high-order tensor for the document, and applying the SVM classifier to the generated core tensor for the document to determine a category for the document.

2. The computer-readable medium of claim 1 wherein the documents are derived from new groups.

3. The computer-readable medium of claim 1 wherein the generating of a core tensor for a document includes unfolding the generated high-order tensor for the document, applying a singular value decomposition to the unfolded tensor to generate a unfolded core tensor, and folding the unfolded core tensor.

4. A method performed by a computer system to categorize a document, the method performed by a processor of the computer system comprising:

for each of a plurality of categories, storing documents within that category, each document having words with characters;

for each document,

generating a high-order tensor having an order of at least three, each order represented by a coordinate with characters as dimensions of the coordinate, each element of the high-order tensor representing a sequence of at least three characters and being set to a weight based on number of occurrences of that sequence of at least three characters within the document; and

generating a core tensor by reducing dimensionality of the generated high-order tensor;

training a classifier using the generated core tensors for the documents and the categories of the documents; and

categorizing a document by generating a high-order tensor for the document, generating a core tensor for the generated high-order tensor for the document, and applying the classifier to the generated core tensor for the document to determine a category for the document.

5. The method of claim 4 wherein the training includes for each category, generating an average of the core tensors of documents within that category.

6. The method of claim 5 wherein the applying of the classifier to the generated core tensor for the document includes selecting as the category for the document the category whose average core tensor is most similar to the generated core tensor for the document.

7. The method of claim 4 wherein the classifier is a support vector machine (“SVM”) classifier.

8. The method of claim 4 wherein the weight is based on a term frequency by inverse document frequency metric.

9. The method of claim 4 wherein the generating a core tensor for a document includes unfolding the generated high-order tensor for the document, applying a singular value decomposition to the unfolded tensor to generate a unfolded core tensor, and folding the unfolded core tensor.

10. A computer system that categorizes a document, comprising:

a processor; and

a memory storing:

a corpus of documents, each document having words and a category;

a tensor space model module that generates a high-order tensor having an order of at least three, each order represented by a coordinate with characters as dimensions of the coordinate, each element of the high-order tensor representing a sequence of at least three characters and being set to a weight based on number of occurrences of that sequence of at least three characters within the document;

an analyzing module that generates a core tensor by reducing dimensionality of the generated high-order tensor;

a training module that trains a classifier using the generated core tensors for the documents and the categories of the documents; and

a categorization module that categorizes a document by generating a high-order tensor for the document, generates a core tensor for the generated high-order tensor for the document, and applies the classifier to the generated core tensor for the document to determine a category for the document.

11. The computer system of claim 10 wherein the training module, for each category, generates an average of the core tensors of documents within that category.

12. The computer system of claim 11 wherein the categorization module includes selecting as the category for the document the category whose average core tensor is most similar to the generated core tensor for the document.

13. The computer system of claim 12 wherein the weight is based on a term frequency by inverse document frequency metric.

14. The computer system of claim 13 wherein the analyzing module generates a core tensor for a document by unfolding the generated high-order tensor for the document, applying a singular value decomposition to the unfolded tensor to generate a unfolded core tensor, and folding the unfolded core tensor.

15. The computer system of claim 10 wherein the classifier is a support vector machine (“SVM”) classifier.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2015
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034766/0509 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2006
From: LIU, NING; ZHANG, BENYU; YAN, JUN; CHEN, ZHENG; ZENG, HUA-JUN; WANG, JIAN
To: MICROSOFT CORPORATION
Reel/Frame 017602/0862 →