IP Library Granted Patent US 8,375,061
Granted Patent B2
US 8,375,061 · App. 12/796,266 · Granted Feb 12, 2013

Graphical models for representing text documents for computer analysis

Inventor: Charu Aggarwal (Hawthorne, NY)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,375,061
App. No.
12/796,266
Granted
Feb 12, 2013
Kind
B2
Abstract

In a method for representing a text document with a graphical model, a document including a plurality of ordered words is received and a graph data structure for the document is created. The graph data structure includes a plurality of nodes and edges, with each node representing a distinct word in the document and each edge identifying a number of times two nodes occur within a predetermined distance from each other. The graph data structure is stored in an information repository.

Claims (19)

1. A method, comprising:

receiving a document including a plurality of ordered words;

creating a graph data structure for the document, wherein the graph data structure includes a plurality of nodes and edges, each node representing a distinct word in the document and each edge identifying a number of times two nodes occur within at least one word from each other;

storing the graph data structure in an information repository;

receiving a request to perform text analysis on the document; and

performing text analysis on the graph data structure and providing a result that is responsive to the request.

2. The method of claim 1 , wherein an edge is a directed edge or an undirected edge.

3. The method of claim 1 , wherein after receiving the document and before creating the graph data structure, the method further comprises pruning stop words from the document, wherein the graph data structure is created from the pruned document.

4. The method of claim 1 , wherein the text analysis is a text search.

5. The method of claim 1 , wherein at least one edge identifies a number of times two nodes occur within two words from each other.

6. A computer program product, comprising:

a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code comprising:

computer readable program code configured to receive a document including a plurality of ordered words; and

computer readable program code configured to create a graph data structure for the document, wherein the graph data structure includes a plurality of nodes and edges, each node representing a distinct word in the document and each edge identifying a number of times two nodes occur within at least one word from each other.

7. The computer program product of claim 6 , wherein an edge is a directed edge or an undirected edge.

8. The computer program product of claim 6 , wherein after receiving the document and before creating the graph data structure, the computer readable program code of the computer readable storage medium further comprises computer readable program code configured to prune stop words from the document, wherein the graph data structure is created from the pruned document.

9. The computer program product of claim 6 , wherein the computer readable program code of the computer readable storage medium further comprises:

computer readable program code configured to receive a request to perform text analysis on the document; and

computer readable program code configured to perform text analysis on the graph data structure to obtain a result that is responsive to the request.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2010
From: AGGARWAL, CHARU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 024502/0377 →
Continuity (1)
Related Publication 20110302168A1 · Dec 8, 2011