IP Library Granted Patent US 7,720,799
Granted Patent B2
US 7,720,799 · App. 11/707,394 · Granted May 18, 2010

Systems and methods for employing an orthogonal corpus for document indexing

Assignee: IndraWeb, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,720,799
App. No.
11/707,394
Granted
May 18, 2010
Kind
B2
Abstract

Methods and systems for processing a body of reference material to generate a directory for accessing information from a database.

Claims (16)

1. A method for extending a body of textual reference material to generate a directory for accessing documents from a document collection that is initially unconnected with said body of textual reference material, comprising

processing the body of textual reference material into a plurality of text portions, each text portion being associated with a single topic from a plurality of topics,

processing said plurality of text portions, generating keywords for each, assigning a weight to each keyword in a text portion, associating a keyword with a corresponding text portion if the weight of said keyword in said corresponding text portion is equal to or greater than a weight of said keyword in the text portions other than the corresponding text portion, or is equal to or greater than a predetermined threshold value, and forming first keyword-weight pairs of said associated keywords, and

applying the associated keywords to at least one document from the initially unconnected document collection, and forming second keyword-weight pairs associated with the at least one document, forming a numeric score between the first and second keyword-weight pairs, and associating based on said score the at least one document from the initially unconnected document collection and the single topic from the plurality of topics.

2. A method according to claim 1 comprising creating a graphical interface representative of an identified orthogonal organization of topics for allowing a user to access information retrieved from the document collection and having an association with a topic.

3. A method according to claim 2 , wherein the orthogonal organization of topics includes a hierarchical organization of substantially orthogonal topics.

4. A method according to claim 1 , wherein the body of textual reference material is selected from the group consisting of an encyclopedia, a dictionary, a text book, a novel, a newspaper, a web site, and a glossary.

5. A method according to claim 1 , wherein processing the body of textual reference material includes identifying a table of contents and chapter headings for the body of reference material.

6. A method according to claim 1 , wherein processing the body of textual reference material includes identifying a keyword index for the body of reference material.

7. A method according to claim 1 , wherein processing the body of textual reference material includes identifying definition entries in a dictionary or glossary.

8. A method according to claim 1 , wherein processing the body of textual reference material includes normalizing the identified orthogonal organization of topics.

9. A method according to claim 1 , wherein processing the plurality of text portions includes generating a word map representative of a statistical analysis of words contained in at least one text portion.

10. A computer readable medium having computer readable program codes embodied therein for causing a computer to extend a body of textual reference material to generate a directory for accessing documents from a document collection that is initially unconnected with said body of textual reference material, the computer readable medium program codes performing functions comprising:

processing the body of textual reference material into a plurality of text portions, each text portion being associated with a single topic from a plurality of topics,

processing said plurality of text portions, generating keywords for each, assigning a weight to each keyword in a text portion, associating a keyword with a corresponding text portion if the weight of said keyword in said corresponding text portion is equal to or greater than a weight of said keyword in the text portions other than the corresponding text portion, or is equal to or greater than a predetermined threshold value, and forming first keyword-weight pairs of said associated keywords, and

applying the associated keywords to at least one document from the initially unconnected document collection, and forming second keyword-weight pairs associated with the at least one document, forming a numeric score between the first and second keyword-weight pairs, and associating based on said score the at least one document from the initially unconnected document collection and the single topic from the plurality of topics.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2015
From: SEMMX, INC
To: LINKAPEDIA,INC
Reel/Frame 037277/0889 →
CHANGE OF NAME Recorded Jun 20, 2012
From: INDRAWEB.COM, INC.
To: SEMMX, INC.
Reel/Frame 028415/0244 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2008
From: KON, HENRY; BURCH, GEORGE
To: INDRAWEB.COM, INC.
Reel/Frame 020476/0382 →
Continuity (3)
Continuation 0954879600 · Apr 13, 2000
Provisional Application 6012910300 · Apr 13, 1999
Related Publication 20080010311A1 · Jan 10, 2008