IP Library Patent Application 12029259
Patent Application
App. No. 12/029,259

FULL TEXT QUERY AND SEARCH SYSTEMS AND METHODS OF USE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
12/029,259
Abstract

The invention is a method for textual searching of text-based databases including databases of compiled internet content, scientific literature, abstracts for books and articles, newspapers, journals, and the like. Specifically, the algorithm supports searches using full-text or webpage as query and keyword searches allowing multiple entries and an information-content based ranking system (Shannon Information score) that uses p-values to represent the likelihood that a hit is due to random matches. Additionally, users can specify the parameters that determine hits and their ranking with scoring based on phrase matches and sentence similarities.

Claims (11)

1 - 28 . (canceled)

29 . A data processing system comprising

1) a database of string entries,

2) a routine for processing the string entries, the routine selected from the group consisting of calculating a frequency distribution of string entries, associating an external frequency distribution with string entries in the database, and associating an external probability distribution with a collection of string entries in the database,

and 3) a routine for analyzing the database using the distribution.

30 . The data processing system of claim 29 wherein the routine for analyzing the database is selected from the group consisting of searching the database, querying the database, clustering the content of the database, and classifying the content of the database.

31 . The data processing system of claim 30 wherein a search query is selected from the group consisting of a keyword, a plurality of keywords, a title, an abstract, a full text query, a webpage, a webpage URL address, a highlighted segment of a webpage, and any part thereof.

32 . The data processing system of claim 29 further comprising a routine for calculating an information measure using the distribution.

33 . The data processing system of claim 32 , wherein the information measure comprises a negative log of the frequency or a negative log of the probability.

34 . The data processing system of claim 29 , wherein a string associated with the distribution defines an Infotom, the string comprising contiguous digitized text, the digitized text selected from the group consisting of letters, spaces, numbers, keywords, binary code, symbols, glyphs, and hieroglyphs.

35 . The data processing system of claim 32 , wherein the information measure is calculated using a Shannon information function.