IP Library Patent Application 11420681
Patent Application
App. No. 11/420,681

Method and Apparatus for Focused Crawling

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/420,681
Abstract

The present invention pertains to the field of computer software. More specifically, the present invention relates to dynamic discovery of documents or information through a focused crawler or search engine.

Claims (28)

1 . A method of ranking documents, comprising:

accessing a first plurality of documents from a database of a plurality of received documents, the first plurality of documents to be ranked;

generating a graph of the first plurality of documents;

expanding the graph with a second plurality of one or more documents from the database, such that a third plurality includes a union of the first plurality of documents and the second plurality of documents, and the third plurality of documents is smaller than the plurality of received documents, the second plurality including one or more of: 1) one or more documents connected within a first specified number of links in a forward direction from one or more documents of the first plurality of documents, the forward direction being forward from the first plurality of documents, and 2) one or more documents connected within a second specified number of links in a backward direction from one or more documents of the first plurality of documents, the backward direction being backward from the first plurality of documents;

assigning weights to one or more nodes of the graph;

finding an assignment of weights to one or more nodes of the graph, by propagating weights through the graph, the assignment of weight to a node based at least in part on calculating a weighted sum of weights propagated from neighboring nodes; and

generating a ranked list of at least the first plurality of documents, the ranked list at least partly generated from the graph.

2 . The method of claim 1 , further comprising:

shrinking the graph by removing one or more nodes of the graph.

3 . The method of claim 1 , further comprising:

shrinking the graph by combining one or more sets of one or more nodes of the graph.

4 . The method of claim 3 , wherein the combining is based on common characteristics of the nodes or relationships between the nodes.

5 . The method of claim 1 , wherein the propagating weights through the graph occurs up to a limited node distance.

6 . The method of claim 1 , wherein weights assigned to a node include at least one of relevance of the document to a query input and importance of the document independent of the query input.

7 . A method of finding starting points, comprising:

accessing a plurality of sample documents, the sample documents including links to other documents;

generating a graph representation of the plurality of sample documents and a plurality of referring documents, each of the referring documents referring to at least one of the plurality of sample documents, such that nodes of the graph represent the sample documents and the referring documents, and edges of the graph represent links linking the sample documents and the referring documents; and

finding starting point documents, at least partly by performing a link structure analysis on the graph.

8 . The method of claim 7 , wherein the link structure analysis includes:

assigning weights at least to nodes representing sample documents; and

propagating weights through the graph in reverse direction, to nodes representing referring pages; and

assigning weights to nodes of the graph, by calculating a weighted sum of weights propagated from neighboring nodes.

9 . The method of claim 7 , wherein at least one of the plurality of referring documents refers to at least one of the plurality of sample documents directly.

10 . The method of claim 7 , wherein at least one of the plurality of referring documents refers to at least one of the plurality of sample documents indirectly through one or more documents.

11 . The method of claim 7 , wherein finding starting point documents further includes one or more of: 1) evaluating relevance of documents using logical expressions of keywords and phrases and 2) evaluating relevance of documents using content located in a specified part of a format.

12 . The method of claim 7 , wherein relevance includes importance.

13 . The method of claim 7 , wherein the starting point documents provide starting points for at least one crawl.

14 . The method of claim 7 , wherein evaluating relevance of documents includes evaluating relevance of at least a first document and a second document, the second document referring to the first document.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded May 13, 2013
From: SILICON VALLEY BANK
To: FIRSTRAIN, INC.
Reel/Frame 030401/0139 →
SECURITY AGREEMENT Recorded Jan 26, 2010
From: FIRSTRAIN, INC.
To: SILICON VALLEY BANK
Reel/Frame 023839/0947 →
RELEASE OF SECURITY INTEREST Recorded Jan 22, 2010
From: VENTURE LENDING & LEASING IV, INC.
To: FIRSTRAIN, INC.
Reel/Frame 023832/0399 →