IP Library › Granted Patent US 7,424,484
Granted Patent B2
US 7,424,484 · App. 10/358,941 · Granted Sep 9, 2008

Path-based ranking of unvisited web pages

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,424,484
App. No.
10/358,941
Granted
Sep 9, 2008
Kind
B2
Abstract

Path-based ranking of unvisited Web pages for WWW crawling is provided, via identifying all the paths beginning with a “seed” URL and leading to visited relevant web pages as “good-path set”, and for each unvisited web page, identifying the paths beginning from the “seed” URL leading to it as “partial-path set”; classifying all the visited web pages and labeling each web Page with the labels of a class or classes it belongs to; training a statistic model for generalizing the common patterns among all ones of “good-path set”; and evaluating the “partial-path set” with the statistic model and ranking the unvisited web pages with the evaluation results.

Claims (28)

1. A computer implemented method for path-based ranking of unvisited Web pages for WWW crawling, comprising:

identifying all the paths beginning with a “seed” URL and leading to visited web pages containing content relevant to a search as “good-path set”, and for each unvisited web page, identifying the paths beginning from the “seed” URL leading to it as “partial-path set”;

classifying all the visited web pages and labeling each web page with the labels of a class or classes it belongs to;

training a statistic model for generalizing the common patterns among all ones of good-path set; and

evaluating the partial-path sets with the statistic model to determine a similarity between the partial-path sets and all ones of good-path set and ranking the unvisited web pages with the evaluation results,

whereby the highest ranked unvisited web pages are those most likely to contain content relevant to the search.

2. The method according to claim 1 , further comprising:

using a directed graph to represent linkage structure among all the visited web pages, wherein nodes are used to represent web pages, and edges are used to represent links between web pages.

3. The method according to claim 1 , further comprising:

adjusting the “good-path set” through Human-computer interaction stage.

4. The method according to claim 1 , wherein:

said statistic model is described by the following formula:

P (L i |(L i-1 ,L i-2 , . . . ,L 0 )), i≧0  (1)

where L i denotes a web page's label and the web page is the i th one in a path whose length is defined as i+1 and (L i-1 , L i-2 , . . . , L 0 ) denotes a certain path having the length of i and whose n th web page is L n where 0≦n<i.

5. A method for deploying computing infrastructure, comprising integrating computer readable code into a computing system, wherein the code in combination with the computing system is capable of performing a method for path-based ranking of unvisited Web pages for WWW crawling, and further comprising the following:

identifying all the paths beginning with a “seed” URL and leading to visited web pages containing content relevant to a search as “good-path set”, and for each unvisited web page, identifying the paths beginning from the “seed” URL leading to it as “partial-path set”;

classifying all the visited web pages and labeling each web page with the labels of a class or classes it belongs to;

training a statistic model for generalizing the common patterns among all ones of good-path set; and

evaluating the partial-path sets with the statistic model to determine a similarity between the partial-path sets and all ones of good-path set and ranking the unvisited web pages with the evaluation results,

whereby the highest ranked unvisited web pages are those most likely to contain content relevant to the search.

6. The method according to claim 5 , further comprising:

using a directed graph to represent linkage structure among all the visited web pages, wherein nodes are used to represent web pages, and edges are used to represent links between web pages.

7. The method according to claim 5 , further comprising:

adjusting the “good-path set” through Human-computer interaction stage.

8. The method according to claim 5 , wherein:

said statistic model is described by the following formula:

P (L i |(L i-1 ,L i-2 , . . . ,L 0 )), i≧0  (1)

where L i denotes a web page's label and the web page is the i th one in a path whose length is defined as i+1 and (L i-1 , L i-2 , . . . , L 0 ) denotes a certain path having the length of i and whose n th web page is L n where 0≦n<i.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2003
From: MA, XIAOCHUNG; PAN, YUE; SU, HUI
To: IBM CORPORATION
Reel/Frame 014765/0475 →
Priority Claims (1)
CN 02 1 03529 · Feb 5, 2002 · national
Continuity (1)
Related Publication 20030149694A1 · Aug 7, 2003