IP Library Granted Patent US 7,379,932
Granted Patent B2
US 7,379,932 · App. 11/314,432 · Granted May 27, 2008

System and a method for focused re-crawling of Web sites

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,379,932
App. No.
11/314,432
Granted
May 27, 2008
Kind
B2
Abstract

A method ( 100 ) of crawling the Web ( 620 ) is disclosed. The method ( 100 ) crawls ( 120 ) Web pages on the Web starting from a given ( 110 ) set of seed Universal Resource Locators (URLs). Crawled Web pages are partitioned ( 140 ) into sets of relevant and irrelevant pages. A set of exclusion and/or inclusion patterns are discovered ( 150 ) from the sets of relevant and irrelevant pages, and subsequent crawling of the Web is restricted through the set of exclusion and/or inclusion patterns.

Claims (57)

1. A method of crawling the Web, said method comprising:

crawling Web pages on the Web starting from a given set of seed Universal Resource Locators (URLs);

partitioning crawled Web pages into sets of relevant and irrelevant pages;

discovering from said sets of relevant and irrelevant pages a set of exclusion and inclusion patterns; and

restricting subsequent crawling of the Web through said set of exclusion and inclusion patterns, wherein said discovering comprises:

re-partitioning irrelevant pages containing a link to at least one relevant page as relevant;

forming a first set containing words appearing in URLs from relevant pages;

forming a second set containing words appearing in URLs from irrelevant pages; and

determining a set of exclusion patterns such that all words from said second set satisfy at least one of said exclusion patterns and no words from said first set satisfy any of said exclusion patterns.

2. A method according to claim 1 , all the limitations of which are incorporated herein by reference, further comprising finding a minimum Steiner tree before performing said re-partitioning.

3. A method of crawling the Web, said method comprising:

crawling Web pages on the Web starting from a given set of seed Universal Resource Locators (URLs);

partitioning crawled Web pages into sets of relevant and irrelevant pages;

discovering from said sets of relevant and irrelevant pages a set of exclusion and inclusion patterns; and

restricting subsequent crawling of the Web through said set of exclusion and inclusion patterns, wherein said discovering comprises:

re-partitioning irrelevant pages containing a link to at least one relevant page as relevant;

forming a first set containing words appearing in URLs from relevant pages;

forming a second set containing words appearing in URLs from irrelevant pages; and

determining a set of inclusion patterns such that all words from said first set satisfy at least one of said exclusion patterns and no words from said second set satisfy any of said exclusion patterns.

4. A program storage device readable by computer, tangibly embodying a program of instructions executable by said computer to perform a method of crawling the Web, said method comprising:

crawling Web pages on the Web starting from a given set of seed Universal Resource Locators (URLs);

partitioning crawled Web pages into sets of relevant and irrelevant pages;

discovering from said sets of relevant and irrelevant pages a set of exclusion and/or inclusion patterns; and

restricting subsequent crawling of the Web through said set of exclusion and/or inclusion patterns, wherein said discovering comprises:

re-partitioning irrelevant pages containing a link to at least one relevant page as relevant;

forming a first set containing words appearing in URLs from relevant pages; and

forming a second set containing words appearing in URLs from irrelevant pages;

determining a set of exclusion patterns such that all words from said second set satisfy at least one of said exclusion patterns and no words from said first set satisfy any of said exclusion patterns.

5. A program storage device readable by computer, tangibly embodying a program of instructions executable by said computer to perform a method of crawling the Web, said method comprising:

crawling Web pages on the Web starting from a given set of seed Universal Resource Locators (URLs);

partitioning crawled Web pages into sets of relevant and irrelevant pages;

discovering from said sets of relevant and irrelevant pages a set of exclusion and/or inclusion patterns; and

restricting subsequent crawling of the Web through said set of exclusion and/or inclusion patterns, wherein said discovering comprises:

re-partitioning irrelevant pages containing a link to at least one relevant page as relevant;

forming a first set containing words appearing in URLs from relevant pages;

forming a second set containing words appearing in URLs from irrelevant pages; and

determining a set of inclusion patterns such that all words from said first set satisfy at least one of said exclusion patterns and no words from said second set satisfy any of said exclusion patterns.

6. A program storage device according to claim 5 , all the limitations of which are incorporated herein by reference, further comprising finding a minimum Steiner tree before performing said re-partitioning.

7. A method of crawling a network, said method comprising:

crawling network pages on the network starting from a given set of seed Universal Resource Locators (URLs);

partitioning crawled network pages into sets of relevant and irrelevant pages;

discovering from said sets of relevant and irrelevant pages a set of exclusion and inclusion patterns; and

restricting subsequent crawling of the network through said set of exclusion and inclusion patterns, wherein said discovering comprises:

re-partitioning irrelevant pages containing a link to at least one relevant page as relevant;

forming a first set containing words appearing in URLs from relevant pages;

forming a second set containing words appearing in URLs from irrelevant pages; and

determining a set of exclusion patterns such that all words from said second set satisfy at least one of said exclusion patterns and no words from said first set satisfy any of said exclusion patterns.

8. A method of crawling a network, said method comprising:

crawling network pages on the network starting from a given set of seed Universal Resource Locators (URLs);

partitioning crawled network pages into sets of relevant and irrelevant pages;

discovering from said sets of relevant and irrelevant pages a set of exclusion and inclusion patterns; and

restricting subsequent crawling of the network through said set of exclusion and inclusion patterns, wherein said discovering comprises:

re-partitioning irrelevant pages containing a link to at least one relevant page as relevant;

forming a first set containing words appearing in URLs from relevant pages; and

forming a second set containing words appearing in URLs from irrelevant pages;

determining a set of inclusion patterns such that all words from said first set satisfy at least one of said exclusion patterns and no words from said second set satisfy any of said exclusion patterns.

9. A method according to claim 8 , all the limitations of which are incorporated herein by reference, further comprising finding a minimum Steiner tree before performing said re-partitioning.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2016
From: MIDWAY TECHNOLOGY COMPANY LLC
To: SERVICENOW, INC.
Reel/Frame 038324/0816 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2016
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MIDWAY TECHNOLOGY COMPANY LLC
Reel/Frame 037704/0257 →