IP Library Granted Patent US 7,805,426
Granted Patent B2
US 7,805,426 · App. 11/763,299 · Granted Sep 28, 2010

Defining a web crawl space

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,805,426
App. No.
11/763,299
Granted
Sep 28, 2010
Kind
B2
Abstract

Provided are techniques for defining a web crawl space to be crawled. A seed list including one or more seed names is received from a user, wherein each seed name represents a website. In response to receiving the seed list, a web crawl space for the received seed list is generated by generating one or more allow rules.

Claims (53)

1. A method for defining a web crawl space to be crawled, comprising:

at a computer including a processor:

receiving a seed list including one or more seed names from a user, wherein each seed name represents a website with a two part name or a three part name, wherein the last part of the seed name is a top level domain name; and

in response to receiving the seed list, generating a web crawl space for the received seed list by generating one or more new allow rules that determine what part of a webspace is to be crawled by:

for each of the one or more seed names,

generating a new allow rule by appending a slash and a wildcard character at an end of the seed name to allow crawling of content of the website;

determining whether the seed name has less than three parts or at least three parts;

in response to determining that the seed name has less than three parts, updating the allow rule for the seed name by appending a wildcard character and a dot before a beginning of the seed name to allow crawling of content of one or more subdomains; and

in response to determining that the seed name has at least three parts, updating the allow rule for the seed name by replacing a first portion of the seed name before a dot with a wildcard character.

2. The method of claim 1 , further comprising:

attempting to fetch a document by crawling the web crawl space;

determining that the document has been moved; and

performing redirect processing by:

identifying a new seed name that represents a new location of the document; and

generating one or more additional allow rules on the new seed name to enlarge the web crawl space.

3. The method of claim 2 , further comprising:

adding the one or more additional allow rules to the set of allow rules; and

crawling the web crawl space defined by the set of allow rules.

4. A computer program product comprising a computer-readable medium storing a computer readable program, wherein the computer readable program when executed by a processor on a computer causes the computer to:

receive a seed list including one or more seed names from a user, wherein each seed name represents a website with a two part name or a three part name, wherein the last part of the seed name is a top level domain name; and

in response to receiving the seed list, generate a web crawl space for the received seed list by generating one or more allow new rules that determine what part of a webspace is to be crawled by:

for each of the one or more seed names,

generating a new allow rule by appending a slash and a wildcard character at an end of the seed name to allow crawling of content of the website;

determining whether the seed name has less than three parts or at least three parts;

in response to determining that the seed name has less than three parts, updating the allow rule for the seed name by appending a wildcard character and a dot before a beginning of the seed name to allow crawling of content of one or more subdomains; and

in response to determining that the seed name has at least three parts, updating the allow rule for the seed name by replacing a first portion of the seed name before a dot with a wildcard character.

5. The computer program product of claim 4 , wherein the computer readable program when executed on a computer causes the computer to:

attempt to fetch a document by crawling the web crawl space;

determine that the document has been moved; and

perform redirect processing by:

identifying a new seed name that represents a new location of the document; and

generating one or more additional allow rules based on the new seed name to enlarge the web crawl space.

6. The computer program product of claim 5 , wherein the computer readable program when executed on a computer causes the computer to:

add the one or more additional allow rules to the set of allow rules; and

crawl the web crawl space defined by the set of allow rules.

7. A system for defining a web crawl space to be crawled, comprising:

hardware logic capable of performing operations, the operations comprising:

receiving a seed list including one or more seed names from a user, wherein each seed name represents a website with a two part name or a three part name, wherein the last part of the seed name is a top level domain name; and

in response to receiving the seed list, generating a web crawl space for the received seed list by generating one or more new allow rules that determine what part of a webspace is to be crawled by:

for each of the one or more seed names,

generating a new allow rule by appending a slash and a wildcard character at an end of the seed name to allow crawling of content of the website;

determining whether the seed name has less than three parts or at least three parts;

in response to determining that the seed name has less than three parts, updating the allow rule for the seed name by appending a wildcard character and a dot before a beginning of the seed name to allow crawling of content of one or more subdomains; and

in response to determining that the seed name has at least three parts, updating the allow rule for the seed name by replacing a first portion of the first seed name before a dot with a wildcard character.

8. The system of claim 7 , wherein the operations further comprise:

attempting to fetch a document by crawling the web crawl space;

determining that the document has been moved; and

performing redirect processing by:

identifying a new seed name that represents a new location of the document; and

generating one or more additional allow rules on the new seed name to enlarge the web crawl space.

9. The system of claim 8 , wherein the operations further comprise:

adding the one or more additional allow rules to the set of allow rules; and

crawling the web crawl space defined by the set of allow rules.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2015
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: LINKEDIN CORPORATION
Reel/Frame 035201/0479 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2007
From: LEUNG, TONY KAI-CHI; TIRAN, MAXIME ROMAIN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019693/0588 →