IP Library Granted Patent US 9,529,911
Granted Patent B2
US 9,529,911 · App. 13/860,923 · Granted Dec 27, 2016

Building of a web corpus with the help of a reference web crawl

Inventors: Sebastien Richard (Paris, FR); Xavier Grehant (Paris, FR); Jim Ferenczi (Paris, FR)
Assignee: Dassault Systemes
G06F17/30864
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,529,911
App. No.
13/860,923
Granted
Dec 27, 2016
Kind
B2
Abstract

Computer-implemented method for building a web corpus (WCD) comprising the steps of: sending by a web crawler (WC) a query to a reference web crawl agent (RWCA), this query containing a least one identifier of a resource, receiving by the web crawler (WC) a response from the reference web crawl agent (RWCA); if this response does not contain the resource identified by the identifier, downloading by the web crawler (WC) the resource from the website (WS) corresponding to the identifier and adding the resource to the web corpus (WCD; and if this response contains the resource identified by the identifier, adding the resource to the web corpus (WCD).

Claims (39)

1. A computer-implemented method for building a web corpus (WCD), the method comprising the steps of:

providing to a web crawler (WC) a query containing a list of several URLs, each URL identifying a respective resource to be included in the web corpus (WCD) to be built, the web corpus being a set of resources that corresponds to a web crawl, the web crawl being a repository of downloaded resources, the resources of the repository being downloaded by visiting corresponding URLs, each URL identifying a respective resource to be downloaded;

providing one or more reference web crawl agents (RWCA), each reference web crawl agent acting as an interface between the web crawler and a respective reference web crawl (RWCD), each reference web crawl being distinct from the web corpus to be built,

for each URL in the query:

sending by the web crawler (WC) the URL to at least one of the one or more reference web crawl agents (RWCA);

by said at least one of the one or more reference web crawl agents (RWCA):

checking if said URL identifies a resource already downloaded in the respective reference web crawl (RWCD);

building a response, where:

if the checking has a positive result, then the response contains the resource identified by said URL, else

if the checking has a negative result, then:

the response does not contain the resource identified by said URL, or

said at least one of the one or more reference web crawl agents (RWCA) initiates the downloading of the resource and said resource is added to the respective reference web crawl (RWCD), and the response contains the resource identified by said URL;

sending to said web crawler (WC) the response;

upon receiving by said web crawler (WC) said responses from said at least one of the one or more reference web crawl agents (RWCA):

if no response contains the resource identified by said URL, downloading by said web crawler (WC) said resource from a website (WS) corresponding to said URL and adding said resource to said web corpus (WCD), otherwise if a response contains the resource identified by said URL, adding said resource to said web corpus (WCD) to be built, the adding being performed without querying a web server.

2. The computer-implemented method according to claim 1 , further comprising steps of:

building a reference index (RID) from said reference web crawl (RWCD),

sending by said web crawler (WC) an index query to said reference index (RID),

receiving by said web crawler (WC) a response from said reference index, and

wherein the sending of said query to said reference web crawl agent (RWCA), is done depending on the content of said response.

3. The computer-implemented method according to claim 2 , wherein said index query contains an identifier of a resource, and wherein if said response contains indexed information related to said resource, deciding on whether to send a query to said reference web crawl agent (RWCA) according to said indexed information.

4. The computer-implemented method according to claim 2 , wherein said index query comprises query criteria and said response of said reference index contains a list of identifiers.

5. The computer-implemented method according to claim 4 , wherein said response of said reference index contains in addition indexed information corresponding to said identifiers.

6. The computer-implemented method according to claim 2 , wherein said index query comprises an identifier, and wherein said reference index sends a response containing a set of identifiers contained in the resource identified by said identifier.

7. A Web Crawler (WC) apparatus adapted to build a web corpus (WCD) having a processor configured to:

receive, at the web crawler, a query containing a list of several URLs, each URL identifying a respective resource to be included in the web corpus (WCD) to be built;

provide one or more reference web crawl agents (RWCA), each reference web crawl agent acting as an interface between the subject web crawler and a respective reference web crawl (RWCD), each reference web crawl being distinct from the web corpus to be built, the web corpus being a set of resources that corresponds to a web crawl, the web crawl being a repository of downloaded resources, the resources of the repository being downloaded by visiting corresponding URLs, each URL identifying a respective resource to be downloaded;

for each URL in the query:

send by the web crawler (WC) the URL to at least one of the one or more reference web crawl agents (RWCA);

by said at least one of the one or more reference web crawl agents (RWCA):

check if said URL identifies a resource already downloaded in the respective reference web crawl (RWCD);

build a response, where:

if the checking has a positive result, then the response contains the resource identified by said URL, else

if the checking has a negative result, then:

the response does not contain the resource identified by said URL, or

said at least one of the one or more reference web crawl agents (RWCA) initiates the downloading of said resource and said resource is added to the respective reference web crawl (RWCD), and the response contains the resource identified by said URL;

send to said web crawler (WC) the response;

upon receiving by said web crawler (WC) said responses from said at least one of the one or more reference web crawl agents (RWCA):

if no response contains the resource identified by said URL, download by said web crawler (WC) said resource from a website (WS) corresponding to said URL and add said resource to said web corpus (WCD), otherwise if a response contains the resource identified by said URL, add said resource to said web corpus (WCD) to be built, the adding being performed without querying a web server.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Dec 15, 2014
From: EXALEAD; DASSAULT SYSTEMES
To: DASSAULT SYSTEMES
Reel/Frame 034631/0570 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2013
From: RICHARD, SEBASTIEN; GREHANT, XAVIER; FERENCZI, JIM
To: EXALEAD
Reel/Frame 030563/0356 →
Priority Claims (1)
EP 12305432 · Apr 12, 2012 · regional
Continuity (1)
Related Publication 20130275406A1 · Oct 17, 2013