IP Library Granted Patent US 10,572,550
Granted Patent B2
US 10,572,550 · App. 15/326,045 · Granted Feb 25, 2020

Method of and system for crawling a web resource

Inventors: Damien Raymond Jean-François Lefortier (Moscow, RU); Liudmila Alexandrovna Ostroumova (Yaroslavl, RU); Egor Aleksandrovich Samosvat (Moscow, RU); Pavel Viktorovich Serdyukov (Moscow, RU); Ivan Semeonovich Bogatyy (Moscow, RU); Arsenii Andreevich Chelnokov (Moscow, RU); Gleb Gennadievich Gusev (Moscow, RU)
Assignee: YANDEX EUROPE AG
G06F16/951G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,572,550
App. No.
15/326,045
Granted
Feb 25, 2020
Kind
B2
Abstract

A method for determining a crawling schedule is disclosed, the method being executable at a crawling server coupled to a first web resource server and a second web resource server. The method comprises: acquiring a first new web page associated with the first web resource server; acquiring a second new web page associated with the second web resource server; determining a first crawling benefit parameter for the first new web page, the first crawling benefit parameter being based on a predicted popularity parameter and a predicted popularity decay parameter thereof; determining a second crawling benefit parameter for the second new web page, the second crawling benefit parameter being based on a predicted popularity parameter and a predicted popularity decay parameter thereof; based on the first crawling benefit parameter and the second crawling benefit parameter, determining a crawling order for the first new web page and the second new web page.

Claims (95)

1. A method of setting up a crawling schedule, the method executable at a crawling server, the crawling server coupled to a communication network, the communication network having coupled thereto a first web resource server and a second web resource server, the method comprising:

acquiring a first new web page associated with the first web resource server;

acquiring a second new web page associated with the second web resource server;

calculating a first crawling benefit parameter associated with the first new web page, the first crawling benefit parameter being based on:

a first predicted popularity parameter being indicative of long-term popularity of the first new web page, wherein the long-term popularity of the first new web page represents total number of visits to the first new web page, and

a first predicted popularity decay parameter of the first new web page, the first predicted popularity decay parameter being indicative of a rate at which the popularity of the first web page increases and decreases with time,

the first predicted popularity decay parameter having been predicted based on short-term popularity of the first web page indicative of a number of visits to the first web page over a predefined time interval after the first web page has been discovered, the long term popularity of the first new web page, and the predefined time interval and current age of the first new web page, wherein

the larger the first predicted popularity parameter is, the larger the first crawling benefit parameter is, and wherein

the first crawling benefit parameter decreases at the rate represented by the first predicted popularity decay parameter;

calculating a second crawling benefit parameter associated with the second new web page, the second crawling benefit parameter being based on:

a second predicted popularity parameter being indicative of long-term popularity of the second new web page, wherein the long-term popularity of the second new web page represents total number of visits to the second new web page, and

a second predicted popularity decay parameter of the second new web page, the second predicted popularity decay parameter being indicative of a rate at which the popularity of the second web page increases and decreases with time,

the second predicted popularity decay parameter having been predicted based on a short-term popularity of the second web page indicative of a number of visits to the second web page over a predefined time interval after the second web page has been discovered, the long term popularity of the second new web page, and the predefined time interval and current age of the second new web page, wherein

the larger the second predicted popularity parameter is, the larger the second crawling benefit parameter is, and wherein

the second crawling benefit parameter decreases at the rate represented by the second predicted popularity decay parameter; and

based on the first crawling benefit parameter and the second crawling benefit parameter, determining a crawling order for the first new web page and the second new web page, such that the first new page and the second new page are ordered in a descending order of an associated one of the first crawling benefit parameter and the second crawling benefit parameter.

2. The method of claim 1 , further comprising acquiring a first old web page associated with one of the first web resource server and the second web resource server, the first old web page having been previously crawled.

3. The method of claim 2 , further comprising calculating a third crawling benefit parameter associated with the first old web-page, the third crawling benefit parameter being based on a predicted popularity parameter and a predicted popularity decay parameter of at least one change associated with the first old web-page.

4. The method of claim 3 , further comprising, based on the first crawling benefit parameter, the second crawling benefit parameter and the third crawling benefit parameter, determining a crawling order for the first new web page, the second new web page and re-crawling of the first old web-page.

5. The method of claim 1 , further comprising estimating respective predicted popularity parameter and predicted popularity decay parameter associated with the first new web page and the second new web page using machine learning algorithm executed by the crawling server.

6. The method of claim 5 , further comprising training the machine learning algorithm.

7. The method of claim 6 , wherein said training is based on at least one feature selected from a list of:

number of transitions to all URLs in the pattern P: V in (P);

average number of transitions to a URL in the pattern V in (P)=|P|, where |P| is the number of URLs in P;

number of transitions to all URL's in the pattern P during the first t hours: V t in (P);

average number of transitions to a URL in the pattern P during the first t hours: V t in (P)=|P|;

fraction of transitions to all URL's in the pattern P during the first t hours: V t in (P)=V in (P).

8. The method of claim 7 , wherein said training is further based on a of the pattern |P|.

9. The method of claim 7 , wherein at least one feature used for said training is weighted.

10. The method of claim 6 , wherein said training is based on at least one feature selected from a list of:

number of times URLs in the pattern act as referrers in browsing V out (P);

average number of times a URL in the pattern acts as a referrer V out (P)=|P|;

number of times URLs in the pattern act as referrers during the first t hours V t out (P);

average number of times a URL in the pattern acts as a referrer during the first t hours V t out (P)=|P|;

fraction of times URLs in the pattern act as referrers during the first t hours V t out (P)=V out (P).

11. The method of claim 1 , wherein each of the first crawling benefit parameter and the second crawling benefit parameter is calculated using equation:

r

(

u

)

=

a

1

e

log

(

1

-

a

2

)

t

Δ

t

,

where a 1 is an estimation of a number of total visits (p);

a 2 is estimation of p t (u)/p(u);

p t (u) is an estimation of a number of visits during a predefined time interval t after the creation of the web resource; and

Δt is a current age of the web resource.

12. The method of claim 1 , wherein said determining a crawling order comprises applying a crawling algorithm.

13. The method of claim 12 , wherein the crawling algorithm is selected from a list of possible crawling algorithms that is configured to take into account the predicted popularity parameter and the predicted popularity decay parameter.

14. The method of claim 1 , wherein the method further comprises using a time when the respective first new web page and the second new web page were acquired by the crawling application as a proxy for the creation day.

15. The method of claim 1 , wherein:

the first predicted popularity decay parameter has been predicted based on short-term popularities of a first plurality of web pages, each of the first plurality of web pages being similar to the first web page, the short-term popularity of a given web page being indicative of a number of visits to the given web page over a predefined time interval after the given web page was discovered;

the second predicted popularity decay parameter has been predicted based on short-term popularities of a second plurality of web pages, each of the second plurality of web pages being similar to the second web page.

16. A server coupled to a communication network, the communication network having coupled thereto a first web resource server and a second web resource server, the server comprising:

a communication interface for communication with an electronic device via a communication network, a processor operationally connected with the communication interface, the processor being configured to:

acquire a first new web page associated with the first web resource server;

acquire a second new web page associated with the second web resource server;

calculate a first crawling benefit parameter associated with the first new web page, the first crawling benefit parameter being based on:

a first predicted popularity parameter being indicative of long-term popularity of the first new web page, wherein the long-term popularity of the first new web page represents total number of visits to the first new web page, and

a first predicted popularity decay parameter of the first new web page, the first predicted popularity decay parameter being indicative of a rate at which the popularity of the first web page increases and decreases with time,

the first predicted popularity decay parameter having been predicted based on a short-term popularity of the first web page indicative of a number of visits to the first web page over a predefined time interval after the first web page has been discovered, the long term popularity of the first new web page, and the predefined time interval and current age of the first new web page, wherein

the larger the first predicted popularity parameter is, the larger the first crawling benefit parameter is, and wherein

the first crawling benefit parameter decreases at the rate represented by the first predicted popularity decay parameter;

calculate a second crawling benefit parameter associated with the second new web page, the second crawling benefit parameter being based on:

a second predicted popularity parameter being indicative of long-term popularity of the second new web page, wherein the long-term popularity of the second new web page represents total number of visits to the second new web page, and

a second predicted popularity decay parameter of the second new web page, the second predicted popularity decay parameter being indicative of a rate at which the popularity of the second web page increases and decreases with time,

the second predicted popularity decay parameter having been predicted based on short-term popularity of the second web page indicative of a number of visits to the second web page over a predefined time interval after the second web page has been discovered, the long term popularity of the second new web page, and the predefined time interval and current age of the second new web page, wherein

the larger the second predicted popularity parameter is, the larger the second crawling benefit parameter is, and wherein

the second crawling benefit parameter decreases at the rate represented by the second predicted popularity decay parameter; and

based on the first crawling benefit parameter and the second crawling benefit parameter, determine a crawling order for the first new web page and the second new web page, such that the first new page and the second new page are ordered in a descending order of an associated one of the first crawling benefit parameter and the second crawling benefit parameter.

17. The server of claim 16 , the processor being further configured to acquire a first old web page associated with one of the first web resource server and the second web resource server, the first old web page having been previously crawled.

18. The server of claim 16 , the processor being further configured to estimate respective predicted popularity parameter and predicted popularity decay parameter associated with the first new web page and the second new web page using machine learning algorithm executed by the crawling server.

19. The server of claim 18 , the processor being further configured to train the machine learning algorithm.

20. The server of claim 16 , wherein:

the first predicted popularity decay parameter has been predicted based on short-term popularities of a first plurality of web pages, each of the first plurality of web pages being similar to the first web page, the short-term popularity of a given web page being indicative of a number of visits to the given web page over a predefined time interval after the given web page was discovered;

the second predicted popularity decay parameter has been predicted based on short-term popularities of a second plurality of web pages, each of the second plurality of web pages being similar to the second web page.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068524/0184 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2018
From: OSTROUMOVA, LIUDMILA ALEXANDROVNA; SAMOSVAT, EGOR ALEKSANDROVICH; SERDYUKOV, PAVEL VIKTOROVICH; BOGATYY, IVAN SEMEONOVICH; CHELNOKOV, ARSENII ANDREEVICH; GUSEV, GLEB GENNADIEVICH
To: YANDEX LLC
Reel/Frame 046814/0782 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2018
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 046814/0917 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2018
From: LEFORTIER, DAMIEN RAYMOND JEAN-FRANÇOIS
To: YANDEX LLC
Reel/Frame 046815/0200 →
Priority Claims (1)
RU 2014130448 · Jul 24, 2014 · national
Continuity (1)
Related Publication 20170206274A1 · Jul 20, 2017