IP Library › Granted Patent US 8,782,031
Granted Patent B2
US 8,782,031 · App. 13/206,256 · Granted Jul 15, 2014

Optimizing web crawling with user history

Inventors: Dean M. Wierman (Bellevue, WA); Fabrice Canel (Redmond, WA); Balaji Shyamkumar (Sammamish, WA); Charles (Xi) Zhang (Sammamish, WA)
Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,782,031
App. No.
13/206,256
Granted
Jul 15, 2014
Kind
B2
Abstract

A politeness manager estimates traffic to the sites based on historical log data generated and sent by plug-ins or toolbars on client web browsers. The historical log data details dates and times the web browsers visit different web sites that is used to understand what timeframes specific web sites are busy and what timeframes the web sites are not busy. Crawl rates for different timeframes for a web site are determined based on the historical log data, and web crawlers are scheduled to crawl the web site according to the crawl rates to minimize the chances that web crawler requests are responsible for the site crashing.

Claims (46)

1. A method for crawling a web site, comprising:

receiving, at a server device, log data from a plurality of web browsers, the log data indicating users accessing the web site through the web browsers;

using, at the service device, the log data to estimate traffic to the web site during a timeframe;

determining, by the server device, a threshold frequency of page requests for the web site during the timeframe based on the estimate of traffic;

determining, at the server device, a crawl rate during the timeframe that is less than the threshold frequency of page requests; and

using the crawl rate to schedule one or more web crawlers to request the web site.

2. The method of claim 1 , wherein the log data indicates a plurality of uniform resource locators (“URLs”) historically accessed by the plurality of web browsers.

3. The method of claim 2 , wherein the log data indicates times and dates that the plurality of web browsers accessed the plurality of URLS.

4. The method of claim 2 , wherein the log data indicates a plurality of referral URLs users accessed before the URLs.

5. The method of claim 1 , further comprising determining one or more peak and non-peak timeframes for accessing the web site based on the estimate of traffic to the web site, wherein the peak timeframe is a time that the web site historically experiences more traffic than the non-peak timeframe.

6. The method of claim 5 , further comprising scheduling the one or more web crawlers to access the web site more frequently during the peak timeframe than during the non-peak timeframe.

7. The method of claim 5 , further comprising scheduling the one or more web crawlers to not access the web site during the peak timeframe.

8. The method of claim 1 , wherein the one or more web crawlers, upon accessing the web site, analyze content on the web site to index the web site.

9. The method of claim 1 , further comprising:

receiving analytics logs from one or more analytics applications that monitor access to the web site; and

using the webmaster data to specify the crawl rate during the timeframe.

10. The method of claim 1 , wherein the log data is received in an HTTP URL message.

11. The method of claim 1 , further comprising:

receiving an HTTP error; and

based on receiving the HTTP error, stopping the one or more web crawlers from requesting the web site.

12. The method of claim 1 , wherein the crawl rate indicates a maximum frequency for accessing the web site.

13. One or more computer-storage media embodied with computer-executable instructions that, when executed by a processor, coordinate web crawling of a web site, comprising:

receiving log data from a web browser indicating historical browsing history;

determining that the web browser accessed the web site at a specific time;

aggregating additional log data from a plurality of other web browsers, the additional log data indicating additional historical browsing history of the other web browsers accessing the web site;

using the log data and the additional log data to estimate traffic to the web site during a timeframe;

determining a threshold frequency of page requests for the web site during the timeframe based on the estimate of traffic; and

assigning one or more web crawlers to access the web site at a rate less than the threshold frequency of page requests.

14. The media of claim 13 , further comprising:

determining a second estimate of traffic to the web site during a second timeframe based on the log data and the additional log data; and

based on the second estimate of traffic indicating less historical traffic to the web site during the second timeframe than the estimate of traffic, scheduling the one or more crawlers to access the web site more frequently during the second timeframe than the timeframe.

15. The media of claim 13 , further comprising:

determining a second estimate of traffic to the web site during a second timeframe based on the log data and the additional log data; and

based on the second estimate of traffic indicating more historical traffic to the web site during the second timeframe than the estimate of traffic, scheduling the one or more crawlers to access the web site less frequently during the second timeframe than the timeframe.

16. The media of claim 13 , wherein the threshold frequency of page requests for the web site comprises a number of requests per a specific quantity of time.

17. The media of claim 13 , further comprising transmitting a notification comprising:

one or more peak timeframes for crawling the web site based on the historical browsing history and the additional historical browsing history,

non-peak timeframes for crawling the web site based on the historical browsing history and the additional historical browsing history, and

a request for approval for crawling the web site at one or more rates during either the peak or non-peak timeframes.

18. A server, comprising:

one or more processors executing a politeness manager to ( 204 and 304 ):

(1) estimate a threshold frequency of page requests for a web site during a first and second timeframe based on historical log data indicating a plurality of users accessing the web site, and

(2) schedule one or more web crawlers to access the web site less than the threshold frequency of page requests; and

one or more web crawlers accessing the web site as scheduled by the politeness manager.

19. The server of claim 18 , the threshold frequency of page requests is also based on historical analytics logs from one or more analytics applications that monitor access to the web site.

20. The server of claim 18 , wherein the politeness manager further transmits a request to crawl the web site at a rate less than the threshold frequency of page requests.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2011
From: WIERMAN, DEAN M.; CANEL, FABRICE; SHYAMKUMAR, BALAJI; ZHANG, CHARLES (XI)
To: MICROSOFT CORPORATION
Reel/Frame 026725/0801 →
Continuity (1)
Related Publication 20130041881A1 · Feb 14, 2013