IP Library Granted Patent US 11,416,291
Granted Patent B1
US 11,416,291 · App. 17/373,482 · Granted Aug 16, 2022

Database server management for proxy scraping jobs

Inventors: Eivydas Vilcinskas (Siauliai, LT); Arnas Petruskevicius (Vilnius, LT); Giedrius Stalioraitis (Vilnius, LT); Martynas Juravicius (Vilnius, LT); Rimantas Stankevicius (Vilnius, LT)
Assignee: Metacluster LT, UAB
G06F9/4881G06F16/951
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,291
App. No.
17/373,482
Granted
Aug 16, 2022
Kind
B1
Abstract

Embodiments disclose a system that allows for improved generation of web requests for scraping that, because of the nature of the requests and time and manner they are sent out, appear more organic, as in human generated, than conventional automated scraping systems. The system then manages how a client request to scrape a target website is made to the site, masking the request in a manner that makes it appear to the Web server as if the request is not generated by an automated system. In this way, by appearing more organic, Web servers may be less likely to block requests from the disclosed system or may take longer to block requests from the disclosed system. By avoiding Web servers blocking requests and extending the lifetime of IP proxies before they are blocked, embodiments can use a limited IP proxy address space more efficiently.

Claims (41)

1. A computer-implemented method for determining which servers are available to process web scraping jobs, comprising:

repeatedly checking health of each of a plurality of database servers;

based on the health checks, determine whether each of the plurality of database servers are to be enabled or disabled in a table, the plurality of database servers operating independently of one another, each of the database servers configured to manage data storage to at least a portion of a job database that stores the status of the web scraping jobs while the web scraping jobs are being executed;

when a web scraping request is received from a client computing device:

selecting one of the plurality of database servers identified as enabled in the table; and

sending a job description specified by the web scraping request to the selected database server for storage in the job database as a pending web scraping job to generate at least one web request to retrieve content as specified by the job description.

2. The method of claim 1 , wherein each of the repeated checking comprises, for each of the plurality of database servers, connecting to the portion of the job database for the respective database server.

3. The method of claim 1 , wherein each of the plurality of database servers comprises a message broker that queues job descriptions to be stored in the job database, and each of the repeatedly checking comprises, for each of the plurality of database servers, checking a connection between a server that receives web scraping requests from client computing devices and the respective database server's message broker.

4. The method of claim 1 , wherein each of the plurality of database servers comprises a message broker that queues job descriptions to be stored in the jobs database, and each of the repeatedly checking comprises, for each of the plurality of database servers, checking a number of messages queued within the respective database server's message broker.

5. The method of claim 1 , wherein each of the plurality of database servers is a shard managing storage in a horizontal partition of the job database.

6. The method of claim 1 , wherein each of the plurality of database servers do not synchronize states to one another.

7. The method of claim 1 , wherein the plurality of database servers are executed by a plurality of different computing devices.

8. The method of claim 1 , further comprising:

determining whether a number of database servers that are disabled in the plurality of database servers exceeds a threshold; and

when the number of database servers that are disabled exceeds the threshold, alerting an administrator.

9. A non-transitory computer-readable device having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations, comprising:

repeatedly checking health of each of a plurality of database servers;

based on the health checks, determining whether each of a plurality of database servers are to be enabled or disabled in a table, the plurality of database servers operating independently of one another, each of the database servers configured to manage data storage to at least a portion of a job database that stores the status of web scraping jobs while the web scraping jobs are being executed;

when a web scraping request is received from a client computing device:

selecting one of the plurality of database servers identified as enabled in the table; and

sending a job description specified by the web scraping request to the selected database server for storage in the job database as a pending web scraping job to generate at least one web request to retrieve content as specified by the job description.

10. The device of claim 9 , wherein each of the repeatedly checking comprises, for each of the plurality of database servers, connecting to the portion of the job database for the respective database server.

11. The device of claim 9 , wherein each of the plurality of database servers comprises a message broker that queues job descriptions to be stored in the job database, and each of the repeatedly checking comprises, for each of the plurality of database servers, checking a connection between a server that receives web scraping requests from client computing devices and the respective database server's message broker.

12. The device of claim 9 , wherein each of the plurality of database servers comprises a message broker that queues job descriptions to be stored in the job database, and each of the repeatedly checking comprises, for each of the plurality of database servers, checking a number of messages queued within the respective database server's message broker.

13. The device of claim 9 , wherein each of the plurality of database servers is a shard managing storage in a horizontal partition of the job database.

14. The device of claim 9 , wherein each of the plurality of database servers do not synchronize states to one another.

15. The device of claim 9 , wherein the plurality of database servers are executed by a plurality of different computing devices.

16. The device of claim 9 , further comprising:

determining whether a number of database servers that are disabled in the plurality of database servers exceeds a threshold; and

when the number of database servers that are disabled exceeds the threshold, alerting an administrator.

17. A system for determining which servers are available to process web scraping jobs, comprising:

a processor;

a job database that stores the status of the web scraping jobs while the web scraping jobs are being executed;

a memory that stores the job database;

a plurality of database servers operating independently of one another, each of the database servers configured to manage data storage to at least a portion of the job database;

a database monitor configured to repeatedly check health of each of the plurality of database servers and, based on the results of the health checks, determine whether each of the plurality of database servers are to be enabled or disabled in a table;

a database server selector configured to, when a web scraping request is received from a client computing device, select one of the database servers identified as enabled in the table; and

a request intake manager configured to send a job description specified by the web scraping request to the selected database server for storage in the job database as a pending web scraping job to generate at least one web request to retrieve content as specified by the job description.

18. The system of claim 17 , wherein the database monitor is configured to, for each of the plurality of database servers, check a connection between the request intake manager and the job database.

19. The system of claim 17 , wherein each of the plurality of database servers comprises a message broker that queues job descriptions to be stored in the job database, and the database monitor is configured to, for each of the plurality of database servers, check a connection between a server that receives web scraping requests from client computing devices and the respective database server's message broker.

20. The system of claim 17 , wherein each of the plurality of database servers comprises a message broker that queues job descriptions to be stored in the job database, and the database monitor is configured to, for each of the plurality of database servers, check a message queued within the respective database server's message broker.

Assignments (3)
MERGER Recorded Mar 20, 2023
From: METACLUSTER LT, UAB
To: TESO LT, UAB
Reel/Frame 063954/0992 →
CHANGE OF NAME Recorded Mar 20, 2023
From: TESO LT, UAB
To: OXYLABS, UAB
Reel/Frame 063955/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2021
From: VILCINSKAS, EIVYDAS; PETRUSKEVICIUS, ARNAS; STALIORAITIS, GIEDRIUS; JURAVICIUS, MARTYNAS; STANKEVICIUS, RIMANTAS
To: METACLUSTER LT, UAB
Reel/Frame 056870/0671 →
Continuity (1)
Provisional Application 63219660 · Jul 8, 2021