IP Library Granted Patent US 11,204,971
Granted Patent B1
US 11,204,971 · App. 17/373,608 · Granted Dec 21, 2021

Token-based authentication for a proxy web scraping service

Inventors: Eivydas Vilcinskas (Siauliai, LT); Arnas Petruskevicius (Vilnius, LT); Giedrius Stalioraitis (Vilnius, LT); Martynas Juravicius (Vilnius, LT); Rimantas Stankevicius (Vilnius, LT)
Assignee: Metacluster LT, UAB
G06F16/951H04L63/083H04L67/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,204,971
App. No.
17/373,608
Granted
Dec 21, 2021
Kind
B1
Abstract

Embodiments disclose a system that allows for improved generation of web requests for scraping that, because of the nature of the requests and time and manner they are sent out, appear more organic, as in human generated, than conventional automated scraping systems. The system then manages how a client request to scrape a target website is made to the site, masking the request in a manner that makes it appear to the Web server as if the request is not generated by an automated system. In this way, by appearing more organic, Web servers may be less likely to block requests from the disclosed system or may take longer to block requests from the disclosed system. By avoiding Web servers blocking requests and extending the lifetime of IP proxies before they are blocked, embodiments can use a limited IP proxy address space more efficiently.

Claims (50)

1. A computer-implemented method for securing a web scraping system, comprising:

at an entry point to the web scraping system, performing the following:

(a) validating credentials received with an API request from a client computing device, the API request asking that the web scraping system scrape content from a target website;

(b) when the credentials are validated, generating a token indicating an identity of a client associated with the credentials;

(c) transmitting the API request along with the token to a server configured to initiate a scraping process on the web scraping system;

at the server configured to initiate the web scraping system:

(d) analyzing the token to determine whether the client is authorized to conduct the request; and

(e) when the client is authorized, causing the web scraping system to scrape the target website.

2. The method of claim 1 , further comprising:

(f) passing the API request between a plurality of servers, each configured to perform a function of the web scraping system, the server configured to initiate the web scraping system being included in the plurality of servers;

at each of the respective servers:

(g) analyzing the token to determine whether the client is authorized to conduct the function performed by the respective server; and

(e) when the client is authorized to conduct the function, performing the function.

3. The method of claim 2 , wherein the plurality of servers includes a server configured to service API requests formatted as a web proxy request.

4. The method of claim 2 , wherein the plurality of servers includes a server configured to service synchronous API requests, leaving a connection between the web scraping system and the client computing device open while the web scraping system scrapes the target web site.

5. The method of claim 2 , wherein the plurality of servers includes a server configured to service asynchronous API requests, closing a connection between the web scraping system and the client computing device before the web scraping system scrapes the target web site.

6. The method of claim 2 , wherein the generating (b) comprises generating the token to include a role of the client.

7. The method of claim 2 , wherein the generating (b) comprises generating the token to include a digital signature that cryptographically guarantees that the identity of the client has not been tampered with.

8. The method of claim 1 , wherein the API request is a first API request, and the token is a first token, further comprising:

(f) validating credentials received with a second API request, the second API request asking to retrieve content that the web scraping system has previously scraped from the target website;

(g) when the credentials are validated, generating a second token indicating an identity of a client associated with the credentials received with the second API request;

(h) determining whether the first and second tokens indicate that the first and second API requests came from the client; and

(i) when the first and second tokens indicate that the first and second API requests came from the client, returning the scraped content in response to the second API request.

9. The method of claim 8 , further comprising, when the first and second tokens do not indicate that the first and second requests came from the client, refusing to return the scraped content.

10. The method of claim 1 , wherein the entry point is a load balancer that selects the server from a plurality of parallel servers.

11. A non-transitory computer-readable device having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations, the operations comprising:

at an entry point to a web scraping system, performing the following:

(a) validating credentials received with an API request from a client computing device, the API request asking that the web scraping system scrape a target web site;

(b) when the credentials are validated, generating a token indicating an identity of a client associated with the credentials;

(c) transmitting the API request along with the token to a server configured to initiate a scraping process on the web scraping system;

at the server configured to initiate the web scraping system:

(d) analyzing the token to determine whether the client is authorized to conduct the request; and

(e) when the client is authorized, causing the web scraping system to scrape the target website.

12. The device of claim 11 , the operations further comprising:

(f) passing the API request between a plurality of servers, each configured to perform a function of the web scraping system, the server configured to initiate the web scraping system being included in the plurality of servers;

at each of the respective servers:

(g) analyzing the token to determine whether the client is authorized to conduct the function performed by the respective server; and

(e) when the client is authorized to conduct the function, performing the function.

13. The device of claim 12 , wherein the plurality of servers includes a server configured to service API requests formatted as a web proxy request.

14. The device of claim 12 , wherein the plurality of servers includes a server configured to service synchronous API requests, leaving a connection between the web scraping system and the client computing device open while the web scraping system scrapes the target web site.

15. The device of claim 12 , wherein the plurality of servers includes a server configured to service asynchronous API requests, closing a connection between the web scraping system and the client computing device before the web scraping system scrapes the target web site.

16. The device of claim 12 , wherein the generating (b) comprises generating the token to include a role of the client.

17. The device of claim 12 , wherein the generating (b) comprises generating the token to include a digital signature that cryptographically guarantees that the identity of the client has not been tampered with.

18. The device of claim 11 , wherein the API request is a first API request, and the token is a first token, the operations further comprising:

(f) validating credentials received with a second API request, the second API request asking to retrieve content that the web scraping system has previously scraped from the target website;

(g) when the credentials are validated, generating a second token indicating an identity of a client associated with the credentials received with the second API request;

(h) determining whether the first and second tokens indicate that the first and second requests came from the client; and

(e) when the first and second tokens indicate that the first and second requests came from the client, returning the scraped content in response to the second API request.

19. The device of claim 18 , the operations further comprising, when the first and second tokens do not indicate that the first and second requests came from the client, refusing to return the scraped content.

20. The device of claim 11 , wherein the entry point is a load balancer that selects the server from a plurality of parallel servers.

Assignments (3)
MERGER Recorded Mar 20, 2023
From: METACLUSTER LT, UAB
To: TESO LT, UAB
Reel/Frame 063954/0992 →
CHANGE OF NAME Recorded Mar 20, 2023
From: TESO LT, UAB
To: OXYLABS, UAB
Reel/Frame 063955/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2021
From: VILCINSKAS, EIVYDAS; PETRUSKEVICIUS, ARNAS; STALIORAITIS, GIEDRIUS; JURAVICIUS, MARTYNAS; STANKEVICIUS, RIMANTAS
To: METACLUSTER LT, UAB
Reel/Frame 056870/0836 →
Continuity (1)
Provisional Application 63219660 · Jul 8, 2021
Cited By (24)
US 12,223,095 US 12,425,492 US 12,445,511 US 12,457,273 US 12,483,635 US 12,517,972 US 12,524,490 US 12,524,491 US 12,536,243 US 12,542,764 US 12,549,645 US 12,563,130 US 12,587,429 US 12,587,430 US 12,587,579 US 12,603,809 US 12,652,330 US 12,659,218 US 12,671,750 US 12,706,984 US 12,711,192 US 12,719,734 US 12,719,735 US 12,719,945