IP Library Granted Patent US 11,681,699
Granted Patent B2
US 11,681,699 · App. 17/364,432 · Granted Jun 20, 2023

Automated extraction of data from web pages

Inventors: Samuel Alison (Austin, TX); Ryan Engle (Austin, TX); Jacob Riesterer (Austin, TX); Jonathan Coon (Austin, TX)
Assignee: Capital One Services, LLC
G06F16/245G06F16/95G06F16/986G06Q30/0629H04L67/01H04L67/02H04L67/10H04L67/564H04L67/567H04L67/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,681,699
App. No.
17/364,432
Granted
Jun 20, 2023
Kind
B2
Abstract

Various embodiments provide techniques for automatically extracting data from web pages. Such extraction can take place without the use of a browser, and without necessarily rendering the entire web page. Thus, data extraction can be performed more efficiently and more quickly, while reducing the computing resources needed to perform such operations. In at least one embodiment, data extraction and translation are performed by automatically parsing structured data from visible and hidden elements of a web page.

Claims (68)

1. A computer-implemented method, comprising:

generating a request based on one or more requests indicated to contain parameters for displaying first information on a website, the one or more requests indicated to contain the parameters for displaying the first information having been determined by:

requesting and retrieving at least one web page from a web server, by transmitting requests and receiving replies, respectively;

extracting structured data of the at least one web page from the replies; and

based on the structured data, determining the one or more requests, of the requests, containing the parameters for displaying the first information; and

transmitting the request to the web server to obtain second information, the second information indicating pricing, descriptions, and/or in-stock information for different variants of products of the website.

2. The computer-implemented method of claim 1 , further comprising:

mapping the second information into a consistent format.

3. The computer-implemented method of claim 1 , wherein the one or more requests indicated to contain the parameters for displaying the first information have been further determined by performing at least one of an extraction operation or a transform operation to obtain the first information.

4. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

analyzing script tag data within script tags included in the replies;

generating an abstract syntax tree based on the script tag data; and

traversing and interpreting the abstract syntax tree to extract the structured data.

5. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

receiving a domain object model for the at least one web page;

analyzing the domain object model; and

based on an analysis result for the domain object model, locating microdata within the domain object model.

6. The computer-implemented method of claim 5 , wherein analyzing the domain object model comprises:

applying a model trained by machine learning, to analyze the domain object model.

7. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

applying a model trained by machine learning, to identify the structured data to be extracted.

8. The computer-implemented method of claim 1 , wherein:

extracting the structured data from the at least one web page comprises:

analyzing a structured format by which the structured data has been organized; and

based on an analysis result for the structured format, generating a format for future requests; and

the request to obtain the second information from the web server is generated using the format.

9. A non-transitory computer-readable medium, comprising instructions stored thereon, that when executed by a processor on a client device, perform operations comprising:

generating a request based on one or more requests indicated to contain parameters for displaying first information on a website, the one or more requests indicated to contain the parameters for displaying the first information having been determined by:

requesting and retrieving at least one web page from a web server, by transmitting requests and receiving replies, respectively;

extracting structured data of the at least one web page from the replies; and

based on the structured data, determining the one or more requests, of the requests, containing the parameters for displaying the first information; and

transmitting the request to the web server to obtain second information, the second information indicating pricing, descriptions, and/or in-stock information for different variants of products of the website.

10. The non-transitory computer-readable medium of claim 9 , further comprising additional instructions stored thereon, that when executed by the processor, perform further operations including:

mapping the second information into a consistent format.

11. The non-transitory computer-readable medium of claim 9 , wherein the one or more requests indicated to contain the parameters for displaying the first information have been further determined by performing at least one of an extraction operation or a transform operation to obtain the first information.

12. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

analyzing script tag data within script tags included in the replies;

generating an abstract syntax tree based on the script tag data; and

traversing and interpreting the abstract syntax tree to extract the structured data.

13. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

receiving a domain object model for the at least one web page;

analyzing the domain object model; and

based on an analysis result for the domain object model, locating microdata within the domain object model.

14. The non-transitory computer-readable medium of claim 13 , wherein analyzing the domain object model comprises:

applying a model trained by machine learning, to automatically analyze the domain object model.

15. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

applying a model trained by machine learning, to identify the structured data to be extracted.

16. The non-transitory computer-readable medium of claim 9 , wherein:

extracting the structured data from the at least one web page comprises:

analyzing a structured format by which the structured data has been organized; and

based on an analysis result for the structured format, generating a format for future requests; and

the request to obtain the second information from the web server is generated using the format.

17. A system, comprising:

a processor, on a client device, configured to perform a process including:

generating a request based on one or more requests indicated to contain parameters for displaying first information on a website, the one or more requests indicated to contain the parameters for displaying the first information having been determined by:

requesting and retrieving at least one web page from a web server, by transmitting requests and receiving replies, respectively;

extracting structured data of the at least one web page from the replies; and

based on the structured data, determining the one or more requests, of the requests, containing the parameters for displaying the first information; and

transmitting the request to the web server to obtain second information, the second information indicating pricing, descriptions, and/or in-stock information for different variants of products of the website.

18. The system of claim 17 , wherein the one or more requests indicated to contain the parameters for displaying the first information have been further determined by performing at least one of an extraction operation or a transform operation to obtain the first information.

19. The system of claim 17 , wherein extracting the structured data from the at least one web page comprises:

analyzing script tags data within script tags;

generating an abstract syntax tree based on the script tags data; and

traversing and interpreting the abstract syntax tree to extract the structured data.

20. The system of claim 17 , wherein extracting the structured data from the at least one web page comprises:

receiving a domain object model for the at least one web page;

analyzing the domain object model; and

based on an analysis result for the domain object model, locating microdata within the domain object model.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2021
From: WIKIBUY HOLDINGS, LLC
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 057169/0506 →
CHANGE OF NAME Recorded Aug 9, 2021
From: CAPITAL ONE WMS LLC
To: WIKIBUY HOLDINGS, LLC
Reel/Frame 057122/0659 →
MERGER AND CHANGE OF NAME Recorded Aug 4, 2021
From: WIKIBUY HOLDINGS, LLC; CAPITAL ONE WMS LLC
To: CAPITAL ONE WMS LLC
Reel/Frame 057079/0279 →
CHANGE OF NAME Recorded Jul 26, 2021
From: IMPOSSIBLE VENTURES, LLC
To: WIKIBUY HOLDINGS, LLC
Reel/Frame 057044/0502 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2021
From: ALISON, SAMUEL; ENGLE, RYAN; RIESTERER, JACOB; COON, JONATHAN
To: IMPOSSIBLE VENTURES, LLC
Reel/Frame 056732/0919 →
Continuity (6)
Continuation 16567367 · Sep 11, 2019
Continuation 15287089 · Oct 6, 2016
Provisional Application 62376243 · Aug 17, 2016
Provisional Application 62238574 · Oct 7, 2015
Provisional Application 62238565 · Oct 7, 2015
Related Publication 20210326338A1 · Oct 21, 2021