IP Library Granted Patent US 11,055,281
Granted Patent B2
US 11,055,281 · App. 16/567,367 · Granted Jul 6, 2021

Automated extraction of data from web pages

Inventors: Samuel Alison (Austin, TX); Ryan Engle (Austin, TX); Jacob Riesterer (Austin, TX); Jonathan Coon (Austin, TX)
Assignee: Capital One Services, LLC
G06F16/245G06F16/95G06F16/986G06Q30/0629H04L67/02H04L67/10H04L67/2819H04L67/2838H04L67/32H04L67/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,055,281
App. No.
16/567,367
Granted
Jul 6, 2021
Kind
B2
Abstract

Various embodiments provide techniques for automatically extracting data from web pages. Such extraction can take place without the use of a browser, and without necessarily rendering the entire web page. Thus, data extraction can be performed more efficiently and more quickly, while reducing the computing resources needed to perform such operations. In at least one embodiment, data extraction and translation are performed by automatically parsing structured data from visible and hidden elements of a web page.

Claims (75)

1. A computer-implemented method for extracting data from a web page, comprising:

requesting and retrieving, by transmitting requests and receiving replies, respectively, at least one web page from a web server;

extracting structured data from the at least one web page;

based on the structured data, determining one or more requests, of the requests, contain parameters for displaying information;

based on a determination that the one or more requests contain the parameters for displaying the information, performing at least one of an extraction operation or a transform operation to obtain the information; and

generating and transmitting at least one request based on the one or more requests to obtain other information.

2. The computer-implemented method of claim 1 , further comprising:

mapping the other information into a consistent format.

3. The computer-implemented method of claim 1 , wherein performing at least one of the extraction operation or the transform operation to obtain the information comprises performing at least one task selected from a group consisting of:

parsing structured information in script tags;

extracting content from received HTML through DOM parsing; or

requesting the content returned in a readable structure.

4. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

analyzing script tag data within script tags;

generating an abstract syntax tree based on the script tag data; and

traversing and interpreting the abstract syntax tree to extract the structured data.

5. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

receiving a domain object model for the at least one web page;

analyzing the domain object model; and

based on an analysis result for the domain object model, locating microdata within the domain object model.

6. The computer-implemented method of claim 5 , wherein analyzing the domain object model comprises:

applying a model trained by machine learning, to analyze the domain object model.

7. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

applying a model trained by machine learning, to identify the structured data to be extracted.

8. The computer-implemented method of claim 1 , wherein extracting the structured data from the at least one web page comprises:

analyzing a structured format by which the structured data has been organized;

based on an analysis result for the structured format, generating a format for future requests; and

requesting the other information from the web server using the format.

9. A non-transitory computer-readable medium for extracting data from a web page, comprising instructions stored thereon, that when executed by a processor on a client device, perform operations comprising:

requesting and retrieving, by transmitting requests and receiving replies, respectively, at least one web page from a web server;

extracting structured data from the at least one web page;

based on the structured data, determining one or more requests, of the requests, contain parameters for displaying information;

based on a determination that the one or more requests contain the parameters for displaying the information, performing at least one of an extraction operation or a transform operation to obtain the information; and

generating and transmitting at least one request to obtain other information based on the one or more requests.

10. The non-transitory computer-readable medium of claim 9 , further comprising additional instructions stored thereon, that when executed by the processor, perform further operations including:

mapping the other information into a consistent format.

11. The non-transitory computer-readable medium of claim 9 , wherein performing at least one of the extraction operation or the transform operation to obtain the information comprises performing at least one task selected from a group consisting of:

parsing structured information in script tags;

extracting content from received HTML through DOM parsing; or

requesting the content returned in a readable structure.

12. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

analyzing script tag data within script tags;

generating an abstract syntax tree based on the script tag data; and

traversing and interpreting the abstract syntax tree to extract the structured data.

13. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

receiving a domain object model for the at least one web page;

analyzing the domain object model; and

based on an analysis result for the domain object model, locating microdata within the domain object model.

14. The non-transitory computer-readable medium of claim 13 , wherein analyzing the domain object model comprises:

applying a model trained by machine learning, to automatically analyze the domain object model.

15. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

applying a model trained by machine learning, to identify the structured data to be extracted.

16. The non-transitory computer-readable medium of claim 9 , wherein extracting the structured data from the at least one web page comprises:

analyzing a structured format by which the structured data has been organized;

based on an analysis result for the structured format, generating a format for future requests; and

requesting the other information from the web server using the format.

17. A system for extracting data from a web page, comprising:

a processor, on a client device, configured to perform a process including:

requesting and retrieving, by transmitting requests and receiving replies, respectively, at least one web page from a web server;

extracting structured data from the at least one web page;

based on the structured data, determining one or more requests, of the requests, contain parameters for displaying information;

based on a determination that the one or more requests contain the parameters for displaying the information, performing at least one of an extraction operation or a transform operation to obtain the information; and

generating and transmitting at least one request to obtain other information based on the one or more requests.

18. The system of claim 17 , wherein performing at least one of the extraction operation or the transform operation to obtain the information comprises performing at least one task selected from a group consisting of:

parsing structured information in script tags;

extracting content from received HTML through DOM parsing; or

requesting the content returned in a readable structure.

19. The system of claim 17 , wherein extracting the structured data from the at least one web page comprises:

analyzing script tags data within script tags;

generating an abstract syntax tree based on the script tags data; and

traversing and interpreting the abstract syntax tree to extract the structured data.

20. The system of claim 17 , wherein extracting the structured data from the at least one web page comprises:

receiving a domain object model for the at least one web page;

analyzing the domain object model; and

based on an analysis result for the domain object model, locating microdata within the domain object model.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2020
From: WIKIBUY HOLDINGS, LLC
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 051844/0085 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2020
From: CAPITAL ONE WMS LLC
To: WIKIBUY HOLDINGS, LLC
Reel/Frame 051804/0375 →
MERGER AND CHANGE OF NAME Recorded Feb 7, 2020
From: WIKIBUY HOLDINGS, LLC; CAPITAL ONE WMS LLC
To: CAPITAL ONE WMS LLC
Reel/Frame 051750/0155 →
CHANGE OF NAME Recorded Dec 10, 2019
From: IMPOSSIBLE VENTURES, LLC
To: WIKIBUY HOLDINGS, LLC
Reel/Frame 051240/0465 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2019
From: ALISON, SAMUEL; ENGLE, RYAN; RIESTERER, JACOB; COON, JONATHAN
To: IMPOSSIBLE VENTURES, LLC
Reel/Frame 051176/0208 →
Continuity (5)
Continuation 15287089 · Oct 6, 2016
Provisional Application 62376243 · Aug 17, 2016
Provisional Application 62238574 · Oct 7, 2015
Provisional Application 62238565 · Oct 7, 2015
Related Publication 20200004740A1 · Jan 2, 2020