IP Library Granted Patent US 12,307,496
Granted Patent B2
US 12,307,496 · App. 18/312,820 · Granted May 20, 2025

Automated extraction of data from web pages

Inventors: Samuel Alison (Austin, TX); Ryan Engle (Austin, TX); Jacob Riesterer (Austin, TX); Jonathan Coon (Austin, TX)
Assignee: Capital One Services, LLC
G06Q30/0629G06F16/245G06F16/95G06F16/986H04L67/01H04L67/02H04L67/10H04L67/564H04L67/567H04L67/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,496
App. No.
18/312,820
Granted
May 20, 2025
Kind
B2
Abstract

Various embodiments provide techniques for automatically extracting data from web pages. Such extraction can take place without the use of a browser, and without necessarily rendering the entire web page. Thus, data extraction can be performed more efficiently and more quickly, while reducing the computing resources needed to perform such operations. In at least one embodiment, data extraction and translation are performed by automatically parsing structured data from visible and hidden elements of a web page.

Claims (68)

1. A system, comprising:

at least one memory storing instructions; and

at least one processor, operatively connected to the at least one memory, and configured to perform a process including:

generating an API that is generalized across a plurality of websites by, for each website:

obtaining at least one reply transmitted by a web server associated with the website in response to a request transmitted to the web server;

based on the at least one reply, extracting data from the website;

applying an abstract syntax parser to the extracted data to generate an abstract syntax tree;

based on the abstract syntax tree, generating a translation between a format of the data from the website and a standardized API format; and

based on the translation, generating at least one API request that is specific to the website and that is executable to retrieve requested data from the website in the standardized API format; and

utilizing the generated API to obtain respective information from each of the plurality of websites by, in each case, transmitting the at least one generated API request specific to the website to the web server associated with the website, the information, in each case, pertaining to at least one aspect of a same product or service available on the plurality of websites.

2. The system of claim 1 , wherein extracting the data from the website includes:

obtaining a domain object model for the website;

applying a model trained by machine learning to identify targeted information from the domain object model; and

extracting the identified targeted information from the website.

3. The system of claim 2 , wherein the targeted information includes information having a predetermined structure.

4. The system of claim 1 , wherein extracting the data from the website includes, in each case:

analyzing script tag data within script tags included in the at least one reply, wherein the abstract syntax tree is generated based on the script tag data; and

traversing and interpreting the abstract syntax tree.

5. The system of claim 1 , wherein extracting the data from the website includes, in each case:

obtaining a domain object model for the website;

analyzing the domain object model to identify microdata of the website; and

extracting the microdata from the website.

6. The system of claim 1 , wherein the data from the website, in each case, is structured data.

7. The system of claim 1 , wherein:

the product or service is displayed on a website visited by a user device; and

the utilizing of the API is in response to the display on the website.

8. A computer-implemented method, comprising:

generating an API that is generalized across a plurality of websites by, for each website:

obtaining at least one reply transmitted by a web server associated with the website in response to a request transmitted to the web server;

based on the at least one reply, extracting data from the website; and

based on the extracted data, generating at least one API request that is specific to the website and that is executable to retrieve requested data from the website in a standardized API format; and

utilizing the generated API to obtain respective information from each of the plurality of websites by, in each case, transmitting the at least one generated API request specific to the website to the web server associated with the website, the information, in each case, pertaining to at least one aspect of a same product or service available on the plurality of websites.

9. The computer-implemented method of claim 8 , wherein extracting the data from the website includes:

analyzing script tag data within script tags included in the at least one reply.

10. The computer-implemented method of claim 8 , wherein extracting the data from the website includes:

obtaining a domain object model for the website;

analyzing the domain object model to identify microdata of the website; and

extracting the microdata from the website.

11. The computer-implemented method of claim 8 , wherein extracting the data from the website includes:

applying a model trained by machine learning to identify data from the website having a predetermined structure; and

extracting information having the predetermined structure from the website.

12. The computer-implemented method of claim 8 , wherein the data from the website, in each case, is structured data.

13. The computer-implemented method of claim 8 , wherein:

the product or service is displayed on a website visited by a user device; and

the utilizing of the API is in response to the display on the website.

14. A computer-implemented method, comprising:

generating an API that is generalized across a plurality of websites by, for each website:

obtaining at least one reply transmitted by a web server associated with the website in response to a request transmitted to the web server;

based on the at least one reply, extracting data from the website;

applying an abstract syntax parser to the extracted data to generate an abstract syntax tree;

traversing and interpreting the generated abstract syntax tree to determine a translation between a format of the data from the website and a standardized API format; and

based on the translation, generating at least one API request that is specific to the website and that is executable to retrieve requested data from the website in the standardized API format.

15. The computer-implemented method of claim 14 , wherein extracting the data from the website includes:

analyzing script tag data within script tags included in the at least one reply, wherein the abstract syntax tree is generated based on the script tag data; and

traversing and interpreting the abstract syntax tree.

16. The computer-implemented method of claim 14 , wherein extracting the data from the website includes:

obtaining a domain object model for the website;

analyzing the domain object model to identify microdata of the website; and

extracting the microdata from the website.

17. The computer-implemented method of claim 14 , further comprising:

utilizing the generated API to obtain respective information from each of the plurality of websites by, in each case, transmitting the at least one generated API request specific to the website to the web server associated with the website, the information, in each case, pertaining to at least one aspect of a same product or service available on the plurality of websites.

18. The computer-implemented method of claim 17 , wherein:

the product or service is displayed on a website visited by a user device; and

the utilizing of the API is in response to the display on the website.

19. The computer-implemented method of claim 14 , wherein the data from the website, in each case, is structured data.

20. The computer-implemented method of claim 14 , wherein extracting the data from the website includes:

applying a model trained by machine learning to identify data from the website having a predetermined structure; and

extracting information having the predetermined structure from the website.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2023
From: WIKIBUY HOLDINGS, LLC
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 063552/0949 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2023
From: ALISON, SAMUEL; ENGLE, RYAN; RIESTERER, JACOB; COON, JONATHAN
To: IMPOSSIBLE VENTURES, LLC
Reel/Frame 063557/0077 →
CHANGE OF NAME Recorded May 5, 2023
From: CAPITAL ONE WMS LLC
To: WIKIBUY HOLDINGS, LLC
Reel/Frame 063557/0081 →
MERGER AND CHANGE OF NAME Recorded May 5, 2023
From: WIKIBUY HOLDINGS, LLC; CAPITAL ONE WMS LLC
To: CAPITAL ONE WMS LLC
Reel/Frame 063557/0087 →
CHANGE OF NAME Recorded May 5, 2023
From: IMPOSSIBLE VENTURES, LLC
To: WIKIBUY HOLDINGS, LLC
Reel/Frame 063557/0928 →
Continuity (7)
Continuation 17364432 · Jun 30, 2021
Continuation 16567367 · Sep 11, 2019
Continuation 15287089 · Oct 6, 2016
Provisional Application 62376243 · Aug 17, 2016
Provisional Application 62238574 · Oct 7, 2015
Provisional Application 62238565 · Oct 7, 2015
Related Publication 20230273920A1 · Aug 31, 2023
References Cited (52)
US 8065195B2 · Tarvydas et al. · 2011 [cited by applicant]
US 8577749B2 · Aliabadi et al. · 2013 [cited by applicant]
US 8600931B1 · Wehrle et al. · 2013 [cited by applicant]
US 8775275B1 · Pope · 2014 [cited by applicant]
US 8881303B2 · Liu et al. · 2014 [cited by applicant]
US RE45371E · Simons · 2015 [cited by applicant]
US 9189811B1 · Bhosle et al. · 2015 [cited by applicant]
US 9626688B2 · King · 2017 [cited by applicant]
US 9639853B2 · Shiffert et al. · 2017 [cited by applicant]
US 9766922B2 · Amershi et al. · 2017 [cited by applicant]
US 9798528B2 · Gao et al. · 2017 [cited by applicant]
US 9892099B2 · Cao · 2018 [cited by applicant]
US 9922327B2 · Johnson et al. · 2018 [cited by applicant]
US 9953335B2 · Shiffert et al. · 2018 [cited by applicant]
US 9965769B1 · Shiffert et al. · 2018 [cited by applicant]
US 20020087883A1 · Wohlgemuth et al. · 2002 [cited by applicant]
US 20050165789A1 · Minton et al. · 2005 [cited by applicant]
US 20060242266A1 · Keezer et al. · 2006 [cited by applicant]
US 20070180380A1 · Khavari et al. · 2007 [cited by applicant]
US 20080005079A1 · Flake et al. · 2008 [cited by applicant]
US 20080098300A1 · Corrales et al. · 2008 [cited by applicant]
US 20080189190A1 · Ferber · 2008 [cited by applicant]
US 20090182643A1 · Holstein et al. · 2009 [cited by applicant]
US 20090313352A1 · Dupont · 2009 [cited by applicant]
US 20100121810A1 · Bromenshenkel et al. · 2010 [cited by applicant]
US 20100241518A1 · McCann · 2010 [cited by applicant]
US 20110088036A1 · Patanella · 2011 [cited by applicant]
US 20110136516A1 · Ellis · 2011 [cited by applicant]
US 20120011431A1 · Mao · 2012 [cited by applicant]
US 20120198342A1 · Mahmud · 2012 [cited by applicant]
US 20120265637A1 · Moeggenberg · 2012 [cited by applicant]
US 20130151381A1 · Klein · 2013 [cited by applicant]
US 20130191723A1 · Pappas et al. · 2013 [cited by applicant]
US 20140100991A1 · Lenahan et al. · 2014 [cited by applicant]
US 20140229258A1 · Seriani · 2014 [cited by applicant]
US 20140229335A1 · Chen · 2014 [cited by applicant]
US 20140244429A1 · Clayton et al. · 2014 [cited by applicant]
US 20140281918A1 · Wei et al. · 2014 [cited by applicant]
US 20140325337A1 · Mcweeney · 2014 [cited by applicant]
US 20150287047A1 · Situ et al. · 2015 [cited by applicant]
US 20160063595A1 · Oral et al. · 2016 [cited by applicant]
US 20160191351A1 · Smith et al. · 2016 [cited by applicant]
US 20170104841A1 · Duke et al. · 2017 [cited by applicant]
US 20170171311A1 · Tennie et al. · 2017 [cited by applicant]
US 20170277764A1 · Osotio · 2017 [cited by applicant]
US 20180089737A1 · Ali et al. · 2018 [cited by applicant]
WO 2008121737A1 · 2008 [cited by applicant]
WO 2009061399A1 · 2009 [cited by applicant]
WO 2017062680A1 · 2017 [cited by applicant]
Schultz, NBA Data Scraping—Game Data, https://bigishdate.com/2015/05/31/nba-data-scraping-game-data/(Year: 2015). [cited by applicant]
Reda, Web Scraping 201: finding the API, http://www.gregreda.com/2015/02/15/we-scraping-finding-the-api/ (Year: 2015). [cited by applicant]
Moore, Nylon Calculus 101: Data Scraping With Python, http://web.archive.org/web/20150910031106/http://nyloncalculus.com/2015/09/07/nylon-calculus-101-data-scraping-with-python/ (Year: 2015). [cited by applicant]