IP Library › Granted Patent US 12,198,789
Granted Patent B2
US 12,198,789 · App. 16/668,524 · Granted Jan 14, 2025

Efficient crawling using path scheduling, and applications thereof

Inventors: Carlos Vera-Ciro (Madison, WI); Robert Raymond Lindner (Fitchburg, WI)
Assignee: VEDA Data Solutions, Inc.
G16H10/60G06F16/21G06F40/14G16H40/20G16H50/20G16H50/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,789
App. No.
16/668,524
Granted
Jan 14, 2025
Kind
B2
Abstract

The present disclosure is directed to systems and methods for extracting unstructured data from a data source in a structure manner. Embodiments provide ways to retrieve unstructured data along from data sources not optimized for automated retrieval. For example, embodiments may generate a branched tree for each data source that maps out paths to individual sites of, for example, a healthcare provider listing the unstructured data. Using this branched tree, tasks can be generated to navigate along a path with the data source to each site and extract the unstructured data from the data source. In this way, embodiments provide the ability to navigate through a site from a base site to a site that has the relevant data.

Claims (49)

1. A computer-implemented method for extracting data from a plurality of data sources, comprising:

generating a decision tree for a data source of the plurality of data sources, wherein the decision tree comprises a base website at a root node and a plurality of respective sites represented as corresponding leaf nodes of the decision tree, wherein the decision tree specifies one or more paths to navigate from the base website to the plurality of respective sites containing unstructured demographic data for a healthcare provider, wherein each of the plurality of respective sites and the corresponding leaf nodes of the decision tree represent a step in the decision tree, wherein a respective site is a website accessible from the base website of the data source;

generating, based on the decision tree, a list of tasks corresponding to each of the plurality of data sources, wherein each task includes instructions for how to extract the unstructured demographic data corresponding to the respective site corresponding to one of the one or more paths;

generating a user interface to be presented on a display, wherein the user interface indicates the list of tasks to be performed for each of the plurality of data sources and a status for the list of tasks;

navigating, based on the decision tree, within a corresponding data source from the base website to the respective site as specified by a specified path, wherein the specified path is a plurality of steps to navigate from the base website to the respective site where the unstructured demographic data is located;

parsing the unstructured demographic data from the respective site into separate categories;

storing the parsed unstructured demographic data in separate databases based on the separate categories; and

generating a report based on the parsed unstructured demographic data that displays the parsed unstructured demographic data in a structured format.

2. The method of claim 1 , wherein the navigating comprises iteratively accessing the respective site for a predetermined number of attempts when the corresponding data source or the respective site is initially inaccessible.

3. The method of claim 2 , further comprising receiving an error notification when the corresponding data source or the respective site is inaccessible after completing the predetermined number of attempts.

4. The method of claim 1 , wherein the user interface further indicates a color code indicator of a priority level of each of the plurality of data sources.

5. The method of claim 1 , further comprising managing a plurality of data extractors performing tasks on each of the plurality of data sources.

6. The method of claim 5 , wherein managing the plurality of data extractors comprises managing a maximum number of data extractors performing tasks on each of the plurality of data sources.

7. The method of claim 6 , wherein when the maximum number of data extractors for a first data source of the plurality of data sources is reached, the method further comprises assigning tasks of a second data source of the plurality of data sources having a same priority level as the first data source.

8. The method of claim 6 , wherein when the maximum number of data extractors for a first data source is reached, the method further comprises assigning tasks of a second data source of the plurality of data sources having a different priority level as the first data source.

9. The method of claim 5 , wherein managing the plurality of data extractors comprises periodically adjusting the plurality of data extractors performing tasks on the corresponding data source.

10. A non-transitory program storage device having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform a method, the method comprising:

generating a decision tree for a data source of a plurality of data sources, wherein the decision tree comprises a base website at a root node and a plurality of respective sites represented as corresponding leaf nodes of the decision tree, wherein the decision tree specifies one or more paths to navigate from the base website to the plurality of respective sites containing unstructured demographic data for a healthcare provider, wherein each of the plurality of respective sites and the corresponding leaf nodes of the decision tree represent a step in the decision tree, wherein a respective site is a website accessible from the base website of the data source;

generating, based on the decision tree, a list of tasks corresponding to each of the plurality of data sources, wherein each task includes instructions for how to extract the unstructured demographic data corresponding to the respective site corresponding to one of the one or more paths;

generating a user interface to be presented on a display, wherein the user interface indicates the list of tasks to be performed for each of the plurality of data sources and a status for the list of tasks;

navigating, based on the decision tree, within a corresponding data source from the base website to the respective site as specified by a specified path, wherein the specified path is a plurality of steps to navigate from the base website to the respective site where the unstructured demographic data is located;

parsing the unstructured demographic data from the respective site into separate categories;

storing the parsed unstructured demographic data in separate databases based on the separate categories; and

generating a report based on the parsed unstructured demographic data that displays the parsed unstructured demographic data in a structured format.

11. The program storage device of claim 10 , wherein the navigating comprises iteratively accessing the respective site for a predetermined number of attempts when the corresponding data source or the respective site is initially inaccessible.

12. The program storage device of claim 11 , the instructions further comprising receiving an error notification when the corresponding data source or the respective site is inaccessible after completing the predetermined number of attempts.

13. The program storage device of claim 10 , wherein the user interface further indicates a color code indicator of a priority level of each of the plurality of data sources.

14. The program storage device of claim 10 , the instructions further comprising managing a plurality of data extractors performing tasks on each of the plurality of data sources.

15. The program storage device of claim 14 , wherein managing the plurality of data extractors comprises managing a maximum number of data extractors performing tasks on each of the plurality of data sources.

16. The program storage device of claim 15 , wherein when the maximum number of data extractors for a first data source of the plurality of data sources is reached, the method further comprises assigning tasks of a second data source of the plurality of data sources having a same priority level as the first data source.

17. The program storage device of claim 15 , wherein when the maximum number of data extractors for a first data source is reached, the method further comprises assigning tasks of a second data source of the plurality of data sources having a different priority level as the first data source.

18. The program storage device of claim 14 , wherein managing the plurality of data extractors comprises periodically adjusting the plurality of data extractors performing tasks on the corresponding data source.

19. A system comprising:

a first computing device comprising:

a first memory; and

a first processor communicatively coupled to the first memory and configured to:

generate a decision tree for a data source of a plurality of data sources, wherein the decision tree comprises a base website at a root node and a plurality of respective sites represented as corresponding leaf nodes of the decision tree, wherein the decision tree specifies one or more paths to navigate from the base website to the plurality of respective sites containing unstructured demographic data for a healthcare provider, wherein each of the plurality of respective sites and the corresponding leaf nodes of the decision tree represent a step in the decision tree, wherein a respective site is a website accessible from the base website of the data source;

generate a list of tasks for each of the plurality of data sources based on the decision tree, wherein each task includes instructions for how to extract the unstructured demographic data corresponding to the respective site of the one or more paths and comprises instructions for extracting the unstructured demographic data from the respective site;

assign a task from the list of tasks to a second computing device based on a priority level of a corresponding data source;

generate a user interface to be presented on a display, wherein the user interface indicates the list of tasks to be performed for each of the plurality of data sources and a status for the list of tasks; and

transmit the assigned task to the second computing device; the second computing device comprising:

a second memory; and

a second processor communicatively coupled to the second memory and configured to:

execute the assigned task to navigate, based on the decision tree, the corresponding data source to the respective site and extract the unstructured demographic data from the respective site based on the assigned task; and

transmit the extracted unstructured demographic data to the first computing device, wherein upon receipt of the extracted unstructured demographic data, the first processor is further configured to:

parse the extracted unstructured demographic data into separate categories;

store the parsed unstructured demographic data in separate databases based on the separate categories; and

generate a report based on the parsed unstructured demographic data that displays the parsed unstructured demographic data in a structured format.

20. The system of claim 19 , wherein the user interface further indicates a color code indicator of a priority level of each of the plurality of data sources.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2025
From: COMERICA BANK
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071309/0392 →
SECURITY INTEREST Recorded Nov 27, 2023
From: VEDA DATA SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 065668/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2019
From: VERA-CIRO, CARLOS; LINDNER, ROBERT RAYMOND
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 050883/0259 →
Continuity (1)
Related Publication 20210134407A1 · May 6, 2021
References Cited (90)
US 5664109A · Johnson · 1997 [cited by examiner]
US 6606625B1 · Muslea · 2003 [cited by examiner]
US 6714941B1 · Lerman · 2004 [cited by examiner]
US 6725425B1 · Rajan · 2004 [cited by examiner]
US 6732102B1 · Khandekar · 2004 [cited by examiner]
US 6941318B1 · Tamayo et al. · 2005 [cited by applicant]
US 7103838B1 · Krishnamurthy · 2006 [cited by examiner]
US 7200804B1 · Khavari · 2007 [cited by examiner]
US 7523129B1 · Bent · 2009 [cited by examiner]
US 7774742B2 · Gupta · 2010 [cited by examiner]
US 8042112B1 · Zhu · 2011 [cited by examiner]
US 9171080B2 · Song · 2015 [cited by examiner]
US 11030691B2 · Parmar · 2021 [cited by examiner]
US 20030195889A1 · Yao et al. · 2003 [cited by applicant]
US 20040015784A1 · Chidlovskii · 2004 [cited by examiner]
US 20040172307A1 · Gruber · 2004 [cited by examiner]
US 20050022115A1 · Baumgartner · 2005 [cited by examiner]
US 20050229151A1 · Gupta · 2005 [cited by examiner]
US 20060242180A1 · Graf · 2006 [cited by examiner]
US 20070094060A1 · Apps · 2007 [cited by examiner]
US 20070198727A1 · Guan · 2007 [cited by examiner]
US 20080027895A1 · Combaz · 2008 [cited by examiner]
US 20080091663A1 · Inala · 2008 [cited by examiner]
US 20090055727A1 · Hansen · 2009 [cited by examiner]
US 20090132524A1 · Stouffer · 2009 [cited by examiner]
US 20100162097A1 · Dalvi · 2010 [cited by examiner]
US 20110145218A1 · Meyerzon · 2011 [cited by examiner]
US 20110185273A1 · DaCosta · 2011 [cited by examiner]
US 20110258195A1 · Welling · 2011 [cited by examiner]
US 20110307479A1 · Yin · 2011 [cited by examiner]
US 20120124086A1 · Song · 2012 [cited by examiner]
US 20130166207A1 · Shao et al. · 2013 [cited by applicant]
US 20130236111A1 · Pintsov · 2013 [cited by applicant]
US 20140172754A1 · He et al. · 2014 [cited by applicant]
US 20140359411A1 · Botta · 2014 [cited by examiner]
US 20150161257A1 · Shivaswamy · 2015 [cited by examiner]
US 20150199744A1 · Tolvanen · 2015 [cited by examiner]
US 20160092458A1 · Gottlob · 2016 [cited by examiner]
US 20160092730A1 · Smirnov et al. · 2016 [cited by applicant]
US 20160188717A1 · Rosenberg et al. · 2016 [cited by applicant]
US 20160371603A1 · A V · 2016 [cited by examiner]
US 20170104841A1 · Duke · 2017 [cited by examiner]
US 20180018429A1 · Rice · 2018 [cited by examiner]
US 20180060495A1 · Mahapatra · 2018 [cited by examiner]
US 20180083901A1 · McGregor, Jr. · 2018 [cited by examiner]
US 20180150562A1 · Gundimeda et al. · 2018 [cited by applicant]
US 20190172586A1 · Choksi · 2019 [cited by examiner]
US 20190287171A1 · Parmar · 2019 [cited by examiner]
US 20200089712A1 · Tripathi · 2020 [cited by examiner]
US 20200210511A1 · Korobov · 2020 [cited by examiner]
Hong-ye, C., “Method of Web Information Extraction Based on Decision Tree,” 2009, IEEE, pp. 664-666. [cited by examiner]
Crescenzi, V. et al., “RoadRunner: Towards Automatic Data Extraction from Large Web Sites,” 2001, 19 pages. [cited by examiner]
Hsu, C-N. et al., “Generating Finite-State Transducers for Semi-Structured Data Extraction From the Web,” (1998), Elsevier Science pp. 521-538. [cited by examiner]
Arasu, A. et al., “Extracting Structured Data From Web Pages,” (2003), ACM, pp. 337-348. [cited by examiner]
Cho, J. et al., “Efficient Crawling Through URL Ordering,” Computer Networks and ISDN Systems 30 (1998) 161-172. [cited by examiner]
Hersovici, M. et al., “The Shark-Search Algorithm: An Application: Tailored Web Site Mapping,” Computer Networks and ISDN Systems 30 (1998) 317-326. [cited by examiner]
McCallum, A. et al., “Buidling Domain-Specific Search Engines with Machine Learning Techniques,” 2009, 12 pages. [cited by examiner]
Diligenti, M. et al., “Focused Crawling using Context Graphs,” 26th International Conference on Very Large Databases, VLDB 2000, Cairo, Egypt, pp. 527-534, 2000. [cited by examiner]
Shen, W. et al., “An Algorithm on Web Article Automatic Extraction Based on DOM Structure,” International Journal of Hybrid Information Technology vol. 8, No. 3 (2015), pp. 243-254. [cited by examiner]
Meng, X. et al., “A Supervised Visual Wrapper Generator for Web-Data Extraction,” Proceedings of the 27th Annual International Computer Software and Applications Conference (Compsac'03) 2003, 6 pages. [cited by examiner]
Papadakis, N. K. et al., “Stavies: A System for Information Extraction from Unknown Web Data Sources Through Automatic Web Wrapper Generation Using Clustering Techniques,” IEEE Transactions on Knowledge and Data Enginee… [cited by examiner]
Liu, L. et al., “XRWAP: An XML-enabled Wrapper Construction System for Web Information Sources,” (2000), 11 pages. [cited by examiner]
Lin, L. et al., “Using Structured Tokens to Identify Webpages for Data Extraction,” G. Dong et al. (Eds.): APWeb/WAIM 2007, LNCS 4505, pp. 241-252, 2007, Springer-Verlag Berlin Heidelberg 2007. [cited by examiner]
Shete, D. et al., “Survey Paper on Web Content Extraction & Classification,” 2021 6th International Conference for Convergence in Technology (12CT) Pune, India. Apr. 2-4, 2021, IEEE, pp. 1-6. [cited by examiner]
Uzun, E. et al.,“An Effective and Efficient Web Content Extractor for Optimizing the Crawling Process,” Softw. Pract. Exper. 2014; 44:1181-1199. [cited by examiner]
Liu, Yang et al. “Extracting Patient Demographics and Personal Medical Information from Online Health Forums.” AMIA . . . Annual Symposium proceedings. AMIA Symposium 2014 (2014): 1825-34. [cited by examiner]
Pol K. et al., “A Survey on Web Content Mining and Extraction of Structured and Semistructured Data,” (2008_, IEEE, pp. 543-546. [cited by examiner]
Bhargavi, P. et al., “Knowledge Extraction Using Rule Based Decision Tree Approach,” IJCSNS International Journal of Computer Science and Network 296 Security, vol. 8 No. 7, Jul. 2008. [cited by examiner]
Hong-ye, C., “Method of Web Information Extraction Based on Decision Tree,” (2009), IEEE, 3 pages. [cited by examiner]
Esposito, F. et al., “A Comparative Analysis of Method for Pruning Decision Trees,” May 1997, IEEE, pp. 476-491. [cited by examiner]
Yang, J. et al., “An Interface Agent for Wrapper-Based Information Extraction,” 2005, Springer, pp. 291-302. [cited by examiner]
Shah, P.B. et al., “Survery Based on DOM and Visual Ques for Extracting Structure Data From Web,” (2015), IITRD, pp. 1-8. [cited by examiner]
Kumar, A. et al., “Novel Self-Learning Based Crawling and Data Mining for Automatic Information Extraction,” 2015, IEEE, pp. 732-738. [cited by examiner]
Finch, D.K., “TagLine: Information Extraction for Semi-Structured Text Elements in Medical Progress Notes,” Jan. 2012, Univ. of South Florida, 251 total pages. [cited by examiner]
Safavian, S.R. et al., “A Survey of Decision Tree Classifier Methodology,” (1991), IEEE, pp. 660-674. [cited by examiner]
Taniguchi, K. et al., “Mining Semi-Structured Data by Path Expressions,” (2001), Springer, pp. 378-388. [cited by examiner]
Sigursson, K. et al., “Heritrix User Manual,” (2004), Internet Archive, 57 pages. [cited by examiner]
Girardi, C., Ricca, F. and Tonella, P. (2006), “Web crawlers compared”, International Journal of Web Information Systems, vol. 2 No. 2, pp. 85-94. [cited by examiner]
Shi, S. et al. NEXIR: A Novel Web Extraction Rule Language toward a Three-Stage Web Data Extraction Model. In: Lin, X., Manolopoulos, Y., Srivastava, D., Huang, G. (eds) Web Information Systems Engineering—WISE 2013. WI… [cited by examiner]
Furche, T., Gottlob, G., Grasso, G. et al. Oxpath: A language for scalable data extraction, automation, and crawling on the deep web. The VLDB Journal 22, 47-72 (2013). (Year: 2013). [cited by examiner]
Ferrara, Emilio & De Meo, Pasquale & Fiumara, Giacomo & Baumgartner, Robert. (2012). Web Data Extraction, Applications and Techniques: A Survey. ACM Computing Surveys (Under Review). 70. 10.1016/j.knosys.2014.07.007. (Y… [cited by examiner]
Baumgartner, R., Gatterbauer, W., Gottlob, G. (2009). Web Data Extraction System. In: Liu, L., Özsu, M.T. (eds) Encyclopedia of Database Systems. Springer, Boston, MA (Year: 2009). [cited by examiner]
A. Schulz, J. Lässig and M. Gaedke, “Practical Web Data Extraction: Are We There Yet?—A Short Survey,” 2016 IEEE/WIC/ACM International Conference on Web Intelligence (WI), Omaha, NE, USA, 2016, pp. 562-567 (Year: 2016). [cited by examiner]
S. M. Meystre, G. K. Savova, K. C. Kipper-Schuler, and J. F. Hurdle, “Extracting information from textual documents in the electronic health record: a review of recent research,” Yearb Med Inform, pp. 128-144, 2008. (Ye… [cited by examiner]
Kahn, Asad, Text-Mining and Analysis of the Doctor's Meta-data and Text-Reviews Using Topic-Modeling (LDA) Techniques, Sep. 2019, available Oct. 30, 2019, 115 total pages. (Year: 2019). [cited by examiner]
Neustein, Amy & Imambi, S. & Teixeira, António & Ferreira, Liliana & Rodrigues, Mario. (2014). Application of Text Mining to Biomedical Knowledge Extraction: Analyzing Clinical Narratives and Medical Literature. (Year: … [cited by examiner]
Yoon Ho Cho, et al. “A personalized recommender system based on web usage mining and decision tree induction,” Expert Systems with Applications, vol. 23, Iss. 3, 2002; https://doi.org/10.1016/S0957-4174(02)00052-0. (Yea… [cited by examiner]
International Search Report mailed Feb. 2, 2021 for Appl. No. PCT/US2020/58286, 3 pages. [cited by applicant]
The Written Opinion mailed Feb. 2, 2021 for Appl. No. PCT/US2020/58286, 8 pages. [cited by applicant]
Extended European Search Report directed to related European Application No. 20883009.1, mailed on Oct. 5, 2023, 8 pages. [cited by applicant]