IP Library › Patent Application 18962665
Patent Application
App. No. 18/962,665

EFFICIENT CRAWLING USING PATH SCHEDULING, AND APPLICATIONS THEREOF

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/962,665
Abstract

The present disclosure is directed to systems and methods for extracting unstructured data from a data source in a structure manner. Embodiments provide ways to retrieve unstructured data along from data sources not optimized for automated retrieval. For example, embodiments may generate a branched tree for each data source that maps out paths to individual sites of, for example, a healthcare provider listing the unstructured data. Using this branched tree, tasks can be generated to navigate along a path with the data source to each site and extract the unstructured data from the data source. In this way, embodiments provide the ability to navigate through a site from a base site to a site that has the relevant data.

Claims (58)

1 . (canceled)

2 . A computed-implemented method, comprising:

generating a decision tree for a plurality of data sources, wherein the decision tree comprises a data source at a root node of the decision tree, another data source at a leaf node of the decision tree, and a path from the data source at the root node to the other data source at the leaf node;

generating, based on the decision tree, a plurality of tasks associated with the plurality of data sources;

selecting, based on a priority level of the corresponding data source, a task from the plurality of tasks, wherein the task comprises an instruction corresponding to the path for extracting demographic data located at the other data source; and

parsing, based on the instruction of the task, the demographic data from the other data source into a category.

3 . The computer-implemented method of claim 2 , further comprising:

generating a user interface to be presented on a display, wherein the user interface indicates the plurality of tasks to be performed for each of the plurality of data sources and a status for the plurality of tasks.

4 . The computer-implemented method of claim 2 , further comprising:

navigating, based on the decision tree, from the data source at the root node to the other data source at the leaf node as specified by the path, wherein the path comprises a plurality of steps to navigate from the root node to the leaf node.

5 . The computer-implemented method of claim 2 , further comprising:

iteratively accessing the other data source for a predetermined number of attempts when the other data source is initially inaccessible; and

receiving an error notification when the other data source is inaccessible after completing the predetermined number of attempts.

6 . The computer-implemented method of claim 2 , further comprising:

managing a plurality of data extractors performing the plurality tasks for each of the plurality of data sources; and

in response to a maximum number of the plurality of data extractors for a first data source of the plurality of data sources being reached, assigning tasks of a second data source of the plurality of data sources having a same priority level as the first data source.

7 . The computer-implemented method of claim 2 , further comprising:

storing the parsed demographic data in a database based on the category.

8 . The computer-implemented method of claim 2 , further comprising:

generating a report based on the parsed demographic data that displays the parsed demographic data in a structured format.

9 . A system, comprising:

a memory configured to store operations; and

one or more processors configured to perform the operations, the operations comprising:

generating a decision tree for a plurality of data sources, wherein the decision tree comprises a data source at a root node of the decision tree, another data source at a leaf node of the decision tree, and a path from the data source at the root node to the other data source at the leaf node;

generating, based on the decision tree, a plurality of tasks associated with the plurality of data sources;

selecting, based on a priority level of the corresponding data source, a task from the plurality of tasks, wherein the task comprises an instruction corresponding to the path for extracting demographic data located at the other data source; and

parsing, based on the instruction of the task, the demographic data from the other data source into a category.

10 . The system of claim 9 , wherein the operations further comprise:

generating a user interface to be presented on a display, wherein the user interface indicates the plurality of tasks to be performed for each of the plurality of data sources and a status for the plurality of tasks.

11 . The system of claim 9 , wherein the operations further comprise:

navigating, based on the decision tree, from the data source at the root node to the other data source at the leaf node as specified by the path, wherein the path comprises a plurality of steps to navigate from the root node to the leaf node.

12 . The system of claim 9 , wherein the operations further comprise:

iteratively accessing the other data source for a predetermined number of attempts when the other data source is initially inaccessible; and

receiving an error notification when the other data source is inaccessible after completing the predetermined number of attempts.

13 . The system of claim 9 , wherein the operations further comprise:

managing a plurality of data extractors performing the plurality tasks for each of the plurality of data sources; and

in response to a maximum number of the plurality of data extractors for a first data source of the plurality of data sources being reached, assigning tasks of a second data source of the plurality of data sources having a same priority level as the first data source.

14 . The system of claim 9 , wherein the operations further comprise:

storing the parsed demographic data in a database based on the category.

15 . The system of claim 9 , wherein the operations further comprise:

generating a report based on the parsed demographic data that displays the parsed demographic data in a structured format.

16 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating a decision tree for a plurality of data sources, wherein the decision tree comprises a data source at a root node of the decision tree, another data source at a leaf node of the decision tree, and a path from the data source at the root node to the other data source at the leaf node;

generating, based on the decision tree, a plurality of tasks associated with the plurality of data sources;

selecting, based on a priority level of the corresponding data source, a task from the plurality of tasks, wherein the task comprises an instruction corresponding to the path for extracting demographic data located at the other data source; and

parsing, based on the instruction of the task, the demographic data from the other data source into a category.

17 . The non-transitory computer-readable medium according to claim 16 , wherein the operations further comprise:

generating a user interface to be presented on a display, wherein the user interface indicates the plurality of tasks to be performed for each of the plurality of data sources and a status for the plurality of tasks.

18 . The non-transitory computer-readable medium according to claim 16 , wherein the operations further comprise:

navigating, based on the decision tree, from the data source at the root node to the other data source at the leaf node as specified by the path, wherein the path comprises a plurality of steps to navigate from the root node to the leaf node.

19 . The non-transitory computer-readable medium according to claim 16 , wherein the operations further comprise:

iteratively accessing the other data source for a predetermined number of attempts when the other data source is initially inaccessible; and

receiving an error notification when the other data source is inaccessible after completing the predetermined number of attempts.

20 . The non-transitory computer-readable medium according to claim 16 , wherein the operations further comprise:

managing a plurality of data extractors performing the plurality tasks for each of the plurality of data sources; and

in response to a maximum number of the plurality of data extractors for a first data source of the plurality of data sources being reached, assigning tasks of a second data source of the plurality of data sources having a same priority level as the first data source.

21 . The non-transitory computer-readable medium according to claim 16 , wherein the operations further comprise:

generating a report based on the parsed demographic data that displays the parsed demographic data in a structured format.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2025
From: VERA-CIRO, CARLOS; LINDNER, ROBERT RAYMOND
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071855/0268 →