IP Library › Granted Patent US 12,602,392
Granted Patent B2
US 12,602,392 · App. 17/670,896 · Granted Apr 14, 2026

Scalable metadata-driven data ingestion pipeline

Inventor: Dirk Renick (Albany, CA)
Assignee: Insight Direct USA, Inc.
G06F16/254G06F16/2358G06F16/258G06F21/604G06F21/6218
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,392
App. No.
17/670,896
Granted
Apr 14, 2026
Kind
B2
Abstract

A method of ingesting data into a data lake includes initiating a data pipeline by performing a lookup activity that includes querying a data structure to retrieve metadata corresponding to data sources, and performing a sequence of activities for each of the data sources. The metadata stored in the data structure includes information that enables identifying and connecting to each of the data sources. The sequence of activities includes connecting to a respective one of the data sources using a portion of the metadata that corresponds to the respective one of the data sources; accessing, from the respective one of the data sources, a data object specified by the portion of the metadata that corresponds to the respective one of the data sources; and storing the data object in the data lake, the data lake being remote from each of the data sources.

Claims (106)

1 . A method of ingesting data into a data lake, the method comprising:

consolidating pre-entered metadata enabling an identification of multiple data sources of multiple data pipelines, and a connection to each of the multiple data sources, into a common metadata table for an overarching data pipeline prior to initiating the overarching data pipeline, wherein the pre-entered metadata of the metadata table includes values associated with each of the multiple data sources that specify, for each of the multiple data sources, at least one of:

a data source identification;

a data source name;

whether a respective one of the multiple data sources is enabled; and

a key value that is associated with access credentials for the respective one of the multiple data sources;

initiating, by a computer device, the overarching data pipeline by performing a lookup activity, the lookup activity including querying the metadata table that includes the pre-entered metadata, wherein the querying specifies a plurality of data sources from among the multiple data sources;

creating an overarching pipeline start log corresponding to a start of the overarching data pipeline when the overarching data pipeline is initiated;

retrieving, based on the querying, respective metadata from the metadata table that corresponds to the specified plurality of data sources of the multiple data sources; and

performing, by the computer device, a sequence of activities for each data source of the specified plurality of data sources, the sequence of activities including:

creating a data source pipeline start log corresponding to a start of the sequence of activities for a respective one of the specified plurality of data sources before forming a connection to the respective one of the specified plurality of data sources;

connecting to the respective one of the specified plurality of data sources using a portion of the metadata from the metadata table that corresponds to the respective one of the specified plurality of data sources to form the connection to the respective one of the specified plurality of data sources by:

providing connection information to a linked services module based on the portion of the metadata from the metadata table that corresponds to the respective one of the specified plurality of data sources;

forming the connection to the respective one of the specified plurality of data sources via the linked services module;

connecting, via the linked services module, to a secure storage that contains the associated access credentials for the respective one of the specified plurality of data sources;

obtaining the access credentials for the respective one of the specified plurality of data sources from the secure storage by passing the key value to the secure storage via the linked services module; and

using the access credentials for the respective one of the specified plurality of data sources in forming the connection to the respective one of the specified plurality of data sources via the linked services module;

accessing, from the respective one of the specified plurality of data sources via the connection, a data object specified by the portion of the metadata that corresponds to the respective one of the specified plurality of data sources; and

storing the data object in the data lake, the data lake being remote from each of the multiple data sources;

creating a data source pipeline end log corresponding to an end of the sequence of activities for the respective one of the specified plurality of data sources when the data object is stored in the data lake; and

creating an overarching data pipeline end log corresponding to an end of the overarching data pipeline after the sequence of activities has been performed for each of the specified plurality of data sources.

2 . The method of claim 1 , wherein initiating the overarching data pipeline further comprises initiating the overarching data pipeline manually by passing arguments to parameters of the overarching data pipeline at a runtime of the overarching data pipeline.

3 . The method of claim 1 , wherein initiating the overarching data pipeline further comprises initiating the overarching data pipeline by scheduling an instance of the overarching data pipeline to occur automatically after a defined period elapses or after an occurrence of a trigger event.

4 . The method of claim 1 , wherein connecting to the respective one of the specified plurality of data sources further comprises:

connecting to a secure storage that contains access credentials for the respective one of the specified plurality of data sources;

obtaining the access credentials for the respective one of the specified data sources from the secure storage; and

using the access credentials for the respective one of the specified plurality of data sources in forming the connection to the respective one of the specified plurality of data sources.

5 . The method of claim 1 and further comprising:

transforming the data object from the data lake into a different format such that the data object is compatible with a different database.

6 . The method of claim 1 , wherein the data structure is stored in an SQL database.

7 . The method of claim 1 , wherein the pre-entered metadata of the metadata table further includes values associated with each of the multiple data sources that specify, for each of the multiple data sources, at least one of:

a data object name; and

a last execution date.

8 . The method of claim 1 , wherein storing the data object in the data lake further comprises incorporating values from the portion of the metadata that corresponds to the respective one of the specified plurality of data sources into a file name of the data object to enable identifying the data object in the data lake.

9 . The method of claim 8 , wherein the file name of the data object includes at least one of:

a data source name that corresponds to the respective one of the specified plurality of data sources; and

a data object name that corresponds to the data object in the respective one of the specified plurality of data sources.

10 . The method of claim 1 , wherein storing the data object in the data lake further comprises storing the data object in a Parquet format.

11 . The method of claim 10 and further comprising:

transforming the data object from the Parquet format into a different format such that the data object is compatible with a different database that is separate from the data lake.

12 . The method of claim 1 , wherein storing the data object in the data lake further comprises creating a copy of the data object from the respective one of the specified plurality of data sources and storing the copy of the data object in the data lake.

13 . The method of claim 12 , wherein storing the copy of the data object in the data lake further comprises writing the copy of the data object to a container within the data lake, the container being a destination for respective copies of data objects accessed from each of the specified plurality of data sources.

14 . The method of claim 13 , wherein writing the copy of the data object to the container within the data lake further comprises incorporating values from the portion of the metadata that corresponds to the respective one of the specified plurality of data sources into a file name of the copy of the data object to enable identifying the copy of the data object in the data lake.

15 . The method of claim 1 , wherein storing the data object in the data lake further comprises writing the data object to a container within the data lake, the container being a destination for respective data objects accessed from each of the specified plurality of data sources.

16 . The method of claim 1 and further comprising:

storing the overarching data pipeline start log and the overarching data pipeline end log in a log table on a database.

17 . The method of claim 16 and further comprising sending a notification corresponding to the end of the overarching data pipeline after the sequence of activities has been performed for each of the specified plurality of data sources, wherein the notification is an email.

18 . The method of claim 1 and further comprising:

storing the data source pipeline start log and the data source pipeline end log for each of the specified plurality of data sources in a log table on a database.

19 . The method of claim 1 and further comprising:

creating an error log and sending a notification if an error occurs while initiating the overarching data pipeline or while performing the sequence of activities, wherein the notification is an email; and

storing the error log in a log table on a database.

20 . The method of claim 1 , wherein the data object comprises a new or updated portion of data stored in the respective one of the specified plurality of data sources.

21 . The method of claim 20 , wherein storing the data object in the data lake further comprises creating a copy of the new or updated portion of the data from the respective one of the specified plurality of data sources and storing the copy in the data lake.

22 . The method of claim 21 , wherein storing the copy in the data lake further comprises writing the copy to a container within the data lake, the container being a destination for respective copies of new or updated portions of data accessed from each of the specified plurality of data sources.

23 . The method of claim 22 , wherein writing the copy to the container within the data lake further comprises incorporating values from the portion of the metadata that corresponds to the respective one of the specified plurality of data sources into a file name of the copy to enable identifying the copy in the data lake.

24 . A method of ingesting data into a data lake, the method comprising:

consolidating pre-entered metadata enabling an identification of multiple data sources of multiple data pipelines, and a connection to each of the multiple data sources, into a common metadata table for an overarching data pipeline prior to initiating the overarching data pipeline, wherein the pre-entered metadata of the metadata table includes values associated with each of the multiple data sources that specify, for each of the multiple data sources, at least one of:

a data source identification;

a data source name;

whether a respective one of the multiple data sources is enabled; and

a key value that is associated with access credentials for the respective one of the multiple data sources;

initiating, by a computer device, the overarching data pipeline by performing a lookup activity, the lookup activity including querying the metadata table that includes the pre-entered metadata, wherein the querying specifies at least one attribute of a plurality of data sources of the multiple data sources;

creating an overarching pipeline start log corresponding to a start of the overarching data pipeline when the overarching data pipeline is initiated;

retrieving, based on the querying, respective metadata from the metadata table that corresponds to the specified plurality of data sources of the multiple data sources corresponding to the at least one attribute; and

performing, by the computer device, a sequence of activities for each data source of the specified plurality of data sources, the sequence of activities including:

creating a data source pipeline start log corresponding to a start of the sequence of activities for a respective one of the specified plurality of data sources before forming a connection to the respective one of the specified plurality of data sources;

connecting to the respective one of the specified plurality of data sources using a portion of the metadata from the metadata table that corresponds to the respective one of the specified plurality of data sources to form the connection to the respective one of the specified plurality of data sources by:

providing connection information to a linked services module based on the portion of the metadata that corresponds to the respective one of the specified plurality of data sources;

forming the connection to the respective one of the specified plurality of data sources via the linked services module;

connecting, via the linked services module, to a secure storage that contains the associated access credentials for the respective one of the specified plurality of data sources;

obtaining the access credentials for the respective one of the specified plurality of data sources from the secure storage by passing the key value to the secure storage via the linked services module; and

using the access credentials for the respective one of the specified plurality of data sources in forming the connection to the respective one of the specified plurality of data sources via the linked services module;

accessing, from the respective one of the specified plurality of data sources via the connection, a data object specified by the portion of the metadata that corresponds to the respective one of the specified plurality of data sources; and

storing the data object in the data lake, the data lake being remote from each of the multiple data sources;

creating a data source pipeline end log corresponding to an end of the sequence of activities for the respective one of the specified plurality of data sources when the data object is stored in the data lake; and

creating an overarching pipeline end log corresponding to an end of the overarching data pipeline after the sequence of activities has been performed for each of the specified plurality of data sources.

25 . The method of claim 24 and further comprising:

configuring the overarching data pipeline to be initiated manually by passing arguments to parameters of the overarching data pipeline at a runtime of the overarching data pipeline; or

configuring the overarching data pipeline to be initiated by scheduling an instance of the overarching data pipeline to occur automatically after a defined period elapses or after an occurrence of a trigger event.

26 . The method of claim 24 and further comprising:

defining a first connection to a database that stores the metadata table via a linked services module; and

defining a second connection to the data lake via the linked services module.

27 . A data pipeline system for ingesting data into cloud storage, the system comprising:

one or more processors; and

computer-readable memory encoded with instructions that, when executed by the one or more processors, cause the data pipeline system to:

consolidate pre-entered metadata enabling an identification of multiple data sources of multiple data pipelines, and a connection to each of the multiple data sources, into a common metadata table for an overarching data pipeline prior to initiating the overarching data pipeline, wherein the pre-entered metadata of the metadata table includes values associated with each of the multiple data sources that specify, for each of the multiple data sources, at least one of:

a data source identification;

a data source name;

whether a respective one of the multiple data sources is enabled; and

a key value that is associated with access credentials for the respective one of the multiple data sources;

initiate the overarching data pipeline by performing a lookup activity, the lookup activity including querying the metadata table that includes the pre-entered metadata, wherein the querying specifies a plurality of data sources of the multiple data sources;

create an overarching pipeline start log corresponding to a start of the overarching data pipeline when the overarching data pipeline is initiated;

retrieve, based on the querying, metadata that corresponds to the specified plurality of data sources of the multiple data sources; and

perform a sequence of activities for each data source of the specified data sources, the sequence of activities including:

creating a data source pipeline start log corresponding to a start of the sequence of activities for a respective one of the specified plurality of data sources before forming a connection to the respective one of the specified plurality of data sources;

connecting to the respective one of the specified plurality of data sources using a portion of the metadata from the metadata table that corresponds to the respective one of the specified plurality of data sources to form the connection to the respective one of the specified data sources by:

providing connection information to a linked services module based on the portion of the metadata that corresponds to the respective one of the specified data sources;

forming the connection to the respective one of the specified data sources via the linked services module;

connecting, via the linked services module, to a secure storage that contains the associated access credentials for the respective one of the specified plurality of data sources;

obtaining the access credentials for the respective one of the specified plurality of data sources from the secure storage by passing the key value to the secure storage via the linked services module; and

using the access credentials for the respective one of the specified plurality of data sources in forming the connection to the respective one of the specified plurality of data sources via the linked services module;

accessing, from the respective one of the specified plurality of data sources via the connection, a data object specified by the portion of the metadata that corresponds to the respective one of the specified plurality of data sources; and

storing the data object in the data lake, the data lake being remote from each of the multiple data sources;

creating a data source pipeline end log corresponding to an end of the sequence of activities for the respective one of the specified plurality of data sources when the data object is stored in the data lake; and

create an overarching pipeline end log corresponding to an end of the overarching data pipeline after the sequence of activities has been performed for each of the specified plurality of data sources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2022
From: RENICK, DIRK
To: INSIGHT DIRECT USA, INC.
Reel/Frame 059001/0408 →
Continuity (1)
Related Publication 20230259518A1 · Aug 17, 2023
References Cited (123)
US 6161103A · Rauer et al. · 2000 [cited by applicant]
US 6189004B1 · Rassen et al. · 2001 [cited by applicant]
US 6212524B1 · Weissman et al. · 2001 [cited by applicant]
US 7461076B1 · Weissman et al. · 2008 [cited by applicant]
US 7475080B2 · Chowdhary et al. · 2009 [cited by applicant]
US 7681185B2 · Kapoor et al. · 2010 [cited by applicant]
US 7720804B2 · Fazal et al. · 2010 [cited by applicant]
US 7739224B1 · Weissman et al. · 2010 [cited by applicant]
US 7747563B2 · Gehring · 2010 [cited by applicant]
US 7792921B2 · Taylor et al. · 2010 [cited by applicant]
US 7844570B2 · Netz et al. · 2010 [cited by applicant]
US 8086583B2 · Crutchfield et al. · 2011 [cited by applicant]
US 8311975B1 · Gonsalves · 2012 [cited by applicant]
US 8315972B2 · Chkodrov et al. · 2012 [cited by applicant]
US 9466037B2 · Reed et al. · 2016 [cited by applicant]
US 9477451B1 · O'Farrell · 2016 [cited by applicant]
US 9727591B1 · Sharma et al. · 2017 [cited by applicant]
US 9727604B2 · Jin et al. · 2017 [cited by applicant]
US 9740992B2 · Fazal et al. · 2017 [cited by applicant]
US 10025622B2 · Rijhsinghani et al. · 2018 [cited by applicant]
US 10095766B2 · Kapoor et al. · 2018 [cited by applicant]
US 10122783B2 · Qiao et al. · 2018 [cited by applicant]
US 10360239B2 · Kapoor et al. · 2019 [cited by applicant]
US 10469585B2 · Zhao et al. · 2019 [cited by applicant]
US 10599678B2 · Kapoor et al. · 2020 [cited by applicant]
US 10642854B2 · Pattnaik et al. · 2020 [cited by applicant]
US 10719480B1 · Todd · 2020 [cited by applicant]
US 10853338B2 · Meacham et al. · 2020 [cited by applicant]
US 10860599B2 · Mccluskey et al. · 2020 [cited by applicant]
US 10885051B1 · Peters et al. · 2021 [cited by applicant]
US 10997243B1 · Paulus et al. · 2021 [cited by applicant]
US 11030166B2 · Swamy et al. · 2021 [cited by applicant]
US 11113294B1 · Bourbie et al. · 2021 [cited by applicant]
US 11119980B2 · Szczepanik et al. · 2021 [cited by applicant]
US 11137987B2 · Namarvar et al. · 2021 [cited by applicant]
US 11269913B1 · Dervay · 2022 [cited by examiner]
US 11487777B1 · Bose · 2022 [cited by examiner]
US 20060080156A1 · Baughn et al. · 2006 [cited by applicant]
US 20060253830A1 · Rajanala et al. · 2006 [cited by applicant]
US 20070203933A1 · Iversen et al. · 2007 [cited by applicant]
US 20070282889A1 · Ruan et al. · 2007 [cited by applicant]
US 20080046803A1 · Beauchamp · 2008 [cited by examiner]
US 20080183744A1 · Adendorff et al. · 2008 [cited by applicant]
US 20080288448A1 · Agredano et al. · 2008 [cited by applicant]
US 20090055439A1 · Pai et al. · 2009 [cited by applicant]
US 20090265335A1 · Hoffman et al. · 2009 [cited by applicant]
US 20090281985A1 · Aggarwal · 2009 [cited by applicant]
US 20100106747A1 · Honzal et al. · 2010 [cited by applicant]
US 20100122258A1 · Reed et al. · 2010 [cited by applicant]
US 20100211539A1 · Ho · 2010 [cited by applicant]
US 20110029478A1 · Broeker · 2011 [cited by applicant]
US 20110113005A1 · He et al. · 2011 [cited by applicant]
US 20110264618A1 · Potdar et al. · 2011 [cited by applicant]
US 20120005151A1 · Vasudevan et al. · 2012 [cited by applicant]
US 20120173478A1 · Jensen et al. · 2012 [cited by applicant]
US 20120254103A1 · Cottle et al. · 2012 [cited by applicant]
US 20130275365A1 · Wang et al. · 2013 [cited by applicant]
US 20140095249A1 · Tarakad et al. · 2014 [cited by applicant]
US 20140244573A1 · Gonsalves · 2014 [cited by applicant]
US 20140304217A1 · Nowakowski et al. · 2014 [cited by applicant]
US 20150169602A1 · Shankar et al. · 2015 [cited by applicant]
US 20150317350A1 · Roy-Faderman · 2015 [cited by applicant]
US 20160147850A1 · Winkler et al. · 2016 [cited by applicant]
US 20160241579A1 · Roosenraad et al. · 2016 [cited by applicant]
US 20170011087A1 · Hyde et al. · 2017 [cited by applicant]
US 20170116305A1 · Kapoor et al. · 2017 [cited by applicant]
US 20170116306A1 · Kapoor et al. · 2017 [cited by applicant]
US 20170116307A1 · Kapoor et al. · 2017 [cited by applicant]
US 20170132357A1 · Brewerton · 2017 [cited by examiner]
US 20170286255A1 · Kinnear et al. · 2017 [cited by applicant]
US 20180025035A1 · Xia et al. · 2018 [cited by applicant]
US 20180081641A1 · Ron et al. · 2018 [cited by applicant]
US 20180095952A1 · Rehal · 2018 [cited by applicant]
US 20180246944A1 · Yelisetti et al. · 2018 [cited by applicant]
US 20180357237A1 · Balasubrahmanian et al. · 2018 [cited by applicant]
US 20190138319A1 · Haupt et al. · 2019 [cited by applicant]
US 20190243836A1 · Nanda et al. · 2019 [cited by applicant]
US 20190332495A1 · Fair et al. · 2019 [cited by applicant]
US 20190340103A1 · Nelson et al. · 2019 [cited by applicant]
US 20200019558A1 · Okorafor · 2020 [cited by examiner]
US 20200067789A1 · Khuti et al. · 2020 [cited by applicant]
US 20200097851A1 · Alvarez et al. · 2020 [cited by applicant]
US 20200151198A1 · Wang · 2020 [cited by applicant]
US 20200379985A1 · Athavale et al. · 2020 [cited by applicant]
US 20210056084A1 · Guha · 2021 [cited by examiner]
US 20210081848A1 · Polleri et al. · 2021 [cited by applicant]
US 20210117437A1 · Gibson · 2021 [cited by applicant]
US 20210133189A1 · Prado · 2021 [cited by examiner]
US 20210135772A1 · Brown et al. · 2021 [cited by applicant]
US 20210173846A1 · Verma et al. · 2021 [cited by applicant]
US 20210232603A1 · Sundaram et al. · 2021 [cited by applicant]
US 20210232604A1 · Sundaram et al. · 2021 [cited by applicant]
US 20210271568A1 · Abdul Rasheed et al. · 2021 [cited by applicant]
US 20210326717A1 · Mueller et al. · 2021 [cited by applicant]
US 20210351979A1 · Delay et al. · 2021 [cited by applicant]
US 20220171738A1 · Lee et al. · 2022 [cited by applicant]
US 20220207418A1 · Cardoso et al. · 2022 [cited by applicant]
US 20220239759A1 · Grey et al. · 2022 [cited by applicant]
US 20220309324A1 · Nama et al. · 2022 [cited by applicant]
US 20220350674A1 · Velasco · 2022 [cited by applicant]
US 20230015688A1 · Durakovic et al. · 2023 [cited by applicant]
US 20230082010A1 · Clifford et al. · 2023 [cited by applicant]
US 20230168895A1 · Shah et al. · 2023 [cited by applicant]
US 20230236835A1 · Österlund et al. · 2023 [cited by applicant]
AU 2017224831A1 · 2018 [cited by applicant]
CA 2795757A1 · 2013 [cited by applicant]
CN 112597218A · 2021 [cited by applicant]
CN 112883091A · 2021 [cited by applicant]
EP 2079020A1 · 2009 [cited by applicant]
IN 2606MUM2009A · 2012 [cited by applicant]
IN 201914044380A · 2019 [cited by applicant]
WO 2020092291A1 · 2020 [cited by applicant]
International Search and Written Opinion for International Application No. PCT/US2023/016919, dated Sep. 1, 2023, 19 pages. [cited by applicant]
Dominik Jens Elias Waibel et al.: “InstantDL: an easy-to-use deep learning pipeline for image segmentation and classification,” BMC Bioinformatics, Biomed Central LTD, London, UK, vol. 22, No. 1, Mar. 2, 2021 (Mar. 2, 2… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2023/016922, dated Jun. 28, 2023, 10 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2023/016925, dated Sep. 20, 2023, 9 pages. [cited by applicant]
Partial International Search and Provisional Opinion for International Application No. PCT/US2023/016919, dated Jul. 11, 2023, 14 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2023/016915, dated Jul. 4, 2023, 14 pages. [cited by applicant]
Xu Ye: “Data Migration from on-premise relational Data Warehouse to Azure using Azure Data Factory,” Oct. 2018 (Oct. 2018), XP093057148 retrieved Jun. 23, 2023. [cited by applicant]
Yan Zhao et al: “Data Lake Ingestion Management,” arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14856, Jul. 5, 2021 (Jul. 5, 2021), XP091008365, section 2. [cited by applicant]
Malhotra Gaurav: “Get started quickly using templates in Azure Data Factory,” Feb. 11, 2019 (Feb. 11, 2019), XP093057259, retrieved Jun. 23, 2023. [cited by applicant]
Anonymous: “Slowly changing dimension,” Feb. 6, 2022 (Feb. 6, 2022), XP093057274, retrieved Jun. 23, 2023. [cited by applicant]
Masino et al. “A SCALA DSL for Rapid ETL Configuration and Execution,” The Third Annual Scala Workshop, 2012, pp. 1-12 (Year: 2012). [cited by applicant]