IP Library › Granted Patent US 11,321,337
Granted Patent B2
US 11,321,337 · App. 16/203,725 · Granted May 3, 2022

Crowdsourcing data into a data lake

Inventors: Antonio Nucci (San Jose, CA); Ahmed Khattab (San Jose, CA); Carlos M. Pignataro (Cary, NC); Ravi K. Papisetti (Leander, TX); Prasad Potipireddi (San Ramon, CA); Richard M. Plane (Wake Forest, NC)
Assignee: CISCO TECHNOLOGY, INC.
G06F16/254G06F16/211G06F16/2471G06F16/24578G06F16/258G06F16/383G06F21/6254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,321,337
App. No.
16/203,725
Granted
May 3, 2022
Kind
B2
Abstract

A Services Delivery Platform (SDP) architecture is provided that is configured to onboard new data sets into an SDP data lake. The SDP enables the crowdsourcing of data on-boarding by configuring this process into an interactive, intuitive, step-by-step guided workflow while governing/controlling key functions like verification, acceptance and execution.

Claims (84)

1. A method performed at a central compute node in a distributed computing system that includes a plurality of remote compute nodes whose computing resources and software functions are made available in a platform agnostic matter to users of the distributed computing system, the method comprising:

obtaining, from a remote compute node, a request to onboard at least one data set to a data lake of the distributed computing system;

obtaining, from the remote compute node, information about the at least one data set available from a source repository, the information including database type of the source repository and schema of the at least one data set;

obtaining information about a destination repository for the at least one data set in the data lake managed by the distributed computing system;

based on the information about at least one data set and the information about the destination repository, determining a mapping scheme that describes a transformation from the schema of the at least one data set to a unified data schema for the distributed computing system;

copying data for the at least one data set to the destination repository in the distributed computing system;

transforming the data of the at least one data set according to the mapping scheme;

storing a plurality of data catalogs associated with data from a plurality of enterprises, wherein data of at least a first enterprise is separated into sensitive data for internal use only and anonymized data for external use;

obtaining, from a second enterprise, a request for the data of the first enterprise;

initiating a workload migration from the second enterprise to the first enterprise, wherein a workload is executed at the first enterprise on the anonymized data of the first enterprise; and

returning results of the workload executed at the first enterprise to the second enterprise.

2. The method of claim 1 , further comprising:

obtaining information about quality rules for data in the at least one data set;

applying the quality rules on data for the at least one data set that is copied; and

generating quality metrics, every time the data is onboarded according to the quality rules.

3. The method of claim 1 , further comprising:

obtaining information about a taxonomy associated with data for the at least one data set; and

managing storage of the data for the data set to the destination in the data lake based on the taxonomy.

4. The method of claim 1 , further comprising generating for display a graphical element that indicates availability of the at least one data set in a catalog of data sets in the data lake available for use.

5. The method of claim 1 , wherein copying includes:

automatically creating an end-to-end data flow between the source repository and the destination repository; and

executing one or more operations on data in the data set, including reading, writing, schema transforming, and filtering based on criteria or identifiers included in the data.

6. The method of claim 1 , further comprising:

serving a catalog portal that includes a publish/discovery/subscriber service with respect to a plurality of external data lakes;

publishing data sets into a catalog managed by the catalog portal; and

parsing schemas associated with the data sets into metadata stored in the catalog.

7. The method of claim 1 , further comprising:

obtaining advanced information including one or more of quality rule definitions, taxonomy enforcement definitions, and masking definitions.

8. An apparatus comprising:

a communication interface configured to enable network communications between a central compute node in a distributed computing system and a plurality of remote compute nodes in whose computing resources and software functions are made available in a platform agnostic matter to users of the distributed computing system; and

one or more processors coupled to the communication interface, wherein the one or more processors are configured to perform operations including:

obtaining, from a remote compute node, a request to onboard at least one data set to a data lake of the distributed computing system;

obtaining, from the remote compute node, information about the at least one data set available from a source repository, the information including database type of the source repository and schema of the at least one data set;

obtaining information about a destination repository for the at least one data set in the data lake managed by the distributed computing system;

based on the information about at least one data set and the information about the destination repository, determining a mapping scheme that describes a transformation from the schema of the at least one data set to a unified data schema for the distributed computing system;

copying data for the at least one data set to the destination repository in the distributed computing system;

transforming the data of the at least one data set according to the mapping scheme;

storing a plurality of data catalogs associated with data from a plurality of enterprises, wherein data of at least a first enterprise is separated into sensitive data for internal use only and anonymized data for external use;

obtaining, from a second enterprise, a request for the data of the first enterprise;

initiating a workload migration from the second enterprise to the first enterprise, wherein a workload is executed at the first enterprise on the anonymized data of the first enterprise; and

returning results of the workload executed at the first enterprise to the second enterprise.

9. The apparatus of claim 8 , wherein the one or more processors are configured to further perform operations including:

obtaining information about quality rules for data in the at least one data set;

applying the quality rules on data for the at least one data set that is copied; and

generating quality metrics, every time the data is onboarded according to the quality rules.

10. The apparatus of claim 8 , wherein the one or more processors are configured to further perform operations including:

obtaining information about a taxonomy associated with data for the at least one data set; and

managing storage of the data for the data set to the destination in the data lake based on the taxonomy.

11. The apparatus of claim 8 , wherein the one or more processors are configured to further perform operations including generating for display a graphical element that indicates availability of the at least one data set in a catalog of data sets in the data lake available for use.

12. The apparatus of claim 8 , wherein the one or more processors are configured to perform the copying by:

automatically creating an end-to-end data flow between the source repository and the destination repository; and

executing one or more operations on data in the data set, including reading, writing, schema transforming, and filtering based on criteria or identifiers included in the data.

13. The apparatus of claim 8 , wherein the one or more processors are configured to further perform operations including:

serving a catalog portal that includes a publish/discovery/subscriber service with respect to a plurality of external data lakes;

publishing data sets into a catalog managed by the catalog portal; and

parsing schemas associated with the data sets into metadata stored in the catalog.

14. The apparatus of claim 8 , wherein the one or more processors are configured to further perform operations including:

obtaining information including one or more of quality rule definitions, taxonomy enforcement definitions, and masking definitions.

15. One or more computer readable storage media encoded with software comprising computer executable instructions and when the software is executed operable to perform operations at a central compute node in a distributed computing system that includes a plurality of remote compute nodes whose computing resources and software functions are made available in a platform agnostic matter to users of the distributed computing system, the operations including:

obtaining, from a remote compute node, a request to onboard at least one data set to a data lake of the distributed computing system;

obtaining, from the remote compute node, information about the at least one data set available from a source repository, the information including database type of the source repository and schema of the at least one data set;

obtaining information about a destination repository for the at least one data set in the data lake managed by the distributed computing system;

based on the information about at least one data set and the information about the destination repository, determining a mapping scheme that describes a transformation from the schema of the at least one data set to a unified data schema for the distributed computing system;

copying data for the at least one data set to the destination repository in the distributed computing system;

transforming the data of the at least one data set according to the mapping scheme;

storing a plurality of data catalogs associated with data from a plurality of enterprises, wherein data of at least a first enterprise is separated into sensitive data for internal use only and anonymized data for external use;

obtaining from a second enterprise a request for the data of the first enterprise;

initiating a workload migration from the second enterprise to the first enterprise, wherein a workload is executed at the first enterprise on the anonymized data of the first enterprise; and

returning results of the workload executed at the first enterprise to the second enterprise.

16. The one or more computer readable storage media of claim 15 , further comprising instructions operable for:

obtaining information about quality rules for data in the at least one data set;

applying the quality rules on data for the at least one data set that is copied; and

generating quality metrics, every time the data is onboarded according to the quality rules.

17. The one or more computer readable storage media of claim 15 , further comprising instructions operable for:

obtaining information about a taxonomy associated with data for the at least one data set; and

managing storage of the data for the data set to the destination in the data lake based on the taxonomy.

18. The one or more computer readable storage media of claim 15 , further comprising instructions operable for:

serving a catalog portal that includes a publish/discovery/subscriber service with respect to a plurality of external data lakes;

publishing data sets into a catalog managed by the catalog portal; and

parsing schemas associated with the data sets into metadata stored in the catalog.

19. The one or more computer readable storage media of claim 15 , further comprising instructions operable for:

obtaining information including one or more of quality rule definitions, taxonomy enforcement definitions, and masking definitions.

20. The one or more computer readable storage media of claim 15 , further comprising instructions operable for:

generating for display a graphical element that indicates availability of the at least one data set in a catalog of data sets in the data lake available for use.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2018
From: NUCCI, ANTONIO; KHATTAB, AHMED; PIGNATARO, CARLOS M.; PAPISETTI, RAVI K.; POTIPIREDDI, PRASAD; PLANE, RICHARD M.
To: CISCO TECHNOLOGY, INC.
Reel/Frame 047618/0367 →
Continuity (3)
Provisional Application 62732763 · Sep 18, 2018
Provisional Application 62680074 · Jun 4, 2018
Related Publication 20190370263A1 · Dec 5, 2019
Cited By (2)
US 12,411,972 US 12,413,485