Prioritizing curation targets for data curation based on downstream impact
Methods and systems for curating data by a data manager are disclosed. Data may be curated from various data sources before being provided to downstream consumers that may rely on the trustworthiness of the curated data in order to provide desired computer-implemented services. During the data curation process, data curation resources are used to improve the trustworthiness and/or value of the collected data. However, data curation resources (e.g., data curators, computing resources) may be limited and/or insufficient to perform the data curation process as desired, which may result in unusable and/or uncurated (e.g., untrustworthy) data. Thus, portions of the data (e.g., curation targets) may be prioritized (e.g., relative to other curation targets). The curation targets may be curated with the available data curation resources based on their relative priority in order to reduce the likelihood of providing untrustworthy data to entities (e.g., downstream consumers) that facilitate the computer-implemented services.
1 . A method for curating data by a data manager, comprising:
obtaining at least a portion of the data from a data source;
identifying curation targets of the data;
making a determination regarding whether sufficient data curation resources are available to perform a data curation process for the curation targets within a target period of time;
in an instance of the determination where there are insufficient data curation resources available:
identifying a data curation resource of the data curation resources that has available curation bandwidth;
obtaining, based on scoring criteria, impact scores for the curation targets, wherein each impact score of the impact scores is based on at least one of:
a frequency of use of a corresponding curation target by an inference model that ingests at least a second portion of the data to generate an inference;
a measure of relative contribution of the corresponding curation target to the inference;
a measure of confidence in the inference; and
a measure of importance of the corresponding curation target or the inference to a downstream consumer;
obtaining a rank for at least one curation target based on the impact score, the rank being usable to order the at least one curation target;
selecting a curation target of the curation targets for the data curation resource based on a rank ordering of the curation targets, the rank ordering being based on the impact scores for the curation targets;
assigning the curation target to the data curation resource in order to complete the data curation process for a portion of the curation targets within the target period of time; and
curating the curation target using the data curation resource to obtain at least partially curated data.
2 . The method of claim 1 , wherein each impact score is based, at least in part, on a number of occurrences of the curation target in downstream use of the data.
3 . The method of claim 2 , wherein each impact score is further based, at least in part, on an attribution score for the curation target, the attribution score indicating a relative level of contribution to a future outcome in which the curation target is usable in the downstream use of the data.
4 . The method of claim 3 , wherein each impact score is further based, at least in part, on a level of confidence in predicting the future outcome through the downstream use of the data.
5 . The method of claim 4 , wherein each impact score is further based, at least in part, on the measure of importance of the curation target to a downstream consumer.
6 . The method of claim 5 , wherein each impact score is further based, at least in part, on a measure of dependence that the downstream consumer has on the predicting of the future outcome.
7 . The method of claim 6 , wherein the portion of the curation targets excludes at least one of the curation targets.
8 . The method of claim 1 ,
wherein the at least partially curated data complying with a schema for downstream use of the curation target.
9 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for curating data by a data manager, the operations comprising:
obtaining at least a portion of the data from a data source;
identifying curation targets of the data;
making a determination regarding whether sufficient data curation resources are available to perform a data curation process for the curation targets within a target period of time;
in an instance of the determination where there are insufficient data curation resources available:
identifying a data curation resource of the data curation resources that has available curation bandwidth;
obtaining, based on scoring criteria, impact scores for the curation targets, wherein each impact score of the impact scores is based on at least one of:
a frequency of use of a corresponding curation target by an inference model that ingests at least a second portion of the data to generate an inference;
a measure of relative contribution of the corresponding curation target to the inference;
a measure of confidence in the inference; and
a measure of importance of the corresponding curation target or the inference to a downstream consumer;
obtaining a rank for at least one curation target based on the impact score, the rank being usable to order the at least one curation target;
selecting a curation target of the curation targets for the data curation resource based on a rank ordering of the curation targets, the rank ordering being based on the impact scores for the curation targets;
assigning the curation target to the data curation resource in order to complete the data curation process for a portion of the curation targets within the target period of time; and
curating the curation target using the data curation resource to obtain at least partially curated data.
10 . The non-transitory machine-readable medium of claim 9 , wherein each impact score is based at least in part, on a number of occurrences of the curation target in downstream use of the data.
11 . The non-transitory machine-readable medium of claim 10 , wherein each impact score is further based, at least in part, on an attribution score for the curation target, the attribution score indicating a relative level of contribution to a future outcome in which the curation target is usable in the downstream use of the data.
12 . The non-transitory machine-readable medium of claim 9 , wherein the at least partially curated data complying with a schema for downstream use of the curation target.
13 . A data processing system, comprising:
a processor; and
a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for curating data by a data manager, the operations comprising:
obtaining at least a portion of the data from a data source;
identifying curation targets of the data;
making a determination regarding whether sufficient data curation resources are available to perform a data curation process for the curation targets within a target period of time;
in an instance of the determination where there are insufficient data curation resources available:
identifying a data curation resource of the data curation resources that has available curation bandwidth;
obtaining, based on scoring criteria, impact scores for the curation targets, wherein each impact score of the impact scores is based on at least one of:
a frequency of use of a corresponding curation target by an inference model that ingests at least a second portion of the data to generate an inference;
a measure of relative contribution of the corresponding curation target to the inference;
a measure of confidence in the inference; and
a measure of importance of the corresponding curation target or the inference to a downstream consumer;
obtaining a rank for at least one curation target based on the impact score, the rank being usable to order the at least one curation target;
selecting a curation target of the curation targets for the data curation resource based on a rank ordering of the curation targets, the rank ordering being based on the impact scores for the curation targets;
assigning the curation target to the data curation resource in order to complete the data curation process for a portion of the curation targets within the target period of time; and
curating the curation target using the data curation resource to obtain at least partially curated data.
14 . The data processing system of claim 13 , wherein each impact score is based at least in part, on a number of occurrences of the curation target in downstream use of the data.
15 . The data processing system of claim 14 , wherein each impact score is further based, at least in part, on an attribution score for the curation target, the attribution score indicating a relative level of contribution to a future outcome in which the curation target is usable in the downstream use of the data.
16 . The data processing system of claim 15 , wherein each impact score is further based, at least in part, on a level of confidence in predicting the future outcome through the downstream use of the data.
17 . The data processing system of claim 16 , wherein each impact score is further based, at least in part, on the measure of importance of the curation target to a downstream consumer.
18 . The data processing system of claim 17 , wherein each impact score is further based, at least in part, on a measure of dependence that the downstream consumer has on the predicting of the future outcome.
19 . The data processing system of claim 18 , wherein the portion of the curation targets excludes at least one of the curation targets.
20 . The data processing system of claim 13 , wherein the at least partially curated data complying with a schema for downstream use of the curation target.