Ranking datasets based on data attributes
Ranking a group of datasets using a computer includes determining a set of target data fields from a set of process documents that indicate user data field preferences. A set of target dataset attributes from a set of data use documents indicate user data scope preferences. A plurality of metadata sets for an associated plurality of datasets the computer determines having a field suitability value exceeding a predetermined suitability threshold value. The FSV represents a degree of similarity between a set of fields associated with said dataset and the set of target data fields. The computer assesses metadata sets with regard to the target attributes and generates a compared attribute score for each candidate dataset. A degree of likelihood is indicated that an associated dataset will have content exhibiting said target dataset attributes. The computer candidate datasets is based on the compared attribute score.
1. A computer implemented method to sort a plurality of datasets according to dataset attributes, comprising:
identifying, by a computer, a set of target data fields from a set of process documents, said process documents indicating data field preferences of a user;
identifying, by said computer, a set of target dataset attributes from a set of data use documents, said data use documents indicating data scope preferences for said user, attributes include a data property either derived using automation or added to the set of the target dataset attributes as part of an input provide by a domain expert;
generating, by a computer, a plurality of metadata sets for an associated plurality of datasets;
determining, by said computer, candidate datasets having a field suitability value that exceeds a predetermined suitability threshold value, said field suitability value representing a degree of similarity between a set of fields associated with said dataset and the set of target data fields;
assessing, by said computer, the associated metadata set for each candidate dataset, with regard to the target attributes and generating, by said computer, a compared attribute score for each candidate dataset, indicating a degree of likelihood that an associated dataset will have content exhibiting said target dataset attributes; and
generating, by said computer, a list of said candidate datasets sorted by said compared attribute scores.
2. The method of claim 1 , wherein said data use documents include information in a format selected from a list consisting of Business Process Execution Language (BEPL), and Unified Modeling Language (UML).
3. The method of claim 1 , wherein said data target attributes are extracted from elements of said process documents, selected from a list consisting of class diagrams, activity diagrams, sequence diagrams, and component diagrams.
4. The method of claim 1 , further including designating a candidate dataset having a highest compared attribute score as a selected dataset.
5. The method of claim 4 , further including establishing a set of search parameters for a search to be conducted on said selected dataset; and updating a historic use field in the metadata set associated with a dataset selected for searching with a search context value that represent aspects of the search parameters.
6. The method of claim 5 , wherein said ranking is based, at least in part, on the historic use field values.
7. The method of claim 1 , wherein said compared attribute scores are based, at least in part on an associated desirability value associated with each of said target dataset attributes.
8. The method of claim 1 , wherein said sets of metadata include information selected from a list consisting of: domain, gender, age group, geographic distribution, demographic distribution, statistical ranges of numerical values, and context of applicability.
9. system to sort a plurality of datasets according to dataset attributes, which comprises:
a computer system comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:
identify a set of target data fields from a set of process documents, said process documents indicating data field preferences of a user;
identify a set of target dataset attributes from a set of data use documents, said data use documents indicating data scope preferences for said user, attributes include a data property either derived using automation or added to the set of the target dataset attributes as part of an input provide by a domain expert;
generate a plurality of metadata sets for an associated plurality of datasets;
determine candidate datasets having a field suitability value that exceeds a predetermined suitability threshold value, said field suitability value representing a degree of similarity between a set of fields associated with said dataset and the set of target data fields;
assess the associated metadata set for each candidate dataset, with regard to the target attributes and generating, by said computer, a compared attribute score for each candidate dataset, indicating a degree of likelihood that an associated dataset will have content exhibiting said target dataset attributes; and
generate a list of said candidate datasets sorted by said compared attribute scores.
10. The system of claim 9 , wherein said data use documents include information in a format selected from a list consisting of Business Process Execution Language (BEPL), and Unified Modeling Language (UML).
11. The system of claim 9 , wherein said data target attributes are extracted from elements of said process documents, selected from a list consisting of class diagrams, activity diagrams, sequence diagrams, and component diagrams.
12. The system of claim 9 , further including instructions for the computer to designate a candidate dataset having a highest compared attribute score as a selected dataset.
13. The system of claim 12 , further including instructions for the computer to establish a set of search parameters for a search to be conducted on said selected dataset; and to update a historic use field in the metadata set associated with a dataset selected for searching with a search context value that represent aspects of the search parameters.
14. The system of claim 13 , wherein said ranking is based, at least in part, on the historic use field values.
15. The system of claim 9 , wherein said compared attribute scores are based, at least in part on an associated desirability value associated with each of said target dataset attributes.
16. The system of claim 9 , wherein said sets of metadata include information selected from a list consisting of: domain, gender, age group, geographic distribution, demographic distribution, statistical ranges of numerical values, and context of applicability.
17. computer program product to sort a plurality of datasets according to dataset attributes, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:
identify, using a computer, a set of target data fields from a set of process documents, said process documents indicating data field preferences of a user;
identify, using a computer, a set of target dataset attributes from a set of data use documents, said data use documents indicating data scope preferences for said user, attributes include a data property either derived using automation or added to the set of the target dataset attributes as part of an input provide by a domain expert;
generate, using a computer, a plurality of metadata sets for an associated plurality of datasets;
determine, using a computer, candidate datasets having a field suitability value that exceeds a predetermined suitability threshold value, said field suitability value representing a degree of similarity between a set of fields associated with said dataset and the set of target data fields;
assess, using a computer, the associated metadata set for each candidate dataset, with regard to the target attributes and generating, by said computer, a compared attribute score for each candidate dataset, indicating a degree of likelihood that an associated dataset will have content exhibiting said target dataset attributes; and
generate, using a computer, a list of said candidate datasets sorted by said compared attribute scores.
18. The computer program product of claim 17 , wherein said data use documents include information in a format selected from a list consisting of Business Process Execution Language (BEPL), and Unified Modeling Language (UML).
19. The computer program product of claim 17 , wherein said data target attributes are extracted from elements of said process documents, selected from a list consisting of class diagrams, activity diagrams, sequence diagrams, and component diagrams.
20. The computer program product of claim 17 , further including instructions for the computer to designate a candidate dataset having a highest compared attribute score as a selected dataset.