Methods and systems to classify software components based on multiple information sources
Systems and methods classifying software components based on multiple information sources are provided. An exemplary method includes retrieving a number of sources including a project documentation file, source code, and dependent project list associated with a software component, extracting a number of entities from the number of sources, processing the number of entities based on a machine learning model, mapping the number of entities to a set of rules, generating a number of categorizations based on the mapping of the number of entities to the set of rules, and ranking the number of categorizations based on the set of rules.
1. A method for classifying software components based on multiple information sources, the method comprising:
retrieving a plurality of sources comprising a project documentation file, source code, and dependent project list associated with a software component;
extracting contextual information for the software component from the plurality of sources;
pre-processing the contextual information using natural language processing;
fetching, based on the pre-processed contextual information, a set of rules for prioritizing classification results;
generating a plurality of categorizations for the software component using the pre-processed contextual information; and
ranking the plurality of categorizations for the software component based on the set of rules
wherein:
generating the plurality of categorizations comprises providing a first categorization associated with a direct match between contextual information and the set of rules and providing a second categorization associated with an indirect match between the contextual information and the set of rules, the indirect match associated with a similarity score, the similarity score identified as equal to or greater than a threshold score; and
the contextual information comprises a short description, a full description, features, code comments, project tags, and dependent libraries.
2. The method of claim 1 , wherein pre-processing the contextual information using natural language processing comprises removing hyperlinks, stopwords, and version information.
3. The method of claim 1 , wherein the natural language processing uses a machine learning model, the method comprising training the machine learning model using training data extracted from a plurality of project documentation files associated with the dependent project list.
4. The method of claim 1 , wherein ranking the plurality of categorizations based on the set of rules comprises determining whether a categorization matches a name of the project documentation file.
5. The method of claim 1 , further comprising presenting a user with the ranked categorizations.
6. A system for classifying software components based on multiple information sources, the system comprising:
one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
retrieving a plurality of sources comprising a project documentation file, source code, and dependent project list associated with a software component;
extracting contextual information for the software component from the plurality of sources;
pre-processing the contextual information using natural language processing;
fetching, based on the pre-processed contextual information, a set of rules for prioritizing classification results;
generating a plurality of categorizations for the software component using the pre-processed contextual information; and
ranking the plurality of categorizations for the software component based on the set of rules
wherein:
generating the plurality of categorizations comprises providing a first categorization associated with a direct match between contextual information and the set of rules and providing a second categorization associated with an indirect match between the contextual information and the set of rules, the indirect match associated with a similarity score, the similarity score identified as equal to or greater than a threshold score; and
the contextual information comprises a short description, a full description, features, code comments, project tags, and dependent libraries.
7. The system of claim 6 , wherein pre-preprocessing the contextual information using the natural language processing comprises removing unnecessary information comprising hyperlinks, stopwords, and version information.
8. The system of claim 6 , wherein the natural language processing uses a machine learning model generated based on training data extracted from a plurality of project documentation files associated with the dependent project list.
9. The system of claim 6 , wherein ranking the plurality of categorizations based on the set of rules comprises determining whether a categorization matches a name of the project documentation file.
10. The system of claim 6 , the operations further comprising presenting a user with the ranked categorizations.
11. One or more non-transitory computer-readable media for classifying software components based on multiple information sources, the non-transitory computer-readable media storing instructions thereon, wherein the instructions when executed by one or more processors cause the one or more processors to:
retrieve a plurality of sources comprising a project documentation file, source code, and dependent project list associated with a software component;
extract contextual information for the software component from the plurality of sources, wherein the contextual information comprises a short description, a full description, features, code comments, project tags, and dependent libraries;
pre-process the contextual information using natural language processing;
fetch, based on the pre-processed contextual information, a set of rules for prioritizing classification results;
generate a plurality of categorizations for the software component using the pre-processed contextual information by providing a first categorization associated with a direct match between the contextual information and the set of rules and providing a second categorization associated with an indirect match between the contextual information and the set of rules, the indirect match associated with a similarity score, the similarity score identified as equal to or greater than a threshold score; and
rank the plurality of categorizations for the software component based on the set of rules.
12. The non-transitory computer-readable media of claim 11 , wherein pre-processing the contextual information using the natural language processing comprises removing hyperlinks.
13. The non-transitory computer-readable media of claim 11 , wherein the natural language processing uses a machine learning model generated based on training data extracted from a plurality of project documentation files associated with the dependent project list.
14. The non-transitory computer-readable media of claim 11 , wherein ranking the plurality of categorizations based on the set of rules comprises determining whether a categorization matches a name of the project documentation file.