Data driven ranking of competing entities in a marketplace
A method, computer system, and a computer program product for competitive analysis is provided. The present invention may include identifying one or more potential competitors by searching a knowledge corpus using one or more see terms. The present invention may include determining one or more competitors by eliminating at least one potential competitor. The present invention may include generating a competitive analyst report.
1 . A method for competitive analysis, the method comprising:
receiving data related to a target business, wherein the target business is identified based on a user selection within a user interface;
identifying a plurality of seed terms associated with the target business using a machine learning model with Natural Language Processing to analyze the data related to the target business using one or more text analysis techniques;
performing a semantic relatedness analysis of the plurality of seed terms to determine whether each of the plurality of seed terms are above or below a similarity threshold, wherein one or more of the plurality of seed terms are determined to be above or below the similarity threshold based on a cosine similarity distance and requiring manual verification or elimination by the user within the user interface, and wherein the semantic relatedness analysis further includes analyzing the plurality of seed terms based on their source and using string similarity, Latent Semantic Indexing (LSI), and averaging unigram word embeddings;
training the machine learning model with Natural Language Processing based on one or more seed terms of the plurality of seed terms which are manually verified by the user within the user interface;
identifying a plurality of related businesses to the target business by searching a knowledge corpus using the one or more seed terms, wherein the knowledge corpus is comprised of external documentation and publicly available information related to a plurality of businesses;
determining one or more competitors by eliminating at least one of the plurality of related businesses identified, wherein the at least one of the plurality of related businesses identified are eliminated based on a weighting of mentions of the one or more seed terms, wherein the weighting of the mentions is determined based on at least a recency of the mention, a form of the external documentation in which the mention is derived, and an engagement with the external documentation in which the mention is derived; and
generating a competitive analysis report, wherein the competitive analysis report includes at least a favorability of the target business in comparison to the one or more competitors and an aggregate competitive score over time which is continuously updated in real time based on additional seed terms identified in the external documentation related to the one or more competitors and manual changes to the plurality of seed terms by the user, wherein the competitive score over time is displayed to the user within the user interface utilizing visual insights.
2 . The method of claim 1 , wherein performing the semantic relatedness analysis further comprises:
tokenizing the plurality of seed terms using n-Gram tokenization and converting a plurality of tokenized seed terms into a plurality of vectors, wherein each of the plurality of vectors are compared using cosine similarity;
eliminating, by the user, the one or more seed terms below the similarity threshold;
verifying, by the user, the one or more seed terms above the similarity threshold;
storing the one or more verified seed terms in a target business knowledge corpus; and
training the machine learning model with Natural Language Processing based on one or more eliminated seed terms and one or more verified seed terms.
3 . The method of claim 2 , wherein more than one similarity threshold is utilized, and wherein at least one of the similarity thresholds is automatically verified.
4 . The method of claim 1 , wherein the text analysis techniques include at least keyword extraction and the data related to the target business includes internal documentation, the external documentation, and manual input received from the user within the user interface, wherein the manual input is ontology terms relating to a taxonomy of the target business.
5 . The method of claim 1 , wherein the plurality of related businesses are listed according to an aggregate number of identifications, wherein the aggregate number of identifications may be categorized according to categories of the target business identified by the user in the user interface, wherein the competitive analysis report includes insights for each of the one or more competitors according to the categories of the target business identified by the user in the user interface.
6 . The method of claim 1 , wherein the favorability is determined based on a sentiment analysis of the target business as compared to the one or more competitors.
7 . The method of claim 1 , wherein the visual insights include a breakdown of the external documentation related to each of the one or more competitors.
8 . The method of claim 1 , wherein the determining of the one or more competitors further comprises:
characterizing, utilizing one or more Natural Language Processing (NLP) techniques, a manner in which the plurality of related businesses are mentioned within the external documentation; and
assigning a greater weighting to the one or more seed terms based on the manner in which the one or more competitors are mentioned, wherein the manner in which the one or more competitors are mentioned corresponds to one or more goals of the target business identified by the user within the user interface.
9 . The method of claim 1 , wherein the weighting of the mentions are adjusted from default weighting settings based on the target business, wherein the target business is a product or service offered by a company.
10 . The method of claim 1 , wherein the weighting of the mentions are adjusted from default weighting settings based on input from the user, wherein the user assigns a greater weighting to a first form of external documentation and a lesser weighting to a second form of external documentation, and wherein the external documentation is continuously added to the knowledge corpus as it is made publicly available.
11 . The method of claim 1 , wherein the engagement with the external documentation in which the mention is derived is determined based on viewership, wherein the weighting of the one or more seed terms in the external documentation corresponds to the viewership of the external documentation.
12 . A computer system for competitive analysis, comprising:
one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:
receiving data related to a target business, wherein the target business is identified based on a user selection within a user interface;
identifying a plurality of seed terms associated with the target business using a machine learning model with Natural Language Processing to analyze the data related to the target business using one or more text analysis techniques;
performing a semantic relatedness analysis of the plurality of seed terms to determine whether each of the plurality of seed terms are above or below a similarity threshold, wherein one or more of the plurality of seed terms are determined to be above or below the similarity threshold based on a cosine similarity distance and requiring manual verification or elimination by the user within the user interface, and wherein the semantic relatedness analysis further includes analyzing the plurality of seed terms based on their source and using string similarity, Latent Semantic Indexing (LSI), and averaging unigram word embeddings;
training the machine learning model with Natural Language Processing based on one or more seed terms of the plurality of seed terms which are manually verified by the user within the user interface;
identifying a plurality of related businesses to the target business by searching a knowledge corpus using the one or more seed terms, wherein the knowledge corpus is comprised of external documentation and publicly available information related to a plurality of businesses;
determining one or more competitors by eliminating at least one of the plurality of related businesses identified, wherein the at least one of the plurality of related businesses identified are eliminated based on a weighting of mentions of the one or more seed terms, wherein the weighting of the mentions is determined based on at least a recency of the mention, a form of the external documentation in which the mention is derived, and an engagement with the external documentation in which the mention is derived; and
generating a competitive analysis report, wherein the competitive analysis report includes at least a favorability of the target business in comparison to the one or more competitors and an aggregate competitive score over time which is continuously updated in real time based on additional seed terms identified in the external documentation related to the one or more competitors and manual changes to the plurality of seed terms by the user, wherein the competitive score over time is displayed to the user within the user interface utilizing visual insights.
13 . The computer system of claim 12 , wherein performing the semantic relatedness analysis further comprises:
tokenizing the plurality of seed terms using n-Gram tokenization and converting a plurality of tokenized seed terms into a plurality of vectors, wherein each of the plurality of vectors are compared using cosine similarity;
eliminating, by the user, the one or more seed terms below the similarity threshold;
verifying, by the user, the one or more seed terms above the similarity threshold;
storing the one or more verified seed terms in a target business knowledge corpus; and
training the machine learning model with Natural Language Processing based on one or more eliminated seed terms and one or more verified seed terms.
14 . The computer system of claim 13 , wherein more than one similarity threshold is utilized, and wherein at least one of the similarity thresholds is automatically verified.
15 . The computer system of claim 12 , wherein the text analysis techniques include at least keyword extraction and the data related to the target business includes internal documentation, the external documentation, and manual input received from the user within the user interface, wherein the manual input is ontology terms relating to a taxonomy of the target business.
16 . A computer program product for competitive analysis, comprising:
one or more non-transitory computer-readable storage media and program instructions stored on at least one of the one or more tangible storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:
receiving data related to a target business, wherein the target business is identified based on a user selection within a user interface;
identifying a plurality of seed terms associated with the target business using a machine learning model with Natural Language Processing to analyze the data related to the target business using one or more text analysis techniques;
performing a semantic relatedness analysis of the plurality of seed terms to determine whether each of the plurality of seed terms are above or below a similarity threshold, wherein one or more of the plurality of seed terms are determined to be above or below the similarity threshold based on a cosine similarity distance and requiring manual verification or elimination by the user within the user interface, and wherein the semantic relatedness analysis further includes analyzing the plurality of seed terms based on their source and using string similarity, Latent Semantic Indexing (LSI), and averaging unigram word embeddings;
training the machine learning model with Natural Language Processing based on one or more seed terms of the plurality of seed terms which are manually verified by the user within the user interface;
identifying a plurality of related businesses to the target business by searching a knowledge corpus using the one or more seed terms, wherein the knowledge corpus is comprised of external documentation and publicly available information related to a plurality of businesses;
determining one or more competitors by eliminating at least one of the plurality of related businesses identified, wherein the at least one of the plurality of related businesses identified are eliminated based on a weighting of mentions of the one or more seed terms, wherein the weighting of the mentions is determined based on at least a recency of the mention, a form of the external documentation in which the mention is derived, and an engagement with the external documentation in which the mention is derived; and
generating a competitive analysis report, wherein the competitive analysis report includes at least a favorability of the target business in comparison to the one or more competitors and an aggregate competitive score over time which is continuously updated in real time based on additional seed terms identified in the external documentation related to the one or more competitors and manual changes to the plurality of seed terms by the user, wherein the competitive score over time is displayed to the user within the user interface utilizing visual insights.
17 . The computer program product of claim 16 , wherein performing the semantic relatedness analysis further comprises:
tokenizing the plurality of seed terms using n-Gram tokenization and converting a plurality of tokenized seed terms into a plurality of vectors, wherein each of the plurality of vectors are compared using cosine similarity;
eliminating, by the user, the one or more seed terms below the similarity threshold;
verifying, by the user, the one or more seed terms above the similarity threshold;
storing the one or more verified seed terms in a target business knowledge corpus; and
training the machine learning model with Natural Language Processing based on one or more eliminated seed terms and one or more verified seed terms.
18 . The computer program product of claim 17 , wherein more than one similarity threshold is utilized, and wherein at least one of the similarity thresholds is automatically verified.
19 . The computer program product of claim 16 , wherein the text analysis techniques include at least keyword extraction and the data related to the target business includes internal documentation, the external documentation, and manual input received from the user within the user interface, wherein the manual input is ontology terms relating to a taxonomy of the target business.