IP Library Granted Patent US 12,737,642
Granted Patent B2
US 12,737,642 · App. 17/247,448 · Granted Sep 15, 2026

Data driven ranking of competing entities in a marketplace

Inventors: Jonathan M. Smith (White Plains, NY); Sheema Usmani (White Plains, NY); Alexander Shypula (New York, NY); Phillip Werner Simplicio (West Hartford, CT); Biplav Srivastava (Rye, NY); Amir Sabet Sarvestani (New York, NY)
Assignee: International Business Machines Corporation
G06N5/022G06F18/214G06F18/22G06N20/00G06Q10/0637
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,642
App. No.
17/247,448
Granted
Sep 15, 2026
Kind
B2
Abstract

A method, computer system, and a computer program product for competitive analysis is provided. The present invention may include identifying one or more potential competitors by searching a knowledge corpus using one or more see terms. The present invention may include determining one or more competitors by eliminating at least one potential competitor. The present invention may include generating a competitive analyst report.

Claims (59)

1 . A method for competitive analysis, the method comprising:

receiving data related to a target business, wherein the target business is identified based on a user selection within a user interface;

identifying a plurality of seed terms associated with the target business using a machine learning model with Natural Language Processing to analyze the data related to the target business using one or more text analysis techniques;

performing a semantic relatedness analysis of the plurality of seed terms to determine whether each of the plurality of seed terms are above or below a similarity threshold, wherein one or more of the plurality of seed terms are determined to be above or below the similarity threshold based on a cosine similarity distance and requiring manual verification or elimination by the user within the user interface, and wherein the semantic relatedness analysis further includes analyzing the plurality of seed terms based on their source and using string similarity, Latent Semantic Indexing (LSI), and averaging unigram word embeddings;

training the machine learning model with Natural Language Processing based on one or more seed terms of the plurality of seed terms which are manually verified by the user within the user interface;

identifying a plurality of related businesses to the target business by searching a knowledge corpus using the one or more seed terms, wherein the knowledge corpus is comprised of external documentation and publicly available information related to a plurality of businesses;

determining one or more competitors by eliminating at least one of the plurality of related businesses identified, wherein the at least one of the plurality of related businesses identified are eliminated based on a weighting of mentions of the one or more seed terms, wherein the weighting of the mentions is determined based on at least a recency of the mention, a form of the external documentation in which the mention is derived, and an engagement with the external documentation in which the mention is derived; and

generating a competitive analysis report, wherein the competitive analysis report includes at least a favorability of the target business in comparison to the one or more competitors and an aggregate competitive score over time which is continuously updated in real time based on additional seed terms identified in the external documentation related to the one or more competitors and manual changes to the plurality of seed terms by the user, wherein the competitive score over time is displayed to the user within the user interface utilizing visual insights.

2 . The method of claim 1 , wherein performing the semantic relatedness analysis further comprises:

tokenizing the plurality of seed terms using n-Gram tokenization and converting a plurality of tokenized seed terms into a plurality of vectors, wherein each of the plurality of vectors are compared using cosine similarity;

eliminating, by the user, the one or more seed terms below the similarity threshold;

verifying, by the user, the one or more seed terms above the similarity threshold;

storing the one or more verified seed terms in a target business knowledge corpus; and

training the machine learning model with Natural Language Processing based on one or more eliminated seed terms and one or more verified seed terms.

3 . The method of claim 2 , wherein more than one similarity threshold is utilized, and wherein at least one of the similarity thresholds is automatically verified.

4 . The method of claim 1 , wherein the text analysis techniques include at least keyword extraction and the data related to the target business includes internal documentation, the external documentation, and manual input received from the user within the user interface, wherein the manual input is ontology terms relating to a taxonomy of the target business.

5 . The method of claim 1 , wherein the plurality of related businesses are listed according to an aggregate number of identifications, wherein the aggregate number of identifications may be categorized according to categories of the target business identified by the user in the user interface, wherein the competitive analysis report includes insights for each of the one or more competitors according to the categories of the target business identified by the user in the user interface.

6 . The method of claim 1 , wherein the favorability is determined based on a sentiment analysis of the target business as compared to the one or more competitors.

7 . The method of claim 1 , wherein the visual insights include a breakdown of the external documentation related to each of the one or more competitors.

8 . The method of claim 1 , wherein the determining of the one or more competitors further comprises:

characterizing, utilizing one or more Natural Language Processing (NLP) techniques, a manner in which the plurality of related businesses are mentioned within the external documentation; and

assigning a greater weighting to the one or more seed terms based on the manner in which the one or more competitors are mentioned, wherein the manner in which the one or more competitors are mentioned corresponds to one or more goals of the target business identified by the user within the user interface.

9 . The method of claim 1 , wherein the weighting of the mentions are adjusted from default weighting settings based on the target business, wherein the target business is a product or service offered by a company.

10 . The method of claim 1 , wherein the weighting of the mentions are adjusted from default weighting settings based on input from the user, wherein the user assigns a greater weighting to a first form of external documentation and a lesser weighting to a second form of external documentation, and wherein the external documentation is continuously added to the knowledge corpus as it is made publicly available.

11 . The method of claim 1 , wherein the engagement with the external documentation in which the mention is derived is determined based on viewership, wherein the weighting of the one or more seed terms in the external documentation corresponds to the viewership of the external documentation.

12 . A computer system for competitive analysis, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

receiving data related to a target business, wherein the target business is identified based on a user selection within a user interface;

identifying a plurality of seed terms associated with the target business using a machine learning model with Natural Language Processing to analyze the data related to the target business using one or more text analysis techniques;

performing a semantic relatedness analysis of the plurality of seed terms to determine whether each of the plurality of seed terms are above or below a similarity threshold, wherein one or more of the plurality of seed terms are determined to be above or below the similarity threshold based on a cosine similarity distance and requiring manual verification or elimination by the user within the user interface, and wherein the semantic relatedness analysis further includes analyzing the plurality of seed terms based on their source and using string similarity, Latent Semantic Indexing (LSI), and averaging unigram word embeddings;

training the machine learning model with Natural Language Processing based on one or more seed terms of the plurality of seed terms which are manually verified by the user within the user interface;

identifying a plurality of related businesses to the target business by searching a knowledge corpus using the one or more seed terms, wherein the knowledge corpus is comprised of external documentation and publicly available information related to a plurality of businesses;

determining one or more competitors by eliminating at least one of the plurality of related businesses identified, wherein the at least one of the plurality of related businesses identified are eliminated based on a weighting of mentions of the one or more seed terms, wherein the weighting of the mentions is determined based on at least a recency of the mention, a form of the external documentation in which the mention is derived, and an engagement with the external documentation in which the mention is derived; and

generating a competitive analysis report, wherein the competitive analysis report includes at least a favorability of the target business in comparison to the one or more competitors and an aggregate competitive score over time which is continuously updated in real time based on additional seed terms identified in the external documentation related to the one or more competitors and manual changes to the plurality of seed terms by the user, wherein the competitive score over time is displayed to the user within the user interface utilizing visual insights.

13 . The computer system of claim 12 , wherein performing the semantic relatedness analysis further comprises:

tokenizing the plurality of seed terms using n-Gram tokenization and converting a plurality of tokenized seed terms into a plurality of vectors, wherein each of the plurality of vectors are compared using cosine similarity;

eliminating, by the user, the one or more seed terms below the similarity threshold;

verifying, by the user, the one or more seed terms above the similarity threshold;

storing the one or more verified seed terms in a target business knowledge corpus; and

training the machine learning model with Natural Language Processing based on one or more eliminated seed terms and one or more verified seed terms.

14 . The computer system of claim 13 , wherein more than one similarity threshold is utilized, and wherein at least one of the similarity thresholds is automatically verified.

15 . The computer system of claim 12 , wherein the text analysis techniques include at least keyword extraction and the data related to the target business includes internal documentation, the external documentation, and manual input received from the user within the user interface, wherein the manual input is ontology terms relating to a taxonomy of the target business.

16 . A computer program product for competitive analysis, comprising:

one or more non-transitory computer-readable storage media and program instructions stored on at least one of the one or more tangible storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

receiving data related to a target business, wherein the target business is identified based on a user selection within a user interface;

identifying a plurality of seed terms associated with the target business using a machine learning model with Natural Language Processing to analyze the data related to the target business using one or more text analysis techniques;

performing a semantic relatedness analysis of the plurality of seed terms to determine whether each of the plurality of seed terms are above or below a similarity threshold, wherein one or more of the plurality of seed terms are determined to be above or below the similarity threshold based on a cosine similarity distance and requiring manual verification or elimination by the user within the user interface, and wherein the semantic relatedness analysis further includes analyzing the plurality of seed terms based on their source and using string similarity, Latent Semantic Indexing (LSI), and averaging unigram word embeddings;

training the machine learning model with Natural Language Processing based on one or more seed terms of the plurality of seed terms which are manually verified by the user within the user interface;

identifying a plurality of related businesses to the target business by searching a knowledge corpus using the one or more seed terms, wherein the knowledge corpus is comprised of external documentation and publicly available information related to a plurality of businesses;

determining one or more competitors by eliminating at least one of the plurality of related businesses identified, wherein the at least one of the plurality of related businesses identified are eliminated based on a weighting of mentions of the one or more seed terms, wherein the weighting of the mentions is determined based on at least a recency of the mention, a form of the external documentation in which the mention is derived, and an engagement with the external documentation in which the mention is derived; and

generating a competitive analysis report, wherein the competitive analysis report includes at least a favorability of the target business in comparison to the one or more competitors and an aggregate competitive score over time which is continuously updated in real time based on additional seed terms identified in the external documentation related to the one or more competitors and manual changes to the plurality of seed terms by the user, wherein the competitive score over time is displayed to the user within the user interface utilizing visual insights.

17 . The computer program product of claim 16 , wherein performing the semantic relatedness analysis further comprises:

tokenizing the plurality of seed terms using n-Gram tokenization and converting a plurality of tokenized seed terms into a plurality of vectors, wherein each of the plurality of vectors are compared using cosine similarity;

eliminating, by the user, the one or more seed terms below the similarity threshold;

verifying, by the user, the one or more seed terms above the similarity threshold;

storing the one or more verified seed terms in a target business knowledge corpus; and

training the machine learning model with Natural Language Processing based on one or more eliminated seed terms and one or more verified seed terms.

18 . The computer program product of claim 17 , wherein more than one similarity threshold is utilized, and wherein at least one of the similarity thresholds is automatically verified.

19 . The computer program product of claim 16 , wherein the text analysis techniques include at least keyword extraction and the data related to the target business includes internal documentation, the external documentation, and manual input received from the user within the user interface, wherein the manual input is ontology terms relating to a taxonomy of the target business.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2020
From: SMITH, JONATHAN M.; USMANI, SHEEMA; SHYPULA, ALEXANDER; SIMPLICIO, PHILLIP WERNER; SRIVASTAVA, BIPLAV; SABET SARVESTANI, AMIR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054614/0713 →
Continuity (1)
Related Publication 20220188653A1 · Jun 16, 2022
References Cited (28)
US 6446061B1 · Doerre · 2002 [cited by applicant]
US 6711589B2 · Dietz · 2004 [cited by applicant]
US 7673234B2 · Kao · 2010 [cited by applicant]
US 7974984B2 · Reuther · 2011 [cited by applicant]
US 10354188B2 · Chalabi · 2019 [cited by applicant]
US 10394955B2 · Fauceglia · 2019 [cited by applicant]
US 10552468B2 · Ciulla · 2020 [cited by applicant]
US 11392651B1 · McClusky · 2022 [cited by examiner]
US 20100153183A1 · Ulwick · 2010 [cited by examiner]
US 20140006415A1 · Rubchinsky · 2014 [cited by examiner]
US 20150278731A1 · Schwaber · 2015 [cited by examiner]
US 20190026280A1 · Aviyam · 2019 [cited by examiner]
US 20190221080A1 · Reetz · 2019 [cited by examiner]
US 20210287261A1 · Medalion · 2021 [cited by examiner]
Baum, et al., “Hits and Misses: Managers' (Mis)Categorization of Competitors in the Manhattan Hotel Industry”, Geography and Strategy (Advances in Strategic management) Abstract, 2003, 1 page, vol. 20, Emerald Group Pub… [cited by applicant]
Disclosed Anonymously, “Context-Based Concept Resolution with Structured and Unstructured Sources,” IP.com, May 17, 2016, 7 pages, IP.com No. IPCOM000246223D. [cited by applicant]
Disclosed Anonymously, “Method and System for Creating Dynamic Ontologies and Knowledge Graphs for Search Engines,” IP.com, Jan. 24, 2018, 4 pages, IP.com No. IPCOM000252560D. [cited by applicant]
Lederman, et al., “Identifying Competitors in Markets with Fixed Product Offerings,” Columbia Business School Research Paper No. 14-10, Jan. 2014, 39 pages, Retrieved from the Internet: <URL: https://papers.ssm.com/sol3… [cited by applicant]
Meij, et al., “Method and System for Automatically Explaining Entity Relationships in a Knowledge Graph,” IP.com, Jun. 15, 2015, 3 pages, IP.com No. IPCOM000242025D. [cited by applicant]
Mell, et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Morales, et al., “Computerized Competitiveness Analysis,” Application and Drawings, Filed on Aug. 6, 2020, 29 Pages, CN Patent Application Serial No. 202010782699.7. [cited by applicant]
Morales, et al., “Computerized Competitiveness Analysis,” Application and Drawings, Filed on Sep. 9, 2019, 41 Pages, U.S. Appl. No. 16/564,252. [cited by applicant]
Reynolds, “The Organization Ontology,” W3C via Wayback Machine, Jan. 16, 2014 [accessed on Dec. 11, 2020], 20 pages, Retrieved from the Internet: <URL: https://web.archive.org/web/20201121143504/http://www.w3.org/ TR/vo… [cited by applicant]
Santini, et al., “Designing an Extensible Domain-Specific Web Corpus for “Layfication”: A Case Study in eCare at Home,” Cyber-Physical Systems for Social Applications, 2019, 61 pages, IGI Global, DOI: 10.4018/978-1-5225… [cited by applicant]
Screen capture from The FOAF Project entitled “FOAF (2000-2015+),” [accessed on Jul. 8, 2020], 1 page, Retrieved from the Internet: <URL: http://www.foaf-project.org/>. [cited by applicant]
Undisclosed, “Skills-ML,” Data at Work, [accessed on Jul. 8, 2020], 2 pages, Alfred P. Sloan Foundation, Retrieved from the Internet: <URL: http://dataatwork.org/skills-ml/>. [cited by applicant]
Xu, “Bootstrapping Relation Extraction from Semantic Seeds,” Dissertation, 176 pages, Saarland University, DE. [cited by applicant]
Zhang, et al., “SemRe-Rank: Improving Automatic Term Extraction by Incorporating Semantic Relatedness With Personalised PageRank,” ACM Trans. Knowl. Discov. Data, Nov. 2017 [Mar. 28, 2018], 40 pages, vol. 9, Issue 4, Ar… [cited by applicant]