IP Library › Granted Patent US 12,169,508
Granted Patent B2
US 12,169,508 · App. 17/481,866 · Granted Dec 17, 2024

System and method for entity disambiguation for customer relationship management

Inventors: Johannes Julien Frederik Erett (London, GB); James A Hodson (El Cerrito, CA)
Assignee: Cognism Limited
G06F16/285G06F7/14G06F16/245G06N5/01G06N20/20G06Q10/063112G06Q30/01G06Q30/015
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,169,508
App. No.
17/481,866
Granted
Dec 17, 2024
Kind
B2
Abstract

Systems and methods for disambiguating company profiles are disclosed. The system builds a database of candidate companies with timed attributes. The system further ingests timed company metadata using a PostgreSQL database. The system disambiguates location and geocoding by matching text patterns and cross-referencing one or more identified location components against one or more geocode databases and classifies and disambiguates company name component from a company name associated with the company using a conditional random field (CRF) model. Further, the system disambiguates employee attributes using a Latent Dirichlet Allocation (LDA) topic model algorithm and train a tree model for pre-selection of candidate companies for the company. The system trains a similarity model for comparison of the candidate companies and in response to a determination that two given candidate companies are same merge the company profiles associated with the given candidate companies.

Claims (10)

1. A system for disambiguating company profiles, the system comprising:

an entity disambiguation computer comprising a memory, a processor, and a plurality of programming instructions, the plurality of programming instructions when executed by the processor cause the processor to:

receive a plurality of candidate company profiles with timed attributes from a database;

monitor the plurality of received candidate company profiles, and extract and ingest at dynamic intervals, a plurality of timed metadata from the plurality of candidate company profiles, the timed metadata comprising individual company data, each timed metadata, of the plurality of timed company profiles, linked to a company profile, of the plurality of the candidate company profiles, associated to a time frame, wherein the individual company data for a company profile comprises, at least location, geocodes, company name, employee attributes, average company headcount reporting data over a given time frame, company website and URL information, and company employment records;

disambiguate, from at least a portion of timed metadata, a plurality of locations and geocodes, for the time frame, using regular expressions to match components of unstructured text and cross-reference one or more identified location components against one or more geocode databases;

classify and disambiguate, from at least a portion of timed metadata, a company name component, for the time frame, from a company name associated with at least one company profile, of the plurality of candidate company profiles, to identify at least a base name, a connector, a function and/or industry, and a legal identifier associated with the company name, wherein a conditional random field (CRF) is used to disambiguate the company name component;

disambiguate, from at least a portion of timed metadata, employee attributes, for the time frame, by mapping skills to skill topics implemented with a Latent Dirichlet Allocation (LDA) topic model algorithm;

train a tree model for pre-selection of a plurality of candidate companies for the first company profile, wherein one or more manually annotated training examples are used to train the tree model, so as to facilitate identification of a pre-selection of the plurality of candidate companies, wherein one or more manually annotated training examples are used to train the tree model, so as to facilitate identification of a pre-selection of the plurality of candidate companies, and wherein one or more algorithms used to train the tree model include Random Forrest algorithm, Gradient Boosting algorithm, Decision Tree algorithm, or a combination thereof;

train a similarity model for comparison of the plurality of candidate companies based at least on the pre-selection of the plurality of candidate companies, wherein the similarity model is trained to determine whether two given candidate companies are the same, wherein the similarity model is trained using one of a Regression Algorithm, a Neural Network, or a Vector Similarity Algorithm paired with a learned threshold model; and

in response to a determination that at least two given candidate companies, of the plurality of candidate companies are matched, merge the timed metadata associated with the matched candidate companies.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2024
From: ERETT, JOHANNES JULIEN FREDERIK; HODSON, JAMES A
To: COGNISM LIMITED
Reel/Frame 068970/0460 →
Continuity (2)
Provisional Application 63081761 · Sep 22, 2020
Related Publication 20220114198A1 · Apr 14, 2022