IP Library › Granted Patent US 12,182,083
Granted Patent B2
US 12,182,083 · App. 18/210,084 · Granted Dec 31, 2024

System and method for entity disambiguation for customer relationship management

Inventors: Johannes Julien Frederik Erett (London, GB); James A Hodson (El Cerrito, CA)
Assignee: Cognism Limited
G06F16/214G06F16/2365
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,182,083
App. No.
18/210,084
Granted
Dec 31, 2024
Kind
B2
Abstract

A system and method for disambiguating entities for managing customer relationships are described. An entity disambiguation computer receives information associated with candidate entities in an entity database. The received information comprises multiple versions of attributes related to one or more entities. Attributes are disambiguated and extracted from the information. A set of timeslice objects representing the multiple versions of each attribute is created. A subset of timeslice objects is selected for comparison based on an overlap between durations in respective timeslice objects. The system and method use a similarity model comprising weight and biases assigned to sets of previously used overlapping durations to predict if the subset of timeslice objects corresponds to the same entity. The subset of timeslice objects is merged if predicted to correspond to the same entity. This merging of timeslice objects disambiguates the information present in the entity database.

Claims (50)

1. A system for disambiguating attributes associated with one or more entities, the system comprising:

an entity disambiguation computer comprising a memory, a processor, and a plurality of programming instructions, the plurality of programming instructions when executed by the processor cause the processor to:

receive electronic information associated with a candidate entity among the one or more entities in an entity database at pre-defined intervals, wherein the received electronic information comprises multiple versions of data associated with the one or more entities received from one or more external data sources over a network;

extract and store at the entity database, one or more attributes associated with the candidate entity from the electronic information;

for each of the one or more attributes, call and run software functions to:

create and store a set of timeslice objects, wherein the set of timeslice objects are associated with respective durations;

select and store a subset of timeslice objects from the set timeslice objects for candidate comparison based on an overlap between durations in respective timeslice objects;

predict whether the subset of timeslice objects corresponds to a same entity by comparing the overlapping durations in the subset of timeslice objects using a machine-learning similarity model comprising weight and biases assigned to sets of previously used overlapping durations to:

compute, for each attribute, distance vectors between the subset of timeslice objects, wherein a vectorizer converts the overlapping durations to distance vectors;

predict if the subset of timeslice objects represented by distance vectors correspond to the same entity by comparing the distance vectors with a similarity model comprising weight and biases assigned to sets of previous distance vectors;

responsive to predicting that the subset of timeslice objects correspond to the same entity, combine the subset of timeslice objects by merging the selected timeslice objects into a single entity identity record, and

store the prediction; and

responsive to determining, based on the prediction, that the subset of timeslice objects correspond to the same entity, store and merge the subset of timeslice objects to generate an unambiguous entity database.

2. The system of claim 1 , wherein the one or more attributes comprises a location, a geocode, an entity name, a stock symbol, a registered entity identity, an entity classification code, an entity uniform resource links (URLs), employee data, an entity event, a technology domain, an entity group connection, an entity brand, and a competitor.

3. The system of claim 2 , wherein to extract the one or more attributes, the plurality of instructions when executed by the processor, further cause the processor to:

tokenize the electronic information;

responsive to identifying that the electronic information has multiple components based on one or more tokens:

determine that the one or more attributes in the electronic information are related to an entity name based on the multiple components; and disambiguate and classify the multiple components into at least a base name, a connector, a function and/or industry, and a legal identifier associated with the entity name.

4. The system of claim 3 , wherein to extract the one or more attributes, the plurality of instructions when executed by the processor, further cause the processor to:

responsive to identifying that a first attribute, of the one or more attributes, is a location:

disambiguate and compare the one or more tokens associated with the location with a plurality of known locations;

responsive to determining that there is a match between the one or more tokens associated with the location and a first known location of the plurality of locations, assign a geocode to the location.

5. The system of claim 3 , wherein the plurality of instructions when executed by the processor, further cause the processor to:

responsive to determining that the one or more tokens are related to the employee data, disambiguate and classify employee attributes from the one or more tokens, wherein the employee attribute comprises an employee skill, an employee job title, a location of employee, a gender, and an educational qualifications.

6. The system of claim 3 , wherein the disambiguation and classification of the multiple components of an entity name is performed using at least one of fingerprinting, semantic embedding or a conditional random fields (CRF) classifier model.

7. A method for disambiguating attributes associated with one or more entities, the method comprising:

receiving, at an entity disambiguation computer, from one or more external data sources over a network, electronic information associated with a candidate entity among the one or more entities in an entity database at pre-defined intervals, wherein the received electronic information comprises multiple versions of the one or more entities;

extracting and storing at the entity database, by the entity disambiguation computer one or more attributes associated with the candidate entity from the electronic information;

for each of the one or more attributes, call and run software functions for:

creating and storing a set of timeslice objects, wherein the set of timeslice objects are associated with respective durations;

selecting and storing a subset of timeslice objects from the set timeslice objects for candidate comparison based on an overlap between durations in respective timeslice objects;

predicting whether the subset of timeslice objects corresponds to a same entity by comparing the overlapping durations in the subset of timeslice objects using a machine learning similarity model comprising weight and biases assigned to sets of previously used overlapping durations by:

computing, for each attribute, distance vectors between selected set of timeslice objects, wherein a vectorizer converts the overlapping durations to distance vectors;

predicting if the selected timeslice objects represented by distance vectors correspond to the same entity by comparing the distance vectors with a similarity model comprising weight and biases assigned to sets of previous distance vectors;

responsive to predicting that the selected timeslice objects correspond to the same entity, merging the selected timeslice objects into a single entity identity record; and generating an unambiguous entity database by merging of the subset of timeslice objects, and

storing the prediction; and

responsive to determining that the subset of timeslice objects correspond to the same entity, based on the prediction, storing and merging the subset of timeslice objects to generate an unambiguous entity database.

8. The method of claim 7 , wherein the one or more attributes comprises a location, a geocode, an entity name, a stock symbol, a registered entity identity, an entity classification code, an entity uniform resource links (URLs), employee data, an entity event, a technology domain, an entity group connection, an entity brand, and a competitor.

9. The method of claim 8 , wherein extracting the one or more attributes further comprises the steps of:

tokenizing the information;

responsive to identifying that information has multiple components based on one or more tokens:

determining that attribute in the received information is related to an entity name based on the multiple components; and

disambiguating and classifying the multiple components into at least a base name, a connector, a function and/or industry, and a legal identifier associated with the entity name.

10. The method of claim 9 , wherein extracting the one or more attributes further comprises the steps of:

responsive to identifying that a first attribute, of the one or more attributes, is a location:

disambiguating and comparing one or more tokens associated with the location with a plurality of known locations;

responsive to determining that there is match between the one or more tokens associated with the location and a first known location of the plurality of locations, assigning a geocode to the location.

11. The method of claim 10 , wherein extracting the one or more attributes further comprises the steps of:

responsive to determining that the one or more tokens are related to the employee data, disambiguating and classifying employee attributes from the one or more tokens, wherein the employee attribute comprises an employee skill, an employee job title, a location of employee, a gender, and an educational qualification.

12. The method of claim 9 , wherein the disambiguation and classification of the multiple components of an entity name is performed using at least one of fingerprinting, semantic embedding or a conditional random fields (CRF) classifier model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2024
From: ERETT, JOHANNES JULIEN FREDERIK; HODSON, JAMES A
To: COGNISM LIMITED
Reel/Frame 068977/0348 →
SECURITY INTEREST Recorded Aug 2, 2024
From: COGNISM INC.; COGNISM LIMITED
To: HSBC INNOVATION BANK LIMITED
Reel/Frame 068162/0061 →
Continuity (3)
Continuation In Part 17481866 · Sep 22, 2021
Provisional Application 63081761 · Sep 22, 2020
Related Publication 20230325366A1 · Oct 12, 2023