IP Library › Granted Patent US 11,436,241
Granted Patent B2
US 11,436,241 · App. 16/506,792 · Granted Sep 6, 2022

Entity resolution based on character string frequency analysis

Inventor: Girish Kunjur (Houston, TX)
Assignee: Fair Isaac Corporation
G06F16/2462G06F16/24578G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,436,241
App. No.
16/506,792
Filed
Jul 9, 2019
Granted
Sep 6, 2022
Kind
B2
Examiner
VO, CECILE H
Art Unit
2153
USPC
707/737
Abstract

Computer-implemented methods, systems and products for character string frequency analysis. The method includes a set of operations or steps, including parsing a plurality of character strings into one or more tokens, categorizing the one or more tokens into one or more token frequency categories, and generating a first similarity score between one or more pairs of character strings of the plurality of character strings. The method further includes calculating one or more degrees of commonality or rarity of the plurality of character strings based on the categorizing, generating one or more penalties for token pairs of the one or more pairs of character strings associated with the first similarity score based on the one or more degrees of commonality or rarity and the categorizing, and generating a second similarity score based the first similarity score and the one or more penalties.

Claims (67)

1. A computer-implemented method comprising:

receiving, using at least one processor, a request to resolve an entity represented by one or more first character strings;

querying, using the at least one processor, via a communication network, in response to the received request, a plurality of second character strings stored in one or more disparate data sources, identifying, using a federated search, one or more second character strings in the plurality of second character strings being similar to the one or more first character strings, retrieving the one or more second character strings from the one or more disparate data sources, and clustering the one or more second character strings based on at least one similarity parameter to cause the at least one processor to execute a resolution of the entity;

executing, using the at least one processor, the resolution of the entity by

parsing the one or more first character strings and the one or more clustered second character strings into one or more tokens;

categorizing the one or more tokens into at least one of one or more token frequency categories;

generating a first similarity score between one or more pairs of character strings selected from the one or more first character strings and the one or more clustered second character strings based on one or more of lexical, synonymic and sound similarities or differences;

calculating one or more degrees of commonality or rarity of the one or more first character strings and the one or more clustered second character strings based on the categorizing of the one or more tokens;

generating one or more penalties for token pairs of the one or more pairs of character strings associated with the first similarity score based on the one or more degrees of commonality or rarity and the categorizing of the one or more tokens; and

generating a second similarity score based the first similarity score and the one or more penalties;

resolving, using the at least one processor, the entity based on the generated second similarity score; and

transmitting, using the at least one processor, via the communication network, the resolved entity for processing by at least one application.

2. The computer-implemented method in accordance with claim 1 , wherein generating the first similarity score between the one or more pairs of character strings comprises:

identifying the token pairs of the one or pairs of character strings; and

identifying one or more similarities between the identified token pairs, the one or more similarities including one or more of lexical similarities, synonymic properties and sound closeness.

3. The computer-implemented method in accordance with claim 1 , wherein generating the first similarity score between the one or more pairs of character strings comprises:

identifying the token pairs of the one or pairs of character strings; and

identifying one or more non-token variations between the identified token pairs, the one or more non-token variations including one or more of missing tokens and added tokens.

4. The computer-implemented method in accordance with claim 1 , wherein categorizing the one or more tokens uses a locale-specific lookup table.

5. The computer-implemented method in accordance with claim 1 , wherein calculating the one or more degrees of commonality or rarity uses a sigmoid function.

6. The computer implemented method in accordance with claim 1 , wherein the one or more penalties comprise token penalties associated with differences between the tokens in the token pairs.

7. The computer-implemented method in accordance with claim 6 , wherein the token penalties are generated based on both the one or more degrees of commonality and rarity.

8. The computer implemented method in accordance with claim 1 , wherein the one or more penalties comprise non-token penalties associated with added or missing tokens in the one or more pairs of character strings.

9. The computer-implemented method in accordance with claim 8 , wherein the non-token penalties are generated based on both the one or more degrees of commonality and rarity.

10. The computer-implemented method in accordance with claim 1 , further comprising storing, by the one or more computer processors, the second similarity score in a similarity index of a database.

11. A system comprising:

at least one programmable processor;

a non-transitory machine-readable medium storing instructions that, when executed by the at least one programmable processor, cause the at least one programmable processor to perform operations comprising:

receiving a request to resolve an entity represented by one or more first character strings;

querying, via a communication network, in response to the received request, a plurality of second character strings stored in one or more disparate data sources, identifying, using a federated search, one or more second character strings in the plurality of second character strings being similar to the one or more first character strings, retrieving the one or more second character strings from the one or more disparate data sources, and clustering the one or more second character strings based on at least one similarity parameter to cause the at least one programmable processor to execute a resolution of the entity;

executing the resolution of the entity by

parsing the one or more first character strings and the one or more clustered second character strings into one or more tokens;

categorizing the one or more tokens into at least one of one or more token frequency categories;

generating a first similarity score between one or more pairs of character strings selected from the one or more first character strings and the one or more clustered second character strings based on one or more of lexical, synonymic and sound similarities or differences;

calculating one or more degrees of commonality or rarity of the one or more first character strings and the one or more clustered second character strings based on the categorizing of the one or more tokens;

generating one or more penalties for token pairs of the one or more pairs of character strings associated with the first similarity score based on the one or more degrees of commonality or rarity and the categorizing of the one or more tokens; and

generating a second similarity score based the first similarity score and the one or more penalties;

resolving the entity based on the generated second similarity score; and

transmitting via the communication network, the resolved entity for processing by at least one application.

12. The system in accordance with claim 11 , wherein generating the first similarity score between the one or more pairs of character strings comprises:

identifying the token pairs of the one or pairs of character strings; and

identifying one or more similarities between the identified token pairs, the one or more similarities including one or more of lexical similarities, synonymic properties and sound closeness.

13. The system in accordance with claim 11 , wherein generating the first similarity score between the one or more pairs of character strings comprises:

identifying the token pairs of the one or pairs of character strings; and

identifying one or more non-token variations between the identified token pairs, the one or more non-token variations including one or more of missing tokens and added tokens.

14. The system in accordance with claim 11 , wherein categorizing the one or more tokens uses a locale-specific lookup table.

15. The system in accordance with claim 11 , wherein calculating the one or more degrees of commonality or rarity uses a sigmoid function.

16. A computer program product comprising a non-transitory machine-readable medium storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to perform operations comprising:

receiving a request to resolve an entity represented by one or more first character strings;

querying, via a communication network, in response to the received request, a plurality of second character strings stored in one or more disparate data sources, identifying, using a federated search, one or more second character strings in the plurality of second character strings being similar to the one or more first character strings, retrieving the one or more second character strings from the one or more disparate data sources, and clustering the one or more second character strings based on at least one similarity parameter to cause the at least one programmable processor to execute a resolution of the entity;

executing the resolution of the entity by

parsing the one or more first character strings and the one or more clustered second character strings into one or more tokens;

categorizing the one or more tokens into at least one of one or more token frequency categories;

generating a first similarity score between one or more pairs of character strings selected from the one or more first character strings and the one or more clustered second character strings based on one or more of lexical, synonymic and sound similarities or differences;

calculating one or more degrees of commonality or rarity of the one or more first character strings and the one or more clustered second character strings based on the categorizing of the one or more tokens;

generating one or more penalties for token pairs of the one or more pairs of character strings associated with the first similarity score based on the one or more degrees of commonality or rarity and the categorizing of the one or more tokens; and

generating a second similarity score based the first similarity score and the one or more penalties;

resolving the entity based on the generated second similarity score; and

transmitting, via the communication network, the resolved entity for processing by at least one application.

17. The computer program product in accordance with claim 16 , wherein generating the first similarity score between the one or more pairs of character strings comprises:

identifying the token pairs of the one or pairs of character strings; and

identifying one or more similarities between the identified token pairs, the one or more similarities including one or more of lexical similarities, synonymic properties and sound closeness.

18. The computer program product in accordance with claim 16 , wherein generating the first similarity score between the one or more pairs of character strings comprises:

identifying the token pairs of the one or pairs of character strings; and

identifying one or more non-token variations between the identified token pairs, the one or more non-token variations including one or more of missing tokens and added tokens.

19. The computer program product in accordance with claim 16 , wherein categorizing the one or more tokens uses a locale-specific lookup table.

20. The computer program product in accordance with claim 16 , wherein calculating the one or more degrees of commonality or rarity uses a sigmoid function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2019
From: KUNJUR, GIRISH
To: FAIR ISAAC CORPORATION
Reel/Frame 049709/0184 →
Continuity (1)
Related Publication 20210011909A1 · Jan 14, 2021
Cited By (3)
US 12,205,138 US 12,354,159 US 12,585,970