IP Library Granted Patent US 11,580,320
Granted Patent B2
US 11,580,320 · App. 17/186,998 · Granted Feb 14, 2023

Algorithm for scoring partial matches between words

Inventors: Rushik Upadhyay (Milpitas, CA); Dhamodharan Lakshmipathy (San Jose, CA); Nandhini Ramesh (San Jose, CA); Aditya Kaulagi (San Francisco, CA)
Assignee: PAYPAL, INC.
G06K9/6215G06F40/216G06F40/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,320
App. No.
17/186,998
Granted
Feb 14, 2023
Kind
B2
Abstract

Techniques are disclosed relating to scoring partial matches between words. In certain embodiments, a method may include receiving a request to determine a similarity between an input text data and a stored text data. The method also includes determining, based on comparing one or more words included in the input text data with one or more words included in the stored text data, a set of word pairs and a set of unpaired words. Further, in response to determining that the set of unpaired words passes elimination criteria, the method includes calculating a base similarity score between the input text data and the stored text data based on the set of word pairs. The method also includes determining a scoring penalty based on the set of unpaired words and generating a final similarity score between the input text data and the stored text data by modifying the base similarity score based on the scoring penalty.

Claims (55)

1. A system comprising:

one or more hardware processors; and

a memory storing computer-executable instructions, that in response to execution by the one or more hardware processors, causes the system to perform operations comprising:

receiving a request to determine a similarity score between an input set of words and a stored set of words;

determining a base similarity score using at least a distance score associated with pairs of words between the input set of words and the stored set of words;

normalizing the base similarity score determined using an average word length from the pairs between the input set of words and the stored set of words;

calculating a base penalty score on a number of word pairs included in the pairs between the input set of words and the stored set of words, wherein the base penalty score is further calculated using at least a difference between a normal base score and the base similarity score; and

determining a final similarity score based at least in part on the base penalty score calculated.

2. The system of claim 1 , wherein the operations further comprise:

updating the base penalty score based at least on a predetermined weight, wherein the predetermined weight is greater than a first predetermined weight.

3. The system of claim 1 , wherein the operations further comprise:

determining a total penalty score via a summation of base penalty scores for each of the pairs; and

updating the base penalty score with the total penalty score.

4. The system of claim 1 , wherein the operations further comprise:

determining phonetic codes for each word in each of the input set of words and the stored set of words;

identifying a set of unpaired words between the input set of words and the stored set of words using the phonetic codes and the distance scores determined; and

adjusting the base similarity score based at least in part on the set of unpaired words identified.

5. The system of claim 4 , wherein the normalizing includes using an average word length from the set of unpaired words.

6. The system of claim 1 , wherein determining the final similarity score based in part on the base penalty score calculated includes reducing the base similarity score based at least on the base penalty score calculated.

7. The system of claim 1 , wherein the determining the base similarity score is further based on an index file having a set of entity identifiers associated with past transactions processed by the entity identifiers.

8. A method comprising:

receiving a request to determine a similarity score between an input set of words and a stored set of words;

determining a base similarity score using at least a distance score associated with pairs of words between the input set of words and the stored set of words;

determining phonetic codes for each word in each of the input set of words and the stored set of words;

identifying a set of unpaired words between the input set of words and the stored set of words using the phonetic codes and the distance scores;

adjusting the base similarity score based at least in part on the set of unpaired words identified;

normalizing the adjusted base similarity score using an average word length from the pairs between the input set of words and the stored set of words;

calculating a base penalty score on a number of word pairs included in the pairs between the input set of words and the stored set of words; and

determining a final similarity score based at least in part on the base penalty score calculated.

9. The method of claim 8 , wherein the base penalty score is determined using at least a difference between a normal base score and the base similarity score.

10. The method of claim 8 , further comprising:

updating the base penalty score based at least on a predetermined weight, wherein the predetermined weight is greater than a first predetermined weight.

11. The method of claim 8 , further comprising:

determining a total penalty score via a summation of base penalty scores for each of the pairs; and

updating the base penalty score with the total penalty score.

12. The method of claim 8 , wherein the normalizing includes using an average word length from the set of unpaired words.

13. The method of claim 8 , wherein determining the final similarity score based at least in part on the base penalty score calculated includes reducing the base similarity score based on the base penalty score calculated.

14. The method of claim 8 , wherein the determining the base similarity score is further based on a blacklist of entities associated with fraudulent transactions processed by the entities.

15. A non-transitory computer readable medium storing computer-executable instructions that in response to execution by one or more hardware processors, causes a payment provider system to perform operations comprising:

receiving a request to determine a similarity score between an input set of words and a stored set of words;

determining a base similarity score using at least a distance score associated with pairs of words between the input set of words and the stored set of words;

normalizing the base similarity score determined using an average word length from the pairs between the input set of words and the stored set of words;

calculating a base penalty score on a number of word pairs included in the pairs between the input set of words and the stored set of words, wherein the base penalty score is further calculated using at least a difference between a normal base score and the base similarity score; and

determining a final similarity score based at least in part on the base penalty score calculated.

16. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

updating the base penalty score based on a predetermined weight, wherein the predetermined weight is greater than a first predetermined weight.

17. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

determining a total penalty score via a summation of base penalty scores for each of the pairs; and

updating the base penalty score with the total penalty score.

18. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

determining phonetic codes for each word in each of the input set of words and the stored set of words;

identifying a set of unpaired words between the input set of words and the stored set of words using the phonetic codes and the distance scores determined; and

adjusting the base similarity score based at least in part on the set of unpaired words identified.

19. The non-transitory computer readable medium of claim 18 , wherein the normalizing includes using the average word length from the set of unpaired words.

20. The non-transitory computer readable medium of claim 15 , wherein the determining the base similarity score is further based on an index file having a set of entity identifiers associated with past transactions processed by the entity identifiers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2021
From: UPADHYAY, RUSHIK; KAULAGI, ADITYA; LAKSHMIPATHY, DHAMODHARAN; RAMESH, NANDHINI
To: PAYPAL, INC.
Reel/Frame 056316/0102 →
Continuity (2)
Continuation 16234855 · Dec 28, 2018
Related Publication 20210232852A1 · Jul 29, 2021