IP Library Granted Patent US 10,339,118
Granted Patent B1
US 10,339,118 · App. 15/674,993 · Granted Jul 2, 2019

Data normalization system

Inventor: Luke Davis (Winchcombe, GB)
Assignee: Palantir Technologies Inc.
G06F16/215G06F7/02G06F16/2468G06F16/2471G06F16/3329G06F16/3343G06F16/3344G06F16/3346
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,339,118
App. No.
15/674,993
Granted
Jul 2, 2019
Kind
B1
Abstract

A data normalization system receives a first string and a second string that are ordered according to an initial string ordering. The data normalization system analyzes, the first string and the second string based on a list of known character sets included in surnames, yielding an analysis, and determines, based on the analysis, that a set of characters in the second string matches a known character set included in the list of known character sets included in surnames. In response to determining that the set of characters in the second string matches a known character set included in the list of known character sets included in surname, the data normalization system orders the first string and the second string according to an updated string ordering.

Claims (61)

1. A data normalization system comprising:

one or more computer processors; and

one or more computer-readable mediums storing instructions that, when executed by the one or more computer processors, causes the data normalization system to perform operations comprising:

receiving a first string and a second string, the first string and the second string being ordered according to an initial string ordering;

converting, using a metaphone algorithm, the first string into a first metaphone string, and the second string into a second metaphone string;

searching, based on the first metaphone string and the second metaphone string, a name index including a listing of metaphone strings representing common names and probabilities that the common names are either a given name or a surname;

determining, based on searching the name index, a confidence score indicating a confidence level that the first string represents the given name and that the second string represents the surname;

determining that the confidence score does not meet or exceed a threshold confidence score;

in response to determining that the confidence score does not meet or exceed the threshold confidence score, analyzing, the first string and the second string based on a list of known character sets included in surnames, yielding an analysis;

determining, based on the analysis, that a set of characters in the second string matches a known character set included in the list of known character sets included in surnames; and

in response to determining that the set of characters in the second string matches a known character set included in the list of known character sets included in surname, ordering the first string and the second string according to an updated string ordering.

2. The data normalization system of claim 1 , the operations further comprising:

calculating an updated confidence score based on the analysis; and

determining that the updated confidence score meets or exceeds the threshold confidence score.

3. The data normalization system of claim 1 , the operations further comprising:

determining, based on searching the name index, a signal strength that the first metaphone string is a given name, a signal strength that the first metaphone string is a surname, a signal strength that the second metaphone string is a given name, and a signal strength that the second metaphone string is a surname;

determining an absolute difference between the signal strength that the first metaphone string is a surname and the signal strength that the first metaphone string is a given name, yielding an absolute signal difference for the first string;

determining an absolute difference between the signal strength that the second metaphone string is a surname and the signal strength that the second metaphone string is a given name, yielding an absolute signal difference for the second string; and

determining that the absolute signal difference for the first string is greater than the absolute signal difference for the second string.

4. The data normalization system of claim 3 , wherein the confidence score is further based on determining that the absolute signal difference for the first string is greater than the absolute signal difference for the second string.

5. The data normalization system of claim 1 wherein the name index includes at least a first set of metaphone strings corresponding to a first language set, and a second set of metaphone strings corresponding to a second language set.

6. The data normalization system of claim 5 , wherein a metaphone string from the first set of metaphone strings and a metaphone string from the second set of metaphone strings both represent an identical name.

7. A method comprising:

receiving a first string and a second string, the first string and the second string being ordered according to an initial string ordering;

converting, using a metaphone algorithm, the first string into a first metaphone string, and the second string into a second metaphone string;

searching, based on the first metaphone string and the second metaphone string, a name index including a listing of metaphone strings representing common names and probabilities that the common names are either a given name or a surname;

determining, based on searching the name index, a confidence score indicating a confidence level that the first string represents the given name and that the second string represents the surname;

determining that the confidence score does not meet or exceed a threshold confidence score;

in response to determining that the confidence score does not meet or exceed the threshold confidence score, analyzing, the first string and the second string based on a list of known character sets included in surnames, yielding an analysis;

determining, based on the analysis, that a set of characters in the second string matches a known character set included in the list of known character sets included in surnames; and

in response to determining that the set of characters in the second string matches a known character set included in the list of known character sets included in surname, ordering the first string and the second string according to an updated string ordering.

8. The method of claim 7 , further comprising:

calculating an updated confidence score based on the analysis; and

determining that the updated confidence score meets or exceeds the threshold confidence score.

9. The method of claim 7 , further comprising:

determining, based on searching the name index, a signal strength that the first metaphone string is a given name, a signal strength that the first metaphone string is a surname, a signal strength that the second metaphone string is a given name, and a signal strength that the second metaphone string is a surname;

determining an absolute difference between the signal strength that the first metaphone string is a surname and the signal strength that the first metaphone string is a given name, yielding an absolute signal difference for the first string;

determining an absolute difference between the signal strength that the second metaphone string is a surname and the signal strength that the second metaphone string is a given name, yielding an absolute signal difference for the second string; and

determining that the absolute signal difference for the first string is greater than the absolute signal difference for the second string.

10. The method of claim 9 , wherein the confidence score is further based on determining that the absolute signal difference for the first string is greater than the absolute signal difference for the second string.

11. The method of claim 7 , wherein the name index includes at least a first set of metaphone strings corresponding to a first language set, and a second set of metaphone strings corresponding to a second language set.

12. The method of claim 11 , wherein a metaphone string from the first set of metaphone strings and a metaphone string from the second set of metaphone strings both represent an identical name.

13. A non-transitory computer-readable medium storing instructions that, when executed by one or more computer processors of a data normalization system, causes the data normalization system to perform operations comprising:

receiving a first string and a second string, the first string and the second string being ordered according to an initial string ordering;

converting, using a metaphone algorithm, the first string into a first metaphone string, and the second string into a second metaphone string;

searching, based on the first metaphone string and the second metaphone string, a name index including a listing of metaphone strings representing common names and probabilities that the common names are either a given name or a surname;

determining, based on searching the name index, a confidence score indicating a confidence level that the first string represents the given name and that the second string represents the surname;

determining that the confidence score does not meet or exceed a threshold confidence score;

in response to determining that the confidence score does not meet or exceed the threshold confidence score, analyzing, the first string and the second string based on a list of known character sets included in surnames, yielding an analysis;

determining, based on the analysis, that a set of characters in the second string matches a known character set included in the list of known character sets included in surnames; and

in response to determining that the set of characters in the second string matches a known character set included in the list of known character sets included in surname, ordering the first string and the second string according to an updated string ordering.

14. The non-transitory computer-readable medium of claim 13 , the operations further comprising:

calculating an updated confidence score based on the analysis; and

determining that the updated confidence score meets or exceeds the threshold confidence score.

15. The non-transitory computer-readable medium of claim 13 , the operations further comprising:

determining, based on searching the name index, a signal strength that the first metaphone string is a given name, a signal strength that the first metaphone string is a surname, a signal strength that the second metaphone string is a given name, and a signal strength that the second metaphone string is a surname;

determining an absolute difference between the signal strength that the first metaphone string is a surname and the signal strength that the first metaphone string is a given name, yielding an absolute signal difference for the first string;

determining an absolute difference between the signal strength that the second metaphone string is a surname and the signal strength that the second metaphone string is a given name, yielding an absolute signal difference for the second string; and

determining that the absolute signal difference for the first string is greater than the absolute signal difference for the second string.

16. The non-transitory computer-readable medium of claim 15 , wherein the confidence score is further based on determining that the absolute signal difference for the first string is greater than the absolute signal difference for the second string.

17. The non-transitory computer-readable medium of claim 13 , wherein the name index includes at least a first set of metaphone strings corresponding to a first language set, and a second set of metaphone strings corresponding to a second language set, and a metaphone string from the first set of metaphone strings and a metaphone string from the second set of metaphone strings both represent an identical name.

Assignments (8)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2017
From: DAVIS, LUKE
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 043270/0001 →
Continuity (1)
Continuation 15391728 · Dec 27, 2016