IP Library › Granted Patent US 11,288,241
Granted Patent B1
US 11,288,241 · App. 16/580,955 · Granted Mar 29, 2022

Systems and methods for integration and analysis of data records

Inventors: Sears Merritt (Groton, MA); Thom Neale (Bedford, MA)
Assignee: MASSACHUSETTS MUTUAL LIFE INSURANCE COMPANY
G06F16/212G06F16/215G06F16/258G06F16/285G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,288,241
App. No.
16/580,955
Granted
Mar 29, 2022
Kind
B1
Abstract

Methods and systems for determining relationships between two or more nominally unrelated data sources utilizing a combination of probabilistic modeling and graphical clustering are described. The systems and methods for utilizing probabilistic model functions as a way of determining and judging the likelihood that two records from different systems are related to the same entity.

Claims (33)

1. A server-implemented method comprising:

preprocessing, by a server, a plurality of datasets by converting the plurality of datasets in a plurality of data formats to the plurality of datasets having a common format;

grouping, by the server, the plurality of datasets into one or more pairs based on one or more blocking keys associated with each dataset, wherein each dataset has a blocking key configured to generate hash values for each dataset based on content of each dataset;

determining, by the server, a score for each pair based on the hash values for each respective dataset in the one or more pairs;

classifying, by the server, each pair based on the score computed for the pair satisfying a predetermined threshold, wherein the server is configured to link datasets of the plurality of datasets within a same classification;

generating, by the server, based on respective edge identifiers generated for each of the one or more pairs according to the classification of the one or more pairs, a graph structure representing a cluster of linked datasets within the plurality of datasets that are determined to be related to each other, wherein the graph structure is generated based on a join operation performed according to the respective edge identifiers generated for each of the one or more pairs; and

in response to receiving a request to search a data record from a computing device of a user, querying, by the server, the cluster of linked datasets to search the data record.

2. The server-implemented method of claim 1 , wherein the plurality of datasets are arranged in a plurality of fields.

3. The server-implemented method of claim 2 , further comprising assigning, by the server, a unique identifier to each field associated with the plurality of datasets.

4. The server-implemented method of claim 3 , wherein the plurality of fields are mapped to a common set of fields selected from a library of fields.

5. The server-implemented method of claim 1 , wherein each of the plurality of datasets has a dataset_name.

6. The server-implemented method of claim 1 , wherein the common format is a single text encoding format.

7. The server-implemented method of claim 1 , further comprising normalizing, by the server, each of the plurality of datasets.

8. The server-implemented method of claim 7 , wherein normalization is performed by executing a master schema.

9. The server-implemented method of claim 8 , wherein the master schema is configured to identify each semantically equivalent field associated with each of the plurality of datasets.

10. The server-implemented method of claim 1 , wherein the score is a floating-point ratio computed based on the hash values associated with each dataset within each pair.

11. A system comprising:

a server configured to:

preprocess a plurality of datasets by converting the plurality of datasets in a plurality of data formats to the plurality of datasets having a common format;

group the plurality of datasets into one or more pairs based on one or more blocking keys associated with each dataset, wherein each dataset has a blocking key configured to generate hash values for each dataset based on content of each dataset;

determine a score for each pair based on the hash values for each respective dataset in the one or more pairs;

classify each pair based on the score computed for the pair satisfying a predetermined threshold, wherein the server is configured to link datasets of the plurality of datasets within a same classification;

generate, based on respective edge identifiers generated for each of the one or more pairs according to the classification of the one or more pairs, a graph structure representing a cluster of linked datasets within the plurality of datasets that are determined to be related to each other, wherein the graph structure is generated based on a join operation performed according to the respective edge identifiers generated for each of the one or more pairs; and

in response to receiving a request to search a data record from a computing device of a user, query the cluster of linked datasets to search the data record.

12. The system of claim 11 , wherein the plurality of datasets are arranged in a plurality of fields.

13. The system of claim 12 , wherein the server is further configured to assign a unique identifier to each field associated with the plurality of datasets.

14. The system of claim 13 , wherein the plurality of fields are mapped to a common set of fields selected from a library of fields.

15. The system of claim 11 , wherein each of the plurality of datasets has a dataset name.

16. The system of claim 11 , wherein the common format is a single text encoding format.

17. The system of claim 11 , wherein the server is further configured to normalize each of the plurality of datasets.

18. The system of claim 17 , wherein normalization is performed by executing a master schema.

19. The system of claim 18 , wherein the master schema is configured to identify each semantically equivalent field associated with each of the plurality of datasets.

20. The system of claim 11 , wherein the score is a floating-point ratio computed based on the hash values associated with each dataset within each pair.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2019
From: MERRITT, SEARS; NEALE, THOM
To: MASSACHUSETTS MUTUAL LIFE INSURANCE COMPANY
Reel/Frame 050477/0559 →
Continuity (2)
Continuation 15380925 · Dec 15, 2016
Provisional Application 62387164 · Dec 23, 2015
Cited By (3)
US 12,248,485 US 12,625,851 US 12,717,991