IP Library Granted Patent US 10,121,557
Granted Patent B2
US 10,121,557 · App. 14/466,909 · Granted Nov 6, 2018

System and method for dynamic document matching and merging

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,121,557
App. No.
14/466,909
Granted
Nov 6, 2018
Kind
B2
Abstract

A system and method for matching and merging documents from disparate data sources into a single data store for a particular entity are provided. The system and method may be particularly useful for a healthcare system to match and merge data from disparate data sources about a healthcare provider.

Claims (36)

1. A system for matching and merging data for an entity from disparate data sources, comprising:

one or more data sources for disparate sources, each data source having a raw file about an entity, each raw file having a plurality of data fields and each field has a value associated with the entity;

a computer system having a processor, the computer system coupled to each of the one or more data sources;

the processor of the computer system being configured to receive the one or more raw files about the entity from the one or more data sources;

the processor of the computer system being configured to perform each of a plurality of matching processes against each of the one or more raw files about the entity to generate a match set for each raw file matched with one of the plurality of matching processes based on results from each of the plurality of matching processes with each of the raw files wherein the match set for each raw file matched by each one of the plurality of matching processes a reference to a document storage location of the raw file, a unique identifier for the raw file and provenance metadata for the one matching process including an identifier of the matching process and parameters associated with the matching process;

the processor of the computer system being configured to merge the one or more raw files represented into a merged document that has the at least one matched data field in each of the one or more raw files; and

the processor of the computer system being configured to identify conflicting values in each data field of the merged document.

2. The system of claim 1 , wherein the processor is configured to transform, before performing the plurality of matching processes, the one or more raw files to initial data cleanse the one or more raw files and generate one or more transformed files about the entity.

3. The system of claim 1 , wherein the processor is configured to rank each conflicting value for a data field based on a confidence of a correctness of each conflicting value.

4. The system of claim 1 , wherein the processor is configured to perform a sequence of queries on the one or more raw files, wherein each query generates the match set.

5. The system of claim 1 , wherein the processor is configured to use a Bayesian identity resolution process and uses an ElasticSearch process.

6. The system of claim 1 , wherein the entity is one of a healthcare provider, subject, an idea, a professional, a person, a corporation and a business entity.

7. The system of claim 1 , wherein the entity is a healthcare provider and the one or more raw files are a Centers for Medicare and Medicaid Services' (CMS) National Plan and Provider Enumeration System (NPPES) data file and an American Medical Association (AMA) file.

8. A method for matching and merging data for an entity from disparate data sources, comprising:

receiving, by a computer system, one or more raw files from disparate sources about an entity, each raw file having a plurality of data fields and each field has a value associated with the entity;

performing, by the computer system, each of a plurality of matching processes against each of the one or more raw files about the entity to generate a match set for each raw file matched with one of the plurality of matching processes based on results from each of the plurality of matching processes with each of the raw files wherein the match set for each raw file matched by each one of the plurality of matching processes a reference to a document storage location of the raw file, a unique identifier for the raw file and provenance metadata for the one matching process including an identifier of the matching process and parameters associated with the matching process;

merging, by the computer system, the one or more raw files represented into a merged document that has the at least one matched data field in each of the one or more raw files; and

identifying, by the computer system, conflicting values in each data field of the merged document.

9. The method of claim 8 further comprising transforming, by the computer system before performing the plurality of matching processes, the one or more raw files to initial data cleanse the one or more raw files and generate one or more transformed files about the entity.

10. The method of claim 8 , wherein identifying the conflicting values further comprises ranking each conflicting value for a data field based on a confidence of a correctness of each conflicting value.

11. The method of claim 8 , wherein performing the plurality of matching processes further comprises performing a sequence of queries on the one or more raw files, wherein each query generates the match set.

12. The method of claim 11 , wherein the match set for a particular raw file has a reference to a storage location of the particular raw file and a unique identifier.

13. The method of claim 8 , wherein performing the plurality of matching processes further comprises using a Bayesian identity resolution process and using an ElasticSearch process.

14. The method of claim 8 , wherein the entity is one of a healthcare provider, subject, an idea, a professional, a person, a corporation and a business entity.

15. The method of claim 8 , wherein the entity is a healthcare provider and the one or more raw files are a Centers for Medicare and Medicaid Services' (CMS) National Plan and Provider Enumeration System (NPPES) data file and an American Medical Association (AMA) file.

16. A system for matching and merging data for an entity from disparate data sources, comprising:

one or more data sources for disparate sources, each data source having a raw file about an entity, each raw file having a plurality of data fields and each field has a value associated with the entity;

a computer system having a processor, the computer system coupled to each of the one or more data sources;

the processor of the computer system being configured to perform each of a plurality of matching processes against each of the one or more raw files about the entity to generate a match set for each raw file matched with one of the plurality of matching processes wherein the match set for each raw file matched by each one of the plurality of matching processes a reference to a document storage location of the raw file, a unique identifier for the raw file and provenance metadata for the one matching process including an identifier of the matching process and parameters associated with the matching process, the plurality of matching processes including a strict matcher process that utilizes statistically significant combinations of biographic identifiers including each component of a person's full name, a birth date and a birth place and a loose matcher process that allows for a variation in a person's name;

the processor of the computer system being configured to merge the one or more raw files represented into a merged document that has the at least one matched data field in each of the one or more raw files; and

the processor of the computer system being configured to identify conflicting values in each data field of the merged document.

17. A method for matching and merging data for an entity from disparate data sources, comprising:

receiving, by a computer system, one or more raw files from disparate sources about an entity, each raw file having a plurality of data fields and each field has a value associated with the entity;

performing, by the computer system, a plurality of matching processes against the one or more raw files about the entity to generate a match set for each raw file matched with one of the plurality of matching processes wherein the match set for each raw file matched by each one of the plurality of matching processes a reference to a document storage location of the raw file, a unique identifier for the raw file and provenance metadata for the one matching process including an identifier of the matching process and parameters associated with the matching process, the plurality of matching processes including a strict matcher process that utilizes statistically significant combinations of biographic identifiers including each component of a person's full name, a birth date and a birth place and a loose matcher process that allows for a variation in a person's name;

merging, by the computer system, the one or more raw files represented into a merged document that has the at least one matched data field in each of the one or more raw files; and

identifying, by the computer system, conflicting values in each data field of the merged document.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Oct 5, 2022
From: BANK OF AMERICA, N.A.
To: CHANGE HEALTHCARE HOLDINGS, LLC
Reel/Frame 061620/0032 →
SECURITY INTEREST Recorded Dec 12, 2019
From: CHANGE HEALTHCARE HOLDINGS, LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 051279/0614 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2019
From: POKITDOK, INC.
To: CHANGE HEALTHCARE HOLDINGS, LLC
Reel/Frame 048195/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2014
From: ALDRIDGE, MATTHEW LEE; HOEKMAN, JEFFREY SCOTT; ALSTAD, COLIN ERIK; TANNER, THEODORE CALHOUN, JR.
To: POKITDOK, INC.
Reel/Frame 033611/0279 →
Cited By (1)
US 12,633,400