IP Library Granted Patent US 11,133,088
Granted Patent B2
US 11,133,088 · App. 15/355,186 · Granted Sep 28, 2021

Resolving conflicting data among data objects associated with a common entity

Inventors: Peter A. Fellowes (Cleveland, OH); Jacob O. Miller (Cleveland, OH); Matthew M. Pohlman (University Heights, OH)
Assignee: International Business Machines Corporation
G16H10/60G16H40/67G16H50/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,133,088
App. No.
15/355,186
Granted
Sep 28, 2021
Kind
B2
Abstract

Resolving conflicting data among data objects associated with a common entity includes assigning an ordered sequence of data analysis processes to a corresponding data field of a plurality of data objects associated with a common entity. At least two of the data objects include different values for the corresponding data field, and each data analysis process performs a different technique to resolve conflicts between different values of data. The ordered sequence of data analysis processes is executed to determine a consensus value to serve as a value for the corresponding data field of each of the plurality of data objects. The data analysis processes are successively executed in the ordered sequence until the consensus value is determined for the corresponding data field.

Claims (33)

1. A computer-implemented method of resolving conflicting data among data objects associated with a common entity comprising:

assigning, via at least one processor, an ordered sequence of a plurality of data analysis processes executable by the at least one processor to a corresponding data field of a plurality of data objects associated with a common entity in a repository, wherein the data objects include electronic records and the repository stores the electronic records, wherein the corresponding data field is the same data field across the plurality of data objects and at least two of the data objects include different values for the corresponding data field, wherein each data analysis process performs a different technique on the corresponding data field of the plurality of data objects to resolve conflicts between the different values for the corresponding data field, and wherein the plurality of data objects each includes a plurality of data fields and at least two of the plurality of data fields are assigned a different ordered sequence of the data analysis processes;

calling and executing, via the at least one processor, the ordered sequence of data analysis processes on the corresponding data field of the plurality of data objects to determine a consensus value from the corresponding data field of the plurality of data objects to serve as a value for the corresponding data field of each of the plurality of data objects, wherein two or more of the data analysis processes are successively executed in the ordered sequence by executing a successive data analysis process after a prior data analysis process fails to determine the consensus value for the corresponding data field, and wherein the ordered sequence includes a field specific data analysis process that determines the consensus value based on statistical information of occurrences of a combination of the corresponding data field and at least one other data field of the plurality of data objects in data from external data sources;

generating and storing, via the at least one processor, a new data object in the repository for the common entity including the consensus value for the corresponding data field of each of the plurality of data objects and an indicator to distinguish the new data object from the plurality of data objects as containing the consensus value; and

processing, via the at least one processor, a query for the repository including data objects of entities by identifying the new data object based on the indicator and processing the query against the new data object including the consensus value instead of the plurality of data objects including different values to produce search results.

2. The method of claim 1 , wherein one or more of the data analysis processes apply probabilistic estimates to determine the consensus value.

3. The method of claim 2 , wherein the probabilistic estimates are based on frequency of co-occurrences of values of the corresponding data field across the plurality of data objects.

4. The method of claim 1 , wherein the plurality of data objects includes demographic records comprising the corresponding data field, and the common entity includes a patient.

5. The method of claim 4 , wherein one of the data analysis processes uses clinical information of the patient to determine a likelihood of a preference for one of the different values for the corresponding data field to be the consensus value.

6. The method of claim 1 , further comprising:

storing information pertaining to the consensus value, the data analysis process of the ordered sequence determining the consensus value, and the data object containing the consensus value as a value for the corresponding data field.

7. A system for resolving conflicting data among data objects associated with a common entity comprising:

at least one processor configured to:

assign an ordered sequence of a plurality of data analysis processes executable by the at least one processor to a corresponding data field of a plurality of data objects associated with a common entity in a repository, wherein the data objects include electronic records and the repository stores the electronic records, wherein the corresponding data field is the same data field across the plurality of data objects and at least two of the data objects include different values for the corresponding data field, wherein each data analysis process performs a different technique on the corresponding data field of the plurality of data objects to resolve conflicts between the different values for the corresponding data field, and wherein the plurality of data objects each includes a plurality of data fields and at least two of the plurality of data fields are assigned a different ordered sequence of the data analysis processes;

call and execute the ordered sequence of data analysis processes on the corresponding data field of the plurality of data objects to determine a consensus value from the corresponding data field of the plurality of data objects to serve as a value for the corresponding data field of each of the plurality of data objects, wherein two or more of the data analysis processes are successively executed in the ordered sequence by executing a successive data analysis process after a prior data analysis process fails to determine the consensus value for the corresponding data field, and wherein the ordered sequence includes a field specific data analysis process that determines the consensus value based on statistical information of occurrences of a combination of the corresponding data field and at least one other data field of the plurality of data objects in data from external data sources;

generate and store a new data object in the repository for the common entity including the consensus value for the corresponding data field of each of the plurality of data objects and an indicator to distinguish the new data object from the plurality of data objects as containing the consensus value; and

process a query for the repository including data objects of entities by identifying the new data object based on the indicator and processing the query against the new data object including the consensus value instead of the plurality of data objects including different values to produce search results.

8. The system of claim 7 , wherein one or more of the data analysis processes apply probabilistic estimates to determine the consensus value.

9. The system of claim 8 , wherein the probabilistic estimates are based on frequency of co-occurrences of values of the corresponding data field across the plurality of data objects.

10. The system of claim 7 , wherein the plurality of data objects includes demographic records comprising the corresponding data field, and the common entity includes a patient.

11. The system of claim 10 , wherein one of the data analysis processes uses clinical information of the patient to determine a likelihood of a preference for one of the different values for the corresponding data field to be the consensus value.

12. The system of claim 7 , wherein the at least one processor is further configured to:

store information pertaining to the consensus value, the data analysis process of the ordered sequence determining the consensus value, and the data object containing the consensus value as a value for the corresponding data field.

13. A computer program product for resolving conflicting data among data objects associated with a common entity, the computer program product comprising one or more computer readable storage media collectively having program instructions embodied therewith, the program instructions executable by at least one processor to cause the at least one processor to:

assign an ordered sequence of a plurality of data analysis processes executable by the at least one processor to a corresponding data field of a plurality of data objects associated with a common entity in a repository, wherein the data objects include electronic records and the repository stores the electronic records, wherein the corresponding data field is the same data field across the plurality of data objects and at least two of the data objects include different values for the corresponding data field, wherein each data analysis process performs a different technique on the corresponding data field of the plurality of data objects to resolve conflicts between the different values for the corresponding data field, and wherein the plurality of data objects each includes a plurality of data fields and at least two of the plurality of data fields are assigned a different ordered sequence of the data analysis processes;

call and execute the ordered sequence of data analysis processes on the corresponding data field of the plurality of data objects to determine a consensus value from the corresponding data field of the plurality of data objects to serve as a value for the corresponding data field of each of the plurality of data objects, wherein two or more of the data analysis processes are successively executed in the ordered sequence by executing a successive data analysis process after a prior data analysis process fails to determine the consensus value for the corresponding data field, and wherein the ordered sequence includes a field specific data analysis process that determines the consensus value based on statistical information of occurrences of a combination of the corresponding data field and at least one other data field of the plurality of data objects in data from external data sources;

generate and store a new data object in the repository for the common entity including the consensus value for the corresponding data field of each of the plurality of data objects and an indicator to distinguish the new data object from the plurality of data objects as containing the consensus value; and

process a query for the repository including data objects of entities by identifying the new data object based on the indicator and processing the query against the new data object including the consensus value instead of the plurality of data objects including different values to produce search results.

14. The computer program product of claim 13 , wherein one or more of the data analysis processes apply probabilistic estimates based on frequency of co-occurrences of values of the corresponding data field across the plurality of data objects to determine the consensus value.

15. The computer program product of claim 13 , wherein the plurality of data objects includes demographic records comprising the corresponding data field, and the common entity includes a patient.

16. The computer program product of claim 15 , wherein one of the data analysis processes uses clinical information of the patient to determine a likelihood of a preference for one of the different values for the corresponding data field to be the consensus value.

17. The computer program product of claim 13 , further comprising program instructions executable by the at least one processor to cause the at least one processor to:

store information pertaining to the consensus value, the data analysis process of the ordered sequence determining the consensus value, and the data object containing the consensus value as a value for the corresponding data field.

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2016
From: FELLOWES, PETER A.; MILLER, JACOB O.; POHLMAN, MATTHEW M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 040366/0489 →