IP Library Granted Patent US 11,294,938
Granted Patent B2
US 11,294,938 · App. 16/238,974 · Granted Apr 5, 2022

Generalized distributed framework for parallel search and retrieval of unstructured and structured patient data across zones with hierarchical ranking

Inventors: David Beymer (San Jose, CA); Tanveer Syeda-Mahmood (San Jose, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/288G06F7/14G06F16/2246G06F16/248G16H10/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,294,938
App. No.
16/238,974
Granted
Apr 5, 2022
Kind
B2
Abstract

A generalized distributed framework is provided for parallel search and retrieval of unstructured and structured patient data across zones with hierarchical ranking. In various embodiments, patient data is ingested from a plurality of data sources. A plurality of data models is populated based on the ingested patient data, each data model comprising an abstract data type. The plurality of data models is stored in an index. A search request is processed against the index, the search request comprising one or more attribute of the abstract data type.

Claims (36)

1. A method comprising:

ingesting a collection of structured and unstructured patient data, the patient data comprising a plurality of modalities, the patient data comprising a longitudinal clinical history of one or more patients, the patient data from a plurality of data sources;

populating a plurality of data models based on the ingested patient data to provide for parallel search, retrieval, and hierarchical ranking of the patient data from the plurality of data sources, each data model comprising an abstract data type having a primary key configured to uniquely designate a row or a document, a search key configured to separately designate a row or a document, a grouping key configured to count number of matches to a query, and an untokenized key configured to preserve one or more predetermined fields of the document definition without tokenization, wherein the plurality of data models comprise one or more domain-specific subclasses;

storing the plurality of data models as one or more indexed documents in an index, each data model being associated with one or more schemas, each data model representing patient data according to one or more data types within the respective patient data, and each data model conforming to at least one of the one or more schemas;

processing a search request against the index, the search request comprising one or more attribute of the abstract data type, wherein the search request is automatically expanded to retrieve one or more of the plurality of data models based on one or more additional attributes associated with the one or more attribute; and

providing a search result by collecting one or more shards based on the search request.

2. The method of claim 1 , further comprising outputting a result of the search request.

3. The method of claim 2 , wherein the output of the search request is sorted based on quality of match.

4. The method of claim 1 , wherein populating the plurality of data models comprises generating and merging a data model for ingested patient data.

5. The method of claim 1 , further comprising:

outputting the data models in a time series.

6. The method of claim 1 , the abstract data type reflecting a class hierarchy.

7. A system comprising:

a computing node comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor of the computing node to cause the processor to perform a method comprising:

ingesting a collection of structured and unstructured patient data, the patient data comprising a plurality of modalities, the patient data comprising a longitudinal clinical history of one or more patients, the patient data from a plurality of data sources;

populating a plurality of data models based on the ingested patient data to provide for parallel search, retrieval, and hierarchical ranking of the patient data from the plurality of data sources, each data model comprising an abstract data type having a primary key configured to uniquely designate a row or a document, a search key configured to separately designate a row or a document, a grouping key configured to count number of matches to a query, and an untokenized key configured to preserve one or more predetermined fields of the document definition without tokenization, wherein the plurality of data models comprise one or more domain-specific subclasses;

storing the plurality of data models as one or more indexed documents in an index, each data model being associated with one or more schemas, each data model representing patient data according to one or more data types within the respective patient data, and each data model conforming to at least one of the one or more schemas;

processing a search request against the index, the search request comprising one or more attribute of the abstract data type, wherein the search request is automatically expanded to retrieve one or more of the plurality of data models based on one or more additional attributes associated with the one or more attribute; and

providing a search result by collecting one or more shards based on the search request.

8. The system of claim 7 , further comprising outputting a result of the search request.

9. The system of claim 8 , wherein the output of the search request is sorted based on quality of match.

10. The system of claim 7 , wherein populating the plurality of data models comprises generating and merging a data model for ingested patient data.

11. The system of claim 7 , further comprising:

outputting the data models in a time series.

12. The system of claim 7 , the abstract data type reflecting a class hierarchy.

13. A computer program product for modeling patient data, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform a method comprising:

ingesting a collection of structured and unstructured patient data, the patient data comprising a plurality of modalities, the patient data comprising a longitudinal clinical history of one or more patients, the patient data from a plurality of data sources;

populating a plurality of data models based on the ingested patient data to provide for parallel search, retrieval, and hierarchical ranking of the patient data from the plurality of data sources, each data model comprising an abstract data type having a primary key configured to uniquely designate a row or a document, a search key configured to separately designate a row or a document, a grouping key configured to count number of matches to a query, and an untokenized key configured to preserve one or more predetermined fields of the document definition without tokenization, wherein the plurality of data models comprise one or more domain-specific subclasses;

storing the plurality of data models as one or more indexed documents in an index, each data model being associated with one or more schemas, each data model representing patient data according to one or more data types within the respective patient data, and each data model conforming to at least one of the one or more schemas;

processing a search request against the index, the search request comprising one or more attribute of the abstract data type, wherein the search request is automatically expanded to retrieve one or more of the plurality of data models based on one or more additional attributes associated with the one or more attribute; and

providing a search result by collecting one or more shards based on the search request.

14. The computer program product of claim 13 , further comprising outputting a result of the search request.

15. The computer program product of claim 13 , wherein populating the plurality of data models comprises generating and merging a data model for ingested patient data.

16. The computer program product of claim 13 , further comprising:

outputting the data models in a time series.

17. The computer program product of claim 13 , the abstract data type reflecting a class hierarchy.

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2019
From: BEYMER, DAVID; SYEDA-MAHMOOD, TANVEER
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047905/0325 →