IP Library Granted Patent US 7,707,169
Granted Patent B2
US 7,707,169 · App. 11/143,317 · Granted Apr 27, 2010

Specification-based automation methods for medical content extraction, data aggregation and enrichment

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,707,169
App. No.
11/143,317
Granted
Apr 27, 2010
Kind
B2
Abstract

A method for knowledge generation from raw medical records uses XML-based specifications. The method includes content extraction, data aggregation and data enrichment. The method operates on various sources of medical data including financial data, clinical documents and medical images.

Claims (56)

1. A method for structuring diverse medical data as a database comprising:

extracting content from the diverse medical data comprising:

generating Extensible Markup Language (XML)-based extraction rules for each type of medical data based on predetermined information from the medical data document type;

converting each medical data document type into one or more standard document types, wherein each standard document type includes text and image object relevant content;

extracting the relevant content from each standard document type using the XML-based extraction rules associated with the medical data document type as an XML syntax;

indexing the extracted relevant content; and

storing the indexed relevant content;

aggregating the stored indexed relevant content comprising:

creating data migration language extraction rules;

extracting the stored indexed relevant content using the data extraction rules;

creating data migration language constraint differential rules;

normalizing the extracted indexed relevant content using the constraint differential rules;

storing the normalized relevant content;

creating data migration language record/rectify merge rules;

consolidating the normalized relevant content using the record/rectify merge, rules; and

storing the consolidated relevant content;

enriching the stored consolidated relevant content comprising:

creating one or more XML templates; and

sequentially processing the consolidated relevant content using the one or more XML templates into models.

2. The method according to claim 1 wherein the standard document types include pdf and Digital Imaging and Communications in Medicine (DICOM) standards.

3. The method according to claim 1 wherein the extraction rules specify relevant content location, string patterns, and contexts where to extract the text and/or image object content.

4. The method according to claim 1 wherein aggregating further comprises mapping data from each standard format to provide for data extraction and data loading.

5. The method according to claim 1 wherein the data extraction rules are created using Data Migration Specification Language (DMSL).

6. The method according to claim 1 wherein aggregating further comprises performing a functional dependence analysis on the indexed relevant content.

7. The method according to claim 1 wherein aggregating further comprises performing an integrity constraint comparison on the indexed relevant content.

8. The method according to claim 1 wherein aggregating further comprises performing a record linkage analysis on the indexed relevant content.

9. The method according to claim 1 wherein the data migration language uses XML tags.

10. The method according to claim 1 wherein the models include associated models, functional models, structural models and behavioral models.

11. A computer program product comprising a computer readable recording medium having recorded thereon a computer program comprising code means for, when executed on a computer, instructing the computer to control steps in a method for structuring diverse medical data as a database comprising:

extracting content from the diverse medical data comprising:

generating Extensible Markup Language (XML)-based extraction rules for each type of medical data based on predetermined information from the medical data document type;

converting each medical data document type into one or more standard document types, wherein each standard document type includes text and image object relevant content;

extracting the relevant content from each standard document type using the XML-based extraction rules associated with the medical data document type as an XML syntax;

indexing the extracted relevant content; and

storing the indexed relevant content;

aggregating the stored indexed relevant content comprising:

creating data migration language extraction rules;

extracting the stored indexed relevant content using the data extraction rules;

creating data migration language constraint differential rules;

normalizing the extracted indexed relevant content using the constraint differential rules;

storing the normalized relevant content;

creating data migration language record/rectify merge rules;

consolidating the normalized relevant content using the record/rectify merge rules; and

storing the consolidated relevant content;

enriching the stored consolidated relevant content comprising:

creating one or more XML templates; and

sequentially processing the consolidated relevant content using the one or more XML templates into models.

12. The computer program product according to claim 11 wherein the standard document types include pdf and Digital Imaging and Communications in Medicine (DICOM) standards.

13. The computer program product according to claim 11 wherein the extraction rules specify relevant content location, string patterns, and contexts where to extract the text and/or image object content.

14. The computer program product according to claim 11 wherein aggregating further comprises mapping data from each standard format to provide for data extraction and data loading.

15. The computer program product according to claim 11 wherein the data extraction rules are created using Data Migration Specification Language (DMSL).

16. The computer program product according to claim 11 wherein aggregating further comprises performing a functional dependence analysis on the indexed relevant content.

17. The computer program product according to claim 11 wherein aggregating further comprises performing an integrity constraint comparison on the indexed relevant content.

18. The computer program product according to claim 11 wherein aggregating further comprises performing a record linkage analysis on the indexed relevant content.

19. The computer program product according to claim 11 wherein the data migration language uses XML tags.

20. The computer program product according to claim 11 wherein the models include associated models, functional models, structural models and behavioral models.

Assignments (2)
MERGER Recorded Apr 12, 2010
From: SIEMENS CORPORATE RESEARCH, INC.
To: SIEMENS CORPORATION
Reel/Frame 024216/0434 →
MERGER Recorded Mar 8, 2010
From: SIEMENS CORPORATE RESEARCH, INC.
To: SIEMENS CORPORATION
Reel/Frame 024042/0242 →