IP Library Granted Patent US 10,025,828
Granted Patent B2
US 10,025,828 · App. 14/543,414 · Granted Jul 17, 2018

Method and system for generating a unified database from data sets

Inventors: Glen de Vries (New York, NY); Michelle Marlborough (Brooklyn, NY)
Assignee: Medidata Solutions, Inc.
G06F17/30528G06F17/30091G06F17/30569G16H10/20G16H10/40G16H10/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,025,828
App. No.
14/543,414
Granted
Jul 17, 2018
Kind
B2
Abstract

A method for generating a unified database includes receiving a structured set of data, where each set is made up of records having fields, aggregating values within a first field of the records, automatically applying a set of rules to the first field values to determine correlations among the first field values, calculating a confidence level regarding a label for the first field, providing the label to the first field, storing the first field values in the first field in the unified database, and receiving more information to increase the confidence level. A system for generating a clinical database and a method for using the database are also described.

Claims (74)

1. A computer-implemented method for generating a unified clinical database, comprising:

receiving a structured set of clinical data, the set comprising one or more records, each record having two or more fields;

aggregating values taken from a first field of the records;

aggregating values taken from a second field of the records;

automatically applying, by a processor, a set of rules to the aggregated first and second field values to calculate statistical correlations between the aggregated first and second field values;

calculating a confidence level regarding a label for the first field based on the rules;

if the confidence level meets or exceeds a pre-determined threshold,

applying the label to the first field;

storing the first field values in a first unified field in the unified database; and

if the confidence level does not meet or exceed the pre-determined threshold, receiving information regarding the correlations between the aggregated first and second field values and recalculating the confidence level.

2. The method of claim 1 , further comprising:

calculating a second confidence level regarding a label for the second field;

if the second confidence level meets or exceeds a pre-determined threshold,

applying the label to the second field; and

storing the second field values in a second unified field in the unified database; and

if the second confidence level does not meet or exceed the pre-determined threshold, receiving information regarding the correlations between the aggregated first and second field values and recalculating the second confidence level.

3. The method of claim 1 , further comprising:

calculating statistical distributions of the aggregated first and second field values; and

comparing the statistical distributions of the aggregated first and second field values with statistical distributions of stored data measures.

4. The method of claim 3 , wherein determining labels for the aggregated first and second field values comprises determining closest data measures having statistical distributions substantially the same as the statistical distributions of the aggregated first and second field values.

5. The method of claim 1 , wherein applying the set of rules to the aggregated first and second field values comprises determining a structure of the records.

6. The method of claim 5 , wherein the structure of the records of a CHEM-7 test comprises seven fields.

7. A computer-implemented method for generating a unified clinical database, comprising:

receiving a structured set of clinical data, the set comprising one or more records, each record having one or more fields;

aggregating values taken from a first field of the records;

automatically applying, by a processor, one or more rules to the aggregated first field values,

wherein applying said rules includes comparing the aggregated first field values to:

stored statistics of known clinical measures, if the aggregated first field values comprise numerical values;

known text entries from a clinical dictionary, if the aggregated first field values comprise alphabetical and/or alphanumeric values;

stored calendar information, if the aggregated first field values comprise date values; and

stored information concerning a clinical trial, including trial design information, if the aggregated first field values comprise alphabetical and/or alphanumeric values;

calculating a confidence level regarding a label for the first field;

if the confidence level meets or exceeds a pre-determined threshold,

applying the label to the first field; and

storing the first field values in a first unified field in the unified database; and

if the confidence level does not meet or exceed the pre-determined threshold, receiving more information and recalculating the confidence level.

8. The method of claim 7 , wherein the rules include eligibility criteria comprising inclusion criteria, exclusion criteria, or both.

9. The method of claim 7 , wherein the rules are refined based on the received data.

10. The method of claim 7 , wherein the stored statistics comprise statistical distributions of known clinical measures.

11. The method of claim 7 , wherein if the aggregated first field values comprise date values, and the date values are after the beginning of a clinical trial, then the first field label refers to testing during the clinical trial.

12. The method of claim 7 , wherein if the aggregated first field values comprise date values, and the date values are before the beginning of a clinical trial, then the first field label refers to historic data.

13. The method of claim 7 , further comprising:

aggregating values taken from a second field of the records;

automatically applying the one or more rules to the aggregated second field values to determine correlations between the aggregated first and second field values; and

determining a label for the second field values.

14. The method of claim 13 , wherein said correlations increase the confidence level regarding the label for the first field values.

15. A computer-implemented method for generating a unified database, comprising:

receiving a structured set of clinical data, the set comprising one or more records, each record having one or more fields;

aggregating values within a first field of the records;

automatically applying, by a processor, a set of rules to the first field values to calculate statistical correlations among the first field values, wherein applying the set of rules includes comparing the aggregated first field values to stored statistics of known data measures, if the aggregated first field values comprise numerical values;

calculating a confidence level regarding a label for the first field;

if the confidence level meets or exceeds a pre-determined threshold,

applying the label to the first field;

storing the first field values in the first field in the unified database; and

receiving more information to increase the confidence level; and

if the confidence level does not meet or exceed the pre-determined threshold, receiving more information and recalculating the confidence level.

16. The method of claim 15 , wherein applying the set of rules to the first field values comprises using inclusion and exclusion criteria.

17. The method of claim 15 , further comprising:

calculating a statistical distribution of the first field values; and

comparing the statistical distribution of the first field values with statistical distributions of stored data measures.

18. The method of claim 17 , further comprising:

determining a closest data measure having a statistical distribution substantially the same as the statistical distribution of the first field values.

19. The method of claim 17 , further comprising calculating the statistical distributions of the stored data measures.

20. The method of claim 15 , wherein applying the set of rules includes comparing the first field values to known text entries from a clinical dictionary, if the aggregated first field values comprise alphabetical and/or alphanumeric values.

21. The method of claim 15 , further comprising:

aggregating values within a second field of the records;

automatically applying the set of rules to the second field values to determine correlations among the second field values and first field values; and

determining a label for the second field values.

22. The method of claim 21 , wherein said determined correlations increase the confidence level regarding the label for the first field values.

23. The method of claim 21 , further comprising:

calculating a statistical distribution of the second field values; and

comparing the statistical distribution of the second field values with statistical distributions of stored data measures.

24. The method of claim 23 , further comprising:

determining a closest data measure having a statistical distribution substantially the same as the statistical distribution of the second field values.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Oct 30, 2019
From: HSBC BANK USA
To: MEDIDATA SOLUTIONS, INC.; CHITA INC.
Reel/Frame 050875/0776 →
SECURITY INTEREST Recorded Jan 2, 2018
From: MEDIDATA SOLUTIONS, INC.
To: HSBC BANK USA, NATIONAL ASSOCIATION
Reel/Frame 044979/0571 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2015
From: DE VRIES, GLEN, MR.; MARLBOROUGH, MICHELLE, MS.
To: MEDIDATA SOLUTIONS, INC.
Reel/Frame 034757/0799 →
Continuity (3)
Continuation 14450197 · Aug 1, 2014
Continuation 13974294 · Aug 23, 2013
Related Publication 20150074133A1 · Mar 12, 2015