IP Library Granted Patent US 8,799,331
Granted Patent B1
US 8,799,331 · App. 13/974,294 · Granted Aug 5, 2014

Generating a unified database from data sets

Inventors: Glen de Vries (New York, NY); Michelle Marlborough (Brooklyn, NY)
Assignee: Medidata Solutions, Inc.
G06F17/30979
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,799,331
App. No.
13/974,294
Granted
Aug 5, 2014
Kind
B1
Abstract

A method for generating a unified database includes receiving a structured set of data, where each set is made up of records having fields, aggregating values within a first field of the records, automatically applying a set of rules to the first field values to determine correlations among the first field values, calculating a confidence level regarding a label for the first field, providing the label to the first field, storing the first field values in the first field in the unified database, and receiving more information to increase the confidence level. A system for generating a clinical database and a method for using the database are also described.

Claims (67)

1. A computer-implemented method for generating a unified clinical database, comprising:

receiving a structured set of clinical data, the set comprising one or more records, each record having one or more fields;

aggregating values within a first field of the records;

automatically applying, by a processor, a set of rules to the first field values to determine correlations among the first field values;

wherein applying the set of rules includes comparing the first field values to known text entries from a clinical dictionary;

calculating a confidence level regarding a label for the first field;

if the confidence level meets or exceeds a pre-determined threshold,

providing the label to the first field;

storing the first field values in the first field in the unified database; and

receiving more information to increase the confidence level; and

if the confidence level does not exceed the pre-determined threshold, receiving more information and recalculating the confidence level.

2. The method of claim 1 , wherein applying the set of rules to the first field values comprises using inclusion and exclusion criteria.

3. The method of claim 1 , further comprising:

calculating a statistical distribution of the first field values; and

comparing the statistical distribution of the first field values with statistical distributions of stored data measures.

4. The method of claim 3 , further comprising:

determining a closest data measure having a statistical distribution substantially the same as the statistical distribution of the first field values.

5. The method of claim 3 , further comprising calculating the statistical distributions of the stored data measures.

6. The method of claim 1 , further comprising:

aggregating values within a second field of the records;

automatically applying the set of rules to the second field values to determine correlations among the second field values and first field values; and

determining a label for the second field values.

7. The method of claim 6 , wherein said determined correlations increase the confidence level regarding the label for the first field values.

8. The method of claim 6 , further comprising:

calculating a statistical distribution of the second field values; and

comparing the statistical distribution of the second field values with statistical distributions of stored data measures.

9. The method of claim 8 , further comprising:

determining a closest data measure having a statistical distribution substantially the same as the statistical distribution of the second field values.

10. A system for generating a clinical database, comprising:

a processor;

a database generator; and

a rules generator for generating rules to be used in the database generator to determine correlations among data input to the database generator,

wherein the database generator:

receives a structured set of clinical data, each set made up of records having one or more fields;

aggregates values within a first field of the records;

automatically applies the rules to the first field values, which includes comparing the first field values to known text entries from a clinical dictionary;

calculates a confidence level regarding a label for the first field;

if the confidence level meets or exceeds a pre-determined threshold,

provides the label to the first field;

stores the first field values in the first field in the clinical database; and

receives more information to increase the confidence level; and

if the confidence level does not exceed the pre-determined threshold, receives more information and recalculates the confidence level.

11. The system of claim 10 , wherein the rules are refined based on the received data.

12. The system of claim 10 , wherein the rules comprise inclusion and exclusion criteria.

13. The system of claim 10 , wherein the rules comprise statistical distributions of known clinical measures.

14. The system of claim 10 , wherein if the records in each set comprise a second field, the database generator:

aggregates values within the second field of the records;

applies the rules to the second field values to determine correlations among the second field values and first field values; and

determines a label for the second field values.

15. A computer-implemented method for generating and using a unified clinical database, comprising:

automatically generating the unified clinical database by:

receiving a structured set of clinical data, the set comprising one or more records, each record having one or more fields;

aggregating values within a first field of the records;

automatically applying, by a processor, a set of rules to the first field values, the rules comprising inclusion and exclusion criteria from at least one prior clinical trial;

wherein applying the set of rules includes comparing the first field values to known text entries from a clinical dictionary;

designing a placebo arm of a clinical trial using the inclusion and exclusion criteria and the clinical data in the unified clinical database; and

calculating a confidence level regarding a label for the first field;

if the confidence level meets or exceeds a pre-determined threshold,

providing the label to the first field;

storing the first field values in the first field in the unified database; and

receiving more information to increase the confidence level; and

if the confidence level does not exceed the pre-determined threshold, receiving more information and recalculating the confidence level.

16. The method of claim 15 , wherein said inclusion and exclusion criteria come from protocol databases.

17. The method of claim 15 , further comprising:

aggregating values within a second field of the records; and

automatically applying the inclusion and exclusion criteria to the second field values.

18. The method of claim 15 , wherein the rules further determine correlations among the first field values to determine a label for the first field values.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Oct 30, 2019
From: HSBC BANK USA
To: MEDIDATA SOLUTIONS, INC.; CHITA INC.
Reel/Frame 050875/0776 →
SECURITY INTEREST Recorded Jan 2, 2018
From: MEDIDATA SOLUTIONS, INC.
To: HSBC BANK USA, NATIONAL ASSOCIATION
Reel/Frame 044979/0571 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2013
From: DE VRIES, GLEN; MARLBOROUGH, MICHELLE
To: MEDIDATA SOLUTIONS, INC.
Reel/Frame 031120/0120 →