IP Library › Granted Patent US 8,392,126
Granted Patent B2
US 8,392,126 · App. 12/565,341 · Granted Mar 5, 2013

Method and system for determining the accuracy of DNA base identifications

Inventor: Tobias Mann (Carlsbad, CA)
Assignee: Illumina, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,392,126
App. No.
12/565,341
Granted
Mar 5, 2013
Kind
B2
Abstract

A method for determining the quality of predicted nucleotide base identifications by receiving training data sets of predicted base identifications; defining subsets within the training data sets; comparing the predicted base identifications with actual base identifications within each subset; determining one or more sampling characteristics for each subset; and determining quality characterizations based on the comparison and the determined sampling characteristics.

Claims (81)

1. A method in a computer for determining the quality of predicted DNA base identifications, the method comprising:

receiving a training data set, the training data set comprising a plurality of predicted DNA base identifications;

defining a group of subsets;

comparing the predicted DNA base identifications with actual DNA base identifications for training data within each subset of the group;

determining a sampling characteristic for each subset of the group based on training data within the respective subset; and

determining a quality characterization for predicted DNA base identifications within at least one of subset of the group based on the comparison and determined sampling characteristics;

wherein the sampling characteristic comprises a confidence value comprising a binomial proportion confidence interval value.

2. The method of claim 1 , further comprising receiving parameter values associated with the training data set.

3. The method of claim 2 , wherein defining the group of subsets is based on the parameter values.

4. The method of claim 2 , wherein defining the group of subsets comprises partitioning the parameter values into a plurality of bins.

5. The method of claim 2 , further comprising

determining the predicted DNA base identifications based on the received parameter values.

6. The method of claim 2 , wherein the parameter values comprise at least one of

a first peak height ratio for a current peak based on a first plurality of called peaks centered at a current peak;

a second peak height ratio for the current peak based on a second plurality of called peaks centered at the current peak;

a peak spacing ratio for the current peak based on a largest peak spacing; and

a smallest peak spacing of the second plurality of called peaks centered at the current peak and a peak resolution.

7. The method of claim 1 , further comprising determining an accuracy of the base prediction based on the determined quality characterization.

8. The method of claim 7 , wherein determining the accuracy of the base prediction comprises referencing a look-up table.

9. The method of claim 1 , wherein determining the quality characterization comprises generating an accuracy prediction look-up table.

10. The method of claim 9 , wherein generating the look-up table comprises:

defining the plurality of subsets by partitioning parameter values associated with data of the training data set into a plurality of bins;

populating the plurality of bins with the predicted DNA base identifications;

iteratively performing the following process:

computing a quality characterization for each subset of the group;

selecting an extreme quality characteristic subset as the considered subset having the largest or smallest quality characteristic of the group;

storing the largest or smallest quality characteristic and corresponding threshold values in the look-up table; and

adjusting the quality characteristic for the group of considered subsets such that the quality characteristic no longer depends on data within the extreme quality characteristic subset.

11. The method of claim 10 , wherein the process further comprises

deleting the extreme quality characteristic subset from the group.

12. The method of claim 10 , wherein the process is iteratively performed until all of the data from the training data set has been within at least one extreme characteristic subset.

13. The method of claim 1 , further comprising:

determining a plurality of quality characterizations, each quality characterization being associated with at least one threshold parameter value; and

storing the quality characterizations and the corresponding threshold parameter values in a look-up table.

14. The method of claim 13 , further comprising:

receiving at least one parameter value associated with data of a non-training data set; and

selecting a quality characterization from the look-up table, the selected quality characterization being that the at least one corresponding threshold parameter values exceed the at least one parameter value associated with data of the non-training data set.

15. The method of claim 1 , wherein the comparing comprises calculating an error value E.

16. The method of claim 1 , wherein the binomial proportion confidence interval value comprises a Clopper-Pearson interval value.

17. The method of claim 1 , wherein the comparing comprises calculating an error value E,

wherein the sampling characteristic comprises a confidence value C, and

wherein the quality characterization comprises a value equal to 100%-E-C.

18. A system performed on a processor for determining the quality of DNA base identifications, the system comprising:

a processor;

a predicted identity input component configured to receive a plurality of predicted DNA base identifications associated with a training data set;

a subset generator configured to define a group of subsets; an identity comparison component configured to compare the predicted DNA base identifications with actual DNA base identifications for training data within each subset of the group;

a sampling determination component configured to determine a sampling characteristic for each subset of the group based on training data within the respective subset; and

a quality characterization determination component configured to determine a quality characterization for predicted DNA base identifications within at least one of subset of the group based on the comparison and determined sampling characteristic

wherein the sampling characteristic comprises a confidence value comprising a binomial proportion confidence interval value.

19. The system of claim 18 , further comprising a training data parameter input component configured to receive parameter values associated with data of the training data set.

20. The system of claim 19 , wherein the subset generator is configured to partition the parameter values into a plurality of bins.

21. The system of claim 19 , further comprising an identity prediction component configured to predict DNA base identifications based on the received parameter values.

22. The system of claim 19 , wherein the parameter values characterize at least a portion of gene array data.

23. The system of claim 19 , wherein the parameter values characterize intrinsic peak characteristics of at least a portion of gene array data.

24. The system of claim 19 , wherein the parameter values comprise at least one of

a first peak height ratio for a current peak based on a first plurality of called peaks centered at a current peak;

a second peak height ratio for the current peak based on a second plurality of called peaks centered at the current peak;

a peak spacing ratio for the current peak based on a largest peak spacing; and

a smallest peak spacing of the second plurality of called peaks centered at the current peak and a peak resolution.

25. The system of claim 18 , further comprising a component configured to generate an accuracy determination based on the determined quality characterization.

26. The system of claim 18 , wherein the quality characterization determination component is configured to generate an accuracy prediction look-up table.

27. The system of claim 18 , wherein the quality characterization determination component is configured to:

define the plurality of subsets by partitioning parameter values associated with data of the training data set into a plurality of bins;

populate the plurality of bins with the predicted DNA base identifications;

iteratively performing the following process:

computing a quality characterization for each subset of the group;

selecting an extreme characteristic subset as the considered subset

having the largest or smallest quality characteristic of the group;

storing the largest or smallest quality characteristic and corresponding threshold values in the look-up table; and

adjusting the quality characteristic for the group of considered subsets such that the quality characteristic no longer depends on data within the extreme quality characteristic subset.

28. The system of claim 27 , wherein the quality characterization determination component is configured to iteratively perform the process until all of the data from the training data set has been within at least one largest quality characteristic subset.

29. The system of claim 18 , wherein the quality characterization determination component is configured to determine a plurality of quality characterizations, each quality characterization being associated with at least one threshold parameter value and to store the quality characterizations and the corresponding threshold parameter values in a look-up table.

30. The system of claim 29 , further comprising:

a data parameter input component configured to receive at least one parameter value associated with data of a non-training data set; and

an accuracy prediction component configured to select a quality characterization from the look-up table, the selected quality characterization being that the at least one corresponding threshold parameter values exceed the at least one parameter value associated with data of the non-training data set.

31. The system of claim 18 , wherein the identity comparison component is configured to calculate an error value E.

32. The system of claim 18 , wherein the binomial proportion confidence interval value comprises a Clopper-Pearson interval value.

33. The system of claim 18 , wherein the identity comparison component is configured to calculate an error value E,

wherein the sampling characteristic comprises a confidence value C, and

wherein the quality characterization comprises a value equal to 100%-E-C.

34. The system of claim 18 , wherein the quality characterization comprises a percentage.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2009
From: MANN, TOBIAS
To: ILLUMINA, INC.
Reel/Frame 023644/0344 →
Continuity (2)
Provisional Application 61102719 · Oct 3, 2008
Related Publication 20100088255A1 · Apr 8, 2010