IP Library Patent Application 14200942
Patent Application
App. No. 14/200,942

Methods, Systems, and Computer Readable Media for Evaluating Variant Likelihood

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/200,942
Abstract

A method for evaluating variant likelihood includes: providing a plurality of template polynucleotide strands, sequencing primers, and polymerase in a plurality of defined spaces disposed on a sensor array; exposing the plurality of template polynucleotide strands, sequencing primers, and polymerase to a series of flows of nucleotide species according to a predetermined order; obtaining measured values corresponding to an ensemble of sequencing reads for at least some of the template polynucleotide strands in at least one of the defined spaces; and evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising: determining a measurement confidence value for each read in the ensemble of sequencing reads and modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands.

Claims (76)

1 . A method for evaluating variant likelihood in nucleic acid sequencing, comprising:

(a) providing a plurality of template polynucleotide strands, sequencing primers, and polymerase in a plurality of defined spaces disposed on a sensor array;

(b) exposing the plurality of template polynucleotide strands, sequencing primers, and polymerase to a series of flows of nucleotide species according to a predetermined order;

(c) obtaining measured values corresponding to an ensemble of sequencing reads for at least some of the template polynucleotide strands in at least one of the defined spaces; and

(d) evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising:

(i) determining a measurement confidence value for each read in the ensemble of sequencing reads, wherein the determining is based on variations between the measured values and model-predicted values for hypothesized sequences obtained using a predictive model of nucleotide incorporations responsive to flows of nucleotide species; and

(ii) modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations responsive to flows of nucleotide species.

2 . The method of claim 1 , wherein modifying the at least some model-predicted values comprises applying a transformation including a product of (i) one of the first and second biases and (ii) a discriminant vector representing a difference between model-predicted values corresponding to different hypothesized sequences.

3 . The method of claim 1 , wherein evaluating the likelihood further comprises assigning a first frequency to a variant sequence and a second frequency to a non-variant sequence, and calculating a likelihood of having observed the ensemble of sequencing reads conditioned on the first frequency as a function of a product of the likelihoods of having observed each of the sequencing reads given the first frequency.

4 . The method of claim 1 , wherein evaluating the likelihood further comprises assigning a first frequency to a variant sequence, a second frequency to a non-variant sequence, and a third frequency to an outlier event.

5 . The method of claim 4 , wherein the outlier event has a flat density across all sequencing reads in the ensemble.

6 . The method of claim 4 , wherein evaluating the likelihood further comprises calculating a likelihood of having observed the ensemble of sequencing reads conditioned on the third frequency as a function of a product of the likelihoods of having observed each of the sequencing reads given the third frequency.

7 . The method of claim 1 , wherein the measurement confidence values are estimated using a function comprising a sum of log-likelihood of values measured for a given flow given a hypothesized sequence.

8 . The method of claim 7 , wherein the measurement confidence values are estimated using a function comprising differences between the measured values and the model-predicted values.

9 . The method of claim 8 , wherein the differences between measured and model-predicted values at each nucleotide flow are assumed to follow independent normal distributions each having a mean and a variance.

10 . The method of claim 9 , wherein the differences between measured and model-predicted values at each nucleotide flow are assumed to follow independent t-distributions.

11 . The method of claim 1 , wherein the measurement confidence values are estimated using an expression comprising

ε

i

=

1

1

+

exp

(

LL

yi

-

LL

xi

)

where LL yi and LL xi are log-likelihoods of values measured for a given sequencing read under hypothesized sequences y and x, respectively.

12 . The method of claim 11 , wherein the measurement confidence values are estimated using an expression for responsibility comprising

ρ

i

=

π

π

+

(

1

-

π

)

*

exp

(

LL

yi

-

LL

xi

)

where π represents a first frequency assigned to a variant sequence, 1−π represents a second frequency assigned to a non-variant sequence, and ρ i represents a measure of responsibility for each of the sequencing reads in the ensemble.

13 . The method of claim 11 , where the variance is estimated by decomposition of the variance in a flow and sequencing read into underlying latent components.

14 . The method of claim 13 , wherein each latent component corresponds to a homopolymer having an integer length.

15 . The method of claim 13 , wherein the latent components include a null variance component representing contribution to a flow regardless of any nucleotide incorporation, a residual variance component representing contribution for nucleotide incorporations not explicitly modeled, and one or more additional variance components.

16 . The method of claim 15 , wherein the one or more additional variance components comprise variance components associated with homopolymers having an integer length.

17 . The method of claim 13 , wherein the latent components are estimated using an EM methodology and a method of moments approximation.

18 . The method of claim 1 , wherein evaluating the likelihood comprises estimating (i) a first frequency π assigned to a variant sequence, (ii) at least one of a measurement confidence value ε i for each of the sequencing reads in the ensemble and a measure of responsibility ρ i for each of the sequencing reads in the ensemble, and (iii) a variance σ ij 2 for each of the flows and sequencing reads in the ensemble.

19 . A non-transitory machine-readable storage medium comprising instructions which, when executed by a processor, cause the processor to perform a method for evaluating variant likelihood in nucleic acid sequencing comprising:

(a) obtaining measured values corresponding to an ensemble of sequencing reads for at least some template polynucleotide strands in at least one defined space, wherein a plurality of template polynucleotide strands, sequencing primers, and polymerase have been provided in a plurality of defined spaces disposed on a sensor array, and wherein the plurality of template polynucleotide strands, sequencing primers, and polymerase have been exposed to a series of flows of nucleotide species according to a predetermined order; and

(b) evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising:

(i) determining a measurement confidence value for each read in the ensemble of sequencing reads, wherein the determining is based on variations between the measured values and model-predicted values for hypothesized sequences obtained using a predictive model of nucleotide incorporations responsive to flows of nucleotide species; and

(ii) modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations responsive to flows of nucleotide species.

20 . A system for evaluating variant likelihood in nucleic acid sequencing, including:

a plurality of template polynucleotide strands, sequencing primers, and polymerase provided in a plurality of defined spaces disposed on a sensor array;

an apparatus configured to expose the plurality of template polynucleotide strands, sequencing primers, and polymerase to a series of flows of nucleotide species according to a predetermined order;

a machine-readable memory; and

a processor configured to execute machine-readable instructions, which, when executed by the processor, cause the system to perform a method for evaluating variant likelihood, comprising:

(a) obtaining measured values corresponding to an ensemble of sequencing reads for at least some of the template polynucleotide strands in at least one of the defined spaces; and

(b) evaluating a likelihood that a variant sequence is present given the measured values corresponding to the ensemble of sequencing reads, the evaluating comprising:

(i) determining a measurement confidence value for each read in the ensemble of sequencing reads, wherein the determining is based on variations between the measured values and model-predicted values for hypothesized sequences obtained using a predictive model of nucleotide incorporations responsive to flows of nucleotide species; and

(ii) modifying at least some model-predicted values using a first bias for forward strands and a second bias for reverse strands, wherein the modifying is based on variations between model-predicted values for different hypothesized sequences obtained using the predictive model of nucleotide incorporations responsive to flows of nucleotide species.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2014
From: HUBBELL, EARL; UTIRAMERUR, SOWMI
To: LIFE TECHNOLOGIES CORPORATION
Reel/Frame 033072/0681 →