IP Library Granted Patent US 8,200,440
Granted Patent B2
US 8,200,440 · App. 12/123,463 · Granted Jun 12, 2012

System, method, and computer software product for genotype determination using probe array data

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,200,440
App. No.
12/123,463
Granted
Jun 12, 2012
Kind
B2
Abstract

An embodiment of a method of analyzing data from processed images of biological probe arrays is described that comprises receiving a plurality of files comprising a plurality of intensity values associated with a probe on a biological probe array; normalizing the intensity values in each of the data files; determining an initial assignment for a plurality of genotypes using one or more of the intensity values from each file for each assignment; estimating a distribution of cluster centers using the plurality of initial assignments; combining the normalized intensity values with the cluster centers to determine a posterior estimate for each cluster center; and assigning a plurality of genotype calls using a distance of the one or more intensity values from the posterior estimate.

Claims (21)

1. A method for determining the genotype of a plurality of nucleic acid samples at a plurality of SNPs, comprising:

(a) hybridizing each nucleic acid sample in said plurality of nucleic acid samples to an array of allele specific probes to obtain a plurality of raw probe intensity measurements;

(b) normalizing said raw probe intensity measurements to obtain normalized probe intensities;

(c) summarizing the normalized probe intensities to obtain an allele signal estimate for each allele of each SNP in each nucleic acid sample, wherein the allele signal estimate for each allele of each SNP in each nucleic acid sample comprises a first value S A and a second value S B where S A is the summary value for a first allele of the SNP and S B is the summary value for a second allele of the SNP;

(d) transforming said allele signal estimates in clustering space in one-dimension to obtain a transformed allele signal estimate for each SNP in each nucleic acid sample using the following equation: transformed allele signal estimate =asinh(K(S A −S B )/(S A +S B ))/asinh(K), where K is a tuning constant;

(e) obtaining a prior distribution of genotype cluster characteristics;

(f) evaluating all possible assignments of the transformed allele signal estimates to the genotype clusters of the prior distribution;

(g) calculating an optimal assignment from the possible assignments of step (f) for each transformed allele signal estimate, to one or more genotype clusters using a Gaussian cluster model;

(h) updating the prior distribution with optimal assignments calculated in (g) to obtain a posterior distribution of genotype cluster characteristics for each SNP; and

(i) using the posterior distribution of genotype cluster characteristics for each SNP to make genotype calls for that SNP in each nucleic acid sample.

2. The method of claim 1 wherein the array comprises more than 1 million different sequence probes and wherein each probe is perfectly complementary to an allele of a SNP that has a minor allele frequency of at least 2% in a population.

3. The method of claim 1 further comprising obtaining a confidence value for each of the genotype calls.

4. The method of claim 1 wherein the step of normalizing said raw probe intensity measurements comprises a quantile normalization step, a log-scale transformation and a median polish step.

5. The method of claim 1 wherein K is 2.

6. The method of claim 1 wherein K is 1 to 4.

7. The method of claim 1 wherein the prior distribution of genotype cluster characteristics is a generic prior derived from a plurality of different SNPs.

8. The method of claim 1 wherein the prior distribution of genotype cluster characteristics is a specific prior for each SNP that has been derived using data for that SNP in a plurality of samples.

9. The method of claim 1 wherein there are three possible genotype clusters including in the prior distribution: a first cluster of samples that are homozygous for the first allele, a second cluster of samples that are homozygous for the second allele and a third cluster of samples that are heterozygous and wherein the third cluster is between the first cluster and the second cluster.

10. The method of claim 1 wherein the posterior distribution of genotype cluster characteristics for each SNP comprises a cluster center for each of three possible genotypes, AA, AB and BB, and covariance matrices for each of the three possible genotypes.

11. The method of claim 10 wherein the genotype call is made by obtaining relative probabilities p(AA), p(AB) and p(BB) from the cluster centers and covariance matrices and calling the genotype with the highest probability for each SNP.

12. The method of claim 10 wherein an ordered relationship between the centers of the genotype clusters is required during calculating step (g).

Assignments (3)
NOTICE OF RELEASE Recorded Apr 5, 2016
From: BANK OF AMERICA, N.A.
To: AFFYMETRIX, INC.
Reel/Frame 038361/0891 →
SECURITY INTEREST Recorded Oct 28, 2015
From: AFFYMETRIX, INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 036988/0166 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2008
From: HUBBELL, EARL A.; CAWLEY, SIMON
To: AFFYMETRIX, INC.
Reel/Frame 021240/0488 →