IP Library Granted Patent US 9,721,062
Granted Patent B2
US 9,721,062 · App. 15/167,530 · Granted Aug 1, 2017

BamBam: parallel comparative analysis of high-throughput sequencing data

Inventors: John Zachary Sanborn (Santa Cruz, CA); David Haussler (Santa Cruz, CA)
Assignee: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
G06F19/22C12Q1/6886G06F3/04845G06F17/241G06F19/345G06N7/005G06T11/206C12Q2600/106C12Q2600/118C12Q2600/156G06F2203/04806G06Q50/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,721,062
App. No.
15/167,530
Granted
Aug 1, 2017
Kind
B2
Abstract

The present invention relates to methods for evaluating and/or predicting the outcome of a clinical condition, such as cancer, metastasis, AIDS, autism, Alzheimer's, and/or Parkinson's disorder. The methods can also be used to monitor and track changes in a patient's DNA and/or RNA during and following a clinical treatment regime. The methods may also be used to evaluate protein and/or metabolite levels that correlate with such clinical conditions. The methods are also of use to ascertain the probability outcome for a patient's particular prognosis.

Claims (28)

1. A computer-based sequence analysis system comprising:

a computer readable memory configured to store at least a first and a second genomic sequence datasets, the sequence datasets comprising genomic reads associated with respective first and second tissues; and

a sequence analysis engine having a processor coupled with the computer readable memory and configured to:

determine a common genomic location in the first and second genomic sequence datasets;

generate at least a pair of pileups by:

reading a first set of pileups that includes genomic reads from the first genomic sequence dataset and that overlap the common genomic location; and

reading a second set of pileups that includes genomic reads from the second genomic sequence dataset and that also overlap the common genomic location;

infer at least a pair of genotypes for the common genomic location based on the at least the pair of pileups, the at least the pair of genotypes including a first genotype associated with the first tissue and a second genotype associated with the second tissue;

identify a genomic difference between the first genotype and the second genotype in the at least the pair of genotypes;

filter false positives based on a skewing from a random distribution; and

store the genomic difference in a device memory.

2. The system of claim 1 , wherein the sequence analysis engine is further configured to infer the at least the pair of genotypes based on a joint probability derived based on the at least the pair of pileups.

3. The system of claim 2 , wherein the sequence analysis engine is further configured to select the first and second genotype based on maximizing the joint probability.

4. The system of claim 1 , wherein the sequence analysis engine is further configured infer the at least the pair of genotypes based on reads in the pair of pileups exceeding mapping quality thresholds.

5. The system of claim 1 , wherein the sequence analysis engine is further configured infer the at least the pair of genotypes based on reads in the pair of pileups exceeding a user-defined base.

6. The system of claim 1 , wherein the genomic difference is selected from the group consisting of: a somatic mutation, a copy number alteration, an allele-specific copy number, a sequence variant, and a sequence loss of heterozygosity.

7. The system of claim 1 , wherein the first tissue and the second tissue are from the same patient.

8. The system of claim 1 , wherein the first tissue comprises a tumor tissue and the second tissue comprises a matched normal tissue.

9. The system of claim 1 , wherein at least one of the first and the second genomic sequence datasets comprises data associated with at least one of the following: DNA, RNA, mRNA, tRNA, rRNA, miRNA, and asRNA.

10. The system of claim 1 , wherein the sequence analysis engine is further configured to keep the first and the second sequence datasets synchronized with respect to a genome.

11. The system of claim 1 , wherein the sequence analysis engine is further configured to read the at least the pair of pileups at the same time.

12. The system of claim 1 , wherein the at least one of the first set of pileups and the second set of pileups include short reads.

13. The system of claim 1 , wherein the at least a pair of pileups includes at least three sets of pileups that include a third set of pileups representing a third genome.

14. The system of claim 13 , wherein the third set of pileups represent a relapsed sequence.

15. The system of claim 1 , wherein the at least one of the first and the second genomic sequence datasets comprises at least one of a BAM file and a SAM file.

16. The system of claim 1 , wherein the common genomic location is relative to a reference genome.

17. The system of claim 1 , wherein the sequence analysis engine is further configured to determine the common genomic location by incrementally moving to a next position in the reference genome.

18. The system of claim 17 , where in the common genomic location comprises a next common genomic location within the reference genome.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2020
From: SANBORN, JOHN ZACHARY; HAUSSLER, DAVID
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 052329/0584 →
Continuity (3)
Continuation 13134047 · May 25, 2011
Provisional Application 61396356 · May 25, 2010
Related Publication 20160275257A1 · Sep 22, 2016