IP Library Granted Patent US 10,268,800
Granted Patent B2
US 10,268,800 · App. 15/634,919 · Granted Apr 23, 2019

BAMBAM: parallel comparative analysis of high-throughput sequencing data

Inventors: John Zachary Sanborn (Santa Cruz, CA); David Haussler (Santa Cruz, CA)
Assignee: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
G06F19/22C12Q1/6886G06F3/04845G06F17/241G06F19/00G06N7/005G06T11/206G16H50/20C12Q2600/106C12Q2600/118C12Q2600/156G06F2203/04806G06Q50/24G16H10/40Y02A90/22Y02A90/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,268,800
App. No.
15/634,919
Granted
Apr 23, 2019
Kind
B2
Abstract

The present invention relates to methods for evaluating and/or predicting the outcome of a clinical condition, such as cancer, metastasis, AIDS, autism, Alzheimer's, and/or Parkinson's disorder. The methods can also be used to monitor and track changes in a patient's DNA and/or RNA during and following a clinical treatment regime. The methods may also be used to evaluate protein and/or metabolite levels that correlate with such clinical conditions. The methods are also of use to ascertain the probability outcome for a patient's particular prognosis.

Claims (29)

1. A computer-based sequence analysis system comprising:

a computer readable memory configured to store at least a first and second digital sequence datasets associated with a healthy tissue and a suspected diseased tissue, respectively; and

a sequence analysis engine having at least one processor coupled with the computer readable memory and configured, upon execution of software instructions, to:

determine a common sequence location within the first and second digital sequence datasets;

generate at least one pair of sequence pileups from the first and second digital sequence datasets and that overlap at the common sequence location;

infer genotypes for the healthy tissue and suspected diseased tissue at the common sequence location as a function of the at least one pair of sequence pileups;

identify a sequence difference at the common sequence location and between genotypes;

filter false positives based on reads in the at least one pair of sequence pileups; and

store the sequence difference as a differential sequence object in a device memory.

2. The system of claim 1 , wherein the first and second digital sequences are derived from tissues of the same person.

3. The system of claim 1 , wherein at least one of the first and second digital sequences is derived from blood.

4. The system of claim 1 , wherein the second digital sequence is derived from a tumor tissue.

5. The system of claim 1 , wherein at least one of the first and second digital sequence datasets comprises DNA sequences.

6. The system of claim 1 , wherein at least one of the first and second digital sequence datasets comprises RNA sequences.

7. The system of claim 6 , wherein the RNA sequences represent at least one of the following types of RNA: mRNA, rRNA, tRNA, miRNA, and asRNA.

8. The system of claim 1 , wherein at least one of the first and second digital sequence datasets comprises genomic sequences.

9. The system of claim 1 , wherein at least one of the first and second digital sequence datasets comprises a proteomic sequence.

10. The system of claim 9 , wherein the proteomic sequence comprises a polypetide sequence.

11. The system of claim 1 , wherein the first and the second digital sequence datasets comprise short reads.

12. The system of claim 1 , wherein the at least one pair of sequence pileups includes sequence reads having up to 100 positions.

13. The system of claim 1 , wherein the at least one pair of sequence pileups includes more than 30 reads.

14. The system of claim 1 , wherein one of the at least one of first and second digital sequence datasets includes at least a billion reads.

15. The system of claim 1 , wherein the false positives are filtered based on read information of reads within the first and second digital sequence datasets.

16. The system of claim 15 , wherein the read information includes at least one of the following: a strand to which a read maps, an allele position in the reads, and an average quality of alleles.

17. The system of claim 1 , wherein the false positives are filtered based on allele distribution.

18. The system of claim 17 , wherein the allele distribution indicates variant alleles are found near a tail of a read.

19. The system of claim 17 , wherein the false positives are filtered when the allele distribution is skewed from an expected distribution.

20. The system of claim 19 , wherein the expected distribution is a random distribution.

21. The system of claim 1 , wherein the differential sequence object includes at least one variant call.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2018
From: SANBORN, JOHN ZACHARY; HAUSSLER, DAVID
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 047770/0678 →
Continuity (4)
Continuation 15167530 · May 27, 2016
Continuation 13134047 · May 25, 2011
Provisional Application 61396356 · May 25, 2010
Related Publication 20170364634A1 · Dec 21, 2017
Cited By (2)
US 12,347,526 US 12,620,454