IP Library Granted Patent US 10,991,451
Granted Patent B2
US 10,991,451 · App. 16/107,745 · Granted Apr 27, 2021

BamBam: parallel comparative analysis of high-throughput sequencing data

Inventors: John Zachary Sanborn (Santa Cruz, CA); David Haussler (Santa Cruz, CA)
Assignee: The Regents of the University of California
G16B20/20C12Q1/6886G06F3/04845G06F40/169G06N7/005G06T11/206G16B30/00G16B30/10G16B40/00G16H50/20C12Q2600/106C12Q2600/118C12Q2600/156G06F2203/04806G16H10/40G16H10/60G16H70/20Y02A90/10Y02A90/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,991,451
App. No.
16/107,745
Granted
Apr 27, 2021
Kind
B2
Abstract

The present invention relates to methods for evaluating and/or predicting the outcome of a clinical condition, such as cancer, metastasis, AIDS, autism, Alzheimer's, and/or Parkinson's disorder. The methods can also be used to monitor and track changes in a patient's DNA and/or RNA during and following a clinical treatment regime. The methods may also be used to evaluate protein and/or metabolite levels that correlate with such clinical conditions. The methods are also of use to ascertain the probability outcome for a patient's particular prognosis.

Claims (35)

1. A parallel genomic comparative analysis system comprising:

a computer-readable memory; and

an sequence analysis engine having at least one processor coupled with the computer-readable memory and configured to:

identify a genomic position within a reference genome;

access a first file storing tumor sequence data including reads associated with a tumor tissue;

access a second file storing matched normal sequence data including reads associated with a matched normal tissue, wherein the first and second files are read at the same time;

store, in the computer-readable memory, a tumor dataset having tumor read sequences from the first file, where the tumor read sequences overlap the genomic position;

store, in the computer-readable memory, a matched normal dataset having matched normal read sequences from the second file, where the matched normal read sequences overlap the genomic position, wherein the computer-readable memory is configured to store the tumor dataset and the matched normal dataset simultaneously;

select a tumor genotype and a matched normal genotype that maximize a likelihood as a function of at least one of the tumor read sequences and the matched normal read sequences at the genomic position; and

store a difference associated with at least one of the tumor genotype and the matched normal genotype in a device memory.

2. The system of claim 1 , wherein the reads associated with at least one of the tumor tissue and the matched normal tissue comprise short reads.

3. The system of claim 1 , wherein the reads comprise at least 30 times coverage.

4. The system of claim 1 , wherein the sequence analysis engine is further configured to move to a next genomic location as a new genomic position upon processing the reads at the genomic position.

5. The system of claim 1 , wherein the sequence analysis engine is further configured to synchronize the first and the second file.

6. The system of claim 5 , wherein the reads from the first file and the second file are synchronized with respect to the genomic position.

7. The system of claim 1 , wherein the genomic position comprises a common genomic location between the first file and the second file relative to the reference genome.

8. The system of claim 1 , wherein the difference is selected from the group consisting of: a somatic variant, a germline variant, a single nucleotide polymorphism, an allele-specific copy number, a loss of heterozygosity, a structural rearrangement, a chromosomal fusion, a mutation, a difference relative to the reference genome, and a breakpoint.

9. The system of claim 1 , wherein the likelihood comprises a joint probability between the tumor genotype and the matched normal genotype.

10. The system of claim 1 , wherein the sequence analysis engine is further configured to calculate a confidence score of the tumor genotype and the matched normal genotype pair.

11. The system of claim 10 , wherein the sequence analysis engine is configured to calculate the confidence score as a posterior probability.

12. The system of claim 10 , wherein the sequence analysis engine is further configured to store the confidence score in the device memory with the difference.

13. The system of claim 1 , wherein at least one of the first file and the second file comprises at least one of a BAM file and a SAM file.

14. The system of claim 1 , wherein the tumor dataset comprises strand information on the tumor read sequences at the genomic position and wherein the sequence analysis engine is further configured to filter tumor variant false positives based on the strand information.

15. The system of claim 1 , wherein the matched normal dataset comprises strand information on the matched normal read sequences at the genomic position, and wherein the sequence analysis engine is further configured to filter matched normal variant false positives based on the strand information.

16. The system of claim 1 , wherein the tumor dataset comprises a tumor pileup at the genomic position.

17. The system of claim 1 , wherein the matched normal dataset comprises a matched normal pileup at the genomic position.

18. The system of claim 1 , wherein the difference comprises at least one difference found in at least one of the first file and the second file.

19. A computer implemented method of comparing genomic sequences in parallel, the method comprising:

identifying a genomic position within a reference genome;

accessing, via an sequence analysis engine, a first file storing tumor sequence data including reads associated with a tumor tissue;

accessing, via the sequence analysis engine, a second file storing matched normal sequence data including reads associated with a matched normal tissue, wherein the first and second file are read at the same time;

storing, in a computer-readable memory, a tumor dataset having tumor read sequences from the first file, where the tumor read sequences overlap the genomic position;

storing, in the computer-readable memory, a matched normal dataset having matched normal read sequences from the second file, where the matched normal read sequences overlap the genomic position, wherein the computer-readable memory is configured to store the tumor dataset and the matched normal dataset simultaneously;

selecting, via the sequence analysis engine, a tumor genotype and a matched normal genotype that maximize a likelihood as a function of at least one of the tumor read sequences and the matched normal read sequences at the genomic position; and

storing a difference associated with at least one of the tumor genotype and the matched normal genotype in a device memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2018
From: SANBORN, JOHN ZACHARY; HAUSSLER, DAVID
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 046653/0686 →
Continuity (5)
Continuation 15634919 · Jun 27, 2017
Continuation 15167530 · May 27, 2016
Continuation 13134047 · May 25, 2011
Provisional Application 61396356 · May 25, 2010
Related Publication 20180357369A1 · Dec 13, 2018
Cited By (2)
US 12,347,526 US 12,620,454