IP Library Granted Patent US 12,431,217
Granted Patent B2
US 12,431,217 · App. 17/095,206 · Granted Sep 30, 2025

Systems and methods for use of known alleles in read mapping

Inventor: Deniz Kural (Charlestown, MA)
Assignee: Seven Bridges Genomics Inc.
G16B30/10G16B30/00G16B20/20G16B30/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,431,217
App. No.
17/095,206
Granted
Sep 30, 2025
Kind
B2
Abstract

The invention generally relates to genomic studies and specifically to improved methods for read mapping using identified nucleotides at known locations. The invention provides methods of using identified nucleotides at known places in a genome to guide the analysis of sequence reads from that genome by excluding potential mappings or assemblies that are not congruent with the identified nucleotides. Information about a plurality of SNPs in the subject's genome is used to identify candidate paths through a genomic directed acyclic graph (DAG). Sequence reads are mapped to the candidate paths.

Claims (39)

1. A method for determining a genomic sequence, the method comprising:

receiving, at a computer system, genetic information identifying one or more single nucleotide polymorphisms (SNPs) at one or more respective positions in a genome of a subject;

receiving a plurality of genomic sequences as a genomic directed acyclic graph (DAG) that represents at least two alternative sequences per position at multiple positions and comprises a list of non-compossible node pairs, wherein at least one path through the genomic DAG corresponds to at least one substantially entire sequence of at least one human chromosome, the genomic DAG stored in the computer system and comprising a plurality of nodes and edges, the nodes representing nucleotide sequences, and the edges connecting pairs of the nodes, wherein each of the one or more SNPs is represented by at least one node in the genomic DAG, wherein the plurality of nodes and edges is stored as objects in a memory of the computer system and an object stores a list of pointers specifying one or more locations in the memory where one or more adjacent objects are stored;

identifying a plurality of candidate paths through the genomic DAG that include the one or more SNPs by identifying a node in the list of non-compossible node pairs using one of the one or more SNPs, identifying a second node paired to the identified node in the list of non-compossible node pairs, and excluding paths containing the second node;

receiving sequence reads obtained by sequencing a biological sample from the subject; and

mapping the sequence reads to the plurality of candidate paths, including one or more of the at least one path through the genomic DAG corresponding to the at least one substantially entire sequence of at least one human chromosome, to identify a nucleotide sequence of at least a portion of the genome.

2. The method of claim 1 wherein mapping the sequence reads comprises aligning the sequence reads to the plurality of candidate paths.

3. The method of claim 2 , wherein aligning the sequence reads to the plurality of candidate paths comprises determining at least one score for one or more traces through a multi-dimensional matrix.

4. The method of claim 3 , wherein determining the at least one score for the one or more traces comprises using the computer system to calculate match scores between the sequence reads and the nucleotide sequences represented by at least some of the nodes, and looking backwards to predecessor nodes in the genomic DAG.

5. The method of claim 1 , further comprising:

obtaining probabilities for additional SNPs based on the one or more SNPs, wherein the probabilities are obtained from measures of linkage disequilibrium between ones of the additional SNPs and ones of the one or more SNPs; and

using the obtained probabilities in mapping the sequence reads to the plurality of candidate paths.

6. The method of claim 1 , wherein the genetic information identifying the one or more SNPs is received as results from a microarray assay.

7. A system for determining a genomic sequence, the system comprising:

a computer system comprising a processor coupled to memory and operable to:

receive genetic information identifying one or more single nucleotide polymorphisms (SNPs) at one or more respective positions in a genome of a subject;

receive a plurality of genomic sequences as a genomic directed acyclic graph (DAG) that represents at least two alternative sequences per position at multiple positions and comprises a list of non-compossible node pairs, wherein at least one path through the genomic DAG corresponds to at least one substantially entire sequence of at least one human chromosome, the genomic DAG stored in the computer system and comprising a plurality of nodes and edges, the nodes representing nucleotide sequences, and the edges connecting pairs of the nodes, wherein each of the one or more SNPs is represented by at least one node in the genomic DAG, wherein the plurality of nodes and edges is stored as objects in the memory and an object stores a list of pointers specifying one or more locations in the memory where one or more adjacent objects are stored;

identify a plurality of candidate paths through the genomic DAG that include the one or more SNPs by identifying a node in the list of non-compossible node pairs using one of the one or more SNPs, identifying a second node paired to the identified node in the list of non-compossible node pairs, and excluding paths containing the second node;

receive sequence reads obtained by sequencing a biological sample from the subject; and

map the sequence reads to the plurality of candidate paths, including one or more of the at least one path through the genomic DAG corresponding to at least one substantially entire sequence of at least one human chromosome, to identify a nucleotide sequence of at least a portion of the genome.

8. The system of claim 7 , wherein each of the plurality of candidate paths defines a path through the genomic DAG.

9. The system of claim 7 , wherein mapping the sequence reads comprises aligning the sequence reads to the plurality of candidate paths.

10. The system of claim 9 , wherein aligning the sequence reads to the plurality of candidate paths comprises determining at least one score for one or more traces through a multi-dimensional matrix.

11. The system of claim 9 , further operable to: obtain probabilities for additional SNPs based on the one or more SNPs; and use the obtained probabilities in aligning the sequence reads to the plurality of candidate paths.

12. The system of claim 11 , wherein the probabilities are obtained from measures of linkage disequilibrium between ones of the additional SNPs and ones of the one or more SNPs.

13. The system of claim 7 , wherein each of the plurality of candidate paths defines a path through the genomic DAG and wherein the system is operable to map the sequence reads by aligning the sequence reads to the plurality of candidate paths.

14. The system of claim 13 , wherein aligning the sequence reads to the plurality of candidate paths comprises determining one or more traces through the genomic DAG by:

calculating match scores between the sequence reads and the nucleotide sequences represented by at least some of the nodes in the genomic DAG; and

dereferencing pointers to read predecessor objects storing prior nodes in the genomic DAG from their referenced locations in the memory, wherein a path through the genomic DAG includes one or more of the prior nodes and the match scores are used in determining the one or more traces.

15. At least one non-transitory computer readable storage device storing instructions that, when executed, cause at least one computer system to perform:

receive genetic information identifying one or more single nucleotide polymorphisms (SNPs) at one or more respective positions in a genome of a subject;

receive a plurality of genomic sequences as a genomic directed acyclic graph (DAG) that represents at least two alternative sequences per position at multiple positions and comprises a list of non-compossible node pairs, wherein at least one path through the genomic DAG corresponds to at least one substantially entire sequence of at least one human chromosome, the genomic DAG stored in the at least one computer system and comprising a plurality of nodes and edges, the nodes representing nucleotide sequences, and the edges connecting pairs of the nodes, wherein each of the one or more SNPs is represented by at least one node in the genomic DAG, wherein the plurality of nodes and edges is stored as objects in memory and an object stores a list of pointers specifying one or more locations in the memory where one or more adjacent objects are stored;

identify a plurality of candidate paths through the genomic DAG that include the one or more SNPs by identifying a node in the list of non-compossible node pairs using one of the one or more SNPs, identifying a second node paired to the identified node in the list of non-compossible node pairs, and excluding paths containing the second node;

receive sequence reads obtained by sequencing a biological sample from the subject; and

map the sequence reads to the plurality of candidate paths, including one or more of the at least one path through the genomic DAG corresponding to at least one substantially entire sequence of at least one chromosome, to identify a nucleotide sequence of at least a portion of the genome.

16. The at least one non-transitory computer readable storage device of claim 15 , wherein filtering paths through the genomic DAG comprises using conditional information to exclude paths in the genomic DAG containing the second node.

17. The at least one non-transitory computer readable storage device of claim 16 , wherein the conditional information comprises probabilistic information.

18. The at least one non-transitory computer readable storage device of claim 15 , wherein the instructions further cause the at least one computer system to perform generating a reduced DAG having the plurality of candidate paths, wherein the reduced DAG has fewer nodes than the genomic DAG, and storing the reduced DAG in the at least one computer system.

19. The at least one non-transitory computer readable storage device of claim 15 , wherein mapping the sequence reads to the plurality of candidate paths further comprises dereferencing at least one pointer to access data stored at one or more locations in the memory associated with the at least one pointer.

Assignments (6)
SECURITY INTEREST Recorded Aug 4, 2022
From: PIERIANDX, INC.; SEVEN BRIDGES GENOMICS INC.
To: ORBIMED ROYALTY & CREDIT OPPORTUNITIES III, LP
Reel/Frame 061084/0786 →
RELEASE OF SECURITY INTEREST Recorded Aug 2, 2022
From: IMPERIAL FINANCIAL SERVICES B.V.
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 061055/0078 →
RELEASE OF SECURITY INTEREST Recorded May 24, 2022
From: IMPERIAL FINANCIAL SERVICES B.V.
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 060173/0792 →
SECURITY INTEREST Recorded May 24, 2022
From: SEVEN BRIDGES GENOMICS INC.
To: IMPERIAL FINANCIAL SERVICES B.V.
Reel/Frame 060173/0803 →
SECURITY INTEREST Recorded Mar 30, 2022
From: SEVEN BRIDGES GENOMICS INC.
To: IMPERIAL FINANCIAL SERVICES B.V.
Reel/Frame 059554/0165 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2021
From: KURAL, DENIZ
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 058029/0167 →
Continuity (3)
Continuation 14592444 · Jan 8, 2015
Provisional Application 61925892 · Jan 10, 2014
Related Publication 20210265012A1 · Aug 26, 2021
References Cited (242)
US 5511158A · Sims · 1996 [cited by applicant]
US 5583024A · McElroy et al. · 1996 [cited by applicant]
US 5674713A · McElroy et al. · 1997 [cited by applicant]
US 5700673A · McElroy et al. · 1997 [cited by applicant]
US 5701256A · Marr et al. · 1997 [cited by applicant]
US 6210891B1 · Nyren et al. · 2001 [cited by applicant]
US 6306597B1 · Macevicz · 2001 [cited by applicant]
US 6818395B1 · Quake et al. · 2004 [cited by applicant]
US 6828100B1 · Ronaghi · 2004 [cited by applicant]
US 6833246B2 · Balasubramanian · 2004 [cited by applicant]
US 6890763B2 · Jackowski et al. · 2005 [cited by applicant]
US 6911345B2 · Quake et al. · 2005 [cited by applicant]
US 6925389B2 · Hitt et al. · 2005 [cited by applicant]
US 6989100B2 · Norton · 2006 [cited by applicant]
US 7169560B2 · Lapidus et al. · 2007 [cited by applicant]
US 7232656B2 · Balasubramanian et al. · 2007 [cited by applicant]
US 7282337B1 · Harris · 2007 [cited by applicant]
US 7577554B2 · Lystad et al. · 2009 [cited by applicant]
US 7598035B2 · Macevicz · 2009 [cited by applicant]
US 7620800B2 · Huppenthal et al. · 2009 [cited by applicant]
US 7835871B2 · Kain et al. · 2010 [cited by applicant]
US 7885840B2 · Sadiq et al. · 2011 [cited by applicant]
US 7917302B2 · Rognes · 2011 [cited by applicant]
US 7960120B2 · Rigatti et al. · 2011 [cited by applicant]
US 8146099B2 · Tkatch et al. · 2012 [cited by applicant]
US 8370079B2 · Sorenson et al. · 2013 [cited by applicant]
US 8639847B2 · Blaszczak et al. · 2014 [cited by applicant]
US 9817944B2 · Kural · 2017 [cited by applicant]
US 10867693B2 · Kural · 2020 [cited by applicant]
US 20020164629A1 · Quake et al. · 2002 [cited by applicant]
US 20060024681A1 · Smith et al. · 2006 [cited by applicant]
US 20060195269A1 · Yeatman et al. · 2006 [cited by applicant]
US 20060292611A1 · Berka et al. · 2006 [cited by applicant]
US 20070114362A1 · Feng et al. · 2007 [cited by applicant]
US 20090026082A1 · Rothberg et al. · 2009 [cited by applicant]
US 20090119313A1 · Pearce · 2009 [cited by applicant]
US 20090127589A1 · Rothberg et al. · 2009 [cited by applicant]
US 20090164135A1 · Brodzik et al. · 2009 [cited by applicant]
US 20090191565A1 · Lapidus et al. · 2009 [cited by applicant]
US 20100035252A1 · Rothberg et al. · 2010 [cited by applicant]
US 20100137143A1 · Rothberg et al. · 2010 [cited by applicant]
US 20100169026A1 · Sorenson et al. · 2010 [cited by applicant]
US 20100188073A1 · Rothberg et al. · 2010 [cited by applicant]
US 20110098193A1 · Kingsmore et al. · 2011 [cited by applicant]
US 20120045771A1 · Beier et al. · 2012 [cited by applicant]
US 20120330566A1 · Chaisson · 2012 [cited by applicant]
US 20130059738A1 · Leamon et al. · 2013 [cited by applicant]
US 20130124100A1 · Drmanac et al. · 2013 [cited by applicant]
US 20130289099A1 · Le Goff et al. · 2013 [cited by applicant]
US 20140025312A1 · Chin et al. · 2014 [cited by applicant]
US 20140051588A9 · Drmanac et al. · 2014 [cited by applicant]
US 20140066317A1 · Talasaz · 2014 [cited by applicant]
US 20150094212A1 · Gottimukkala et al. · 2015 [cited by applicant]
US 20150110754A1 · Bai et al. · 2015 [cited by applicant]
US 20150197815A1 · Kural · 2015 [cited by applicant]
US 20150199472A1 · Kural · 2015 [cited by applicant]
US 20150199473A1 · Kural · 2015 [cited by applicant]
US 20150199474A1 · Kural · 2015 [cited by applicant]
US 20150199475A1 · Kural · 2015 [cited by applicant]
US 20150302145A1 · Kural et al. · 2015 [cited by applicant]
US 20150310167A1 · Kural et al. · 2015 [cited by applicant]
US 20150347678A1 · Kural · 2015 [cited by applicant]
US 20160259880A1 · Semenyuk · 2016 [cited by applicant]
US 20160306921A1 · Kural · 2016 [cited by applicant]
US 20160364523A1 · Locke et al. · 2016 [cited by applicant]
US 20170058320A1 · Locke et al. · 2017 [cited by applicant]
US 20170058341A1 · Locke et al. · 2017 [cited by applicant]
US 20170058365A1 · Locke et al. · 2017 [cited by applicant]
US 20170198351A1 · Lee et al. · 2017 [cited by applicant]
US 20170199959A1 · Locke · 2017 [cited by applicant]
US 20170199960A1 · Ghose et al. · 2017 [cited by applicant]
US 20170242958A1 · Brown · 2017 [cited by applicant]
KR 20100011825A · 2010 [cited by applicant]
WO WO2007086935A2 · 2007 [cited by applicant]
WO WO2010010992A1 · 2010 [cited by applicant]
WO WO2012098515A1 · 2012 [cited by applicant]
WO WO2012142531A2 · 2012 [cited by applicant]
WO WO2013035904A1 · 2013 [cited by applicant]
WO WO2015027050A1 · 2015 [cited by applicant]
WO WO2015048753A1 · 2015 [cited by applicant]
WO WO2015058093A1 · 2015 [cited by applicant]
WO WO2015058097A1 · 2015 [cited by applicant]
WO WO2015058120A1 · 2015 [cited by applicant]
WO WO2015061099A1 · 2015 [cited by applicant]
WO WO2015061103A1 · 2015 [cited by applicant]
WO WO2015105963A1 · 2015 [cited by applicant]
WO WO2015123269A1 · 2015 [cited by applicant]
WO WO2016141294A1 · 2016 [cited by applicant]
WO WO2016201215A1 · 2016 [cited by applicant]
WO WO2017120128A1 · 2017 [cited by applicant]
WO WO2017123864A1 · 2017 [cited by applicant]
WO WO2017147124A1 · 2017 [cited by applicant]
Examination Report issued in SG 11201601124Y dated Mar. 1, 2018. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2017/018830 mailed Jun. 14, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/052065 mailed Dec. 11, 2014. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/060680 mailed Jan. 27, 2015. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2015/010604 mailed Mar. 31, 2015. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/061158 mailed Feb. 4, 2015. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/058328 mailed Dec. 30, 2014. [cited by applicant]
International Search Report and Written Opinion for International Patent Application No. PCT/US2015/015375 mailed May 11, 2015. [cited by applicant]
Written Opinion issued in SG 11201603044S dated Sep. 10, 2017. [cited by applicant]
Alioto et al., A comprehensive assessment of somatic mutation detection in cancer using whole-genome sequencing, Nature Communications, Dec. 9, 2015, pp. 1-13. [cited by applicant]
Altschul et al., Optimal sequence alignment using affine gap costs. Bulletin of mathematical biology. Sep. 1, 1986;48(5-6):603-16. [cited by applicant]
Bao et al., BRANCH: boosting RNA-Seq assemblies with partial or related genomic sequences. Bioinformatics. May 15, 2013;29(10):1250-9. [cited by applicant]
Barbieri, 2013, Exome sequencing identifies recurrent SPOP, FOXA1 and MED12 mutations in prostate cancer, Nature Genetics 44:6 685-689. [cited by applicant]
Beerenwinkel, 2007, Conjunctive Bayesian Networks, Bernoulli 13(4), 893-909. [cited by applicant]
Bertone et al., 2004, Global identification of human transcribed sequences with genome tiling arrays, Science 306:2242-2246. [cited by applicant]
Bertrand et al., Genetic map refinement using a comparative genomic approach. Journal of Computational Biology. Oct. 1, 2009;16(10):1475-86. [cited by applicant]
Black, A simple answer for a splicing conundrum. Proceedings of the National Academy of Sciences. Apr. 5, 2005;102(14):4927-8. [cited by applicant]
Browning et al, Haplotype phasing: existing methods and new developments, 2011, vol. 12, Nature Reviews Genetics, pp. 1-26. [cited by applicant]
Caboche et al, Comparison of mapping algorithms used in high-throughput sequencing: application to Ion Torrent data, 2014, vol. 15, BMC Genomics, 16 pages. [cited by applicant]
Carrington et al., Polypeptide ligation occurs during post-translational modification of concanavalin A. Nature. Jan. 1985;313(5997):64-7. [cited by applicant]
Cartwright, DNA assembly with gaps (DAWG): simulating sequence evolution, 2005, pp. iii31-iii38, vol. 21, Oxford University Press. [cited by applicant]
Chang et al., 2005, The application of alternative splicing graphs in quantitative analysis of alternative splicing form from EST database, Int J. Comp. Appl. Tech 22(1): 14. [cited by applicant]
Chen et al., Transient hypermutability, chromothripsis and replication-based mechanisms in the generation of concurrent clustered mutations. Mutation Research/Reviews in Mutation Research. Jan. 1, 2012;750(1):52-9. [cited by applicant]
Chen-Shan et al., Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nature methods. Jun. 2013;10(6):563-571. [cited by applicant]
Compeau et al., How to apply de Bruijn graphs to genome assembly. Nature biotechnology. Nov. 2011;29(11):987-91. [cited by applicant]
Costa et al., Uncovering the complexity of transcriptomes with RNA-Seq. Journal of Biomedicine and Biotechnology. Oct. 2010:1-19. [cited by applicant]
Craig, 1990, Ordering of cosmid clones covering the Herpes simplex virus type I (HSV-I) genome: a test case for fingerprinting by hybridisation, Nucleic Acids Research 18:9 pp. 2653-2660. [cited by applicant]
Danecek et al., 2011, The variant call format and VCFtools, Bionformatics 27(15):2156-2158. [cited by applicant]
Delcher et al., 1999, Alignment of whole genomes, Nucleic Acids Research, 27(11):2369-2376. [cited by applicant]
Denoeud, 2004, Identification of polymorphic tandem repeats by direct comparison of genome sequence from different bacterial strains: a web-based resource, BMC Bioinformatics 5:4 pp. 1-12. [cited by applicant]
DePristo, et al., 2011, A framework for variation discovery and genotyping using next-generation DNA sequencing data, Nature Genetics 43:491-498. [cited by applicant]
Dinov et al., 2011, Applications of the pipeline environment for visual informatics and genomic computations, BMC Bioinformatics 12:304. [cited by applicant]
Dudley and Butte, 2009, A quick guide for developing effective bioinformatics programming skills, PLoS Comput Biol 5 (12):e1000589. [cited by applicant]
Duitama et al., Linkage disequilibrium based genotype calling from low-coverage shotgun sequencing reads. BMC bioinformatics. Dec. 2011;12(1):1-1. [cited by applicant]
Durham et al., 2005, EGene: a configurable pipeline system for automated sequence analysis, Bioinformatics 21 (12):2812-2813. [cited by applicant]
Durham et al., EGene: a configurable pipeline generation system for automated sequence analysis. Bioinformatics. Jun. 15, 2005;21(12):2812-3. [cited by applicant]
Endelman, New algorithm improves fine structure of the barley consensus SNP map. BMC genomics. Dec. 1, 2011;12(1):407. [cited by applicant]
Farrar, Striped Smith-Waterman speeds database searches six times over other SIMD implementations. Bioinformatics. Jan. 15, 2007;23(2):156-61. [cited by applicant]
Fitch, Distinguishing homologous from analogous proteins. Systematic zoology. Jun. 1, 1970;19(2):99-113. [cited by applicant]
Florea et al., 2005, Gene and alternative splicing annotation with AIR, Genome Research 15:54-66. [cited by applicant]
Garber et al., 2011, Computational methods for transcriptome annotation and quantification using RNA-Seq, Nat Meth 8(6):469-477. [cited by applicant]
Gerlinger, 2012, Intratumor Heterogeneity and Branched Evolution Revealed by Multiregion Sequencing, 366:10 883-892. [cited by applicant]
Golub, 1999, Molecular classification of cancer: class discovery and class prediction by gene expression monitoring, Science 286, pp. 531-537. [cited by applicant]
Goto et al., 2010, BioRuby: bioinformatics software for the Ruby programming language, Bioinformatics 26 (20):2617-2619. [cited by applicant]
Gotoh, An improved algorithm for matching biological sequences. Journal of molecular biology. Dec. 15, 1982;162(3):705-8. [cited by applicant]
Grabherr et al., Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nature biotechnology. Jul. 2011;29(7):644-652. [cited by applicant]
Guttman et al., Ab initio reconstruction of cell type-specific transcriptomes in mouse reveals the conserved multi-exonic structure of lincRNAs. Nature biotechnology. May 2010;28(5):503-10. [cited by applicant]
Guttman et al., Ab initio reconstruction of transcriptomes of pluripotent and lineage committed cells reveals gene structures of thousands of lincRNAs. Nature biotechnology. May 2010;28(5):22 pages. [cited by applicant]
Haas et al., DAGchainer: a tool for mining segmental genome duplications and synteny. Bioinformatics. Dec. 12, 2004;20(18):3643-6. [cited by applicant]
HapMap International Consortium. A haplotype map of the human genome. Nature. 2005;437:1299-320. [cited by applicant]
Harrow et al., 2012, Gencode: The reference human genome annotation for the Encode Project, Genome Res 22:1760-1774. [cited by applicant]
Heber et al., 2002, Splicing graphs and EST assembly problems, Bioinformatics 18 Suppl: 181-188. [cited by applicant]
Hein, A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given. Molecular Biology and Evolution. Nov. 1, 1989;6(6):649-68. [cited by applicant]
Hein, A tree reconstruction method that is economical in the number of pairwise comparisons used. Molecular biology and evolution. Nov. 1, 1989;6(6):669-84. [cited by applicant]
Holland et al., 2008, BioJava: an open-source framework for bioinformatics, Bioinformatics 24(18):2096-2097. [cited by applicant]
Homer et al., Improved variant discovery through local re-alignment of short-read next-generation sequencing data using SRMA. Genome biology. Oct. 2010;11(10):12 pages. [cited by applicant]
Hoon et al., 2003, Biopipe: A flexible framework for protocol-based bioinformatics analysis, Genome Research 13 (8):1904-1915. [cited by applicant]
Huang, 3: Bio-Sequence Comparison and Alignment, ser. Curr Top Comp Mol Biol. Cambridge, Mass.: The MIT Press. 2002:45-69. [cited by applicant]
Kent, 2002, BLAT—The Blast-Like Alignment Tool, Genome Research 4:656-664. [cited by applicant]
Kim et al., 2005, ECgene: Genome-based EST clustering and gene modeling for alternative splicing, Genome Research 15:566-576. [cited by applicant]
Kim et al., A scaffold analysis tool using mate-pair information in genome sequencing. BioMed Research International. Apr. 3, 2008;1-7. [cited by applicant]
Kumar et al., 2010, Comparing de novo assemblers for 454 transcriptome data, BMC Genomics 11:571. [cited by applicant]
Kurtz et al., 2004, Versatile and open software for comparing large genomes, Genome Biology, 5:R12. [cited by applicant]
Laframboise, Single nucleotide polymorphism arrays: a decade of biological, computational and technological advances. Nucleic acids research. Jul. 1, 2009;37(13):4181-93. [cited by applicant]
Lam et al., Compressed indexing and local alignment of DNA. Bioinformatics. Mar. 15, 2008;24(6):791-7. [cited by applicant]
Langmead et al., 2009, Ultrafast and memory-efficient alignment of short DNA sequences to the human genome, Genome Biology 10:R25. [cited by applicant]
Larkin et al., Clustal W and Clustal X version 2.0. bioinformatics. Nov. 1, 2007;23(21):2947-8. [cited by applicant]
Lecca, 2015, Defining order and timing of mutations during cancer progression: the TO-DAG probabilistic graphical model, Frontiers in Genetics, vol. 6 Article 309 1-17, pp. 1-17. [cited by applicant]
Lee et al., 2005, Bioinformatics analysis of alternative splicing, Brief Bioinf 6(I):23-33. [cited by applicant]
Lee et al., Multiple sequence alignment using partial order graphs. Bioinformatics. Mar. 1, 2002;18(3):452-64. [cited by applicant]
Lee, 2003, Generating consensus sequences from partial order multiple sequence alignment graphs, Bioinformatics 19(8):999-1008. [cited by applicant]
LeGault et al., 2013, Inference of alternative splicing from RNA-Seq data with probabilistic splice graphs, Bioinformatics 29(18):2300-2310. [cited by applicant]
LeGault et al., Learning Probabilistic Splice Graphs from RNA-Seq data. University of Wisconsin-Madison. 2010:1-8. [cited by applicant]
Leipzig et al., 2004, The alternative splicing gallery (ASG): Bridging the gap between genome and transcriptome, Nucleic Acids Res., 23(13):3977-3983. [cited by applicant]
Li et al., 2008, SOAP: short oligonucleotide alignment program, Bioinformatics 24(5):713-14. [cited by applicant]
Li et al., 2009, SOAP2: an improved ultrafast tool for short read alignment, Bioinformatics 25(15): 1966-67. [cited by applicant]
Li et al., 2010, A survey of sequence alignment algorithms for next-generation sequencing, Briefings in Bionformatics 11(5):473-483. [cited by applicant]
Li, et al., 2009, The Sequence Alignment/Map format and SAMtools, Bioinformatics 25(16):2078-9. [cited by applicant]
Lindgreen, 2012, AdapterRemoval: easy cleaning of next-generation sequence reads, BMC Res Notes 5:337. [cited by applicant]
Lipman et al., Rapid and sensitive protein similarity searches. Science. Mar. 22, 1985;227(4693):1435-41. [cited by applicant]
Ma et al., Multiple genome alignment based on longest path in directed acyclic graphs. International journal of bioinformatics research and applications. Jan. 1, 2010;6(4):366-83. [cited by applicant]
Manolio, Genomewide association studies and assessment of the risk of disease. New England journal of medicine. Jul. 8, 2010;363(2):166-76. [cited by applicant]
Mardis, 2010, The $1,000 genome, the $1,000 analysis?, Genome Med 2:84-85. [cited by applicant]
Margulies et al., 2005, Genome sequencing in microfabricated high-density picolitre reactors, Nature 437:376-380. [cited by applicant]
McKenna, et al., 2010, The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data, Genome Res 20:1297-303. [cited by applicant]
Miller et al., Assembly algorithms for next-generation sequencing data. Genomics. Jun. 1, 2010;95(6):315-27. [cited by applicant]
Mount, Multiple Sequence Alignment, Bioinformatics, 2001, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York. 2001:139-204. [cited by applicant]
Mourad et al., A hierarchical Bayesian network approach for linkage disequilibrium modeling and data-dimensionality reduction prior to genome-wide association studies. BMC bioinformatics. Dec. 2011;12(1):1-20. [cited by applicant]
Nagalakshmi et al., RNA-Seq: a method for comprehensive transcriptome analysis. Current protocols in molecular biology. Jan. 2010;89(1):4-11. [cited by applicant]
Nagarajan & Pop, 2013, Sequence assembly demystified, Nat Rev 14:157-167. [cited by applicant]
Nakao et al., 2005, Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals, Nucl Ac Res 33(8):2355-2363. [cited by applicant]
Needleman et al., A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journal of molecular biology. Mar. 28, 1970;48(3):443-53. [cited by applicant]
Oshlack et al., From RNA-seq reads to differential expression results. Genome biology. Dec. 2010;11(12):10 pages. [cited by applicant]
Pabinger et al., A survey of tools for variant analysis of next-generation genome sequencing data. Briefings in bioinformatics. Jan. 21, 2013;15(2):256-78. [cited by applicant]
Pearson et al., 1988, Improved tools for biological sequence comparison, PNAS 85(8):2444-8. [cited by applicant]
Pe'er et al., Evaluating and improving power in whole-genome association studies using fixed marker sets. Nature genetics. Jun. 2006;38(6):663-7. [cited by applicant]
Posada et al., Modeltest: testing the model of DNA substitution. Bioinformatics (Oxford, England). Jan. 1, 1998;14(9):817-8. [cited by applicant]
Potter et al., ASC: An Associative-Computing Paradigm, Computer , 27(11):19-25, 1994. [cited by applicant]
Potter et al., The Ensembl analysis pipeline. Genome research. May 1, 2004;14(5):934-41. [cited by applicant]
Pruesse, 2012, SINA: Accurate high-throughput multiple sequence alignment of ribosomal RNA genes, Bioinformatics 28:14 1823-1829. [cited by applicant]
Quail, et al., 2012, A tale of three next generation sequencing platforms: comparison of Ion Torrent, Pacific Biosciences and Illumina MiSeq sequencers, BMC Genomics 13:341. [cited by applicant]
Rajarm et al., Pearl millet [ [cited by applicant]
Raphael et al., A novel method for multiple alignment of sequences with repeated and shuffled elements. Genome Research. Nov. 1, 2004;14(11):2336-46. [cited by applicant]
Robertson et al., De novo assembly and analysis of RNA-seq data. Nature methods. Nov. 2010;7(11):909-12. [cited by applicant]
Rodelsperger, 2008, Syntenator: Multiple gene order alignments with a gene-specific scoring function, Alg Mol Biol 3:14. [cited by applicant]
Rognes et al., Faster Smith-Waterman database searches with inter-sequence SIMD parallelisation, Bioinformatics 2011, 12:221. [cited by applicant]
Rognes et al., ParAlign: a parallel sequence alignment algorithm for rapid and sensitive database searches, Nucleic Acids Research, 2001, vol. 29, No. 7 1647-1652. [cited by applicant]
Rognes et al., Six-fold speed-up of Smith-Waterman sequence database searching using parallel processing on common microprocessors, Bioinformatics vol. 16 No. 8 2000, pp. 699-706. [cited by applicant]
Ronquist et al., MrBayes 3.2: efficient Bayesian phylogenetic inference and model choice across a large model space. Systematic biology. May 1, 2012;61(3):539-42. [cited by applicant]
Rothberg, et al., 2011, An integrated semiconductor device enabling non-optical genome sequencing, Nature 475:348-352. [cited by applicant]
Saebo et al., Paralign: rapid and sensitive sequence similarity searches powered by parallel computing technology, Nucleic Acids Research, 2005, vol. 33, Web Server issue W535-W539. [cited by applicant]
Sato et al., Directed acyclic graph kernels for structural RNA analysis. BMC bioinformatics. Jul. 22, 2008;9(1):12 pages. [cited by applicant]
Schenk et al., A pipeline for comprehensive and automated processing of electron diffraction data in IPLT. Journal of structural biology. May 1, 2013;182(2):173-85. [cited by applicant]
Schneeberger et al., 2009, Sumaltaneous alignment of short reads against multiple genomes, Genome Biology 10(9): R98.2-R98.12. [cited by applicant]
Schwikowski & Vingron, 2002, Weighted sequence graphs: boosting iterated dynamic programming using locally suboptimal solutions, Disc Appl Mat 127:95-117. [cited by applicant]
Shao et al., Bioinformatic analysis of exon repetition, exon scrambling and trans-splicing in humans. Bioinformatics. Mar. 15, 2006;22(6):692-8. [cited by applicant]
Sievers et al., 2011, Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omeag, Mol Syst Biol 7:539. [cited by applicant]
Slater & Birney, 2005, Automated generation of heuristics for biological sequence comparison, BMC Bioinformatics 6:31. [cited by applicant]
Slater et al., 2005, Automated generation of heuristics for biological sequence comparison, BMC Bioinformatics 6:31. [cited by applicant]
Smith & Waterman, 1981, Identification of common molecular subsequences, J Mol Biol, 147(1):195-197. [cited by applicant]
Smith et al., Identification of Common Molecular Subsequences, J. Mol. Biol. (1981) 147, 195-197. [cited by applicant]
Smith et al., Multiple insert size paired-end sequencing for deconvolution of complex transcriptions, RNA Bio 9:5, 596-609; May 2012. [cited by applicant]
Soni and Meller, 2007, Progress toward ultrafast DNA sequencing using solid-state nanopores, Clin Chem 53 (11):1996-2001. [cited by applicant]
Stephens, et al,. 2001, A new statistical method for haplotype reconstruction from population data, Am J Hum Genet 68:978-989. [cited by applicant]
Stewart, et al., 2011, A comprehensive map of mobile element insertion polymorphisms in humans, PLoS Genetics 7 (8):1-19. [cited by applicant]
Torri et al., Next generation sequence analysis and computational genomics using graphical pipeline workflows. Genes. Sep. 2012;3(3):545-75. [cited by applicant]
Trapnell et al., TopHat: discovering splice junctions with RNA-Seq. Bioinformatics. May 1, 2009;25(9):1105-11. [cited by applicant]
Trapnell et al., Transcript assembly and abundance estimation from RNA-Seq reveals thousands of new transcripts and switching among isoforms. Nature biotechnology. May 2010;28(5):18 pages. [cited by applicant]
Trapnell et al., Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nature biotechnology. May 2010;28(5):511-5. [cited by applicant]
Uchiyama et al., CGAT: a comparative genome analysis tool for visualizing alignments in the analysis of complex evolutionary changes between closely related genomes, 2006, e-pp. 1-17, vol. 7:472; BMC Bioinformatics. [cited by applicant]
Wang et al., RNA-Seq: a revolutionary tool for transcriptomics. Nature reviews genetics. Jan. 2009;10(1):57-63. [cited by applicant]
Wang, et al., 2011, Next generation sequencing has lower sequence coverage and poorer SNP-detection capability in the regulatory regions, Scientific Reports 1:55. [cited by applicant]
Waterman, et al., 1976, Some biological sequence metrics, Adv. in Math. 20(3):367-387. [cited by applicant]
Wellcome Trust Case Control Consortium, 2007, Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls, Nature 447:661-678. [cited by applicant]
Wu et al., Fast and SNP-tolerant detection of complex variants and splicing in short reads, Bioinformatics, vol. 26 No. 7 2010, pp. 873-881. [cited by applicant]
Xing et al., An expectation-maximization algorithm for probabilistic reconstructions of full-length isoforms from splice graphs. Nucleic acids research. Jan. 1, 2006;34(10):3150-60. [cited by applicant]
Yanovsky, et al., 2008, Read mapping algorithms for single molecule sequencing data, Procs of the 8th Int Workshop on Algorithms in Bioinformatics 5251:38-49. [cited by applicant]
Yu et al., A tool for creating and parallelizing bioinformatics pipelines. 2007 DoD High Performance Computing Modernization Program Users Group Conference Jun. 18, 2007:417-420. [cited by applicant]
Yu et al., The construction of a tetraploid cotton genome wide comprehensive reference map, Genomics 95 (2010) 230-240. [cited by applicant]
Zeng, 2013, PyroHMMvar: a sensitive and accurate method to call short indels and SNPs for Ion Torrent and 454 data, Bioinformatics 29:22 2859-2868. [cited by applicant]
Zhang et al., Implementation of the Smith-Waterman algorithm on reconfigurable supercomputing platform. Altera, White Paper ver 1.0. 2007:18 pages. [cited by applicant]
PCT/US2014/061158, Feb. 4, 2015, International Search Report and Written Opinion. [cited by applicant]
PCT/US2014/058328, Dec. 30, 2014, International Search Report and Written Opinion. [cited by applicant]
SG11201601124Y, Mar. 1, 2018, Examination Report. [cited by applicant]
PCT/US2014/052065, Dec. 11, 2014, International Search Report and Written Opinion. [cited by applicant]
PCT/US2014/060680, Jan. 27, 2015, International Search Report and Written Opinion. [cited by applicant]
SG11201603044S, Sep. 10, 2017, Written Opinion. [cited by applicant]
PCT/US2015/010604, Mar. 31, 2015, International Search Report and Written Opinion. [cited by applicant]
PCT/US2015/015375, May 11, 2015, International Search Report and Written Opinion. [cited by applicant]
PCT/US2017/018830, Jun. 14, 2017, International Search Report and Written Opinion. [cited by applicant]