IP Library Granted Patent US 12,412,640
Granted Patent B2
US 12,412,640 · App. 16/793,973 · Granted Sep 9, 2025

Systems and methods for reconciling variants in sequence data relative to reference sequence data

Inventor: Amit Jain (London, GB)
Assignee: Seven Bridges Genomics UK, Ltd.
G16B30/00G16B30/10G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,640
App. No.
16/793,973
Granted
Sep 9, 2025
Kind
B2
Abstract

Techniques for identifying variations in sequence data relative to reference sequence data. The techniques include accessing information specifying multiple sets of variants in the sequence data relative to reference sequence data, each of the multiple sets of variants being generated by using a respective variant identification technique; and determining, using the information specifying the multiple sets of variants in the sequence data, a reconciled set of variants in the sequence data relative to the reference sequence data, the determining comprising: determining whether a first variant is present at a first position in the sequence data based, at least in part, on one or more variants at one or more other positions in the sequence data.

Claims (29)

1. A system for identifying variations in sequence data relative to reference sequence data specifying a reference genome, the system comprising:

at least one computer hardware processor; and

at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform:

aligning sequence data to reference sequence data specifying a reference genome to obtain aligned sequence data;

determining information specifying multiple sets of variants in sequence data relative to the reference sequence data specifying the reference genome at least in part by:

applying a first variant identification technique to the aligned sequence data to obtain a first set of variants of the multiple sets of variants; and

applying a second variant identification technique to the aligned sequence data to obtain a second set of variants of the multiple sets of variants, wherein the first variant identification technique is different from the second variant identification technique;

determining, using the information specifying the multiple sets of variants in the sequence data, a reconciled set of variants in the sequence data relative to the reference sequence data specifying the reference genome, the determining comprising:

generating a data structure specifying a graph of the multiple sets of variants, the graph comprising nodes representing genetic sequences and edges connecting at least some of the nodes;

analyzing the data structure to identify a plurality of linear sequences in the graph, wherein each linear sequence of the plurality of linear sequences corresponds to a respective path through the graph;

calculating, for the plurality of linear sequences, a respective plurality of measures of divergence, the calculating comprising:

for each linear sequence in the plurality of linear sequences, comparing the linear sequence to the reference sequence data specifying the reference genome to determine a respective measure of divergence of the linear sequence from the reference sequence data specifying the reference genome;

selecting a particular linear sequence from the plurality of linear sequences based on the respective plurality of measures of divergence; and

determining the reconciled set of variants based on the particular linear sequence.

2. The system of claim 1 , wherein selecting the particular linear sequence from the plurality of linear sequences based on the respective plurality of measures of divergence comprises using a constrained Smith-Waterman alignment technique to select the particular linear sequence from the plurality of linear sequences based on the respective plurality of measures of divergence.

3. A method for identifying variations in sequence data relative to reference sequence data specifying a reference genome, the method comprising:

using at least one computer hardware processor to perform:

aligning sequence data to reference sequence data specifying a reference genome to obtain aligned sequence data;

determining information specifying multiple sets of variants in sequence data relative to the reference sequence data specifying the reference genome at least in part by:

applying a first variant identification technique to the aligned sequence data to obtain a first set of variants of the multiple sets of variants; and

applying a second variant identification technique to the aligned sequence data to obtain a second set of variants of the multiple sets of variants, wherein the first variant identification technique is different from the second variant identification technique;

determining, using the information specifying the multiple sets of variants in the sequence data, a reconciled set of variants in the sequence data relative to the reference sequence data specifying the reference genome, the determining comprising:

generating a data structure specifying a graph of the multiple sets of variants, the graph comprising nodes representing genetic sequences and edges connecting at least some of the nodes;

analyzing the data structure to identify a plurality of linear sequences in the graph, wherein each linear sequence of the plurality of linear sequences corresponds to a respective path through the graph;

calculating, for the plurality of linear sequences, a respective plurality of measures of divergence, the calculating comprising:

for each linear sequence in the plurality of linear sequences, comparing the linear sequence to the reference sequence data specifying the reference genome to determine a respective measure of divergence of the linear sequence from the reference sequence data specifying the reference genome;

selecting a particular linear sequence from the plurality of linear sequences based on the respective plurality of measures of divergence; and

determining the reconciled set of variants based on the particular linear sequence.

4. The method of claim 3 , wherein selecting the particular linear sequence from the plurality of linear sequences based on the respective plurality of measures of divergence comprises using a constrained Smith-Waterman alignment technique to select the particular linear sequence from the plurality of linear sequences based on the respective plurality of measures of divergence.

Assignments (2)
SECURITY INTEREST Recorded Aug 4, 2022
From: PIERIANDX, INC.; SEVEN BRIDGES GENOMICS INC.
To: ORBIMED ROYALTY & CREDIT OPPORTUNITIES III, LP
Reel/Frame 061084/0786 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: JAIN, AMIT
To: SEVEN BRIDGES GENOMICS UK, LTD.
Reel/Frame 052720/0537 →
Continuity (2)
Continuation 15208656 · Jul 13, 2016
Related Publication 20200286588A1 · Sep 10, 2020
References Cited (264)
US 5242794A · Whiteley et al. · 1993 [cited by applicant]
US 5701256A · Marr et al. · 1997 [cited by applicant]
US 6054278A · Dodge et al. · 2000 [cited by applicant]
US 6223128B1 · Allex et al. · 2001 [cited by applicant]
US 6828100B1 · Ronaghi · 2004 [cited by applicant]
US 6931401B2 · Gibson et al. · 2005 [cited by applicant]
US 7232656B2 · Balasubramanian et al. · 2007 [cited by applicant]
US 7809509B2 · Milosavljevic · 2010 [cited by applicant]
US 7917302B2 · Rognes · 2011 [cited by applicant]
US 7957913B2 · Chinitz et al. · 2011 [cited by applicant]
US 7996157B2 · Zabeau et al. · 2011 [cited by applicant]
US 8165821B2 · Zhang · 2012 [cited by applicant]
US 8209130B1 · Kennedy et al. · 2012 [cited by applicant]
US 8340914B2 · Gatewood et al. · 2012 [cited by applicant]
US 8370079B2 · Sorenson et al. · 2013 [cited by applicant]
US 8428886B2 · Wong et al. · 2013 [cited by applicant]
US 8725422B2 · Halpern et al. · 2014 [cited by applicant]
US 8775092B2 · Colwell et al. · 2014 [cited by applicant]
US 8880456B2 · Kermani et al. · 2014 [cited by applicant]
US 9063914B2 · Kural et al. · 2015 [cited by applicant]
US 9092402B2 · Kural et al. · 2015 [cited by applicant]
US 9116866B2 · Kural · 2015 [cited by applicant]
US 9183349B2 · Kupershmidt et al. · 2015 [cited by applicant]
US 9323888B2 · Rava et al. · 2016 [cited by applicant]
US 10600499B2 · Jain · 2020 [cited by applicant]
US 20040023209A1 · Jonasson · 2004 [cited by applicant]
US 20050089906A1 · Furuta et al. · 2005 [cited by applicant]
US 20090119313A1 · Pearce · 2009 [cited by applicant]
US 20090233809A1 · Faham et al. · 2009 [cited by applicant]
US 20090318310A1 · Liu et al. · 2009 [cited by applicant]
US 20100169026A1 · Sorenson et al. · 2010 [cited by applicant]
US 20110004413A1 · Carnevali et al. · 2011 [cited by applicant]
US 20110009278A1 · Kain et al. · 2011 [cited by applicant]
US 20110098193A1 · Kingsmore et al. · 2011 [cited by applicant]
US 20110257889A1 · Klammer et al. · 2011 [cited by applicant]
US 20120041727A1 · Mishra et al. · 2012 [cited by applicant]
US 20120239706A1 · Steinfadt · 2012 [cited by applicant]
US 20120330566A1 · Chaisson · 2012 [cited by applicant]
US 20130059740A1 · Drmanac et al. · 2013 [cited by applicant]
US 20130073214A1 · Hyland et al. · 2013 [cited by applicant]
US 20130124100A1 · Drmanac et al. · 2013 [cited by applicant]
US 20130173177A1 · Pelleymounter · 2013 [cited by applicant]
US 20130311106A1 · White et al. · 2013 [cited by applicant]
US 20130332081A1 · Reese et al. · 2013 [cited by applicant]
US 20130345066A1 · Brinza et al. · 2013 [cited by applicant]
US 20140012866A1 · Bowman et al. · 2014 [cited by applicant]
US 20140051588A9 · Drmanac et al. · 2014 [cited by applicant]
US 20140129201A1 · Kennedy et al. · 2014 [cited by applicant]
US 20140136120A1 · Colwell et al. · 2014 [cited by applicant]
US 20140143188A1 · Mackey et al. · 2014 [cited by applicant]
US 20140200147A1 · Bartha et al. · 2014 [cited by applicant]
US 20140280360A1 · Webber et al. · 2014 [cited by applicant]
US 20140336999A1 · Kumar et al. · 2014 [cited by applicant]
US 20150056613A1 · Kural · 2015 [cited by applicant]
US 20150057946A1 · Kural · 2015 [cited by applicant]
US 20150112602A1 · Kural et al. · 2015 [cited by applicant]
US 20150112658A1 · Kural et al. · 2015 [cited by applicant]
US 20160004817A1 · Lawrence et al. · 2016 [cited by applicant]
US 20160019341A1 · Harris et al. · 2016 [cited by applicant]
US 20160026757A1 · Li et al. · 2016 [cited by applicant]
US 20160026760A1 · Leong et al. · 2016 [cited by applicant]
US 20160085910A1 · Bruand et al. · 2016 [cited by applicant]
US 20160103954A1 · Conway et al. · 2016 [cited by applicant]
US 20160140289A1 · Gibiansky et al. · 2016 [cited by applicant]
US 20160232291A1 · Kyriazopoulou-Panagiotopoulou et al. · 2016 [cited by applicant]
US 20160283654A1 · Ye et al. · 2016 [cited by applicant]
US 20160300013A1 · Ashutosh et al. · 2016 [cited by applicant]
US 20180018423A1 · Jain · 2018 [cited by applicant]
EP 2990976A1 · 2016 [cited by applicant]
GB 2510725A · 2014 [cited by applicant]
WO WO2007030426A2 · 2007 [cited by applicant]
WO WO2012096579A2 · 2012 [cited by applicant]
WO WO2012098515A1 · 2012 [cited by applicant]
WO WO2012142531A2 · 2012 [cited by applicant]
WO WO2013040583A2 · 2013 [cited by applicant]
WO WO2013043909A1 · 2013 [cited by applicant]
WO WO2013106737A1 · 2013 [cited by applicant]
WO WO2013184643A1 · 2013 [cited by applicant]
WO WO2015027050A1 · 2015 [cited by applicant]
WO WO2015031689A1 · 2015 [cited by applicant]
WO WO2015058093A1 · 2015 [cited by applicant]
WO WO2015058095A1 · 2015 [cited by applicant]
WO WO2015112619A1 · 2015 [cited by applicant]
WO WO2015123269A1 · 2015 [cited by applicant]
WO WO2015123600A1 · 2015 [cited by applicant]
WO WO2015173222A1 · 2015 [cited by applicant]
WO WO2016040287A1 · 2016 [cited by applicant]
WO WO2016122318A1 · 2016 [cited by applicant]
WO WO2016138127A1 · 2016 [cited by applicant]
WO WO2016141077A1 · 2016 [cited by applicant]
WO WO2016154584A1 · 2016 [cited by applicant]
[No Author Listed] International HapMap Consortium. A haplotype map of the human genome. Nature. 2005;437: 1299-1320. [cited by applicant]
1000 Genomes Project Consortium, Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, et al. A map of human genome variation from population-scale sequencing. Nature. Nature Research; 2010;467:14 pages. [cited by applicant]
1000 Genomes Project Consortium, Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, et al. A global reference for human genetic variation. Nature. Nature Research; 2015;526: 68-74. [cited by applicant]
Abouelhoda M, Issa SA, Ghanem M. Tavaxy: integrating Taverna and Galaxy workflows with cloud computing support. BMC Bioinformatics. 2012;13: 77:19 pages. [cited by applicant]
Aguiar D, Istrail S. HapCompass: a fast cycle basis algorithm for accurate haplotype assembly of sequence data. J Comput Biol. 2012;19: 577-590. [cited by applicant]
Aguiar D, Istrail S. Haplotype assembly in polyploid genomes and identical by descent shared tracts. Bioinformatics. 2013;29: 1352-60. [cited by applicant]
Altschul SF, Erickson BW. Optimal sequence alignment using affine gap costs. Bull Math Biol. 1986;48: 603-616. [cited by applicant]
Aniba MR, Poch O, Thompson JD. Issues in bioinformatics benchmarking: the case study of multiple sequence alignment. Nucleic Acids Res. 2010;38: 7353-7363. [cited by applicant]
Bansal V, Halpern AL, Axelrod N, Bafna V. An MCMC algorithm for haplotype assembly from whole-genome sequence data. Genome Res. 2008;18: 1336-1346. [cited by applicant]
Bansal V, Harismendy O, Tewhey R, Murray SS, Schork NJ, Topol EJ, et al. Accurate detection and genotyping of SNPs utilizing population sequencing data. Genome Res. genome.cshlp.org; 2010;20: 537-545. [cited by applicant]
Bertrand D, Blanchette M, El-Mabrouk N. Genetic map refinement using a comparative genomic approach. J Comput Biol. 2009;16: 1475-1486. [cited by applicant]
Boyer RS, Moore JS. A fast string searching algorithm. Commun ACM. ACM; 1977;20: 762-772. [cited by applicant]
Breese MR, Liu Y. NGSUtils: a software suite for analyzing and manipulating next-generation sequencing datasets. Bioinformatics. Oxford Univ Press; 2013;29(4):494-496. [cited by applicant]
Buhler J. Search algorithms for biosequences using random projection. Doctor of Philosophy, University of Washington. 2001. 203 pages. [cited by applicant]
Cantarel BL, Weaver D, McNeill N, Zhang J, Mackey AJ, Reese J. Baysic: a Bayesian method for combining sets of genome variants with improved specificity and sensitivity. BMC Bioinformatics. 2014;15:12 pages. [cited by applicant]
Challis D, Yu J, Evani US, Jackson AR, Paithankar S, Coarfa C, et al. An integrative variant analysis suite for whole exome next-generation sequencing data. BMC Bioinformatics. biomedcentral.com; 2012;13:1-12. [cited by applicant]
Chen J-M, Ferec C, Cooper DN. Transient hypermutability, chromothripsis and replication-based mechanisms in the generation of concurrent clustered mutations. Mutat Res. 2012;750: 52-59. [cited by applicant]
Cheng AY, Teo Y-Y, Ong RT-H. Assessing single nucleotide variant detection and genotype calling on whole-genome sequenced individuals. Bioinformatics. 2014;30: 1707-1713. [cited by applicant]
Chuang JS, Roth D. Gene recognition based on DAG shortest paths. Bioinformatics. 2001;17 Suppl 1: S56-64. [cited by applicant]
Clark L. Illumina announces landmark $1,000 human genome sequencing. Wired. Jan. 15, 2014. Available: http://www.wired.co.uk/article/1000-dollar-genome 3 pages. [cited by applicant]
Cleary JG, Braithwaite R, Gaastra K, Hilbush BS, Inglis S, Irvine SA, et al. Joint variant and de novo mutation identification on pedigrees from high-throughput sequencing data. J Comput Biol. online.liebertpub.com; 201… [cited by applicant]
Cock PJA, Gruning BA, Paszkiewicz K, Pritchard L. Galaxy tools and workflows for sequence analysis with applications in molecular plant pathology. PeerJ. 2013;1: e167:22 pages. [cited by applicant]
Compeau PEC, Pevzner PA, Tesler G. How to apply de Bruijn graphs to genome assembly. Nat Biotechnol. 2011;29: 987-991. [cited by applicant]
Cornish A, Guda C. A Comparison of Variant Calling Pipelines Using Genome in a Bottle as a Reference. Biomed Res Int. hindawi.com; 2015;2015: 456479: 11 pages. [cited by applicant]
Craig DW, Pearson JV, Szelinger S, Sekar A, Redman M, Corneveaux JJ, et al. Identification of genetic variants using barcoded multiplexed sequencing. Nat Methods. 2008;5(10):16 pages. [cited by applicant]
Danecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA, et al. The variant call format and VCFtools. Bioinformatics. 2011;27: 2156-2158. [cited by applicant]
David W. Mount. Multiple Sequence Alignment. In: Cuddihy J, Barker P, editors. Bioinformatics: Sequence and Genome Analysis. Cold Spring Harbor Laboratory Press; 2001. 68 pages. [cited by applicant]
Davies KD, Farooqi MS, Gruidl M, Hill CE, Woolworth-Hirschhorn J, Jones H, et al. Multi-Institutional FASTQ File Exchange as a Means of Proficiency Testing for Next-Generation Sequencing Bioinformatics and Variant Inter… [cited by applicant]
Delcher AL, Kasif S, Fleischmann RD, Peterson J, White O, Salzberg SL. Alignment of whole genomes. Nucleic Acids Res. 1999;27: 2369-2376. [cited by applicant]
DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat Genet. 2011;43: 19 pages. [cited by applicant]
Dolled-Filhart MP, Lee M Jr, Ou-Yang C-W, Haraksingh RR, Lin JC-H. Computational and bioinformatics frameworks for next-generation whole exome and genome sequencing. Scientific WorldJournal. hindawi.com; 2013;2013: 10 p… [cited by applicant]
Edmonson MN, Zhang J, Yan C, Finney RP, Meerzaman DM, Buetow KH. Bambino: a variant detector and alignment viewer for next-generation sequencing data in the SAM/BAM format. Bioinformatics. Oxford Univ Press; 2011;27: 86… [cited by applicant]
Endelman JB. New algorithm improves fine structure of the barley consensus SNP map. BMC Genomics. 2011;12: 9 pages. [cited by applicant]
Farrar M. Striped Smith-Waterman speeds database searches six times over other SIMD implementations. Bioinformatics. 2007;23: 156-161. [cited by applicant]
Farrer RA, Henk DA, MacLean D, Studholme DJ, Fisher MC. Using false discovery rates to benchmark SNP-callers in next-generation sequencing projects. Sci Rep. nature.com; 2013;3:1-6. [cited by applicant]
Flicek P, Birney E. Sense from sequence reads: methods for alignment and assembly. Nat Methods. 2009;6: S6-S12. [cited by applicant]
Garrison E, Marth G. Haplotype-based variant detection from short-read sequencing. arXiv. 1207.3907. Jul. 24, 2012:1-9. [cited by applicant]
Ghoneim DH, Myers JR, Tuttle E, Paciorkowski AR. Comparison of insertion/deletion calling algorithms on human next-generation sequencing data. BMC Res Notes. 2014;7:1-10. [cited by applicant]
Giannoulatou E, Yau C, Colella S, Ragoussis J, Holmes CC. GenoSNP: a variational Bayes within-sample SNP genotyping algorithm that does not require a reference population. Bioinformatics. 2008;24: 2209-2214. [cited by applicant]
Giegerich R. A systematic approach to dynamic programming in bioinformatics. Bioinformatics. Oxford Univ Press; 2000;16: 665-677. [cited by applicant]
Glusman G, Cox HC, Roach JC. Whole-genome haplotyping approaches and genomic medicine. Genome Med. 2014;6:1-16. [cited by applicant]
Gotoh O. An improved algorithm for matching biological sequences. J Mol Biol. 1982;162: 705-708. [cited by applicant]
Gotoh O. Multiple sequence alignment: algorithms and applications. Adv Biophys. 1999;36: 159-206. [cited by applicant]
Grasso C, Lee C. Combining partial order alignment and progressive multiple sequence alignment increases alignment speed and scalability to very large alignment problems. Bioinformatics. 2004;20(10): 1546-1556. [cited by applicant]
Haas BJ, Delcher AL, Wortman JR, Salzberg SL. DAGchainer: a tool for mining segmental genome duplications and synteny. Bioinformatics. 2004;20: 3643-3646. [cited by applicant]
Hamada M, Wijaya E, Frith MC, Asai K. Probabilistic alignments with quality scores: an application to short-read mapping toward accurate SNP/indel detection. Bioinformatics. Oxford Univ Press; 2011;27: 3085-3092. [cited by applicant]
Hansen NF. Variant Calling From Next Generation Sequence Data. In: Mathe E, Davis S, editors. Statistical Genomics. New York, NY: Springer New York; 2016. pp. 209-224. [cited by applicant]
He D, Choi A, Pipatsrisawat K, Darwiche A, Eskin E. Optimal algorithms for haplotype assembly from whole-genome sequence data. Bioinformatics. 2010;26: i183-90. [cited by applicant]
Hein J. A new method that simultaneously aligns and reconstructs ancestral sequences for any No. of homologous sequences, when the phylogeny is given. Mol Biol Evol. 1989;6: 649-668. [cited by applicant]
Highnam G, Wang JJ, Kusler D, Zook J, Vijayan V, Leibovich N, et al. An analytical framework for optimizing variant discovery from personal genomes. Nat Commun. 2015;6:1-6. [cited by applicant]
Homer N, Merriman B, Nelson SF. BFAST: an alignment tool for large scale genome resequencing. PLoS One. 2009;4: e7767:1-12. [cited by applicant]
Homer N, Nelson SF. Improved variant discovery through local re-alignment of short-read next-generation sequencing data using SRMA. Genome Biol. 2010;11: 1-12. [cited by applicant]
Horspool RN. Practical fast searching in strings. Softw Pract Exp. John Wiley & Sons, Ltd.; 1980;10: 501-506. [cited by applicant]
Huang X. Bio-sequence Comparison and Applications. In: Jiang T, Xu Y, Zhang MQ, editors. Current Topics in Computational Molecular Biology. MIT Press; 2002. 29 pages. [cited by applicant]
Hutchinson JN, Raj T, Fagerness J, Stahl E, Viloria FT, Gimelbrant A, et al. Allele-specific methylation occurs at genetic variants associated with complex disease. PLoS One. 2014;9: e98464:1-14. [cited by applicant]
Hwang S, Kim E, Lee I, Marcotte EM. Systematic comparison of variant calling pipelines using gold standard personal exome variants. Sci Rep. 2015;5: 17875:1-8. [cited by applicant]
Ji HP. Improving bioinformatic pipelines for exome variant calling. Genome Med. 2012;4:7:1-4. [cited by applicant]
Katoh K, Kuma K-I, Toh H, Miyata T. Mafft version 5: improvement in accuracy of multiple sequence alignment. Nucleic Acids Res. 2005;33: 511-518. [cited by applicant]
Kehr B, Trappe K, Holtgrewe M, Reinert K. Genome alignment with graph data structures: a comparison. BMC Bioinformatics. 2014;15(99):20 pages. [cited by applicant]
Kent WJ. BLAT—The BLAST-Like Alignment Tool. Genome Res. 2002;12: 656-664. [cited by applicant]
Kim D, Pertea G, Trapnell C, Pimentel H, Kelley R, Salzberg SL. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. Genome Biol. 2013;14: R36:13 pages. [cited by applicant]
Kim P-G, Cho H-G, Park K. A scaffold analysis tool using mate-pair information in genome sequencing. J Biomed Biotechnol. 2008;2008: 675741:7 pages. [cited by applicant]
Kim SY, Jacob L, Speed TP. Combining calls from multiple somatic mutation-callers. BMC Bioinformatics. 2014;15: 154:8 pages. [cited by applicant]
Kim SY, Speed TP. Comparing somatic mutation-callers: beyond Venn diagrams. BMC Bioinformatics. 2013;14: 189:16 pages. [cited by applicant]
Koboldt DC, Larson DE, Wilson RK. Using VarScan 2 for Germline Variant Calling and Somatic Mutation Detection. Curr Protoc Bioinformatics. Wiley Online Library; 2013;44: 15.4. 22 pages. [cited by applicant]
Korbel JO, Abyzov A, Mu XJ, Carriero N, Cayting P, Zhang Z, et al. PEMer: a computational framework with simulation-based error models for inferring genomic structural variants from massive paired-end sequencing data. G… [cited by applicant]
Kosugi S, Natsume S, Yoshida K, MacLean D, Cano L, Kamoun S, et al. Coval: improving alignment quality and variant calling accuracy for next-generation sequencing data. PLoS One. journals.plos.org; 2013;8: e75402:1-11. [cited by applicant]
Krishnan V, Zia A, Utiramerur S, Datta S. Benchmarking variant callers: Towards building a robust exome pipeline [Internet]. med.stanford.edu; 2016. Available: http://med.stanford.edu/content/dam/sm/gbsc/GeneticsRetreat… [cited by applicant]
Kumar P, Al-Shafai M, Al Muftah WA, Chalhoub N, Elsaid MF, Aleem AA, et al. Evaluation of SNP calling using single and multiple-sample calling algorithms by validation against array base genotyping and Mendelian inherit… [cited by applicant]
Kurtz S, Phillippy A, Delcher AL, Smoot M, Shumway M, Antonescu C, et al. Versatile and open software for comparing large genomes. Genome Biol. 2004;5: R12. 9 pages. [cited by applicant]
LaFramboise T. Single nucleotide polymorphism arrays: a decade of biological, computational and technological advances. Nucleic Acids Res. 2009;37: 4181-4193. [cited by applicant]
Lai Z, Markovets A, Ahdesmaki M, Chapman B, Hofmann O, McEwen R, et al. VarDict: a novel and versatile variant caller for next-generation sequencing in cancer research. Nucleic Acids Res. 2016;44: e108:1-11. [cited by applicant]
Lam HYK, Clark MJ, Chen R, Chen R, Natsoulis G, O'Huallachain M, et al. Performance comparison of whole-genome sequencing platforms. Nat Biotechnol. nature.com; 2011;30: 1-16. [cited by applicant]
Lam HYK, Pan C, Clark MJ, Lacroute P, Chen R, Haraksingh R, et al. Detecting and annotating genetic variations using the HugeSeq pipeline. Nat Biotechnol. nature.com; 2012;30: 1-9. [cited by applicant]
Lam TW, Sung WK, Tam SL, Wong CK, Yiu SM. Compressed indexing and local alignment of DNA. Bioinformatics. 2008;24: 791-797. [cited by applicant]
Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. nature.com; 2012;9: 1-8. [cited by applicant]
Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009;10: R25-R25.10. [cited by applicant]
Larkin MA, Blackshields G, Brown NP, Chenna R, McGettigan PA, McWilliam H, et al. Clustal W and Clustal X version 2.0. Bioinformatics. 2007;23: 2947-2948. [cited by applicant]
Layer RM, Chiang C, Quinlan AR, Hall IM. Lumpy: a probabilistic framework for structural variant discovery. Genome Biol. genomebiology.biomedcentral.com; 2014;15: R84:19 pages. [cited by applicant]
Lee C, Grasso C, Sharlow MF. Multiple sequence alignment using partial order graphs. Bioinformatics. 2002;18: 452-464. [cited by applicant]
Lee C. Generating consensus sequences from partial order multiple sequence alignment graphs. Bioinformatics. 2003;19: 999-1008. [cited by applicant]
Lee HC, Lai K, Lorenc MT, Imelfort M, Duran C, Edwards D. Bioinformatics tools and databases for analysis of next-generation sequence data. Brief Funct Genomics. bfg.oxfordjournals.org; 2012;11: 12-24. [cited by applicant]
Lee W-P, Stromberg MP, Ward A, Stewart C, Garrison EP, Marth Gt. Mosaik: a hash-based algorithm for accurate next-generation sequencing short-read mapping. PLoS One. 2014;9: e90581:1-11. [cited by applicant]
Li B, Chen W, Zhan X, Busonero F, Sanna S, Sidore C, et al. A likelihood-based framework for variant calling and de novo mutation detection in families. PLoS Genet. journals.plos.org; 2012;8: e1002944:1-12. [cited by applicant]
Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25: 1754-1760. [cited by applicant]
Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25: 2078-2079. [cited by applicant]
Li H, Homer N. A survey of sequence alignment algorithms for next-generation sequencing. Brief Bioinform. 2010;11: 473-483. [cited by applicant]
Li H, Ruan J, Durbin R. Mapping short DNA sequencing reads and calling variants using mapping quality scores. Genome Res. genome.cshlp.org; 2008;18: 1851-1858. [cited by applicant]
Li H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics. Oxford Univ Press; 2011;27: 2987-2993. [cited by applicant]
Li H. Exploring single-sample SNP and INDEL calling with whole-genome de novo assembly. Bioinformatics. 2012;28: 1838-1844. [cited by applicant]
Li H. Toward better understanding of artifacts in variant calling from high-coverage samples. Bioinformatics. 2014;30: 2843-2851. [cited by applicant]
Li R, Li Y, Kristiansen K, Wang J. Soap: short oligonucleotide alignment program. Bioinformatics. 2008;24: 713-714. [cited by applicant]
Li R, Yu C, Li Y, Lam T-W, Yiu S-M, Kristiansen K, et al. SOAP2: an improved ultrafast tool for short read alignment. Bioinformatics. 2009;25: 1966-1967. [cited by applicant]
Li Y, Chen W, Liu EY, Zhou Y-H. Single Nucleotide Polymorphism (SNP) Detection and Genotype Calling from Massively Parallel Sequencing (MPS) Data. Stat Biosci. Springer; 2013;5: 26 pages. [cited by applicant]
Lin S, Carvalho B, Cutler DJ, Arking DE, Chakravarti A, Irizarry RA. Validation and extension of an empirical Bayes method for SNP calling on Affymetrix microarrays. Genome Biol. 2008;9: R63-R63.12. [cited by applicant]
Liu X, Han S, Wang Z, Gelernter J, Yang B-Z. Variant callers for next-generation sequencing data: a comparison study. PLoS One. journals.plos.org; 2013;8: e75619:1-11. [cited by applicant]
Lucking R, Hodkinson BP, Stamatakis A, Cartwright RA. PICS-Ord: unlimited coding of ambiguous regions by pairwise identity and cost scores ordination. BMC Bioinformatics. 2011;12: 1-15. [cited by applicant]
Ma F, Deogun JS. Multiple genome alignment based on longest path in directed acyclic graphs. IJBRA. 2010;6: 366. [cited by applicant]
Manolio TA. Genomewide association studies and assessment of the risk of disease. N Engl J Med. 2010;363: 166-176. [cited by applicant]
Martin ER, Kinnamon DD, Schmidt MA, Powell EH, Zuchner S, Morris RW. SeqEM: an adaptive genotype-calling approach for next-generation sequencing studies. Bioinformatics. Oxford Univ Press; 2010;26: 2803-2810. [cited by applicant]
Mazrouee S, Wang W. FastHap: fast and accurate single individual haplotype reconstruction using fuzzy conflict graphs. Bioinformatics. 2014;30: 1371-8. [cited by applicant]
McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 2010;20: 1297-1303. [cited by applicant]
Miller JR, Koren S, Sutton G. Assembly algorithms for next-generation sequencing data. Genomics. 2010;95: 31 pages. [cited by applicant]
Misra S, Agrawal A, Liao W-K, Choudhary A. Anatomy of a hash-based long read sequence mapping algorithm for next generation DNA sequencing. Bioinformatics. 2011;27: 189-195. [cited by applicant]
Mohiyuddin M, Mu JC, Li J, Bani Asadi N, Gerstein MB, Abyzov A, et al. MetaSV: an accurate and integrative structural-variant caller for next generation sequencing. Bioinformatics. Oxford Univ Press; 2015;31: 2741-2744. [cited by applicant]
Mu JC, Tootoonchi Afshar P, Mohiyuddin M, Chen X, Li J, Bani Asadi N, et al. Leveraging long read sequencing from a single individual to provide a comprehensive resource for benchmarking variant calling methods. Sci Rep… [cited by applicant]
Nagarajan N, Pop M. Sequence assembly demystified. Nat Rev Genet. Nature Publishing Group; 2013;14: 157-167. [cited by applicant]
Najafi A, Nashta-ali D, Motahari SA, Khani M, Khalaj BH, Rabiee HR. Fundamental Limits of Pooled-DNA Sequencing. 2016. 38 pages. [cited by applicant]
Needleman SB, Wunsch CD. A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol. 1970;48: 443-453. [cited by applicant]
Nekrutenko A, Taylor J. Next-generation sequencing data interpretation: enhancing reproducibility and accessibility. Nat Rev Genet. nature.com; 2012;13: 667-672. [cited by applicant]
Nielsen R, Korneliussen T, Albrechtsen A, Li Y, Wang J. SNP calling, genotype calling, and sample allele frequency estimation from New-Generation Sequencing data. PLoS One. journals.plos.org; 2012;7: e37558:1-11. [cited by applicant]
Nielsen R, Paul JS, Albrechtsen A, Song YS. Genotype and SNP calling from next-generation sequencing data. Nat Rev Genet. Nature Publishing Group; 2011;12: 443-451. [cited by applicant]
Ning Z, Cox AJ, Mullikin JC. SSAHA: a fast search method for large DNA databases. Genome Res. 2001;11: 1725-1729. [cited by applicant]
O'Fallon BD, Wooderchak-Donahue W, Crockett DK. A support vector machine for identification of single-nucleotide polymorphisms from next-generation sequencing data. Bioinformatics. 2013;29: 1361-1366. [cited by applicant]
O'Rawe J, Jiang T, Sun G, Wu Y, Wang W, Hu J, et al. Low concordance of multiple variant-calling pipelines: practical implications for exome and genome sequencing. Genome Med. 2013;5: 28:1-18. [cited by applicant]
O'Rawe JA, Ferson S, Lyon GJ. Accounting for uncertainty in DNA sequencing data. Trends Genet. Elsevier; 2015;31: 61-66. [cited by applicant]
Pabinger S, Dander A, Fischer M, Snajder R, Sperk M, Efremova M, et al. A survey of tools for variant analysis of next-generation genome sequencing data. Brief Bioinform. Oxford University Press; 2014;15: 256-278. [cited by applicant]
Pearson WR, Lipman DJ. Improved tools for biological sequence comparison. Proc Natl Acad Sci U S A. 1988;85: 2444-2448. [cited by applicant]
Pe'er I, de Bakker PIW, Mailer J, Yelensky R, Altshuler D, Daly MJ. Evaluating and improving power in whole-genome association studies using fixed marker sets. Nat Genet. 2006;38: 663-667. [cited by applicant]
Pirooznia M, Kramer M, Parla J, Goes FS, Potash JB, McCombie WR, et al. Validation and assessment of variant calling pipelines for next-generation sequencing. Hum Genomics. 2014;8: 14:1-10. [cited by applicant]
Pope BJ, Nguyen-Dumont T, Hammet F, Park DJ. Rover variant caller: read-pair overlap considerate variant-calling software applied to PCR-based massively parallel sequencing datasets. Source Code Biol Med. 2014;9: 3:1-5. [cited by applicant]
Porter J, Berkhahn J, Zhang L. A Comparative Analysis of Computational Indel Calling Pipelines for Next Generation Sequencing Data. Proceedings of the International Conference on Bioinformatics & Computational Biology (… [cited by applicant]
Quail MA, Smith M, Coupland P, Otto TD, Harris SR, Connor TR, et al. A tale of three next generation sequencing platforms: comparison of Ion Torrent, Pacific Biosciences and Illumina MiSeq sequencers. BMC Genomics. 2012… [cited by applicant]
Raphael B, Zhi D, Tang H, Pevzner P. A novel method for multiple alignment of sequences with repeated and shuffled elements. Genome Res. 2004;14: 2336-2346. [cited by applicant]
Rausch T, Zichner T, Schlattl A, Stutz AM, Benes V, Korbel JO. Delly: structural variant discovery by integrated paired-end and split-read analysis. Bioinformatics. Oxford Univ Press; 2012;28: i333-i339. [cited by applicant]
Reumers J, De Rijk P, Zhao H, Liekens A, Smeets D, Cleary J, et al. Optimized filtering reduces the error rate in detecting genomic variants by short-read sequencing. Nat Biotechnol. nature.com; 2011;30: 10 pages. [cited by applicant]
Rieber N, Zapatka M, Lasitschka B, Jones D, Northcott P, Hutter B, et al. Coverage bias and sensitivity of variant calling for four whole-genome sequencing technologies. PLoS One. journals.plos.org; 2013;8: e66621:11 pa… [cited by applicant]
Rimmer A, Phan H, Mathieson I, Iqbal Z, Twigg SRF, WGS500 Consortium, et al. Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications. Nat Genet. 2014;46: 9… [cited by applicant]
Robertson G, Schein J, Chiu R, Corbett R, Field M, Jackman SD, et al. De novo assembly and analysis of RNA-seq data. Nat Methods. 2010;7: 7 pages. [cited by applicant]
Ronquist F, Teslenko M, van der Mark P, Ayres DL, Darling A, Hohna S, et al. MrBayes 3.2: efficient Bayesian phylogenetic inference and model choice across a large model space. Syst Biol. 2012;61: 539-542. [cited by applicant]
Rothberg JM, Hinz W, Rearick TM, Schultz J, Mileski W, Davey M, et al. An integrated semiconductor device enabling non-optical genome sequencing. Nature. 2011;475: 348-352. [cited by applicant]
Ruffalo M, LaFramboise T, Koyuttirk M. Comparative analysis of algorithms for next-generation sequencing read alignment. Bioinformatics. Oxford Univ Press; 2011;27: 2790-2796. [cited by applicant]
San Lucas FA, Wang G, Scheet P, Peng B. Integrated annotation and analysis of genetic variants from next-generation sequencing studies with variant tools. Bioinformatics. Oxford Univ Press; 2012;28: 421-422. [cited by applicant]
Sato K, Mituyama T, Asai K, Sakakibara Y. Directed acyclic graph kernels for structural RNA analysis. BMC Bioinformatics. 2008;9: 318:12 pages. [cited by applicant]
Saunders CT, Wong WSW, Swamy S, Becq J, Murray LJ, Cheetham RK. Strelka: accurate somatic small-variant calling from sequenced tumor-normal sample pairs. Bioinformatics. Oxford Univ Press; 2012;28: 1811-1817. [cited by applicant]
Schneeberger K, Hagmann J, Ossowski S, Warthmann N, Gesing S, Kohlbacher O, et al. Simultaneous alignment of short reads against multiple genomes. Genome Biol. 2009;10: R98.1-R98.12. [cited by applicant]
Schwikowski B, Vingron M. Weighted sequence graphs: boosting iterated dynamic programming using locally suboptimal solutions. Discrete Appl Math. 2003;127: 95-117. [cited by applicant]
Shen Y, Wan Z, Coarfa C, Drabek R, Chen L, Ostrowski EA, et al. A SNP discovery method to assess variant allele probability from next-generation resequencing data. Genome Res. genome.cshlp.org; 2010;20: 273-280. [cited by applicant]
Shendure J, Ji H. Next-generation DNA sequencing. Nat Biotechnol. 2008;26: 1135-1145. [cited by applicant]
Slater GSC, Birney E. Automated generation of heuristics for biological sequence comparison. BMC Bioinformatics. 2005;6: 31:11 pages. [cited by applicant]
Smith TF, Waterman MS. Identification of common molecular subsequences. J Mol Biol. 1981;147: 4 pages. [cited by applicant]
Soni GV, Meller A. Progress toward ultrafast DNA sequencing using solid-state nanopores. Clin Chem. 2007;53: 1996-2001. [cited by applicant]
Sosa MX, Sivakumar IKA, Maragh S, Veeramachaneni V, Hariharan R, Parulekar M, et al. Next-generation sequencing of human mitochondrial reference genomes uncovers high heteroplasmy frequency. PLoS Comput Biol. 2012;8: e1… [cited by applicant]
Spencer DH, Tyagi M, Vallania F, Bredemeyer AJ, Pfeifer JD, Mitra RD, et al. Performance of common analysis methods for detecting low-frequency single nucleotide variants in targeted next-generation sequence data. J Mol… [cited by applicant]
Stephens M, Smith NJ, Donnelly P. A new statistical method for haplotype reconstruction from population data. Am J Hum Genet. 2001;68: 978-989. [cited by applicant]
Stewart C, Kural D, Stromberg MP, Walker JA, Konkel MK, Stutz AM, et al. A comprehensive map of mobile element insertion polymorphisms in humans. PLoS Genet. 2011;7: e1002236:1-19. [cited by applicant]
Szalkowski AM, Anisimova M. Graph-based modeling of tandem repeats improves global multiple sequence alignment. Nucleic Acids Res. 2013;41: e162:1-11. [cited by applicant]
Szalkowski AM. Fast and robust multiple sequence alignment with phylogeny-aware gap placement. BMC Bioinformatics. 2012;13: 129:1-11. [cited by applicant]
Talwalkar A, Liptrap J, Newcomb J, Hartl C, Terhorst J, Curtis K, et al. SMaSH: a benchmarking toolkit for human genome variant calling. Bioinformatics. Oxford Univ Press; 2014;30: 2787-2795. [cited by applicant]
Tan A, Abecasis GR, Kang HM. Unified representation of genetic variants. Bioinformatics. Oxford Univ Press; 2015;31: 2202-2204. [cited by applicant]
Thomas UG. Community-wide Effort Aims to Better Represent Variation in Human Reference Genome. Genome Web. Dec. 18, 2014. https://www.genomeweb.com/informatics/community-wide-effort-aims-better-r- epresent-variation-hum… [cited by applicant]
Tian S, Yan H, Kalmbach M, Slager SL. Impact of post-alignment processing in variant discovery from whole exome data. BMC Bioinformatics. 2016;17: 403:1-13. [cited by applicant]
Torri F, Dinov ID, Zamanyan A, Hobel S, Genco A, Petrosyan P, et al. Next generation sequence analysis and computational genomics using graphical pipeline workflows. Genes. 2012;3: 185 pages. [cited by applicant]
Vallania FLM, Druley TE, Ramos E, Wang J, Borecki I, Province M, et al. High-throughput discovery of rare insertions and deletions in large cohorts. Genome Res. 2010;20: 1711-1718. [cited by applicant]
Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, Levy-Moonshine A, et al. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline. Curr Protoc Bioinformatics.… [cited by applicant]
Vergara IA, Frech C, Chen N. Coo Var: co-occurring variant analyzer. BMC Res Notes. bmcresnotes.biomedcentral.com; 2012;5: 615:1-7. [cited by applicant]
Wallace IM, Blackshields G, Higgins DG. Multiple sequence alignments. Curr Opin Struct Biol. 2005;15: 261-266. [cited by applicant]
Wang W, Wei Z, Lam T-W, Wang J. Next generation sequencing has lower sequence coverage and poorer SNP-detection capability in the regulatory regions. Sci Rep. 2011;1: 55:1-7. [cited by applicant]
Wang Y, Lu J, Yu J, Gibbs RA, Yu F. An integrative variant analysis pipeline for accurate genotype/haplotype inference in population NGS data. Genome Res. genome.cshlp.org; 2013;23: 833-842. [cited by applicant]
Warden CD, Adamson AW, Neuhausen SL, Wu X. Detailed comparison of two popular variant calling packages for exome and targeted exon studies. PeerJ. peerj.com; 2014;2: e600:1-27. [cited by applicant]
Waterman MS, Smith TF, Beyer WA. Some biological sequence metrics. Adv Math. 1976;20: 367-387. [cited by applicant]
Wei Z, Wang W, Hu P, Lyon GJ, Hakonarson H. SNVer: a statistical tool for variant calling in analysis of pooled or individual next-generation sequencing data. Nucleic Acids Res. Oxford Univ Press; 2011;39: e132:1-13. [cited by applicant]
Wilm A, Aw PPK, Bertrand D, Yeo GHT, Ong SH, Wong CH, et al. LoFreq: a sequence-quality aware, ultra-sensitive variant caller for uncovering cell-population heterogeneity from high-throughput sequencing datasets. Nuclei… [cited by applicant]
Wu TD, Nacu S. Fast and SNP-tolerant detection of complex variants and splicing in short reads. Bioinformatics. 2010;26: 873-881. [cited by applicant]
Yang W-Y, Hormozdiari F, Wang Z, He D, Pasaniuc B, Eskin E. Leveraging reads that span multiple single nucleotide polymorphisms for haplotype inference from sequencing data. Bioinformatics. 2013;29: 2245-2252. [cited by applicant]
Yanovsky V, Rumble SM, Brudno M. Read Mapping Algorithms for Single Molecule Sequencing Data. In: Crandall KA, Lagergren J, editors. Algorithms in Bioinformatics. Berlin, Heidelberg: Springer Berlin Heidelberg; 2008. 12… [cited by applicant]
Yau C, Papaspiliopoulos O, Roberts GO, Holmes C. Bayesian Nonparametric Hidden Markov Models with application to the analysis of copy-number-variation in mammalian genomes. J R Stat Soc Series B Stat Methodol. Wiley Onl… [cited by applicant]
You N, Murillo G, Su X, Zeng X, Xu J, Ning K, et al. SNP calling using genotype model selection on high-throughput sequencing data. Bioinformatics. Oxford Univ Press; 2012;28: 643-650. [cited by applicant]
Yu X, Sun S. Comparing a few SNP calling algorithms using low-coverage sequencing data. BMC Bioinformatics. 2013;14: 274:15 pages. [cited by applicant]
Zhang K, Qin Z, Chen T, Liu JS, Waterman MS, Sun F. HapBlock: haplotype block partitioning and tag SNP selection software using a set of dynamic programming algorithms. Bioinformatics. Oxford Univ Press; 2005;21: 131-13… [cited by applicant]
Zhang P, Tan G, Gao GR. Implementation of the Smith-Waterman algorithm on a reconfigurable supercomputing platform. Proceedings of the 1st international workshop on High-performance reconfigurable computing technology a… [cited by applicant]
Zhao Z, Wang W, Wei Z. An empirical Bayes testing procedure for detecting variants in analysis of next generation sequencing data. Ann Appl Stat. Institute of Mathematical Statistics; 2013;7: 20 pages. [cited by applicant]
Zook JM, Chapman B, Wang J, Mittelman D, Hofmann O, Hide W, et al. Integrating human sequence data sets provides a resource of benchmark SNP and indel genotype calls. Nat Biotechnol. 2014;32: 8 pages. [cited by applicant]