IP Library Granted Patent US 12,633,378
Granted Patent B2
US 12,633,378 · App. 18/812,351 · Granted May 19, 2026

Methods and systems for detecting sequence variants

Inventor: Deniz Kural (Somerville, MA)
Assignee: Seven Bridges Genomics Inc.
G16B30/10G16B30/00G16B30/20G16B50/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,378
App. No.
18/812,351
Granted
May 19, 2026
Kind
B2
Abstract

The invention provides methods for identifying rare variants near a structural variation in a genetic sequence, for example, in a nucleic acid sample taken from a subject. The invention additionally includes methods for aligning reads (e.g., nucleic acid reads) to a reference sequence construct accounting for the structural variation, methods for building a reference sequence construct accounting for the structural variation or the structural variation and the rare variant, and systems that use the alignment methods to identify rare variants. The method is scalable, and can be used to align millions of reads to a construct thousands of bases long, or longer.

Claims (59)

1 . A method for identifying a genetic variant as present in a biological sample from a subject, the method comprising:

using at least one processor to perform:

accessing at least one data structure representing a genomic reference graph, the genomic reference graph representing at least 1,000,000 nucleic acids and comprising nodes and edges connecting the nodes, the nodes including:

a first node representing a first nucleotide sequence stored as a first string of symbols, and

parent nodes of the first node, the parent nodes including a second node representing the genetic variant;

aligning one or more sequence reads to the genomic reference graph using the at least one data structure and a dynamic programming algorithm, the aligning comprising, for each particular sequence read of the one or more sequence reads:

determining scores for entries in a first matrix associated with the first node, the first matrix representing a comparison between the particular sequence read and the first string of symbols, the determining comprising:

determining whether a symbol of the particular sequence read matches a first symbol of the first string of symbols;

accessing a score from matrices associated with the parent nodes of the first node; and

determining a score for an entry in the first matrix based on: (i) a result of determining whether the symbol of the particular sequence read matches the first symbol of the first string of symbols, and (ii) the score from the matrices associated with the parent nodes; and

aligning the particular sequence read to the genomic reference graph based on the determined scores;

identifying the genetic variant as present in the biological sample using results of aligning the one or more sequence reads to the genomic reference graph representing the at least 1,000,000 nucleic acids; and

generating an output indicative of a result of identifying the genetic variant as present in the biological sample.

2 . The method of claim 1 , wherein identifying the genetic variant as present in the biological sample using the results of aligning the one or more sequence reads to the genomic reference graph comprises:

identifying the genetic variant as present in the biological sample when a sequence read of the one or more sequence reads aligns, at least partially, to the second node.

3 . The method of claim 1 , wherein the parent nodes include a third node at a same position in the genomic reference graph as the second node, the third node representing a nucleotide sequence lacking the genetic variant.

4 . The method of claim 1 , wherein the genetic variant comprises a deletion, duplication, copy-number variation, insertion, inversion, and/or translocation.

5 . The method of claim 1 , further comprising identifying a second genetic variant using the result of identifying the genetic variant as present in the biological sample, wherein the second genetic variant is not represented by the genomic reference graph.

6 . The method of claim 5 , wherein the second genetic variant comprises a single nucleotide polymorphism.

7 . The method of claim 1 , further comprising sequencing the biological sample to obtain the one or more sequence reads.

8 . At least one non-transitory storage medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method for identifying a genetic variant as present in a biological sample from a subject, the method comprising:

accessing at least one data structure representing a genomic reference graph, the genomic reference graph representing at least 1,000,000 nucleic acids and comprising nodes and edges connecting the nodes, the nodes including:

a first node representing a first nucleotide sequence stored as a first string of symbols, and

parent nodes of the first node, the parent nodes including a second node representing the genetic variant;

aligning one or more sequence reads to the genomic reference graph using the at least one data structure and a dynamic programming algorithm, the aligning comprising, for each particular sequence read of the one or more sequence reads:

determining scores for entries in a first matrix associated with the first node, the first matrix representing a comparison between the particular sequence read and the first string of symbols, the determining comprising:

determining whether a symbol of the particular sequence read matches a first symbol of the first string of symbols;

accessing a score from matrices associated with the parent nodes of the first node; and

determining a score for an entry in the first matrix based on: (i) a result of determining whether the symbol of the particular sequence read matches the first symbol of the first string of symbols, and (ii) the score from the matrices associated with the parent nodes; and

aligning the particular sequence read to the genomic reference graph based on the determined scores;

identifying the genetic variant as present in the biological sample based on results of aligning the one or more sequence reads to the genomic reference graph representing the at least 1,000,000 nucleic acids; and

generating an output indicative of a result of identifying the genetic variant as present in the biological sample.

9 . The at least one non-transitory storage medium of claim 8 , wherein identifying the genetic variant as present in the biological sample using the results of aligning the one or more sequence reads to the genomic reference graph comprises:

identifying the genetic variant as present in the biological sample when a sequence read of the one or more sequence reads aligns, at least partially, to the second node.

10 . The at least one non-transitory storage medium of claim 8 , wherein the parent nodes include a third node at a same position in the genomic reference graph as the second node, the third node representing a nucleotide sequence lacking the genetic variant.

11 . The at least one non-transitory storage medium of claim 8 , wherein the genetic variant comprises a deletion, duplication, copy-number variation, insertion, inversion, and/or translocation.

12 . The at least one non-transitory storage medium of claim 8 , wherein the method further comprises identifying a second genetic variant using the result of identifying the genetic variant as present in the biological sample, wherein the second genetic variant is not represented by the genomic reference graph.

13 . The at least one non-transitory storage medium of claim 12 , wherein the second genetic variant comprises a single nucleotide polymorphism.

14 . The at least one non-transitory storage medium of claim 8 , wherein the method further comprises sequencing the biological sample to obtain the one or more sequence reads.

15 . A system, comprising:

at least one processor; and

at least one non-transitory storage medium storing processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to perform a method for identifying a genetic variant as present in a biological sample from a subject, the method comprising:

accessing at least one data structure representing a genomic reference graph, the genomic reference graph representing at least 1,000,000 nucleic acids and comprising nodes and edges connecting the nodes, the nodes including:

a first node representing a first nucleotide sequence stored as a first string of symbols, and

parent nodes of the first node, the parent nodes including a second node representing the genetic variant;

aligning one or more sequence reads to the genomic reference graph using the at least one data structure and a dynamic programming algorithm, the aligning comprising, for each particular sequence read of the one or more sequence reads:

determining scores for entries in a first matrix associated with the first node, the first matrix representing a comparison between the particular sequence read and the first string of symbols, the determining comprising:

determining whether a symbol of the particular sequence read matches a first symbol of the first string of symbols;

accessing a score from matrices associated with the parent nodes of the first node; and

determining a score for an entry in the first matrix based on: (i) a result of determining whether the symbol of the particular sequence read matches the first symbol of the first string of symbols, and (ii) the score from the matrices associated with the parent nodes; and

aligning the particular sequence read to the genomic reference graph based on the determined scores;

identifying the genetic variant as present in the biological sample based on results of aligning the one or more sequence reads to the genomic reference graph representing the at least 1,000,000 nucleic acids; and

generating an output indicative of a result of identifying the genetic variant as present in the biological sample.

16 . The system of claim 15 , wherein identifying the genetic variant as present in the biological sample using the results of aligning the one or more sequence reads to the genomic reference graph comprises:

identifying the genetic variant as present in the biological sample when a sequence read of the one or more sequence reads aligns, at least partially, to the second node.

17 . The system of claim 15 , wherein the genetic variant comprises a deletion, duplication, copy-number variation, insertion, inversion, and/or translocation.

18 . The system of claim 15 , wherein the method further comprises identifying a second genetic variant using the result of identifying the genetic variant as present in the biological sample, wherein the second genetic variant is not represented by the genomic reference graph.

19 . The system of claim 18 , wherein the second genetic variant comprises a single nucleotide polymorphism.

20 . The system of claim 15 , wherein the method further comprises sequencing the biological sample to obtain the one or more sequence reads.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2025
From: KURAL, DENIZ
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 070140/0059 →
Continuity (10)
Continuation 18494317 · Oct 25, 2023
Continuation 17933260 · Sep 19, 2022
Continuation 16443402 · Jun 17, 2019
Continuation 15906404 · Feb 27, 2018
Continuation 15196345 · Jun 29, 2016
Continuation 14811057 · Jul 28, 2015
Continuation 14041850 · Sep 30, 2013
Provisional Application 61884380 · Sep 30, 2013
Provisional Application 61868249 · Aug 21, 2013
Related Publication 20250182851A1 · Jun 5, 2025
References Cited (279)
US 5511158A · Sims · 1996 [cited by applicant]
US 5701256A · Marr et al. · 1997 [cited by applicant]
US 6054278A · Dodge et al. · 2000 [cited by applicant]
US 6223128B1 · Allex et al. · 2001 [cited by applicant]
US 7577554B2 · Lystad et al. · 2009 [cited by applicant]
US 7580918B2 · Chang et al. · 2009 [cited by applicant]
US 7809509B2 · Milosavljevic · 2010 [cited by applicant]
US 7885840B2 · Sadiq et al. · 2011 [cited by applicant]
US 7917302B2 · Rognes · 2011 [cited by applicant]
US 8209130B1 · Kennedy et al. · 2012 [cited by applicant]
US 8340914B2 · Gatewood et al. · 2012 [cited by applicant]
US 8370079B2 · Sorenson et al. · 2013 [cited by applicant]
US 8639847B2 · Blaszczak et al. · 2014 [cited by applicant]
US 9063914B2 · Kural et al. · 2015 [cited by applicant]
US 9092402B2 · Kural et al. · 2015 [cited by applicant]
US 9116866B2 · Kural · 2015 [cited by applicant]
US 9390226B2 · Kural · 2016 [cited by applicant]
US 9817944B2 · Kural · 2017 [cited by applicant]
US 9904763B2 · Kural · 2018 [cited by applicant]
US 10325675B2 · Kural · 2019 [cited by applicant]
US 11488688B2 · Kural · 2022 [cited by applicant]
US 11837328B2 · Kural · 2023 [cited by applicant]
US 12106826B2 · Kural · 2024 [cited by applicant]
US 20040023209A1 · Jonasson · 2004 [cited by applicant]
US 20050089906A1 · Furuta et al. · 2005 [cited by applicant]
US 20060292611A1 · Berka et al. · 2006 [cited by applicant]
US 20070166707A1 · Schadt et al. · 2007 [cited by applicant]
US 20080077607A1 · Gatawood et al. · 2008 [cited by applicant]
US 20080294403A1 · Zhu et al. · 2008 [cited by applicant]
US 20090119313A1 · Pearce · 2009 [cited by applicant]
US 20090164135A1 · Brodzik et al. · 2009 [cited by applicant]
US 20090300781A1 · Bancroft et al. · 2009 [cited by applicant]
US 20100041048A1 · Diehl et al. · 2010 [cited by applicant]
US 20100169026A1 · Sorenson et al. · 2010 [cited by applicant]
US 20110004413A1 · Carnevali et al. · 2011 [cited by applicant]
US 20110096193A1 · Egawa · 2011 [cited by applicant]
US 20110098193A1 · Kingsmore et al. · 2011 [cited by applicant]
US 20110257889A1 · Klammer · 2011 [cited by examiner]
US 20110295514A1 · Breu et al. · 2011 [cited by applicant]
US 20120041727A1 · Mishra et al. · 2012 [cited by applicant]
US 20120045771A1 · Beier et al. · 2012 [cited by applicant]
US 20120239706A1 · Steinfadt · 2012 [cited by applicant]
US 20120330566A1 · Chaisson · 2012 [cited by examiner]
US 20130059738A1 · Leamon et al. · 2013 [cited by applicant]
US 20130059740A1 · Drmanac et al. · 2013 [cited by applicant]
US 20130073214A1 · Hyland et al. · 2013 [cited by applicant]
US 20130103320A1 · Dzakula · 2013 [cited by examiner]
US 20130124100A1 · Drmanac et al. · 2013 [cited by applicant]
US 20130289099A1 · Goff et al. · 2013 [cited by applicant]
US 20130311106A1 · White et al. · 2013 [cited by applicant]
US 20140025312A1 · Chin et al. · 2014 [cited by applicant]
US 20140051588A9 · Drmanac et al. · 2014 [cited by applicant]
US 20140052381A1 · Utiramerur · 2014 [cited by examiner]
US 20140066317A1 · Talasaz · 2014 [cited by applicant]
US 20140100792A1 · Deciu et al. · 2014 [cited by applicant]
US 20140121116A1 · Richards · 2014 [cited by examiner]
US 20140136120A1 · Colwell et al. · 2014 [cited by applicant]
US 20140180594A1 · Kim et al. · 2014 [cited by applicant]
US 20140200147A1 · Bartha et al. · 2014 [cited by applicant]
US 20140278590A1 · Abbassi et al. · 2014 [cited by applicant]
US 20140280360A1 · Webber et al. · 2014 [cited by applicant]
US 20140323320A1 · Jia et al. · 2014 [cited by applicant]
US 20150056613A1 · Kural · 2015 [cited by applicant]
US 20150057946A1 · Kural · 2015 [cited by applicant]
US 20150094212A1 · Gottimukkala et al. · 2015 [cited by applicant]
US 20150110754A1 · Bai et al. · 2015 [cited by applicant]
US 20150112602A1 · Kural et al. · 2015 [cited by applicant]
US 20150112658A1 · Kural et al. · 2015 [cited by applicant]
US 20150197815A1 · Kural · 2015 [cited by applicant]
US 20150199472A1 · Kural · 2015 [cited by applicant]
US 20150199473A1 · Kural · 2015 [cited by applicant]
US 20150199474A1 · Kural · 2015 [cited by applicant]
US 20150199475A1 · Kural · 2015 [cited by applicant]
US 20150227685A1 · Kural · 2015 [cited by applicant]
US 20150293994A1 · Kelly · 2015 [cited by applicant]
US 20150302145A1 · Kural et al. · 2015 [cited by applicant]
US 20150310167A1 · Kural et al. · 2015 [cited by applicant]
US 20150344970A1 · Vogelstein et al. · 2015 [cited by applicant]
US 20150347678A1 · Kural · 2015 [cited by applicant]
US 20150356147A1 · Mishra et al. · 2015 [cited by applicant]
US 20160259880A1 · Semenyuk · 2016 [cited by applicant]
US 20160306921A1 · Kural · 2016 [cited by applicant]
US 20160364523A1 · Locke et al. · 2016 [cited by applicant]
US 20170053341A1 · Locke et al. · 2017 [cited by applicant]
US 20170058320A1 · Locke et al. · 2017 [cited by applicant]
US 20170058365A1 · Locke et al. · 2017 [cited by applicant]
US 20170193351A1 · Noyes et al. · 2017 [cited by applicant]
US 20170199959A1 · Locke · 2017 [cited by applicant]
US 20170199960A1 · Ghose et al. · 2017 [cited by applicant]
US 20170242958A1 · Brown · 2017 [cited by applicant]
US 20180336314A1 · Kural · 2018 [cited by applicant]
US 20200168295A1 · Kural · 2020 [cited by applicant]
US 20230044434A1 · Kural · 2023 [cited by applicant]
US 20240062850A1 · Kural · 2024 [cited by applicant]
CA 2869574A1 · 2013 [cited by applicant]
EP 3053073B1 · 2019 [cited by applicant]
WO WO2012096579A2 · 2012 [cited by applicant]
WO WO2012098515A1 · 2012 [cited by applicant]
WO WO2012142531A2 · 2012 [cited by applicant]
WO WO2013151803A1 · 2013 [cited by applicant]
WO WO2015027050A1 · 2015 [cited by applicant]
WO WO2015048753A1 · 2015 [cited by applicant]
WO WO2015058093A1 · 2015 [cited by applicant]
WO WO2015058095A1 · 2015 [cited by applicant]
WO WO2015058097A1 · 2015 [cited by applicant]
WO WO2015058120A1 · 2015 [cited by applicant]
WO WO2015061099A1 · 2015 [cited by applicant]
WO WO2015061103A1 · 2015 [cited by applicant]
WO WO2015105963A1 · 2015 [cited by applicant]
WO WO2015123269A1 · 2015 [cited by applicant]
WO WO2016141294A1 · 2016 [cited by applicant]
WO WO2016201215A1 · 2016 [cited by applicant]
WO WO2017120128A1 · 2017 [cited by applicant]
WO WO2017123864A1 · 2017 [cited by applicant]
WO WO2017147124A1 · 2017 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/061158 mailed Feb. 4, 2015. [cited by applicant]
Extended European Search Report issued in EP 14847490.1 dated May 9, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/058328, mailed Dec. 30, 2014. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/061198, mailed Feb. 4, 2015. [cited by applicant]
Communication pursuant to Article 94(3) EPC for European Application No. 14803268.3 dated Apr. 21, 2017. [cited by applicant]
Written Opinion issued in SG 11201603039P dated Jun. 12, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/061162 mailed Mar. 19, 2015. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2016/057324 mailed Jan. 10, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2016/036873 mailed Sep. 7, 2016. [cited by applicant]
Extended European Search Report for European Application No. 14854801.9 dated Apr. 12, 2017. [cited by applicant]
Written Opinion issued in SG 11201602903X dated May 29, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/061156 mailed Feb. 17, 2015. [cited by applicant]
Extended European Search Report for European Application No. 14837955.5 dated Mar. 29, 2017. [cited by applicant]
Written Opinion issued in SG 11201601124Y dated Dec. 21, 2016. [cited by applicant]
Written Opinion issued in SG 11201601124Y dated Mar. 1, 2018. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/052065 mailed on Dec. 11, 2014. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/060680 mailed Jan. 27, 2015. [cited by applicant]
Written Opinion issued in SG 11201603044S dated Sep. 10, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2014/060690 mailed Feb. 10, 2015. [cited by applicant]
Written Opinion issued in SG 11201605506Q dated Jun. 15, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2015/010604 mailed Mar. 31, 2015. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2015/015375 mailed May 11, 2015. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2016/020899 mailed May 5, 2016. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2017/013329 mailed on Apr. 7, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2017/012015 mailed Apr. 19, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2017/018830 mailed Aug. 31, 2017. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2016/033201 mailed Sep. 2, 2016. [cited by applicant]
Agarwal et al., Sinnet: Social interaction network extractor from text. The Companion Volume of the Proceedings of IJCNLP 2013: System Demonstrations Oct. 2013;33-6. [cited by applicant]
Aguiar et al., HapCompass: a fast cycle basis algorithm for accurate haplotype assembly of sequence data. Journal of Computational Biology. Jun. 1, 2012;19(6):577-90. [cited by applicant]
Aguiar et al., Haplotype assembly in polyploid genomes and identical by descent shared tracts. Bioinformatics. Jul. 1, 2013;29(13):1352-60. [cited by applicant]
Airoldi et al., Mixed membership stochastic blockmodels. Advances in neural information processing systems. 2008;21. [cited by applicant]
Albers et al., Dindel: accurate indel calls from short-read data. Genome research. Jun. 1, 2011;21(6):961-73. [cited by applicant]
Alioto et al., A comprehensive assessment of somatic mutation detection in cancer using whole-genome sequencing. Nature communications. Dec. 9, 2015;6(1):1-3. [cited by applicant]
Altera, Implementation of the Smith-Waterman algorithm on reconfigurable supercomputing platform, White Paper ver 1.0 (18 pages). 2007. [cited by applicant]
Altschul et al., Optimal sequence alignment using affine gap costs. Bulletin of mathematical biology. Sep. 1, 1986;48(5-6):603-16. [cited by applicant]
Bansal et al., An MCMC algorithm for haplotype assembly from whole-genome sequence data. Genome research. Aug. 1, 2008;18(8):1336-46. [cited by applicant]
Bao et al., BRANCH: boosting RNA-Seq assemblies with partial or related genomic sequences. Bioinformatics. May 15, 2013;29(10):1250-9. [cited by applicant]
Barbieri et al., Exome sequencing identifies recurrent SPOP, FOXA1 and MED12 mutations in prostate cancer. Nature genetics. Jun. 2012;44(6):685-9. [cited by applicant]
Beerenwinkel et al., Conjunctive bayesian networks. Bernoulli. Nov. 1, 2007:893-909. [cited by applicant]
Berlin et al., Assembling Large Genomes with Single-Molecule Sequencing and Locality Sensitive Hashing. 2014. bioRxiv: 008003. [cited by applicant]
Bertrand et al., Genetic map refinement using a comparative genomic approach. Journal of Computational Biology. Oct. 1, 2009;16(10):1475-86. [cited by applicant]
Black, A simple answer for a splicing conundrum. Proceedings of the National Academy of Sciences. Apr. 5, 2005;102(14):4927-8. [cited by applicant]
Bondy et al., Graph theory with applications. London: Macmillan; Jun. 1, 1976;1-115. [cited by applicant]
Boyer et al., A fast string searching algorithm. Communications of the ACM. Oct. 1, 1977;20(10):762-72. [cited by applicant]
Browning et al., Haplotype phasing: existing methods and new developments. Nature Reviews Genetics. Oct. 2011;12(10):703-14. [cited by applicant]
Caboche et al., Comparison of mapping algorithms used in high-throughput sequencing: application to Ion Torrent data. BMC genomics. Dec. 2014;15(1):1-6. [cited by applicant]
Cartwright, DNA assembly with gaps (Dawg): simulating sequence evolution. Bioinformatics. Nov. 1, 2005;21(Suppl_3):iii31-8. [cited by applicant]
Chang et al., The application of alternative splicing graphs in quantitative analysis of alternative splicing form from EST database. International Journal of Computer Applications in Technology. Apr. 1, 2005;22(1):14-2… [cited by applicant]
Chen et al., Transient hypermutability, chromothripsis and replication-based mechanisms in the generation of concurrent clustered mutations. Mutation Research/Reviews in Mutation Research. Jan. 1, 2012;750(1):52-9. [cited by applicant]
Chin et al., Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nature methods. Jun. 2013;10(6):563-9. [cited by applicant]
Chuang et al., Gene recognition based on DAG shortest paths. Bioinformatics. Jun. 1, 2001;17(suppl_1):S56-64. [cited by applicant]
Compeau et al., How to apply de Bruijn graphs to genome assembly. Nature biotechnology. Nov. 2011;29(11):987-91. [cited by applicant]
Craig, Ordering of cosmid clones covering the Herpes simplex virus type I (HSV-I) genome: a test case for fingerprinting by hybridisation, Nucleic Acids Research 18:9.1990:2653-2660. [cited by applicant]
Danecek et al., The variant call format and VCFtools. Bioinformatics. Aug. 1, 2011;27(15):2156-8. [cited by applicant]
Denœud et al., Identification of polymorphic tandem repeats by direct comparison of genome sequence from different bacterial strains: a web-based resource. BMC bioinformatics. Dec. 2004;5(1):1-2. [cited by applicant]
Depristo et al., A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature genetics. May 2011;43(5):491-8. [cited by applicant]
Duan et al., Optimizing de novo common wheat transcriptome assembly using short-read RNA-Seq data. BMC genomics. Dec. 2012;13(1):1-2. [cited by applicant]
Dudley et al., A quick guide for developing effective bioinformatics programming skills. PLOS computational biology. Dec. 24, 2009;5(12):e1000589. [cited by applicant]
Durbin, Efficient haplotype matching and storage using the positional Burrows-Wheeler transform (PBWT). Bioinformatics. May 1, 2014;30(9):1266-72. [cited by applicant]
Endelman, New algorithm improves fine structure of the barley consensus SNP map. BMC genomics. Dec. 2011;12(1):1-9. [cited by applicant]
Farrar, Striped Smith-Waterman speeds database searches six times over other SIMD implementations. Bioinformatics. Jan. 15, 2007;23(2):156-61. [cited by applicant]
Fitch, Distinguishing homologous from analogous proteins. Systematic zoology. Jun. 1, 1970;19(2):99-113. [cited by applicant]
Flicek et al., Sense from sequence reads: methods for alignment and assembly. Nature methods. Nov. 2009;6(11):S6-12. [cited by applicant]
Garber et al., Computational methods for transcriptome annotation and quantification using RNA-seq. Nature methods. Jun. 2011;8(6):469-77. [cited by applicant]
Gerlinger et al., Intratumor heterogeneity and branched evolution revealed by multiregion sequencing. N Engl j Med. Mar. 8, 2012;366:883-92. [cited by applicant]
Golub et al., Molecular classification of cancer: class discovery and class prediction by gene expression monitoring. science. Oct. 15, 1999;286(5439):531-7. [cited by applicant]
Gotoh, An improved algorithm for matching biological sequences. Journal of molecular biology. Dec. 15, 1982;162(3):705-8. [cited by applicant]
Gotoh, Multiple sequence alignment: algorithms and applications. Advances in biophysics. Jan. 1, 1999;36:159-206. [cited by applicant]
Grabherr et al., Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nature biotechnology. Jul. 2011;29(7):644-52. [cited by applicant]
Grasso et al., Combining partial order alignment and progressive multiple sequence alignment increases alignment speed and scalability to very large alignment problems. Bioinformatics. Jul. 1, 2004;20(10):1546-56. [cited by applicant]
Guttman et al., Ab initio reconstruction of cell type-specific transcriptomes in mouse reveals the conserved multi-exonic structure of lincRNAs. Nature biotechnology. May 2010;28(5):503-10. [cited by applicant]
Guttman, Ab initio reconstruction of transcriptomes of pluripotent and lineage committed cells reveals gene structures of thousands of lincRNAs. Author Manuscript; available in PMC Nov. 2, 2010. 24 pages. Published in f… [cited by applicant]
Haas et al., DAGchainer: a tool for mining segmental genome duplications and synteny. Bioinformatics. Dec. 12, 2004;20(18):3643-6. [cited by applicant]
Harrow et al., GENCODE: the reference human genome annotation for The ENCODE Project. Genome research. Sep. 1, 2012;22(9):1760-74. [cited by applicant]
He et al., Optimal algorithms for haplotype assembly from whole-genome sequence data. Bioinformatics. Jun. 15, 2010;26(12):i183-90. [cited by applicant]
Heber et al., Splicing graphs and EST assembly problem. Bioinformatics. Jul. 1, 2002;18(suppl_1):S181-8. [cited by applicant]
Hein, A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given. Molecular Biology and Evolution. Nov. 1, 1989;6(6):649-68. [cited by applicant]
Hein, A tree reconstruction method that is economical in the number of pairwise comparisons used. Molecular biology and evolution. Nov. 1, 1989;6(6):649-68. [cited by applicant]
Homer et al., Improved variant discovery through local re-alignment of short-read next-generation sequencing data using SRMA. Genome biology. Oct. 2010;11(10):1-2. [cited by applicant]
Horspool, Practical fast searching in strings. Software: Practice and Experience. Jun. 1980;10(6):501-6. [cited by applicant]
Huang, 3: Bio-Sequence Comparison and Alignment, ser. Curr Top Comp Mol Biol. Cambridge, Mass.: The MIT Press. 2002:45-69. [cited by applicant]
Hutchinson et al., Allele-specific methylation occurs at genetic variants associated with complex disease. PloS one. Jun. 9, 2014;9(6):e98464. [cited by applicant]
Kano et al., Text mining meets workflow: linking U-Compare with Taverna. Bioinformatics. Oct. 1, 2010;26(19):2486-7. [cited by applicant]
Kehr et al., Genome alignment with graph data structures: a comparison. BMC bioinformatics. Dec. 2014;15(1):99. [cited by applicant]
Kent, BLAT—the BLAST-like alignment tool. Genome research. Apr. 1, 2002;12(4):656-64. [cited by applicant]
Kim et al., A scaffold analysis tool using mate-pair information in genome sequencing. Journal of Biomedicine and Biotechnology. Jan. 2008;8(3):195-197. [cited by applicant]
Kim et al., ECgene: genome-based EST clustering and gene modeling for alternative splicing. Genome research. Apr. 1, 2005;15(4):566-76. [cited by applicant]
Kim et al., TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. Genome biology. Apr. 2013;14(4):R36. [cited by applicant]
Koolen et al., Clinical and molecular delineation of the 17q21. 31 microdeletion syndrome. Journal of medical genetics. Nov. 1, 2008;45(11):710-20. [cited by applicant]
Kumar et al., Comparing de novo assemblers for 454 transcriptome data. BMC genomics. Dec. 2010;11(1):571. [cited by applicant]
Kurtz et al., Versatile and open software for comparing large genomes. Genome biology. Jan. 2004;5(2):R12. [cited by applicant]
Lam et al., 2008, Compressed indexing and local alignment of DNA. Bioinformatics 24(6):791-97. [cited by applicant]
Langmead et al., Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome biology. Mar. 2009;10(3):1-0. [cited by applicant]
Larkin et al., Clustal W and Clustal X version 2.0. bioinformatics. Nov. 1, 2007;23(21):2947-8. [cited by applicant]
Lecca et al., Defining order and timing of mutations during cancer progression: the TO-DAG probabilistic graphical model. Frontiers in genetics. Oct. 13, 2015;6:309. [cited by applicant]
Lee et al., Accurate read mapping using a graph-based human pan-genome. American Society of Human Genetics 64th Annual Meeting Platform Abstracts. May 2015; Abstract 41. [cited by applicant]
Lee et al., Bioinformatics analysis of alternative splicing. Briefings in bioinformatics. Mar. 1, 2005;6(1):23-33. [cited by applicant]
Lee et al., Multiple sequence alignment using partial order graphs. Bioinformatics. Mar. 1, 2002;18(3):452-64. [cited by applicant]
Lee, Generating consensus sequences from partial order multiple sequence alignment graphs. Bioinformatics 2003;19(8):999-1008. [cited by applicant]
Legault et al., Inference of alternative splicing from RNA-Seq data with probabilistic splice graphs. Bioinformatics. Sep. 15, 2013;29(18):2300-10. [cited by applicant]
LeGault, 2010, Learning Probalistic Splice Graphs from RNA-Seq data, pages.cs.wisc.edu/.about.legault/cs760_writeup.pdf; retrieved from the internet on Apr. 6, 2014. [cited by applicant]
Leipzig et al., The Alternative Splicing Gallery (ASG): bridging the gap between genome and transcriptome. Nucleic Acids Research. Jan. 1, 2004;32(13):3977-83. [cited by applicant]
Li et al., A survey of sequence alignment algorithms for next-generation sequencing. Briefings in bioinformatics. Sep. 1, 2010;11(5):473-83. [cited by applicant]
Li et al., Fast and accurate short read alignment with Burrows—Wheeler transform. bioinformatics. Jul. 15, 2009;25(14):1754-60. [cited by applicant]
Li et al., Mapping short DNA sequencing reads and calling variants using mapping quality scores. Genome research. Nov. 1, 2008;18(11):1851-8. [cited by applicant]
Lipman et al., Rapid and sensitive protein similarity searches. Science. Mar. 22, 1985;227(4693):1435-41. [cited by applicant]
Lücking et al., PICS-Ord: unlimited coding of ambiguous regions by pairwise identity and cost scores ordination. BMC bioinformatics. Dec. 2011;12(1):1-5. [cited by applicant]
Lupski et al., Genomic disorders: molecular mechanisms for rearrangements and conveyed phenotypes. PLoS genetics. Dec. 2005;1(6):e49. [cited by applicant]
Ma et al., Multiple genome alignment based on longest path in directed acyclic graphs. International journal of bioinformatics research and applications. Jan. 1, 2010;6(4):366-83. [cited by applicant]
Mamoulis et al., Non-contiguous sequence pattern queries. InInternational Conference on Extending Database Technology Mar. 14, 2004;783-800. Springer, Berlin, Heidelberg. [cited by applicant]
Marth et al., A general approach to single-nucleotide polymorphism discovery. Nature genetics. Dec. 1999;23(4):452-6. [cited by applicant]
Mazrouee et al., FastHap: fast and accurate single individual haplotype reconstruction using fuzzy conflict graphs. Bioinformatics. Sep. 1, 2014;30(17):1371-8. [cited by applicant]
Mcsherry, Spectral partitioning of random graphs. Proceedings 42nd IEEE Symposium on Foundations of Computer Science Oct. 8, 2001;529-37. [cited by applicant]
Miller et al., Assembly algorithms for next-generation sequencing data. Genomics. Jun. 1, 2010;95(6):315-27. [cited by applicant]
Mount, Multiple Sequence Alignment. Bioinformatics: Sequence and Genome Analysis. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York. Mar. 2001. Chapter 4;139-204. [cited by applicant]
Mourad et al., A hierarchical Bayesian network approach for linkage disequilibrium modeling and data-dimensionality reduction prior to genome-wide association studies. BMC bioinformatics. Dec. 2011;12(1):1-20. [cited by applicant]
Myers, The fragment assembly string graph. Bioinformatics. Jan. 1, 2005;21(suppl_2):ii79-85. [cited by applicant]
Nagarajan et al., Sequence assembly demystified. Nature Reviews Genetics. Mar. 2013;14(3):157-67. [cited by applicant]
Nakao et al., Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals. Nucleic Acids Research. Jan. 1, 2005;33(8):2355-63. [cited by applicant]
Needleman et al., A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journal of molecular biology. Mar. 28, 1970;48(3):443-53. [cited by applicant]
Newman et al., An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage. Nature medicine. May 2014;20(5):1-11. [cited by applicant]
Newman, Community detection and graph partitioning. arXiv preprint arXiv:1305.4974. May 21, 2013. 5 pages. [cited by applicant]
Olsson et al., Serial monitoring of circulating tumor DNA in patients with primary breast cancer for detection of occult metastatic disease. EMBO molecular medicine. Aug. 2015;7(8):1034-47. [cited by applicant]
Oshlack et al., From RNA-seq reads to differential expression results. Genome biology. Dec. 2010;11(12):1-0. [cited by applicant]
Parks et al., Detecting non-allelic homologous recombination from high-throughput sequencing data. Genome biology. Dec. 2015;16(1):1-9. [cited by applicant]
Peixoto, Efficient Monte Carlo and greedy heuristic for the inference of stochastic block models. Physical Review E. Jan. 13, 2014;89(1):012804. [cited by applicant]
Pop et al., Comparative genome assembly. Briefings in bioinformatics. Sep. 1, 2004;5(3):237-48. [cited by applicant]
Pruesse et al., SINA: accurate high-throughput multiple sequence alignment of ribosomal RNA genes. Bioinformatics. Jul. 15, 2012;28(14):1823-9. [cited by applicant]
Rajaram et al., Pearl millet [ [cited by applicant]
Raphael et al., A novel method for multiple alignment of sequences with repeated and shuffled elements. Genome Research. Nov. 1, 2004;14(11):2336-46. [cited by applicant]
Robertson et al., De novo assembly and analysis of RNA-seq data. Nature methods. Nov. 2010;7(11):909-12. [cited by applicant]
Rödelsperger et al., Syntenator: multiple gene order alignments with a gene-specific scoring function. Algorithms for Molecular Biology. Dec. 2008;3(1):1-2. [cited by applicant]
Rognes et al., Six-fold speed-up of Smith-Waterman sequence database searches using parallel processing on common microprocessors. Bioinformatics. Aug. 1, 2000;16(8):699-706. [cited by applicant]
Rognes, Faster Smith-Waterman database searches with inter-sequence SIMD parallelisation. BMC bioinformatics. Dec. 2011;12(1):1-1. [cited by applicant]
Rognes, ParAlign: a parallel sequence alignment algorithm for rapid and sensitive database searches. Nucleic acids research. Apr. 1, 2001;29(7):1647-52. [cited by applicant]
Ronquist et al., MrBayes 3.2: efficient Bayesian phylogenetic inference and model choice across a large model space. Systematic biology. May 1, 2012;61(3):539-42. [cited by applicant]
Sæbø et al., PARALIGN: rapid and sensitive sequence similarity searches powered by parallel computing technology. Nucleic acids research. Jul. 1, 2005;33(suppl_2):W535-9. [cited by applicant]
Sato et al., Directed acyclic graph kernels for structural RNA analysis. BMC bioinformatics. Dec. 2008;9(1):1-2. [cited by applicant]
Schneeberger et al., Simultaneous alignment of short reads against multiple genomes. Genome biology. Sep. 2009;10(9):R98. [cited by applicant]
Schwikowski et al., Weighted sequence graphs: boosting iterated dynamic programming using locally suboptimal solutions. Discrete Applied Mathematics. Apr. 1, 2003;127(1):95-117. [cited by applicant]
Shao et al., Bioinformatic analysis of exon repetition, exon scrambling and trans-splicing in humans. Bioinformatics. Mar. 15, 2006;22(6):692-8. [cited by applicant]
Sievers et al., Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Molecular systems biology. 2011;7(1):539. [cited by applicant]
Slater et al., Automated generation of heuristics for biological sequence comparison. BMC bioinformatics. Dec. 2005;6(1):31. [cited by applicant]
Smith et al., Multiple insert size paired-end sequencing for deconvolution of complex transcriptomes. RNA biology. May 1, 2012;9(5):596-609. [cited by applicant]
Sosa et al., Next-generation sequencing of human mitochondrial reference genomes uncovers high heteroplasmy frequency. PLoS computational biology. Oct. 25, 2012;8(10):e1002737. [cited by applicant]
Sturgeon et al., Rcda: A highly sensitive and specific alternatively spliced transcript assembly tool featuring upstream consecutive exon structures. Genomics. Dec. 1, 2012;100(6):357-62. [cited by applicant]
Subramanian et al., DIALIGN-TX: greedy and progressive approaches for segment-based multiple sequence alignment. Algorithms for Molecular Biology. Dec. 2008;3(1):1-11. [cited by applicant]
Sudmant et al., An integrated map of structural variation in 2,504 human genomes. Nature. Oct. 2015;526(7571):75-81. [cited by applicant]
Sun, Pairwise comparison between genomic sequences and optical-maps. New York University; 2006. [cited by applicant]
Szalkowski et al., Graph-based modeling of tandem repeats improves global multiple sequence alignment. Nucleic acids research. Sep. 1, 2013;41(17):e162. [cited by applicant]
Szalkowski, Fast and robust multiple sequence alignment with phylogeny-aware gap placement. BMC bioinformatics. Dec. 2012;13(1):1-1. [cited by applicant]
Tarhio et al., Approximate boyer-moore string matching. SIAM Journal on Computing. Apr. 1993;22(2):243-60. [cited by applicant]
Thomas, Community-wide effort aims to better represent variation in human reference genome, Genome Web. 2014. 11 pages. [cited by applicant]
Trapnell et al., TopHat: discovering splice junctions with RNA-Seq. Bioinformatics. May 1, 2009;25(9):1105-11. [cited by applicant]
Trapnell et al., Transcript assembly and abundance estimation from RNA-Seq reveals thousands of new transcripts and switching among isoforms. Nature biotechnology. May 2010;28(5):511. [cited by applicant]
Uchiyama et al., CGAT: a comparative genome analysis tool for visualizing alignments in the analysis of complex evolutionary changes between closely related genomes. BMC bioinformatics. Dec. 2006;7(1):1-7. [cited by applicant]
Wang et al., RNA-Seq: a revolutionary tool for transcriptomics. Nature reviews genetics. Jan. 2009;10(1):57-63. [cited by applicant]
Wu et al., Fast and SNP-tolerant detection of complex variants and splicing in short reads. Bioinformatics. Apr. 1, 2010;26(7):873-81. [cited by applicant]
Xing et al., An expectation-maximization algorithm for probabilistic reconstructions of full-length isoforms from splice graphs. Nucleic acids research. Jan. 1, 2006;34(10):3150-60. [cited by applicant]
Yang et al., Leveraging reads that span multiple single nucleotide polymorphisms for haplotype inference from sequencing data. Bioinformatics. Sep. 15, 2013;29(18):2245-52. [cited by applicant]
Yanovsky et al., Read mapping algorithms for single molecule sequencing data. InInternational Workshop on Algorithms in Bioinformatics. Springer, Berlin, Heidelberg. Sep. 15, 2008;38-49. [cited by applicant]
Yu et al., The construction of a tetraploid cotton genome wide comprehensive reference map. Genomics. Apr. 1, 2010;95(4):230-40. [cited by applicant]
Zeng et al., PyroHMMvar: a sensitive and accurate method to call short indels and SNPs for Ion Torrent and 454 data. Bioinformatics. Nov. 15, 2013;29(22):2859-68. [cited by applicant]
Zhang et al., Construction of a high-density genetic map for sesame based on large scale marker development by specific length amplified fragment (SLAF) sequencing. BMC plant biology. Dec. 2013;13(1):1-2. [cited by applicant]