IP Library Granted Patent US 12674158
Granted Patent B2
US 12674158 · App. 17/438,461 · Granted Jul 7, 2026

Methods for DNA library generation to facilitate the detection and reporting of low frequency variants

Inventors: Morgane Macheret (Saint-Sulpice, CH); Christian Pozzorini (Saint-Sulpice, CH); Adrian Willig (Saint-Sulpice, CH); Jonathan Bieler (Saint-Sulpice, CH); Zhenyu Xu (Saint-Sulpice, CH)
Assignee: Sophia Genetics S.A.
C12N15/1065C12Q1/6869
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12674158
App. No.
17/438,461
Granted
Jul 7, 2026
Kind
B2
Abstract

Methods are disclosed for adding adapters to fragmented nucleic acids for next generation sequencing, including providing numerical codes based on variable adapter molecular barcode lengths on both sides of the fragmented nucleic acids and identifying reads from the same fragment based on both barcodes. The methods and products allow for the amplification of the fragmented nucleic acids when there is a low yield of isolated fragmented nucleic acids and also for efficient and reliable detection of low-frequency mutations including in subpopulations of cells within a subject.

Claims (32)

1 . A method for generating a library of DNA-adaptor products from at least two DNA fragments to facilitate the identification of the fragments in a high throughput sequencing data analysis workflow after amplification and sequencing, said method comprising:

generating a pool of adaptors that comprise a plurality of double-stranded or partially double-stranded polynucleotides comprising a spacer sequence on the double-stranded extremity of the adaptors, wherein the adaptors differ from each other by the total length of their spacer sequence of at least 3 and at most L max nucleotides, wherein each spacer sequence-comprises a constant termination subsequence TS of length L TS , wherein L TS comprises at least 3 nucleotides, concatenated with a variable spacer subsequence, and wherein the variable spacer subsequence is truncated from a common constant, predefined nucleotide sequence(S) having a length of L S nucleotides, wherein L S comprises 5-20 nucleotides, and wherein L max corresponds to the sum of L S and L TS ;

ligating, in a reaction mixture, a first and a second adaptor from the pool of adaptors to each end of a first double-stranded DNA fragment to produce a first DNA-adaptor product, wherein the ligation places the constant termination subsequence TS of each of the first and second adaptor between the double-stranded DNA fragment and the variable spacer subsequence of each of the first and second adaptor, respectively, so that the first DNA-adaptor product may be characterized by a numerical code formed by the respective lengths (L 1 , L 2 ) of the first and the second adaptor spacer sequences (SS 1 , SS 2 ); and

ligating, in the same reaction mixture, a third and a fourth adaptor from the pool of adaptors to each end of a second double-stranded DNA fragment to produce a second DNA-adaptor product, wherein the ligation places the constant termination subsequence TS of each of the third and fourth adaptor between the double-stranded DNA fragment and the variable spacer subsequence of each of the third and fourth adaptor, respectively, so that the second DNA-adaptor product may be characterized by a numerical code formed by the respective lengths (L 3 , L 4 ) of the third and the fourth adaptor spacer sequences (SS 3 , SS 4 ),

wherein determining whether the first DNA-adaptor product and second DNA-adaptor product correspond to different fragments relies on the placement of the constant termination subsequence between the double-stranded DNA fragment and the variable spacer subsequence and the characterization of the numerical code of the first and second DNA-adaptor product.

2 . The method of claim 1 , wherein the constant termination subsequence TS differs from the constant, predefined nucleotide sequence S by an edit distance of at least two.

3 . The method of claim 1 , wherein the spacer subsequence is truncated left to right from the start from said constant nucleotide sequence(S).

4 . The method of claim 1 , wherein the spacer subsequence is truncated right to left from the end from said constant nucleotide sequence(S).

5 . The method of claim 1 , wherein the constant termination subsequence TS is a triplet nucleotide ending with a T overhang to facilitate ligation to the DNA fragments.

6 . The method of claim 1 , wherein the constant termination subsequence TS is a quadruplet nucleotide ending with a T overhang to facilitate ligation to the DNA fragments.

7 . The method of claim 1 , further comprising:

amplifying the DNA-adaptor products to produce PCR duplicates suitable for high-throughput sequencing; and

sequencing the PCR duplicates with a high-throughput sequencer to produce raw sequencing reads.

8 . The method of claim 7 , further comprising:

for each sequencing read Rn,

trimming L max nucleotides from the beginning of the read, to produce a trimmed sequencing read;

recording the trimmed sequencing read in a pre-processed sequencing read file; and

aligning to a reference genome the trimmed sequencing reads from the pre-processed sequencing read file, so as to map each trimmed read to a start position and an end position.

9 . The method of claim 7 , further comprising:

for each sequencing read Rn, searching for the constant termination subsequence TS in the first L max nucleotides of the sequencing read;

measuring the length L n of the spacer sequence SS Rn as the distance, in the number of nucleotides, between the start of the sequencing read Rn and the end of the constant termination subsequence TS;

trimming at least L n nucleotides from the beginning of the read, to produce a trimmed sequencing read;

recording the measured length L n and the trimmed sequencing read in a pre-processed sequencing read file; and

aligning to a reference genome the trimmed sequencing reads from the pre-processed sequencing read file, so as to map each trimmed read to a start position and an end position.

10 . The method of claim 9 , wherein sequencing produces pair-end reads, further comprising:

tagging pair-end reads aligned to the same start and end position relative to the reference genome sequence reading direction and having the same numerical code pair of measured spacer sequence lengths (L 1 , L 2 ), as sequencing reads potentially issued from the two strands of the same original double-stranded DNA fragment; and

further subdividing those pair-end reads in two sub-groups according to their strand of origin, where the numerical code pair of measured spacer sequence lengths (L 1 , L 2 ) is given by {L n(forward) , L n(reverse) } in case of pair-end reads with F1R2 orientation and by {L n(reverse) , L n(forward) } in case of pair-end reads with F2R1 orientation.

11 . The method of claim 10 , further comprising collapsing each group of reads sharing the same start, end and numerical code into a consensus sequence for their parent fragment and identifying, with a variant calling method, variants for this parent fragment into the collapsed consensus sequence.

12 . The method of claim 10 , further comprising identifying for each group of reads sharing the same start, end and numerical code, with a statistical variant calling method, the probability of variants for their parent fragment.

13 . The method of claim 1 , wherein the method further includes a multiplex high throughput sequencing genomic analysis method for identifying genomic variants in at least two patient samples from a pool of samples, wherein the library of adaptors are different across samples.

14 . The method of claim 13 , wherein the library of adaptors differs across samples by the termination subsequence TS.

15 . The method of claim 13 , wherein the library of adaptors differs across samples by the predefined nucleotide sequence(S) used for truncating for the variable spacer subsequence.