Nucleic acid indexing techniques
Presented herein are techniques for indexing of nucleic acid, e.g., for use in conjunction with sequencing. The techniques include generating indexed nucleic acid fragments from an individual sample, whereby the index sequence incorporated into each index site of the nucleic acid fragment is selected from a plurality of distinguishable of index sequences and such that the population of generated nucleic acid fragments represents each index sequence from the plurality. In this manner, the generated indexed nucleic acid fragments from a single sample are indexed with a diverse mix of index sequences that reduce misassignment due to index read errors associated with low sequence diversity.
1 . A method for sequencing nucleic acid molecules, comprising:
providing a plurality of dual-indexed nucleic acid fragments generated from a sample, wherein each individual nucleic acid fragment of the plurality of dual-indexed nucleic acid fragments comprises a 5′ adapter sequence, a 5′ index sequence, a target sequence, a 3′ adapter sequence, and a 3′ index sequence, wherein a plurality of different 5′ index sequences selected from a first set of 5′ index sequences associated with the sample and a plurality of different 3′ index sequences selected from a second set of 3′ index sequences associated with the sample are represented in the plurality of dual-indexed nucleic acid fragments, wherein the plurality of different 5′ index sequences and the plurality of different 3′ index sequences are distinguishable from one another, and wherein the plurality of dual-indexed nucleic acid fragments comprises more than one combination of the 5′ index sequence and the 3′ index sequence such that one or more individual fragments of the plurality of dual-indexed nucleic acid fragments have respective different 5′ index sequences and respective different 3′ index sequences relative to one another, wherein more than one fragment of the plurality has a same combination of an index sequence of the first set and an index sequence of the second set but different target sequences relative to one another;
generating sequencing data representative of sequences of the plurality of dual-indexed nucleic acid fragments; and
associating an individual sequence of the sequences with the sample only when the individual sequence includes both the 5′ index sequence selected from the first set and the 3′ index sequence selected from the second set.
2 . The method of claim 1 , comprising contacting the plurality of dual-indexed nucleic acid fragments with a sequencing substrate comprising immobilized capture molecules that hybridize to the 5′ adapter and the 3′ adapter sequences of the plurality of dual-indexed nucleic acid fragments.
3 . The method of claim 1 , comprising eliminating individual sequences in the sequencing data that include only one of the 5′ index sequence or the 3′ index sequence.
4 . The method of claim 1 , comprising providing dual-indexed nucleic acid fragments from a second sample, wherein the dual-indexed nucleic acid fragments from the second sample are modified with the 5′ adapter sequence and the 3′ adapter sequence and wherein a second plurality of different 5′ index sequences and a second plurality of different 3′ index sequences are represented in the dual-indexed nucleic acid fragments from the second sample and wherein the second plurality of different 5′ index sequences and a second plurality of different 3′ index sequences are distinguishable from one another and from the plurality of different 5′ index sequences and the plurality of different 3′ index sequences.
5 . The method of claim 4 , comprising simultaneously contacting the plurality of dual-indexed nucleic acid fragments from the sample and the second sample with a sequencing substrate comprising immobilized capture molecules that hybridize to the 5′ adapter and the 3′ adapter sequences of the nucleic acid fragments.
6 . A method for sequencing nucleic acid molecules, comprising:
providing a plurality of dual-indexed nucleic acid fragments generated from a sample, wherein each individual nucleic acid fragment of the plurality of dual-indexed nucleic acid fragments comprises a sequence of interest derived from the sample, a 5′ adapter sequence, a 5′ index sequence, a 3′ adapter sequence, and a 3′ index, wherein a plurality of different 5′ index sequences selected from a first set of 5′ index sequences associated with the sample and a plurality of different 3′ index sequences selected from a second set of 3′ index sequences associated with the sample are represented in the plurality of dual-indexed nucleic acid fragments and wherein the plurality of different 5′ index sequences and the plurality of different 3′ index sequences are distinguishable from one another and wherein the plurality of dual-indexed nucleic acid fragments comprises more than one combination of the 5′ index sequence and the 3′ index sequence such that one or more individual fragments of the plurality of dual-indexed nucleic acid fragments have respective different 5′ index sequences and respective different 3′ index sequences relative to one another, wherein more than one fragment of the plurality has a same combination of an index sequence of the first set and an index sequence of the second set but different sequences of interest relative to one another;
generating sequencing data representative of the sequence of interest;
generating sequencing data representative of the 5′ index sequence and the 3′ index sequence; and
assigning an individual sequence of interest to the sample only when the individual sequence of interest is associated with both the 5′ index sequence selected from the first set and the 3′ index sequence selected from the second set.
7 . The method of claim 6 , wherein the individual sequence of interest is associated with both the 5′ index sequence selected from the first set and the 3′ index sequence selected from the second set based on co-location of the sequencing data representative of the 5′ index sequence and the 3′ index sequence with the sequencing data representative of the sequence of interest.
8 . The method of claim 6 , comprising assigning sequence data of another individual sequence of interest to the sample based on complementarity to the individual sequence of interest.
9 . The method of claim 6 , wherein the sequencing data representative of the 5′ index sequence and the 3′ index sequence with the sequencing data representative of the sequence of interest are generated from a same single strand of the dual-index nucleic acid fragments.
10 . The method of claim 6 , wherein the sequencing data representative of the 5′ index sequence and the 3′ index sequence are generated from different single strands of the dual-index nucleic acid fragments.