DNA barcodes for multiplexed sequencing
The present disclosure provides methods for optimizing barcode design for multiplex DNA sequencing. Also disclosed are DNA barcodes optimized for use with particular sequencing technologies.
1. A method for labeling a sample with a plurality of barcodes, comprising:
obtaining a dataset, wherein the dataset comprises data associated with the plurality of barcodes, wherein the barcodes are selected to maximize the edit distance of each barcode relative to the other barcodes in the plurality of barcodes;
determining, by a computer processor, the estimated misidentification error rates of the plurality of barcodes by performing simulated sequencing reads for each barcode and calculating, for each barcode, the fraction of simulated reads for which one or more other barcodes in the plurality of barcodes has edit distance less than or equal to that for the barcode from which the simulated read was generated;
determining if the fraction of simulated reads of the barcodes in the dataset is below an error rate threshold; and
labeling one or more samples with the barcodes having simulated reads below the error rate threshold.
2. The method of claim 1 , further comprising removing from the plurality one or more barcodes with a G:C content above a predetermined threshold value.
3. The method of claim 1 , further comprising removing from the plurality one or more barcodes whose sequences have a G:C content below a predetermined threshold value.
4. The method of claim 1 , further comprising removing from the plurality one or more barcodes capable of forming a hairpin structure.
5. The method of claim 1 , further comprising removing from the plurality one or more barcodes whose sequences include a known restriction site.
6. The method of claim 1 , further comprising removing from the plurality one or more barcodes whose sequences include a start codon.
7. The method of claim 1 , further comprising filtering the dataset to remove one or more barcodes whose sequences include a homopolymer run greater than or equal to 2 nucleotides in length.
8. The method of claim 1 , further comprising removing from the plurality one or more barcodes whose sequences are complementary to an mRNA sequence in an organism.
9. The method of claim 1 , further comprising removing from the plurality one or more barcodes whose sequences are complementary to a genomic sequence in an organism.
10. The method of claim 1 , wherein each barcode is from 5 to 40 nucleotides in length.
11. The method of claim 1 , wherein obtaining the dataset comprises generating the plurality of barcodes using a minimum edit distance, wherein the edit distance of each barcode relative to every other barcode is greater than or equal to the minimum edit distance.
12. The method of claim 1 , wherein obtaining the dataset comprises receiving the dataset from a third party.
13. The method of claim 1 , wherein the estimated misidentification error rate is less than a desired worst-case misidentification error rate of 5%, 1%, 0.1%, or 0.01%.
14. The method of claim 1 , wherein the estimated misidentification error rate is the average error rate of the individual barcode error rates.
15. The method of claim 1 , wherein the estimated misidentification error rate is the maximum error rate of the individual barcode error rates.
16. The method of claim 1 , wherein the estimated misidentification error rate is a specific percentile error rate of the individual barcode error rates.
17. The method of claim 1 , further comprising selecting the plurality of barcodes for labeling the one or more samples, wherein the selection is based upon the determined, estimated misidentification error rate of the plurality of barcodes.
18. The method of claim 1 , further comprising labeling a plurality of samples with the barcodes, wherein each barcode is used to label a distinct sample.
19. A method for labeling a sample with a plurality of barcodes, comprising:
determining, by a computer processor, the estimated misidentification error rates for a set of candidate barcodes by performing simulated sequencing reads for each barcode and calculating, for each barcode, the fraction of simulated reads for which one or more other barcodes in the set of candidate barcodes has edit distance less than or equal to that for the barcode from which the simulated read was generated;
determining the plurality of barcodes whose estimated misidentification error rate of simulated reads of the barcodes is below an error rate threshold; and
labeling one or more samples with the plurality of barcodes having simulated reads below the error rate threshold.