Nucleic acid constructs and methods of use
The present invention provides oligonucleotide constructs, sets of such oligonucleotide constructs, and methods of using such oligonucleotide constructs to provide validated sequences or sets of validated sequences corresponding to desired ROIs. Such validated ROIs and constructs containing these have a wide variety of uses, including in synthetic biology, quantitative nucleic acid analysis, polymorphism and/or mutation screening, and the like.
1. A method to analyze nucleic acid sequences, comprising:
attaching a unique identifier nucleic acid sequence (UID) from a pool of UIDs to a first end of each strand of a plurality of analyte nucleic acid fragments to form a plurality of uniquely identified analyte nucleic acid fragments wherein the pool of UIDs is in excess of the plurality of analyte nucleic acid fragments;
redundantly determining nucleotide sequence of a uniquely identified analyte nucleic acid fragment, wherein determined nucleotide sequences which share a UID form a family of members;
comparing determined nucleotide sequences between members of a family; and
identifying a nucleotide sequence as accurately representing an analyte nucleic acid fragment when the sequence is a consensus sequence of the determined nucleotide sequences of the members of the family and wherein the family contains at least 2 members.
2. The method of claim 1 wherein prior to the step of redundantly determining, the uniquely identified analyte nucleic acid fragments are amplified.
3. The method of claim 1 wherein the nucleotide sequence is identified when 100% of members of the family contain the sequence.
4. The method of claim 1 wherein the step of attaching is performed by polymerase chain reaction.
5. The method of claim 1 wherein a first universal priming site is attached to a second end of each of a plurality of analyte nucleic acid fragments.
6. The method of claim 4 wherein at least two cycles of polymerase chain reaction are performed such that a family is formed of uniquely identified analyte nucleic acid fragments that have a UID on the first end and a first universal priming site on a second end.
7. The method of claim 6 , wherein the UID is covalently linked to a second universal priming site.
8. The method of claim 5 wherein the UID is covalently linked to a second universal priming site.
9. The method of claim 8 wherein prior to the step of redundantly determining, the uniquely identified analyte nucleic acid fragments are amplified using a pair of primers which are complementary to the first and the second universal priming sites, respectively.
10. The method of claim 7 wherein the UID is attached to the 5′ end of an analyte nucleic acid fragment and the second universal priming site is 5′ to the UID.
11. The method of claim 7 wherein the UID is attached to the 3′ end of an analyte nucleic acid fragment and the second universal priming site is 3′ to the UID.
12. The method of claim 1 wherein the analyte nucleic acid fragments are formed by applying a shear force to analyte nucleic acid.
13. The method of claim 2 wherein prior to the amplification, the analyte nucleic acid is treated with bisulfite to convert unmethylated cytosine bases to uracil.
14. The method of claim 1 further comprising the step of comparing number of families representing a first analyte DNA fragment to number of families representing a second analyte DNA fragment to determine a relative concentration of a first analyte DNA fragment to a second analyte DNA fragment in the plurality of analyte nucleic acid fragments.
15. The method of claim 4 wherein the UIDs are in excess of the analyte nucleic acid fragments during the polymerase chain reaction.
16. The method of claim 1 wherein the UIDs are about 10, about 20, about 30 or more bases.
17. The method of claim 1 further comprising the step of: identifying a mutation when the nucleotide sequence that accurately represents the analyte nucleic acid fragment is different from a reference sequence.
18. The method of claim 17 wherein a mutation is identified when the nucleotide sequence that accurately represents the analyte nucleic acid fragment is found in at least two families.
19. The method of claim 17 wherein the mutation is a single base substitution an insertion, or a deletion.
20. A method to analyze nucleic acid sequences, comprising:
attaching a unique identifier nucleic acid sequence (UID) of about 10 to about 30 bases from a pool of UIDs to a first end of each strand of a plurality of analyte nucleic acid fragments using at least two cycles of polymerase chain reaction to form uniquely identified analyte nucleic acid fragments, wherein the pool of UIDs is in excess of the plurality of analyte nucleic acid fragments;
redundantly determining nucleotide sequence of a uniquely identified analyte nucleic acid fragment, wherein determined nucleotide sequences which share a UID form a family of members;
comparing determined nucleotide sequences between members of a family;
identifying a nucleotide sequence as accurately representing an analyte nucleic acid fragment when the sequence is a consensus sequence of the determined nucleotide sequences of the members of the family and wherein the family contains at least 2 members; and
comparing the nucleotide sequence identified as accurately representing an analyte nucleic acid fragment to a reference sequence and identifying a single base substitution, insertion, or deletion in the nucleotide sequence relative to the reference sequence.
21. The method of claim 1 wherein the step of attaching is performed by ligation.