NUCLEIC ACID-BASED DATA STORAGE
Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool. But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).
1 .- 29 . (canceled)
30 . A method for writing information into nucleic acid sequences, comprising:
translating said information into a string of symbols;
mapping said symbols to a plurality of nucleic acid sequences; and
generating said plurality of nucleic acid sequences, wherein each nucleic acid sequence is generated by assembling at least three components, wherein a first, second, and third component of the at least three components is a nucleic acid sequence of a respective one of first, second, and third groups of nucleic acid sequences;
wherein:
each component of the first group is a nucleic acid sequence having a first hybridization region at one end thereof, the first hybridization region being the same for all components in the first group;
each component of the second group is a nucleic acid sequence having a second hybridization region at one end thereof and a third hybridization region at another end thereof, the second hybridization region being the same for all components in the second group, the third hybridization region being the same for all components in the second group, and the second hybridization region being complementary to the first hybridization region; and
each nucleic acid sequence of the third group has a fourth hybridization region at one end thereof, the fourth hybridization region being the same for all components in the third group, and the fourth hybridization region being complementary to the third hybridization region.
31 . The method of claim 30 , wherein the first, second, third, and fourth hybridization regions are different from one another.
32 . The method of claim 30 , wherein the first, second, third, and fourth hybridization regions are overhanging ends.
33 . The method of claim 30 , wherein said plurality of nucleic acid sequences are assembled using polymerase chain reaction, ligation, or recombination.
34 . The method of claim 30 , wherein said plurality of nucleic acid sequences are assembled in a one pot reaction.
35 . The method of claim 34 , wherein a second plurality of nucleic acid sequences is assembled in the one pot reaction.
36 . The method of claim 30 , wherein assembly of the plurality of nucleic acid sequences is performed in a time period that is less than or equal to about 1 day, 12 hours, 10 hours, 8 hours, 7 hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours, or 1 hour.
37 . The method of claim 30 , wherein generating said plurality of nucleic acid sequences comprises introducing at least the first, second, and third components into the same droplet.
38 . The method of claim 37 , wherein generating said plurality of nucleic acid sequences comprises introducing into the same droplet at least two components from the first group, at least two components from the second group, and at least three components from the third group.
39 . A system for writing information into nucleic acid sequences, the system comprising:
one or more computer processors individually or collectively programmed to:
translate said information into a string of symbols;
map said symbols to a plurality of nucleic acid sequences; and
generate said plurality of nucleic acid sequences, wherein each nucleic acid sequence is generated by assembling at least three components, wherein a first, second, and third component of the at least three components is a nucleic acid sequence of a respective one of first, second, and third groups of nucleic acid sequences;
wherein:
each component of the first group is a nucleic acid sequence having a first hybridization region at one end thereof, the first hybridization region being the same for all components in the first group;
each component of the second group is a nucleic acid sequence having a second hybridization region at one end thereof and a third hybridization region at another end thereof, the second hybridization region being the same for all components in the second group, the third hybridization region being the same for all components in the second group, and the second hybridization region being complementary to the first hybridization region; and
each nucleic acid sequence of the third group has a fourth hybridization region at one end thereof, the fourth hybridization region being the same for all components in the third group, and the fourth hybridization region being complementary to the third hybridization region.
40 . The system of claim 39 , wherein the first, second, third, and fourth hybridization regions are different from one another.
41 . The system of claim 39 , wherein the first, second, third, and fourth hybridization regions are overhanging ends.
42 . The system of claim 39 , wherein said one or more computer processors are programmed to assemble the plurality of nucleic acid sequences are assembled using polymerase chain reaction, ligation, or recombination.
43 . The system of claim 39 , comprising a microfluidic device arranged to introduce at least the first, second, and third components into the same droplet.
44 . The system of claim 43 , wherein the microfluidic device is arranged to introduce into the same droplet at least two components from the first group, at least two components from the second group, and at least three components from the third group.
45 . The system of claim 39 , wherein said plurality of nucleic acid sequences are assembled in a one pot reaction.
46 . The system of claim 45 , wherein a second plurality of nucleic acid sequences is assembled in the one pot reaction.
47 . The system of claim 39 , wherein assembly of the plurality of nucleic acid sequences is performed in a time period that is less than or equal to about 1 day, 12 hours, 10 hours, 8 hours, 7 hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours, or 1 hour.