NUCLEIC ACID-BASED DATA STORAGE
Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).
1 .- 29 . (canceled)
30 . A method for writing information into nucleic acid sequences, comprising:
translating said information into a string of symbols, each symbol having a symbol value; and
mapping the string of symbols to a plurality of identifiers, the identifiers comprising a combination of a plurality of components,
wherein each of the components comprises a nucleic acid sequence, said identifiers each comprising at least a first component corresponding to the symbol value of a respective one of the symbols and at least a second component indicating a position in the string of symbols.
31 . The method of claim 30 , further comprising selectively amplifying a subset of said plurality of identifiers.
32 . The method of claim 30 , wherein each symbol value of said string of symbols is one of two or more possible symbol values.
33 . The method of claim 30 , comprising constructing one or more libraries of identifiers.
34 . The method of claim 33 , wherein each of the one or more identifier libraries is associated with a distinct metadata, each individual identifier in said identifier library comprising said distinct metadata.
35 . The method of claim 33 , wherein constructing said individual identifier in said identifier library comprises assembling said one or more components from one or more layers and wherein each layer of said one or more layers comprises a distinct set of components.
36 . The method of claim 35 , wherein said individual identifier from said identifier library comprises one component from each layer of said one or more layers.
37 . The method of claim 30 , wherein the identifiers are assembled using an enzymatic reaction.
38 . The method of claim 30 , wherein the identifiers are assembled using overlap-extension polymerase chain reaction (PCR), polymerase cycling assembly, sticky end ligation, biobricks assembly, golden gate assembly, gibson assembly, recombinase assembly, ligase cycling reaction, or template directed ligation.
39 . The method of claim 30 , wherein constructing said individual identifier in said identifier library comprises deleting, replacing, or inserting at least one component in a parent identifier by applying nucleic acid editing enzymes to said parent identifier.
40 . A system for encoding information into nucleic acid sequences, the system comprising:
one or more computer processors, wherein said one or more computer processors are individually or collectively programmed to:
translate said information into a string of symbols, each symbol having a symbol value; and
map the string of symbols to a plurality of identifiers, the identifiers comprising a combination of a plurality of components, wherein each of the components comprises a nucleic acid sequence, said identifiers each comprising at least a first component corresponding to the symbol value of a respective one of the symbols and at least a second component indicating a position in the string of symbols.
41 . The system of claim 40 , wherein the one or more computer processors are further programmed to selectively amplify a subset of said plurality of identifiers.
42 . The system of claim 40 , wherein each symbol value of said string of symbols is one of two or more possible symbol values.
43 . The system of claim 40 , wherein the one or more computer processors are individually or collectively programmed to construct one or more libraries of identifiers.
44 . The system of claim 43 , wherein each of the one or more identifier libraries is associated with a distinct metadata, each individual identifier in said identifier library comprising said distinct metadata.
45 . The system of claim 40 , wherein constructing said individual identifier in said identifier library comprises assembling said one or more components from one or more layers and wherein each layer of said one or more layers comprises a distinct set of components.
46 . The system of claim 45 , wherein said individual identifier from said identifier library comprises one component from each layer of said one or more layers.
47 . The system of claim 40 , wherein the identifiers are assembled using an enzymatic reaction.
48 . The system of claim 40 , wherein the identifiers are assembled using overlap-extension polymerase chain reaction (PCR), polymerase cycling assembly, sticky end ligation, biobricks assembly, golden gate assembly, gibson assembly, recombinase assembly, ligase cycling reaction, or template directed ligation.
49 . The system of claim 40 , wherein constructing said individual identifier in said identifier library comprises deleting, replacing, or inserting at least one component in a parent identifier by applying nucleic acid editing enzymes to said parent identifier.