IP Library Granted Patent US 12,512,185
Granted Patent B2
US 12,512,185 · App. 16/631,405 · Granted Dec 30, 2025

DNA-based data storage and retrieval

Inventor: Long Fan (Nanjing, CN)
Assignee: NANJING GENSCRIPT BIOTECH CO., LTD.
G16B50/40G06F11/1076G16B25/20G16B30/00H03M7/3086H03M7/70H03M13/1515Y10S977/704
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,512,185
App. No.
16/631,405
Granted
Dec 30, 2025
Kind
B2
Abstract

The present disclosure generally relates to DNA-based data storage. An exemplary method for storing input data on nucleic acid comprises: converting the input data into a set of nucleotide sequences and synthesizing a set of nucleic acids comprising the set of nucleotide sequences. The converting comprises a data processing step comprising converting the input data into a binary string, and a nucleotide encoding step comprising converting the binary string using a 5-bit transcoding framework to obtain the set of nucleotide sequences.

Claims (29)

1 . A method for storing input data on nucleic acid molecules, the method comprising:

a) converting the input data into a set of nucleotide sequences, wherein the converting comprises

i) a data processing step comprising converting the input data into a binary string and dividing the binary string into a sequence of non-overlapping 5-bit binary strings;

ii) a nucleotide encoding step comprising converting the sequence of non-overlapping 5-bit binary strings into the set of nucleotide sequences using a 5-bit transcoding framework to convert each 5-bit binary string into a corresponding 3-mer nucleotide sequence, wherein each corresponding 3-mer nucleotide sequence comprises:

a first base and a second base selected from A, T, C, and G; and

a third base selected from R and Y, wherein:

R is selected from any two of A, T, C, and G;

Y is selected from the corresponding other two of A, T, C, and G;

R and Y are chosen so that they are different from the second base immediately in front of R or Y; and

R and Y are chosen based on a GC content of the set of nucleotide sequences; and

b) synthesizing a set of nucleic acid molecules corresponding to the set of nucleotide sequences generated by converting the input data.

2 . The method of claim 1 , wherein the nucleotide encoding step further comprises converting each 5-bit binary string of the sequence of non-overlapping 5-bit binary strings into an integer ranging from 0 to 31 to obtain a string of integers.

3 . The method of claim 2 , wherein the nucleotide encoding step further comprises dividing the string of integers into a plurality of initial sub-sequence of integers having a predetermined length.

4 . The method of claim 3 , wherein the length of each of the plurality of initial sub-sequence of integers is determined based on an oligo length of a selected synthesis platform, a desired error tolerance, a size of the input data, a selected error correction code, or a combination thereof.

5 . The method of claim 3 , wherein the nucleotide encoding step further comprises adding index information to each of the plurality of the initial sub-sequences of integers to obtain a plurality of integer sub-sequences having indexes.

6 . The method of claim 5 , wherein the index information added to each of the plurality of the initial sub-sequences of integers comprises a sequence of integers, wherein the length of the sequence of integers is based on the size of the input data.

7 . The method of claim 5 , wherein the nucleotide encoding step further comprises, after adding the index information, adding redundancy data to the plurality of integer sub-sequences having indexes, thereby obtaining a plurality of integer sub-sequences having redundancy.

8 . The method of claim 7 , wherein adding redundancy data to the plurality of integer sub-sequences having indexes comprises:

creating an empty matrix, wherein the number of columns in the empty matrix is larger than the size of the plurality of integer sub-sequences having indexes, and wherein the number of rows of the empty matrix is larger than the number of integers in each of the plurality integer sub-sequences having indexes;

filling the empty matrix with the plurality of integer sub-sequences having indexes and data generated by applying an error correction coding; and

obtaining the plurality of sub-sequences having redundancy based on the filled matrix.

9 . The method of claim 8 , wherein the number of columns of the empty matrix is determined based on an oligo length of a selected synthesis platform, the type of the error correction code, a predetermined error tolerance value, a size of the plurality of integer sub-sequences having index, or a combination thereof.

10 . The method of claim 8 , wherein the number of rows of the empty matrix is determined based on an oligo length of a selected synthesis platform, a type of the error correction code, a predetermined error tolerance value, a size of the plurality of integer sub-sequences having indexes, or a combination thereof.

11 . The method of claim 8 , wherein the error correction coding is Reed-Solomon (“RS”) coding.

12 . The method of claim 11 , wherein the data generated by applying an error correction coding is generated by applying string correction of the RS coding and/or block correction of the RS coding.

13 . The method of claim 1 , wherein the input data corresponds to a compressed file.

14 . The method of claim 1 , wherein the input data corresponds to two or more files.

15 . The method of claim 1 , wherein the input data corresponds to a text file.

16 . The method of claim 1 , wherein the data processing step further comprises compressing the input data to obtain a compressed file and converting the compressed file into a binary string.

Assignments (2)
CHANGE OF NAME Recorded Feb 19, 2021
From: NANJINGJINSIRUI SCIENCE & TECHNOLOGY BIOLOGY CORP.
To: NANJING GENSCRIPT BIOTECH CO., LTD.
Reel/Frame 055346/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2020
From: FAN, LONG
To: NANJINGJINSIRUI SCIENCE & TECHNOLOGY BIOLOGY CORP.
Reel/Frame 052067/0090 →
Priority Claims (1)
CN 201710611123.2 · Jul 25, 2017 · national
Continuity (1)
Related Publication 20200211677A1 · Jul 2, 2020
References Cited (32)
US 20110295858A1 · Ahn et al. · 2011 [cited by applicant]
CN 103093121A · 2013 [cited by applicant]
CN 105022935A · 2014 [cited by applicant]
CN 104850760B · 2016 [cited by applicant]
CN 106687966A · 2017 [cited by applicant]
CN 105760706B · 2018 [cited by applicant]
EP 1443449A2 · 2004 [cited by applicant]
EP 3173961A1 · 2017 [cited by applicant]
WO 2003025123A2 · 2003 [cited by applicant]
WO 2003052383A2 · 2003 [cited by applicant]
WO WO03052383A2 · 2003 [cited by applicant]
WO 2003052383A3 · 2003 [cited by applicant]
WO WO03052383A3 · 2003 [cited by applicant]
WO 2004053766A1 · 2004 [cited by applicant]
WO 2004088585A2 · 2004 [cited by applicant]
WO 2004088585A3 · 2005 [cited by applicant]
WO 2007137225A2 · 2007 [cited by applicant]
WO 2013178801A2 · 2013 [cited by applicant]
WO 2014014991A2 · 2014 [cited by applicant]
WO 2013178801A3 · 2014 [cited by applicant]
WO WO2016020682A1 · 2016 [cited by applicant]
WO 2017011492A1 · 2017 [cited by applicant]
Ghoshdastider, Umesh, and Banani Saha. “GenomeCompress: a novel algorithm for DNA compression.” (2005): 0973-6824. (Year: 2005). [cited by examiner]
Grass, Robert N et al. “Robust chemical preservation of digital information on DNA in silica with error-correcting codes.” Angewandte Chemie (International ed. in English) vol. 54,8 (2015): 2552-5. doi:10.1002/anie.2014… [cited by examiner]
Yazdi, S.M.H.T., Gabrys, R. & Milenkovic, O. Portable and Error-Free DNA-Based Data Storage. Sci Rep 7, 5011 (2017). https://doi.org/10.1038/s41598-017-05188-1 Jul. 10, 2017 (Year: 2017). [cited by examiner]
Limbachiya, Dixita, and Manish K. Gupta. “Natural data storage: A review on sending information from now to then via nature.” arXiv preprint arXiv:1505.04890 (2015). (Year: 2015). [cited by examiner]
Yazdi, SM Hossein Tabatabaei, et al. “DNA-based storage: Trends and methods.” IEEE Transactions on Molecular, Biological and Multi-Scale Communications 1.3 (2015): 230-248. [cited by examiner]
International Search Report and Written Opinion, dated Oct. 31, 2018, for PCT Application No. PCT/CN2018/097083, filed Jul. 25, 2018, 6 pages. [cited by applicant]
Maurer, K. et al. (Dec. 20, 2006). “Electrochemically Generated Acid and Its Containment to 100 Micron Reaction Areas for the Production of DNA Microarrays,” PLos One 1(1):e34, 7 pages. [cited by applicant]
Chun, J.Y. et al. (2015). “Passing Go with DNA Sequencing: Delivering Messages in a Covert Transgenic Channel,” IEEE CS Security and Privacy Workshops, pp. 17-26. [cited by applicant]
Tornea, O. (2013). “Contributions to DNA Cryptography: Application to Text and Image Secure Transmission,” HAL Open Science, 170 pages. [cited by applicant]
Delarue, M. (2007). “An Asymmetric Underlying Rule in the Assignment of Codons: Possible Clue to a Quick Early Evolution of the Genetic Code via Successive Binary Choices,” RNA 13(2):161-169. [cited by applicant]