IP Library Granted Patent US 12,658,284
Granted Patent B2
US 12,658,284 · App. 18/750,586 · Granted Jun 16, 2026

Reference marks storing embedded oligo index in DNA data storage

Inventors: Iouri Oboukhov (Rochester, MN); Niranjay Ravindran (Rochester, MN); Richard Galbraith (Rochester, MN); Austin Striegel (Rochester, MN)
Assignee: Western Digital Technologies, Inc.
G16B50/50G16B30/00G11C13/02G16B50/40H03M13/1105H03M13/1148
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,658,284
App. No.
18/750,586
Granted
Jun 16, 2026
Kind
B2
Abstract

Example systems and methods for embedding an oligo index in a series of reference marks along the oligo for DNA data storage are described. A data unit may be encoded in oligos that include reference marks at predetermined intervals along the length of each oligo that encode the oligo index, generally in multiple copies encoded with varying symmetries. During decoding, the reference marks may be analyzed through multiple correlation matrices to determine a most likely set of values for the oligo index. In some configurations, the same reference marks may be processed through another correlation matrix to determine and correct insertions and deletions based on multiple copies of the oligo, as identified by the decoded oligo indices in those copies.

Claims (107)

1 . A system, comprising:

an encoder configured to:

determine an oligo for encoding a data unit, wherein the oligo encodes a number of symbols corresponding to user data in the data unit;

insert a plurality of reference marks at a predetermined interval along a length of the oligo, wherein:

at least one set of reference marks in the plurality of reference marks encodes an oligo index; and

the predetermined interval equals a fixed data length between sequential reference marks; and

output write data for the oligo for synthesis of the oligo.

2 . The system of claim 1 , wherein the predetermined interval of the plurality of reference marks corresponds to a single base pair between sequential reference marks.

3 . The system of claim 1 , wherein the oligo index comprises:

an oligo address; and

a cyclic redundancy check code.

4 . The system of claim 1 , wherein the plurality of reference marks comprises a plurality of copies of the oligo index.

5 . The system of claim 4 , wherein:

each copy of the oligo index comprises a different set of reference marks of the at least one set of reference marks; and

at least two copies of the oligo index comprise different translations of the oligo index that order the corresponding sets of reference marks in different translation orders.

6 . The system of claim 5 , wherein:

each copy of the oligo index:

occupies a predetermined set of positions in the plurality of reference marks; and

has a known translation order for that predetermined set of positions and corresponding set of reference marks; and

the different translation orders are selected from:

mirror symmetry; and

translational symmetry using circular offset values less than a length of the oligo index.

7 . The system of claim 6 , wherein:

the plurality of reference marks comprises a reference number of base pairs in the oligo;

the reference number of base pairs is a multiple of the length of the oligo index;

the multiple is at least equal to a number of the plurality of copies; and

the number of the plurality of copies is at least four.

8 . The system of claim 5 , wherein the different sets of reference marks of the at least two copies are interleaved along the length of the oligo index.

9 . A system comprising:

an index decoder configured to:

receive read data determined from sequencing an oligo, wherein:

the oligo encodes:

a number of symbols corresponding to user data in a data unit; and

a plurality of reference marks at a predetermined interval along a length of the oligo;

at least one set of reference marks in the plurality of reference marks encodes an oligo index; and

the predetermined interval equals a fixed data length between sequential reference marks;

determine, from the read data, a convolutional matrix for the oligo, wherein each column of the convolutional matrix corresponds to a base pair offset of the read data for the oligo;

determine, from the read data, a first reference matrix based on a first sequence of reference positions at an encoded frequency of the plurality of reference marks;

determine, from the read data, a second reference matrix based on a second sequence of reference positions at the encoded frequency of the plurality of reference marks;

determine a first correlation matrix based on a comparison of the convolutional matrix to the first reference matrix;

determine a second correlation matrix based on a comparison of the convolutional matrix and the second reference matrix;

determine a most likely path through a combination of the first correlation matrix and the second correlation matrix based on:

a series of probabilities of moving between corresponding rows in the first correlation matrix and the second correlation matrix; and

loop decoding for different symmetries of a plurality of copies of the oligo index; and

decode the oligo index based on the most likely path and the plurality of copies of the oligo index.

10 . The system of claim 9 , further comprising:

a reference mark decoder configured to:

receive, based on oligos with a same decoded oligo index, read data determined from sequencing at least two copies of the oligo;

populate a reference mark correlation matrix based on a comparison of reference mark positions in a first copy of the oligo to reference mark positions in a second copy of the oligo;

determine a most likely path through the reference mark correlation matrix corresponding to offset values for the plurality of reference marks;

determine, based on at least one offset value for the plurality of reference marks, an insertion or deletion in a data segment between sequential reference marks; and

correct symbol alignment in the data segment to compensate for the insertion or deletion; and

a user data decoder configured to:

decode the user data from the read data using an error correction code and corresponding redundancy data; and

output, based on the decoded user data, the data unit.

11 . A method comprising:

determining an oligo for encoding a data unit, wherein the oligo encodes a number of symbols corresponding to user data in the data unit;

inserting a plurality of reference marks at a predetermined interval along a length of the oligo, wherein:

at least one set of reference marks in the plurality of reference marks encodes an oligo index; and

the predetermined interval equals a fixed data length between sequential reference marks; and

outputting write data for the oligo for synthesis of the oligo.

12 . The method of claim 11 , wherein the predetermined interval of the plurality of reference marks corresponds to a single base pair between sequential reference marks.

13 . The method of claim 11 , wherein the oligo index comprises:

an oligo address; and

a cyclic redundancy check code.

14 . The method of claim 11 , wherein:

the plurality of reference marks comprises a plurality of copies of the oligo index;

each copy of the oligo index comprises a different set of reference marks of the at least one set of reference marks; and

at least two copies of the oligo index comprise different translations of the oligo index that order the corresponding sets of reference marks in different translation orders.

15 . The method of claim 14 wherein:

each copy of the oligo index:

occupies a predetermined set of positions in the plurality of reference marks; and

has a known translation order for that predetermined set of positions and corresponding set of reference marks; and

the different translation orders are selected from:

mirror symmetry; and

translational symmetry using circular offset values less than a length of the oligo index.

16 . The method of claim 15 , wherein:

the plurality of reference marks comprises a reference number of base pairs in the oligo;

the reference number of base pairs is a multiple of the length of the oligo index;

the multiple is at least equal to a number of the plurality of copies; and

the number of the plurality of copies is at least four.

17 . The method of claim 14 , wherein the different sets of reference marks of the at least two copies are interleaved along the length of the oligo index.

18 . The method of claim 14 , further comprising:

receiving read data determined from sequencing the oligo;

determining, from the read data, a convolutional matrix for the oligo, wherein each column of the convolutional matrix corresponds to a base pair offset of the read data for the oligo;

determining, from the read data, a first reference matrix based on a first sequence of reference positions at an encoded frequency of the plurality of reference marks;

determining, from the read data, a second reference matrix based on a second sequence of reference positions at the encoded frequency of the plurality of reference marks;

determining a first correlation matrix based on a comparison of the convolutional matrix to the first reference matrix;

determining a second correlation matrix based on a comparison of the convolutional matrix and the second reference matrix;

determining a most likely path through a combination of the first correlation matrix and the second correlation matrix based on:

a series of probabilities of moving between corresponding rows in the first correlation matrix and the second correlation matrix; and

loop decoding for different symmetries of the plurality of copies of the oligo index; and

decoding the oligo index based on the most likely path and the plurality of copies of the oligo index.

19 . The method of claim 18 , further comprising:

receiving, based on oligos with a same decoded oligo index, read data determined from sequencing at least two copies of the oligo;

populating a reference mark correlation matrix based on a comparison of reference mark positions in a first copy of the oligo to reference mark positions in a second copy of the oligo;

determining a most likely path through the reference mark correlation matrix corresponding to offset values for the plurality of reference marks;

determining, based on at least one offset value for the plurality of reference marks, an insertion or deletion in a data segment between sequential reference marks;

correcting symbol alignment in the data segment to compensate for the insertion or deletion;

decoding the user data from the read data using an error correction code and corresponding redundancy data; and

outputting, based on the decoded user data, the data unit.

20 . A system comprising:

means for determining an oligo for encoding a data unit, wherein the oligo encodes a number of symbols corresponding to user data in the data unit;

means for inserting a plurality of reference marks at a predetermined interval along a length of the oligo, wherein:

at least one set of reference marks in the plurality of reference marks encodes an oligo index; and

the predetermined interval equals a fixed data length between sequential reference marks; and

means for outputting write data for the oligo for synthesis of the oligo.

Assignments (2)
PATENT COLLATERAL AGREEMENT (AR) Recorded Aug 23, 2024
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS THE AGENT
Reel/Frame 068762/0695 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2024
From: OBOUKHOV, IOURI; RAVINDRAN, NIRANJAY; GALBRAITH, RICHARD; STRIEGEL, AUSTIN
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 067805/0432 →
Continuity (1)
Related Publication 20250391513A1 · Dec 25, 2025
References Cited (97)
US 8407554B2 · Kermani · 2013 [cited by applicant]
US 8937564B2 · Aloni · 2015 [cited by applicant]
US 9234239B2 · Sikora · 2016 [cited by applicant]
US 9830553B2 · Chen · 2017 [cited by applicant]
US 10027347B2 · Le Scouarnec · 2018 [cited by applicant]
US 10370246B1 · Milenkovic · 2019 [cited by applicant]
US 10423341B1 · Kermani · 2019 [cited by applicant]
US 10566077B1 · Milenkovic · 2020 [cited by applicant]
US 10650312B2 · Roquet · 2020 [cited by applicant]
US 10742233B2 · Erlich · 2020 [cited by applicant]
US 10810495B2 · Roth · 2020 [cited by applicant]
US 10818378B2 · Hutchison, III · 2020 [cited by applicant]
US 10917109B1 · Dimopoulou · 2021 [cited by applicant]
US 10956806B2 · Masuda · 2021 [cited by applicant]
US 11017170B2 · Yin · 2021 [cited by applicant]
US 11093547B2 · Su · 2021 [cited by applicant]
US 11098303B2 · Zhuang · 2021 [cited by applicant]
US 11227219B2 · Roquet · 2022 [cited by applicant]
US 11249941B2 · Unidad · 2022 [cited by applicant]
US 11435905B1 · Kermani · 2022 [cited by applicant]
US 11755649B2 · Wei · 2023 [cited by applicant]
US 11755922B2 · Milenkovic · 2023 [cited by applicant]
US 11763169B2 · Roquet · 2023 [cited by applicant]
US 20040191788A1 · Gleba · 2004 [cited by applicant]
US 20070042372A1 · Arita · 2007 [cited by applicant]
US 20100199155A1 · Kermani · 2010 [cited by applicant]
US 20110172975A1 · Silva Filho · 2011 [cited by applicant]
US 20110269119A1 · Hutchison · 2011 [cited by applicant]
US 20170109229A1 · Huetter · 2017 [cited by applicant]
US 20170134045A1 · Huetter · 2017 [cited by applicant]
US 20180046921A1 · Chen · 2018 [cited by applicant]
US 20190130280A1 · Erden · 2019 [cited by applicant]
US 20190362814A1 · Roquet · 2019 [cited by applicant]
US 20190363739A1 · Erlich · 2019 [cited by applicant]
US 20200043568A1 · Sikora · 2020 [cited by applicant]
US 20200185057A1 · Leake · 2020 [cited by applicant]
US 20200193301A1 · Roquet · 2020 [cited by applicant]
US 20200211677A1 · Fan · 2020 [cited by applicant]
US 20210024924A1 · Qi · 2021 [cited by applicant]
US 20210050073A1 · Chen · 2021 [cited by applicant]
US 20210074380A1 · Yekhanin · 2021 [cited by applicant]
US 20210079382A1 · Leake · 2021 [cited by applicant]
US 20210098081A1 · Yekhanin · 2021 [cited by applicant]
US 20210108194A1 · Nathaniel · 2021 [cited by applicant]
US 20210166159A1 · Rubenstein · 2021 [cited by applicant]
US 20210202032A1 · Krsticevic · 2021 [cited by applicant]
US 20210210171A1 · Stirparo · 2021 [cited by applicant]
US 20210225461A1 · Merriman · 2021 [cited by applicant]
US 20220002781A1 · Harness · 2022 [cited by applicant]
US 20220044763A1 · Karimi · 2022 [cited by applicant]
US 20220364991A1 · Wanunu · 2022 [cited by applicant]
US 20230027270A1 · Rubenstein · 2023 [cited by applicant]
US 20230257801A1 · Brodin · 2023 [cited by applicant]
US 20230317164A1 · Roquet · 2023 [cited by applicant]
US 20240184666A1 · Oboukhov · 2024 [cited by applicant]
US 20240185959A1 · Oboukhov · 2024 [cited by examiner]
US 20240257147A1 · Owen · 2024 [cited by applicant]
US 20250020657A1 · Graige · 2025 [cited by applicant]
US 20250037039A1 · Rosenstein · 2025 [cited by applicant]
US 20250088203A1 · Oboukhov · 2025 [cited by applicant]
CN 118138060A · 2024 [cited by applicant]
EP 3416076A1 · 2018 [cited by applicant]
WO 2018148260A1 · 2018 [cited by applicant]
WO 2019046768A1 · 2019 [cited by applicant]
WO 2021072398A1 · 2021 [cited by applicant]
WO 2021105974A1 · 2021 [cited by applicant]
WO 2021231493A1 · 2021 [cited by applicant]
Zhang et al., “Soft-Decision Decoding for DNA-Based Data Storage,” 2018 International Symposium on Information Theory and Its Applications (ISITA), Singapore, 2018, pp. 16-20, (Year: 2018). [cited by applicant]
Dimopoulou et al., “Storing Digital Data Into DNA: A Comparative Study Of Quaternary Code Construction,” ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing , pp. 4332-4336, (2020). [cited by applicant]
Hamoum et al. “DNA Data Storage Algorithms and Synchronization”, Information Theory [math.IT]. University Bretagne Sud, (2022). [cited by applicant]
Qin, et al. “Robust Multi-read Reconstruction from Contaminated Custers Using Deep Neural Network for DNA Storage.” arXiv preprint arXiv:2210.11106. Oct. 31, 2022 (Oct. 31, 2022) whole document. [cited by applicant]
Zhang et al. Rbec: a Tool for Analysis of Amplicon Sequencing Data from Synthetic Microbial Communities. ISME communications, 1(1), 73. Jan. 17, 2021 (Jan. 17, 2021) whole document (2021). [cited by applicant]
Lu et al., “Design of Nonbinary Error Correction Codes With a Maximum Run-length Constraint to Correct a Single Insertion or Deletion Error for Dna Storage”, Institute of Electrical and Electronics Engineers (IEEE), vol… [cited by applicant]
Cai et al., “Correcting a Single Indel/Edit for DNA-Based Data Storage: Linear-Time Encoders and Order-Optimality”, EEE Transactions on Information Theory, vol. 67, No. 6, pp. 3438-3451, (2021). [cited by applicant]
Chen et al., “Multiple Interleaved RS Codes for Data Storage Using Up to Mb-scale Synthetic DNA in Living Cells”, Synthetic Biology Journal, vol. 2, Issue 3, pp. 428-443, (2021). [cited by applicant]
Goto et al., “Coding of Insertion-deletion-substitution Channels Without Markers”, 2016 IEEE International Symposium on Information Theory, pp. 635-639, (2016). [cited by applicant]
Khuat et al., “A Quaternary Code Correcting a Burst of at Most Two Deletion or Insertion Errors in DNA Storage”, Entropy, vol. 23, No. 12, pp. 1592, (2021). [cited by applicant]
Kiah et al., “Codes for DNA Sequence Profiles”, IEEE Transactions on Information Theory, vol. 62, No. 6, pp. 3125-3146, (2016). [cited by applicant]
Limbachiya et al., “On Optimal Family of Codes for Archival DNA Storage”, Seventh International Workshop on Signal Design and its Applications in Communications, pp. 123-127, (2015). [cited by applicant]
Nakata et al., “Synchronization and Asymmetric Error Correction for Nanopore Sequencing,” 2021 IEEE International Conference on Consumer Electronics—Taiwan, pp. 1-2, (2021). [cited by applicant]
Yin et al., “Design of Constraint Coding Sets for Archive DNA Storage”, IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 19, No. 6, pp. 3384-3394, (2022). [cited by applicant]
Chauhan et al., “Portable and Error-Free DNA-Based Data Storage”, 2021 4th International Conference on Recent Developments in Control, pp. 418-421, (2021). [cited by applicant]
Deng et al., “Optimized Code Design for Constrained DNA Data Storage With Asymmetric Errors”, IEEE Access, vol. 7, pp. 84107-84121, (2019). [cited by applicant]
Jain et al., “Duplication-Correcting Codes for Data Storage in the DNA of Living Organisms”, IEEE Transactions on Information Theory, vol. 63, No. 8, pp. 4996-5010 (2017). [cited by applicant]
Press et al., “HEDGES Error-correcting Code for DNA Storage Corrects Indels and Allow Sequence Constraints”, Proc Natl Acad Sci USA, vol. 117, No. 31, pp. 18489-18496, (2020). [cited by applicant]
Wei et al., “Improved Coding Over Sets for DNA-Based Data Storage”, IEEE Transactions on Information Theory, vol. 68, No. 1, pp. 118-129, (2022). [cited by applicant]
Wu et al., “HD-Code: End-to-End High Density Code for DNA Storage”, IEEE Transactions on NanoBioscience, vol. 20, No. 4, pp. 455-463, (2021). [cited by applicant]
Yan et al., “New Levenshtein-Marker Code for DNA-based Data Storage Capable of Correcting Multiple Edit Errors”, TechRxiv. (2021). [cited by applicant]
Hamoum et al., “Channel Model with Memory for DNA Data Storage with Nanopore Sequencing”, 2021 11th International Symposium on Topics in Coding , pp. 1-5, (2021). [cited by applicant]
Christner et al., “DNA Data Storage—DNA Data Storage Alliance—Rosetta Stone Initiative”, Storage Networking Industry Association, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
Gervasio et al., “DNA Data Storage—A Decade of Coding and Decoding, How Far Have We Got?”, Storage Networking Industry Association, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
Landsman et al., “DNA Data Storage Alliance: Building a DNA Data Storage Ecosystem”, Storage Networking Industry Association, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
Marelli et al., “DNAssim: A Full System Simulator for DNA Storage”, Storage Networking Industry Association, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
OG US et al., “The Looming Need for Molecular Storage”, Storage Networking Industry Association, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
Piantanida et al., “Nucleic Acid Memory—Super Resolution Microscopy Enhances Novel Approach to DNA Data Storage”, Boise State University, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
Reis et al., “Coding and Decoding, an Experience from a Brazilian Research Center”, Storage Networking Industry Association, Storage Developer Conference, Fremont, CA, Sep. 12-15, 2022. [cited by applicant]
Dimopoulou et al., “Image storage onto synthetic DNA”, Signal Processing: Image Communication, vol. 97, 116331, May 27, 2021, Sections 5.1-5.2; and figure B.3. [cited by applicant]