IP Library Patent Application 19140338
Patent Application
App. No. 19/140,338

DNA-BASED VECTOR DATA STRUCTURE WITH PARALLEL OPERATIONS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/140,338
Abstract

Described in this specification are technologies including a vector data structure for representing data objects with identifiers, with applications in searching for data objects that are similar or identical to a given query data object and in vector arithmetic. The technologies include a DNA data structure and a computing architecture that provides operations on individual bits of each data object in parallel. The described architecture utilizes chemical methods that select, hybridize, concatenate, or amplify DNA, and uses them to implement computational instructions that provide parallel search and arithmetic.

Claims (148)

1 . A method for coding digital information into nucleic acid sequence(s), comprising:

(a) coding said digital information into a sequence of symbols and converting said sequence of symbols into codewords;

(b) parsing said codewords into a coded sequence of symbols;

(c) mapping said coded sequence of symbols to a plurality of identifiers, wherein an individual identifier of said plurality of identifiers comprises one or more nucleic acid sequences; and

(d) enumerating an identifier library wherein each symbol of said coded sequence of symbols is encoded by one or more identifier(s), wherein an individual identifier of said plurality of identifiers comprises a plurality of components from one or more layers, each layer of said one or more layers comprising a distinct set of components, each identifier of the identifier library comprising:

one or more data layers encoding a symbol, and

one or more key layers encoding a key.

2 . The method of claim 1 , wherein a set of identifiers having identical components in the key layers forms a codeword vector of identifiers (CVI).

3 . The method as in any one of claims 1-2 , wherein the digital information comprises data values represented by a sequence of bits, and coding comprises converting the sequence of bits into bit vectors, each bit of the bit vectors mapped to one or more identifiers.

4 . The method of claim 3 , wherein a data value is be encoded with its original bit representation.

5 . The method of claim 3 , wherein a data value is be encoded with a sparser or denser bit vector.

6 . The method as in any one of claims 3-5 , wherein the data layers encode the bit vectors.

7 . The method as in any one of claims 1-6 , comprising performing a data query operation on the identifier library

8 . The method of claim 7 , wherein the query operation is performed on one or more query components form a subset of layers.

9 . The method of claim 8 , comprising performing selective polymerase chain reaction (PCR) using query oligonucleotides to amplify only identifiers that contain the one or more query components.

10 . The method of claim 9 , comprising performing multiple cycles of PCR with different query oligonucleotides in each cycle.

11 . The method as in any one of claims 9-10 , comprising performing a size selection process to isolate a target subset of identifiers.

12 . The method as in any one of claims 8-11 , comprising performing directed digestion to digest the one or more query components in identifiers that contain the one or more query components using a nuclease guided by a query oligonucleotide.

13 . The method of claim 12 , comprising performing multiple cycles of directed digestion.

14 . The method as in any one of claims 12-13 , comprising performing a size selection process to isolate a target subset of identifiers.

15 . The method as in any one of claims 8-14 , comprising performing selective hybridization to hybridize a modified oligonucleotide to one or more query components in single-stranded identifiers that contain the one or more query components.

16 . The method of claim 15 , comprising performing multiple cycles of selective hybridization.

17 . The method as in any one of claims 15-16 , comprising performing removal of hybridized identifiers to extract a target subset of identifiers.

18 . The method as in any one of claims 15-17 , wherein each of the modified oligonucleotides is attached to a bead or a surface.

19 . The method as in any one of claims 8-18 , comprising performing selective hybridization to hybridize a query oligonucleotide to one or more query components in single-stranded identifiers and to hybridize oligonucleotides corresponding to all components in layers other than layers of the query components, and performing digestion of unhybridized single-strand components to isolate a target subset of identifiers.

20 . The method as in any one of claims 7-19 , comprising performing a sampling operation.

21 . The method as in any one of claims 1-20 , comprising performing a plurality of cycles of steps (a)-(d).

22 . The method as in any one of claims 1-21 , comprising modulating a copy number of a subset of identifiers to encode a relationship among a set of data objects.

23 . The method as in any one of claims 2-22 , comprising performing one or more editing processes on a CVI.

24 . The method of claim 23 , wherein the editing process comprises using a guide oligonucleotide to direct a programmable nuclease-recombinase enzyme complex to a target component and replacing the target component with a new template oligonucleotide using the recombinase.

25 . The method as in any one of claims 23-24 , wherein the editing process comprises PCR using as one PCR primer an oligonucleotide that (a) only partially matches the identifier and (b) comprises the sequence for a modified adjacent component.

26 . The method as in any one of claims 23-25 , wherein the editing process comprises adding one or more components by PCR using as one PCR primer an oligonucleotide comprising the sequence for an adjacent component.

27 . The method as in any one of claims 1-26 , comprising performing a key extraction process wherein the key layer of each identifier is separated from the data layer to generate a set of sub-identifiers.

28 . The method of claim 27 , wherein the set of sub-identifiers comprises the key layers.

29 . The method of claim 28 , comprising performing a count operation of the sub-identifiers.

30 . The method as in any one of claims 28-29 , comprising performing a conversion operation wherein a data object is converted from binary to unary form.

31 . The method of claim 30 , comprising performing PCR on an aliquot of the subset of identifiers, wherein the number of PCR cycles corresponds to a bit position of at least one identifier.

32 . The method as in any one of claims 26-31 , comprising performing an addition or subtraction operation.

33 . The method of claim 32 , wherein the addition operation comprises pooling two or more identifier libraries.

34 . The method of claim 32 , wherein the subtraction operation A-B comprises:

(a) converting the binary representations of A by a first CVI and B by a second CVI into unary form;

(b) performing chemical procedure to melt the sub-identifiers into their constituent strands and retaining only the positive strand of each sub-identifier of A and only the negative strand of each sub-identifier of B;

(c) pooling the retained sub-identifiers;

(d) heating the pooled sub-identifiers at a temperature conducive to re-annealing of the positive and negative strands;

(e) discarding any double stranded obtained from process (e); and

(f) isolating and counting the remaining sub-identifiers representing A or B.

35 . The method of claim 34 , comprising, after step (e), converting the remaining single-stranded sub-identifiers of A and/or B into double-stranded sub-identifiers.

36 . The method as in any one of claims 2-35 , comprising writing a CVI encoding a data object into a plurality of separate samples.

37 . The method as in any one of claims 1-36 , wherein the one or more data layers encode a symbol value or a symbol position, or both.

38 . The method as in any one of claims 1-37 , wherein the plurality of components are assembled in a linear order.

39 . The method as in any one of claims 1-38 , wherein said coded sequence of symbols comprises symbols taken from a fixed alphabet of symbols.

40 . The method as in any one of claims 1-38 , further comprising converting said coded sequence into a second sequence of symbols.

41 . The method of claim 40 , wherein said second sequence of symbols comprises a formal data structure.

42 . The method of claim 41 , wherein said formal data structure comprises one or more members selected from the group consisting of a tree structure, a trie structure, a table structure, a key-value dictionary structure, and a set.

43 . The method as in any one of claims 41-42 , wherein said formal data structure is queryable by range queries, rank queries, count queries, membership queries, nearest neighbor queries, match queries, selection queries, or any combination thereof.

44 . The method as in any one of claims 41-43 , further comprising parsing said second sequence of symbols into a sequence of words.

45 . The method of claim 44 , further comprising converting said sequence of words into said sequence of codewords using one or more codebooks.

46 . The method of claim 45 , further comprising converting said sequence of codewords into a third sequence of symbols.

47 . The method of claim 46 , wherein converting said sequence of words into said sequence of codewords minimizes a number of one or more types of symbols in said third sequence of symbols.

48 . The method as in any one of claims 1-47 , wherein said coded sequence of symbols comprises one or more blocks of symbols.

49 . The method of claim 48 , wherein converting said sequence of words into said sequence of codewords generates a fixed number of one or more types of symbols in each block of symbols of said one or more blocks of symbols in said third sequence of symbols.

50 . The method as in any one of claims 1-49 , wherein a codebook appends one or more error protection symbols to individual codewords of said sequence of codewords.

51 . The method of claim 50 , wherein said one or more error protection symbols are computed from one or more words of said sequence of words.

52 . The method as in any one of claims 1-51 , wherein said plurality of identifiers are selected from a combinatorial space of identifiers.

53 . The method as in any one of claims 1-52 , wherein an individual component of said plurality of components comprises a nucleic acid sequence.

54 . The method of claim 53 , wherein said nucleic acid sequence is a distinct sequence.

55 . The method as in any one of claims 1-54 , wherein a presence of said individual identifier in said identifier library corresponds to a first symbol value and an absence of said individual identifier from said identifier library corresponds to a second symbol value.

56 . The method of claim 55 , wherein said first symbol value is ‘1’ and said second symbol value is ‘0’.

57 . The method of claim 55 , wherein said first symbol value is ‘0’ and said second symbol value is ‘1’.

58 . The method as in any one of claims 1-57 , wherein said identifier library comprises supplemental nucleic acid sequences.

59 . The method of claim 58 , wherein said supplemental nucleic acid sequences comprise metadata about said first sequence of symbols or an encoding of said first sequence of symbols.

60 . The method of claim 58 , wherein said supplemental nucleic acid sequences do not correspond to digital information and wherein said supplemental nucleic acid sequences conceal said digital information encoded in said identifier library.

61 . The method as in any one of claims 1-60 , wherein said one or more identifier(s) are generated by combinatorial assembly of one or more components.

62 . The method as in any one of claims 1-61 , further comprising constructing a universal identifier library.

63 . The method of claim 62 , wherein said identifier library is constructed from said universal identifier library by degrading or excluding said individual identifiers that are not present in identifier library.

64 . The method of claim 62 , wherein constructing said universal identifier library comprises using one or more reactions.

65 . The method of claim 64 , wherein said one or more reactions that correspond to said individual identifiers not present in said identifier library are removed, deleted, degraded, or inhibited.

66 . The method of claim 64 , wherein said one or more reactions comprise components, templates and/or reagents and wherein said components, said templates, and/or said reagents are loaded on films, threads, fibers, or other substrates.

67 . The method of claim 66 , wherein said components, said templates, and/or said reagents are dispensed adjacent to one another or collocated by stamping, intertwining, braiding, pinching, or weaving said films, said threads, said fibers, or said other substrates.

68 . An integrated nucleic acid-based storage system comprising:

a data encoding unit configured to write digital information in one or more nucleic acid sequences, wherein said data encoding unit writes said digital information in said one or more nucleic acid sequences in the absence of base-by-base nucleic acid synthesis;

a storage unit configured to store said one or more nucleic acid sequences encoding said digital information;

a reading unit configured to access and read said digital information encoded in said one or more nucleic acid sequences; and

one or more computer processors operatively coupled to said data encoding unit, said storage unit, and said reading unit, wherein said one or more computer processors are individually or collectively programmed to

(i) direct said data encoding unit to encode said digital information into said one or more nucleic acid sequences,

(ii) direct said storage unit to store said digital information encoded into said one or more nucleic acid sequences, and

(iii) direct said reading unit to access and decode said digital information stored in said one or more nucleic acid sequences;

where encoding comprises:

(a) coding said digital information into a sequence of symbols and converting said sequence of symbols into codewords;

(b) parsing said codewords into a coded sequence of symbols;

(c) mapping said coded sequence of symbols to a plurality of identifiers, wherein an individual identifier of said plurality of identifiers comprises one or more nucleic acid sequences; and

(d) enumerating an identifier library wherein each symbol of said coded sequence of symbols is encoded by one or more identifier(s), wherein an individual identifier of said plurality of identifiers comprises a plurality of components from one or more layers, each layer of said one or more layers comprising a distinct set of components, each identifier of the identifier library comprising:

one or more data layers encoding a symbol, and

one or more key layers encoding a key.

69 . The system of claim 68 , wherein a set of identifiers having identical components in the key layers forms a codeword vector of identifiers (CVI).

70 . The system as in any one of claims 68-69 , wherein the digital information comprises data values represented by a sequence of bits, and coding comprises converting the sequence of bits into bit vectors, each bit of the bit vectors mapped to one or more identifiers.

71 . The system of claim 70 , wherein a data value is be encoded with its original bit representation.

72 . The system of claim 70 , wherein a data value is be encoded with a sparser or denser bit vector.

73 . The system as in any one of claims 70-72 , wherein the data layers encode the bit vectors.

74 . The system as in any one of claims 68-73 , wherein encoding or decoding comprises performing a data query operation on the identifier library

75 . The system of claim 74 , wherein the query operation is performed on one or more query components form a subset of layers.

76 . The system of claim 75 , wherein encoding or decoding comprises performing selective polymerase chain reaction (PCR) using query oligonucleotides to amplify only identifiers that contain the one or more query components.

77 . The system of claim 76 , wherein encoding or decoding comprises performing multiple cycles of PCR with different query oligonucleotides in each cycle.

78 . The system as in any one of claims 76-77 , wherein encoding or decoding comprises performing a size selection process to isolate a target subset of identifiers.

79 . The system as in any one of claims 75-78 , wherein encoding or decoding comprises performing directed digestion to digest the one or more query components in identifiers that contain the one or more query components using a nuclease guided by a query oligonucleotide.

80 . The system of claim 79 , wherein encoding or decoding comprises performing multiple cycles of directed digestion.

81 . The system as in any one of claims 79-80 , wherein encoding or decoding comprises performing a size selection process to isolate a target subset of identifiers.

82 . The system as in any one of claims 75-81 , wherein encoding or decoding comprises performing selective hybridization to hybridize a modified oligonucleotide to one or more query components in single-stranded identifiers that contain the one or more query components.

83 . The system of claim 82 , wherein encoding or decoding comprises performing multiple cycles of selective hybridization.

84 . The system as in any one of claims 82-83 , wherein encoding or decoding comprises performing removal of hybridized identifiers to extract a target subset of identifiers.

85 . The system as in any one of claims 82-84 , wherein each of the modified oligonucleotides is attached to a bead or a surface.

86 . The system as in any one of claims 75-85 , wherein encoding or decoding comprises performing selective hybridization to hybridize a query oligonucleotide to one or more query components in single-stranded identifiers and to hybridize oligonucleotides corresponding to all components in layers other than layers of the query components, and performing digestion of unhybridized single-strand components to isolate a target subset of identifiers.

87 . The system as in any one of claims 75-86 , wherein encoding or decoding comprises performing a sampling operation.

88 . The system as in any one of claims 68-87 , wherein encoding comprises performing a plurality of cycles of steps (a)-(d).

89 . The system as in any one of claims 68-88 , wherein encoding comprises modulating a copy number of a subset of identifiers to encode a relationship among a set of data objects.

90 . The system as in any one of claims 69-89 , wherein encoding comprises performing one or more editing processes on a CVI.

91 . The system of claim 90 , wherein the editing process comprises using a guide oligonucleotide to direct a programmable nuclease-recombinase enzyme complex to a target component and replacing the target component with a new template oligonucleotide using the recombinase.

92 . The system as in any one of claims 90-91 , wherein the editing process comprises PCR using as one PCR primer an oligonucleotide that (a) only partially matches the identifier and (b) comprises the sequence for a modified adjacent component.

93 . The system as in any one of claims 90-92 , wherein the editing process comprises adding one or more components by PCR using as one PCR primer an oligonucleotide comprising the sequence for an adjacent component.

94 . The system as in any one of claims 68-93 , comprising performing a key extraction process wherein the key layer of each identifier is separated from the data layer to generate a set of sub-identifiers.

95 . The system of claim 94 , wherein the set of sub-identifiers comprises the key layers.

96 . The system of claim 95 , wherein encoding or decoding comprises performing a count operation of the sub-identifiers.

97 . The system as in any one of claims 95-96 , wherein encoding or decoding comprises performing a conversion operation wherein a data object is converted from binary to unary form.

98 . The system of claim 97 , wherein encoding or decoding comprises performing PCR on an aliquot of the subset of identifiers, wherein the number of PCR cycles corresponds to a bit position of at least one identifier.

99 . The system as in any one of claims 93-98 , wherein encoding or decoding comprises performing an addition or subtraction operation.

100 . The system of claim 99 , wherein the addition operation comprises pooling two or more identifier libraries.

101 . The system of claim 99 , wherein the subtraction operation A-B comprises:

(a) converting the binary representations of A by a first CVI and B by a second CVI into unary form;

(b) performing chemical procedure to melt the sub-identifiers into their constituent strands and retaining only the positive strand of each sub-identifier of A and only the negative strand of each sub-identifier of B;

(c) pooling the retained sub-identifiers;

(d) heating the pooled sub-identifiers at a temperature conducive to re-annealing of the positive and negative strands;

(e) discarding any double stranded obtained from process (e); and

(f) isolating and counting the remaining sub-identifiers representing A or B.

102 . The system of claim 101 , comprising, after step (e), converting the remaining single-stranded sub-identifiers of A and/or B into double-stranded sub-identifiers.

103 . The system as in any one of claims 69-102 , wherein encoding or decoding comprises writing a CVI encoding a data object into a plurality of separate samples.

104 . The system as in any one of claims 68-103 , wherein the one or more data layers encode a symbol value or a symbol position, or both.

105 . The system as in any one of claims 68-104 , wherein the plurality of components are assembled in a linear order.

106 . The system as in any one of claims 68-105 , wherein said system is automated.

107 . The system as in any one of claims 68-106 , wherein said system is networked.

108 . The system as in any one of claims 68-107 , wherein said identifier library generated is a universal library.

109 . The system as in any one of claims 68-108 , further comprising a plurality of modules.

110 . The system of claim 109 , wherein a first module creates an identifier library.

111 . The system as in any one of claims 109-110 , wherein a second module implements deletion of said individual identifiers or of an identifier reaction.

112 . The system as in any one of claims 109-111 , wherein a third module separates said individual identifiers present in said identifier library from said individual identifiers not present in said identifier library.

113 . The system as in any one of claims 109-112 , wherein a fourth module groups or pools said identifier library into one or more partitions.

114 . The system as in any one of claims 68-113 , wherein one or more reaction compartments, vessels, partitions, or substrates are mounted or stored on a disc, a plate, a film, a fiber, a tape, or a thread separate from said system before, after, or both before and after generation of said identifier library or a universal library.

115 . The method of claim 19 , wherein the single-stranded identifiers are immobilized on beads or a surface.

116 . The system of claim 86 , wherein the single-stranded identifiers are immobilized on beads or a surface.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2026
From: CATALOG TECHNOLOGIES, INC.
To: BIOMEMORY AMERICA, LLC
Reel/Frame 075235/0936 →
SECURITY INTEREST Recorded Oct 3, 2025
From: CATALOG TECHNOLOGIES, INC.
To: HANWHA IMPACT NEW TECH LLC
Reel/Frame 072998/0067 →