Object storage and access management systems and methods
A geographically distributed erasure coding system includes multiple computer readable, non-transitory storage memories capable of storing a digital dataset including multiple object blocks, where each storage memory is configured to store one or more of the object blocks of the dataset according to an erasure coding policy. The system includes one or more processors configured to implement the erasure coding policy by distributing the multiple object blocks of the dataset to the multiple storage memories according to distribution criteria of the erasure coding policy, and the distribution criteria include at least one status parameter associated with each storage memory. The multiple storage memories are geographically distributed at different locations from one another.
1 . A distributed digital object storage system comprising:
at least one non-transitory computer readable memory storing a suite of software instructions; and
at least one processor coupled with the at least one non-transitory computer readable memory and that performs the following operations upon execution of the software instructions:
establishing, in the at least one non-transitory computer readable memory, at least one erasure coding policy for a dataset, the at least one erasure coding policy including distribution criteria defined based on location attributes associated with distributed storage node locations of distributed storage nodes and based on a data chunk type attribute associated with the dataset, wherein the location attributes adhere to a common hierarchical namespace including S2 cell identifiers, each of the distributed storage node includes a location attribute having an S2 cell identifier value indicating a location of the distributed storage node according to a hierarchy of S2 cells, and the at least one erasure coding policy includes at least a specified distance between distributed storage node locations in the hierarchy of S2 cells based on S2 cell identifier values of the location attributes corresponding to each of the distributed storage nodes;
creating multiple object blocks from the dataset, each block of the multiple object blocks having a data chunk type attribute;
establishing separated block storage locations among the distributed storage nodes for the multiple object blocks according to the distribution criteria including the S2 cell identifier values of the location attributes of the distributed storage nodes and the data chunk type attributes of the multiple object blocks, wherein establishing separated block storage locations includes determining a minimum separation distance between separated block storage locations based on an S2 cell identifier value associated with each separated block storage location and at least one risk metric, and determining a maximum separation distance between separated block storage locations based on the S2 cell identifier value associated with each separated block storage location and a constraint associated with the at least one risk metric; and
storing, over a network, the multiple object blocks at their respective separated block storage locations on the distributed storage nodes according to the erasure coding policy.
2 . The distributed digital object storage system of claim 1 , wherein the location attributes comprise location identifiers.
3 . The distributed digital object storage system of claim 2 , wherein the location identifiers comprise geographic location identifiers.
4 . The distributed digital object storage system of claim 2 , wherein the location identifiers comprise at least one of the following: a zip code, a longitude, a latitude, or a plus code.
5 . The distributed digital object storage system of claim 1 , wherein the location attributes adhere to location address space of the distribute storage nodes.
6 . The distributed digital object storage system of claim 5 , wherein each of the multiple object blocks have identifiers adhering to the location address space.
7 . The distributed digital object storage system of claim 6 , wherein the location address space comprises a hash address space.
8 . The distributed digital object storage system of claim 1 , wherein the data chunk type attributes adhere to a common namespace.
9 . The distributed digital object storage system of claim 8 , wherein the common namespace comprises an a priori defined namespace.
10 . The distributed digital object storage system of claim 8 , wherein the common namespace comprises attribute-value pair parameters.
11 . The distributed digital object storage system of claim 1 , where in the at least two of the distributed storage nodes are separated by at least 100 miles.
12 . The distributed digital object storage system of claim 11 , wherein the at least two nodes are separated by at least 1000 miles.
13 . The distributed digital object storage system of claim 1 , wherein the distributed storage nodes comprise at least one of the following: a network area storage system, a storage area network system, or a RAID system.
14 . The distributed digital object storage system of claim 1 , wherein the at least one erasure coding policy enforces a code rate representing a ratio (r) of a number of blocks in the dataset (k) to a total number of block in the multiple object blocks (n), wherein n is greater than k.
15 . The distributed digital object storage system of claim 1 , wherein the data chunk type attribute of at least one block in the multiple object blocks represents a data chunk.
16 . The distributed digital object storage system of claim 15 , wherein the data chunk type attribute of the at least one block in the multiple object blocks represents a redundant data chunk.
17 . The distributed digital object storage system of claim 1 , wherein the data chunk type attribute of at least one block in the multiple object blocks represents a parity chunk.
18 . The distributed digital object storage system of claim 1 , wherein the operations further include recording an event related to at least one block in the multiple object blocks on a notarized ledger.
19 . A method of implementing an erasure coding policy for a distributed digital object storage system, the method comprising:
establishing, in at least one non-transitory computer readable memory, at least one erasure coding policy for a dataset, the at least one erasure coding policy including distribution criteria defined based on location attributes associated with distributed storage node locations of distributed storage nodes and based on a data chunk type attribute associated with the dataset, wherein the location attributes adhere to a common hierarchical namespace including S2 cell identifiers, each of the distributed storage node includes a location attribute having an S2 cell identifier value indicating a location of the distributed storage node according to a hierarchy of S2 cells, and the at least one erasure coding policy includes at least a specified distance between distributed storage node locations in the hierarchy of S2 cells based on S2 cell identifier values of the location attributes corresponding to each of the distributed storage nodes;
creating multiple object blocks from the dataset, each block of the multiple object blocks having a data chunk type attribute;
establishing separated block storage locations among the distributed storage nodes for the multiple object blocks according to the distribution criteria including the S2 cell identifier values of the location attributes of the distributed storage nodes and the data chunk type attributes of the multiple object blocks, wherein establishing separated block storage locations includes determining a minimum separation distance between separated block storage locations based on an S2 cell identifier value associated with each separated block storage location and at least one risk metric, and determining a maximum separation distance between separated block storage locations based on the S2 cell identifier value associated with each separated block storage location and a constraint associated with the at least one risk metric; and
storing, over a network, the multiple object blocks at their respective separated block storage locations on the distributed storage nodes according to the erasure coding policy.
20 . A non-transitory computer-readable medium comprising computer-executable instructions configured to, when executed by at least one processor, cause the processor to perform operations including:
establishing, in at least one non-transitory computer readable memory, at least one erasure coding policy for a dataset, the at least one erasure coding policy including distribution criteria defined based on location attributes associated with distributed storage node locations of distributed storage nodes and based on a data chunk type attribute associated with the dataset, wherein the location attributes adhere to a common hierarchical namespace including S2 cell identifiers, each of the distributed storage node includes a location attribute having an S2 cell identifier value indicating a location of the distributed storage node according to a hierarchy of S2 cells, and the at least one erasure coding policy includes at least a specified distance between distributed storage node locations in the hierarchy of S2 cells based on S2 cell identifier values of the location attributes corresponding to each of the distributed storage nodes;
creating multiple object blocks from the dataset, each block of the multiple object blocks having a data chunk type attribute;
establishing separated block storage locations among the distributed storage nodes for the multiple object blocks according to the distribution criteria including the S2 cell identifier values of the location attributes of the distributed storage nodes and the data chunk type attributes of the multiple object blocks, wherein establishing separated block storage locations includes determining a minimum separation distance between separated block storage locations based on an S2 cell identifier value associated with each separated block storage location and at least one risk metric, and determining a maximum separation distance between separated block storage locations based on the S2 cell identifier value associated with each separated block storage location and a constraint associated with the at least one risk metric; and
storing, over a network, the multiple object blocks at their respective separated block storage locations on the distributed storage nodes according to the erasure coding policy.