Horizontally scalable multi-dimensional analysis architecture
A system and method are disclosed including a distributed computer network having one or more application servers and one or more data nodes. The one or more data nodes each include a shard of data representing a measure. The one or more application servers receive a measure model and an edit to the measure. The one or more application servers append the edit to a register and generate a calculation sequence. The one or more application servers further disaggregate one or more values to a shard of data on the one or more data nodes based, at least in part, on the edited measure and calculate one or more unedited measures based, at least in part, on the disaggregation.
1 . A system comprising:
a distributed computer network comprising one or more application servers and one or more data nodes, wherein the one or more data nodes each comprises a data shard representing one or more objects and one or more measures related to the one or more objects, the one or more application servers is configured to:
receive a measure model wherein the measure model comprises one or more rules, one or more measures, and one or more flexibility parameters, wherein a measure comprises one or more properties of data on which a calculation may be made, and further wherein each measure of the one or more measures has associated measure properties, the measure properties comprising a data type and an aggregation type, wherein the one or more rules further comprise one or more formulas that define a relationship between two or more measures, wherein the one or more flexibility parameters comprise an ordered list of one or more measures in a rule, the one or more measures authorized to be calculated in an event the one or more measures in the rule are edited, and wherein the measure model establishes invariant relationships;
receive an edit to a measure of the one or more measures, wherein the received edit is an update to the measure of the one or more measures;
append the edit to one or more queues, wherein the one or more queues comprise a database of edits;
generate a calculation sequence based on the one or more rules;
disaggregate one or more values to the shard of data on the one or more data nodes based, at least in part, on the edited measure, wherein a type of disaggregator used to disaggregate the one or more values is based, at least in part, on a measure property associated with the edited measure; and
calculate one or more unedited measures based on the disaggregation, wherein the application server performs protection processing when a calculation sequence is not able to be generated based on one or more edited measures and the one or more flexibility parameters and wherein the protection processing comprises testing each measure that could be edited to determine if the calculation sequence could be generated, wherein calculating the one or more unedited measures further comprises performing a sequence of calculations to calculate a value of the one or more unedited measures, so that all desired invariant relationships between measures are maintained, as defined in the measure model; and
wherein the one or more application servers generate the calculation sequence by:
avoiding any rule of the one or more rules that calculates an edited measure;
contributing exactly one rule to the calculation sequence from the one or more rules that comprises any measure edited or calculated by a rule in the calculation sequence; and
ordering rules of the calculation sequence so that any measure calculated by a first rule of the calculation sequence will not be used in a second rule of the calculation sequence that appears earlier in the sequence.
2 . The system of claim 1 , wherein disaggregation comprises spreading edits made above a level in a hierarchy where the data shard is stored in the one or more data nodes to a level where the data shard is stored in the one or more data nodes.
3 . The system of claim 2 , wherein the one or more application servers is further configured to:
aggregate a subset of the measures at levels where a subset of measures are not stored.
4 . A method comprising:
receiving a measure model over a computer network from one or more data nodes in a distributed computer network, each of the one or more data nodes comprises a data shard representing one or more objects and one or more measures related to the one or more objects, wherein the measure model comprises one or more rules, one or more measures, and one or more flexibility parameters, wherein a measure comprises one or more properties of data on which a calculation may be made, and further wherein each measure of the one or more measures has associated measure properties, the measure properties comprising a data type and an aggregation type, wherein the one or more rules further comprise one or more formulas that define a relationship between two or more measures, wherein the one or more flexibility parameters comprise an ordered list of one or more measures in a rule, the one or more measures authorized to be calculated in an event the one or more measures in the rule are edited, and wherein the measure model establishes invariant relationships;
receiving an edit to a measure of the one or more measures, wherein the received edit is an update to the measure of the one or more measures;
appending the edit to one or more queues using one or more application servers comprising a processor, wherein the one or more queues comprise a database of edits;
generating a calculation sequence based on the one or more rules using the one or more application servers;
disaggregating, using the one or more application servers, one or more values to the shard of data on the one or more data nodes based, at least in part, on the edited measure, wherein a type of disaggregator used to disaggregate the one or more values is based, at least in part, on a measure property associated with the edited measure; and
calculating, using the one or more application servers, one or more unedited measures based on the disaggregation, wherein when a calculation sequence is not able to be generated based on one or more edited measures and the one or more flexibility parameters, performing protection processing comprising testing each measure that could be edited to determine if the calculation sequence could be generated, wherein calculating the one or more unedited measures further comprises performing a sequence of calculations to calculate a value of the measures that a user has not edited, so that all of desired invariant relationships between measures are maintained, as defined in the measure model; wherein generating the calculation sequence further comprises:
eliminating any rule of the one or more rules that calculates an edited measure;
contributing exactly one rule to the calculation sequence from the one or more rules that comprises any measure edited or calculated by a rule in the calculation sequence; and
ordering rules of the calculation sequence so that a measure calculated by a first rule of the calculation sequence will not be used in a second rule of the calculation sequence that appears earlier in the sequence.
5 . The method of claim 4 , wherein disaggregation comprises spreading edits made above a level in a hierarchy where the data shard is stored in the one or more data nodes to a level where the data shard is stored in the one or more data nodes.
6 . The method of claim 5 , further comprising:
aggregating a subset of the measures at levels where a subset of measures are not stored.
7 . A non-transitory computer-readable medium embodied with software, the software when executed configured to:
receive a measure model over a computer network from one or more data nodes in a distributed computer network, each of the one or more data nodes comprises a data shard representing one or more objects and one or more measures related to the one or more objects, wherein the measure model comprises one or more rules, one or more measures, and one or more flexibility parameters, wherein a measure comprises one or more properties of data on which a calculation may be made, and further wherein each measure of the one or more measures has associated measure properties, the measure properties comprising a data type and an aggregation type, wherein the one or more rules further comprise one or more formulas that define a relationship between two or more measures, wherein the one or more flexibility parameters comprise an ordered list of one or more measures in a rule, the one or more measures authorized to be calculated in an event the one or more measures in the rule are edited, and wherein the measure model establishes invariant relationships;
receive an edit to a measure of the one or more measures, wherein the received edit is an update to the measure of the one or more measures;
append the edit to one or more queues, the one or more queues comprising a database of edits;
generate a calculation sequence based on the one or more rules;
disaggregate one or more values to the shard of data on the one or more data nodes based, at least in part, on the edited measure, wherein a type of disaggregator used to disaggregate the one or more values is based, at least in part, on a measure property associated with the edited measure;
calculate one or more unedited measures based on the disaggregation, wherein when a calculation sequence is not able to be generated based on one or more edited measures and the one or more flexibility parameters, perform protection processing comprising testing each measure that could be edited to determine if the calculation sequence could be generated, wherein calculating the one or more unedited measures further comprises performing a sequence of calculations to calculate a value of the measures that a user has not edited, so that all of desired invariant relationships between measures are maintained, as defined in the measure model;
eliminate any rule of the one or more rules that calculates an edited measure;
contribute exactly one rule to the calculation sequence from the one or more rules that comprises any measure edited or calculated by a rule in the calculation sequence; and
order rules of the calculation sequence so that a measure calculated by a first rule of the calculation sequence will not be used in a second rule of the calculation sequence that appears earlier in the sequence.
8 . The non-transitory computer-readable medium of claim 7 , wherein disaggregation comprises spreading edits made above a level in a hierarchy where the data shard is stored in the one or more data nodes to a level where the data shard is stored in the one or more data nodes.
9 . The non-transitory computer-readable medium of claim 8 , wherein the software is further configured to:
aggregate a subset of the measures at levels where the subset of measures are not stored.