System and method for data compaction utilizing mismatch probability estimation
Codebook data compaction using a universal codebook and mismatch probability estimations to improve entropy encoding methods. Training data sets are analyzed to determine the frequency of occurrence of each sourceblock in the training data sets. A mismatch probability estimate is calculated comprising an estimated frequency at which any given data sourceblock received during encoding will not have a codeword in the codebook. Entropy encoding is used to generate codebooks comprising codewords for data sourceblocks based on the frequency of occurrence of each sourceblock. A “mismatch codeword” is inserted into the codebook based on the mismatch probability estimate to represent those cases when a block of data to be encoded does not have a codeword in the codebook.
1. A system for codebook data compaction using a universal codebook and mismatch probability estimations, comprising:
a computing device comprising a processor, a memory, and a non-volatile data storage device;
a codebook node comprising a first plurality of programming instructions stored in the memory which, when operating on the processor, causes the computing device to:
receive digital data to be compacted using a codebook from a source computing device;
add new sourceblocks to the codebook based on received non-training data to be compacted;
create a behavior codebook from a set of rules, limitations, policies that specify:
prioritization of which pieces of source data should be encoded with which codewords;
limits on types and sizes of source blocks that may be compacted; and
parameters for recursive compaction;
wherein the behavior codebook alters or determines the specific operations of the data compaction using the codebook;
encode and decode data using the codebook and the behavior codebook; and
return the newly encoded or decoded data to the source computing device.
2. The system of claim 1 , wherein the source computing device is the same computing device that operates the codebook node.
3. The system of claim 1 , wherein the source computing device is a mobile device, smartphone, tablet, laptop computer, desktop computer, server, or Internet-Of-Things device.
4. The system of claim 1 , wherein the codebook node uses a machine learning engine to refine or optimize the codebook used for data compaction.
5. The system of claim 1 , wherein the codebook node is a server that communicates with the source computing device over a network.
6. A method for codebook data compaction using a universal codebook and mismatch probability estimations, comprising the steps of:
receiving digital data to be compacted using a codebook, from a source computing device, using a codebook node;
adding new sourceblocks to the codebook based on received non-training data to be compacted, using the codebook node;
creating a behavior codebook from a set of rules, limitations, policies that specify:
prioritization of which pieces of source data should be encoded with which codewords;
limits on types and sizes of source blocks that may be compacted; and
parameters for recursive compaction;
wherein the behavior codebook alters or determines the specific operations of the data compaction using the codebook, using the codebook node;
encoding and decode data using the codebook and the behavior codebook, using the codebook node; and
returning the newly encoded or decoded data to the source computing device, using a codebook node.
7. The method of claim 6 , wherein the source computing device is the same computing device that operates the codebook node.
8. The method of claim 6 , wherein the source computing device is a mobile device, smartphone, tablet, laptop computer, desktop computer, server, or Internet-Of-Things device.
9. The method of claim 6 , wherein the codebook node uses a machine learning engine to refine or optimize the codebook used for data compaction.
10. The method of claim 6 , wherein the codebook node is a server that communicates with the source computing device over a network.