IP Library › Granted Patent US 10,965,315
Granted Patent B2
US 10,965,315 · App. 16/059,633 · Granted Mar 30, 2021

Data compression method

Inventor: Andrew Kamal (Washington Township, MI)
H03M7/60G06F17/18H03M99/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,965,315
App. No.
16/059,633
Granted
Mar 30, 2021
Kind
B2
Abstract

An example method of compressing a data set includes determining whether individual values from a data set correspond to a first category or a second category of values. Based on one of the values corresponding to the first category, the value is added to a compressed data set. Based on one of the values corresponding to the second category, the value is excluded from the compressed data set, and a statistical distribution of values of the second category is updated based on the value. During a first phase, the determining is performed for a plurality of values from a first portion of the data set based on comparison of the values to criteria. During a second phase, the determining is performed for a plurality of values from a second portion of the data set based on the statistical distribution.

Claims (73)

1. A method of compressing a data set, comprising:

obtaining a data set and criteria for determining whether individual values from the data set correspond to a first category or a second category of values;

determining that some values of the data set correspond to the first category, and that other values of the data set correspond to the second category;

based on one of the values corresponding to the first category, adding the value to a compressed data set; and

based on one of the values corresponding to the second category:

excluding the value from the compressed data set; and

updating a statistical distribution of values of the second category in the data set based on the value;

wherein during a first phase, the determining is performed for a plurality of values from a first portion of the data set based on comparison of the values to the criteria; and

wherein during a second phase that is subsequent to the first phase, the determining is performed for a plurality of values from a second portion of the data set that is different from the first portion based on the statistical distribution.

2. The method of claim 1 , wherein values corresponding to the first category of data are more complex than values corresponding to the second category of data.

3. The method of claim 1 , comprising during the second phase:

determining a probability that a particular value from the second portion of the data set corresponds to the second category based on the statistical distribution; and

determining that the particular value corresponds to the second category based on the probability exceeding a predefined threshold.

4. The method of claim 3 , wherein said determining a probability that a particular value from the second portion of the data set corresponds to the second category based on the statistical distribution is performed based on Bayes' theorem.

5. The method of claim 1 , wherein said second phase is initiated in response to a trigger event.

6. The method of claim 5 , wherein:

each determination corresponds to an iteration;

a value from the data set is only added to the statistical distribution based on the value not already being present in the statistical distribution; and

the trigger event comprises no values from the first portion of the data set being added to the statistical distribution for a predefined quantity of consecutive iterations.

7. The method of claim 5 , wherein the trigger event comprises completion of said determining for a predefined portion of the data set.

8. The method of claim 1 , wherein during the first phase, determining whether a value of the data set corresponds to the first category or the second category comprises determining that the value corresponds to the first category based on the value being an irrational number.

9. The method of claim 1 , wherein during the first phase, determining whether a value of the data set corresponds to the first category or the second category comprises determining that the value corresponds to the first category based on the value being a complex number.

10. The method of claim 1 , wherein during the first phase, determining whether a value of the data set corresponds to the first category or the second category comprises determining that the value corresponds to the first category based on the value being a mixed hash that includes both numeric and alphabetical characters.

11. The method of claim 1 , wherein during the first phase, determining whether a value of the data set corresponds to the first category or the second category comprises determining that the value corresponds to the first category based on the value including a non-zero decimal value at or beyond an Xth decimal place, where X is a predefined value that is greater than nine.

12. The method of claim 1 , wherein during the first phase, determining whether a value of the data set corresponds to the first category or the second category comprises determining that the value corresponds to the second category based on the value being an integer.

13. The method of claim 1 , wherein said updating a statistical distribution of values of the second category in the data set based on the value comprises:

adding the value to the statistical distribution based on the value not already being present in the statistical distribution; and

updating the statistical distribution to reflect a quantity of times the value has been found in the data set based on the value already being in the statistical distribution.

14. The method of claim 1 , comprising during the second phase:

determining a redundancy of a particular value from the second portion of the data set within the data set; and

determining that the particular value corresponds to the second category based on the redundancy exceeding a predefined threshold.

15. The method of claim 1 , wherein the compressed data set is stored in a quadtree data structure.

16. The method of claim 15 , wherein the quadtree data structure is a point quadtree data structure.

17. The method of claim 15 , wherein:

values determined to correspond to the first category during first phase are stored in a first quadrant of the quadtree data structure; and

values determined to correspond to the first category during the second phase are stored in one or more other quadrants of the quadtree data structure that are different from the first quadrant.

18. The method of claim 17 , wherein the quadrant in which a given value is stored in the point quadtree data structure is based on which portion of the data set the value was obtained from.

19. The method of claim 15 , wherein:

the quadtree data structure includes four quadrants;

a quantum computing processor includes a plurality of qubits, each corresponding to one of the quadrants; and

the determination of whether a value corresponds to the first category and should be added to a particular quadrant is performed by one or more of the qubits corresponding to the particular quadrant.

20. The method of claim 1 , comprising:

verifying that values corresponding to the second category are not present in the compressed data set based on the Riemann zeta function.

21. The method of claim 20 , wherein said verifying that values corresponding to the second category are not present in the compressed data set based on the Riemann zeta function comprises:

determining a subset of values in the compressed data set that reside within a critical strip of the Riemann zeta function;

verifying whether the subset of values satisfy the criteria; and

based on a value from the subset not satisfying the criteria, excluding the value from the compressed data set.

22. A quantum computer comprising:

processing circuitry including a quantum processor having a plurality of qubits divided into

four groups, each group corresponding to a quadrant of a point quadtree data structure; the processing circuitry configured to:

obtain a data set and criteria for determining whether individual values from the data set correspond to a first category or a second category of values;

determine that some values of the data set correspond to the first category, and that other values of the data set correspond to the second category;

based on one of the values corresponding to the first category, add the value to a compressed data set in the point quadtree data structure; and

based on one of the values corresponding to the second category:

exclude the value from the compressed data set; and

update a statistical distribution of values of the second category in the data set based on the value;

wherein values from the data set corresponding to the first category are stored in multiple quadrants of the point quadtree data structure; and

wherein the determination of whether a value corresponds to the first category and should be added to a particular quadrant is performed by one or more of the qubits corresponding to the particular quadrant.

23. The quantum computer of claim 22 , wherein:

during a first phase, the determination is performed for a plurality of values from a first portion of the data set based on comparison of the values to the criteria, and

during a second phase that is subsequent to the first phase, the determination is performed for a plurality of values from a second portion of the data set that is different from the first portion based on the statistical distribution.

24. The quantum computer of claim 22 , wherein the quadrant in which a given value is stored in the point quadtree data structure is based on which portion of the data set the value was obtained from.

25. A computing device comprising

memory; and

a processing circuit operatively connected to the memory and configured to:

obtain a data set and criteria for determining whether individual values from the data set correspond to a first category or a second category of vales;

determine that some values of the data set correspond to the first category, and that other values of the data set correspond to the second category;

based on one of the values corresponding to the first category, add the value to a compressed data set; and

based on one of the values corresponding to the second category:

exclude the value from the compressed data set; and

update a statistical distribution of values of the second category in the data set based on the value;

wherein during a first phase, the determination is performed for a plurality of first values from a first portion of the data set based on comparison of the values to the criteria; and

wherein during a second phase that is subsequent to the first phase, the determination for a plurality of second values from a second portion of the data set that is different from the first portion is performed based on the statistical distribution.

Continuity (1)
Related Publication 20200052714A1 · Feb 13, 2020
Cited By (3)
US 12,418,310 US 12,499,174 US 12,725,023