Data storage method, and apparatus enabling selective and efficient data reading
This application relates to a data storage method, a reading method, an apparatus, a storage medium, and a program product. The data storage method is applied to a processor and includes: (S 1 ) obtaining event data; (S 2 ) processing the event data, to obtain a plurality of data blocks and a plurality of corresponding metadata blocks, where each data block includes a part of processed event data, and each metadata block includes a start time, an end time, a quantity, and a compression manner of event data in a data block corresponding to the metadata block, and a storage address of the data block; and (S 3 ) storing the plurality of data blocks into a first storage area, and storing the plurality of metadata blocks into a second storage area.
1 . A data storage method, wherein the method comprises:
obtaining event data;
processing the event data, to obtain a plurality of data blocks and a plurality of corresponding metadata blocks, wherein each data block of the plurality of data blocks comprises a part of processed event data, and each metadata block of the plurality of corresponding metadata blocks comprises a start time, an end time, a quantity, and a compression manner of event data in a data block corresponding to the metadata block, and a storage address of the data block; and
storing the plurality of data blocks into a first storage area, and storing the plurality of metadata blocks into a second storage area,
wherein the processing the event data, to obtain the plurality of data blocks and the plurality of corresponding metadata blocks comprises:
performing chunking on the event data based on time information and/or space information in the event data, to obtain a plurality of event chunks;
processing each event chunk of the plurality of event chunks, to obtain a plurality of sequences, wherein the plurality of sequences comprise time information, space information, and event content information of event data in the event chunk, a quantity of elements in each sequence is equal to a quantity of pieces of event data in the event chunk, an element value in each sequence of the plurality of sequences is an integer greater than or equal to 0, and a quantity of elements having element values greater than a first threshold value is greater than a quantity of elements having element values less than or equal to the first threshold value;
for each sequence, of the plurality of sequences selecting, from a preset compression manner, a compression manner matching the sequence, and compressing the sequence based on the selected compression manner, to obtain compressed data;
obtaining one data block based on a plurality of pieces of compressed data obtained based on the plurality of sequences corresponding to each event chunk; and
obtaining, based on time information in event data comprised in each event chunk, a quantity of pieces of event data, a selected compression manner, and a storage address of the data block, to obtain one metadata block corresponding to the data block.
2 . The data storage method according to claim 1 , wherein the plurality of sequences comprise a first sequence, and the first sequence comprises the time information of the event data in the event chunk; and
the processing each event chunk, to obtain a plurality of sequences comprises:
sorting the event data in the event chunk based on a time information order, and performing subtraction between time information of each piece of event data and time information of a previous piece of event data, to obtain a first group of differences; and
using time information of event data ranking first and all differences in the first group of differences respectively as elements in the first sequence, to obtain the first sequence.
3 . The data storage method according to claim 2 , wherein the plurality of sequences further comprise a second sequence, a third sequence, and a fourth sequence, the second sequence comprises first-dimensional space information of the event data in the event chunk, the third sequence comprises second-dimensional space information of the event data in the event chunk, and the fourth sequence comprises the event content information of the event data in the event chunk; and
the processing each event chunk, to obtain a plurality of sequences further comprises:
performing subtraction between first-dimensional space information of each piece of event data and first-dimensional space information of the previous piece of event data, to obtain a second group of differences;
converting, using a preset first function, a negative number in first-dimensional space information of the event data ranking first and the second group of differences into a positive number;
obtaining the second sequence based on converted first-dimensional space information of the event data ranking first and each difference in a converted second group of differences;
performing subtraction between second-dimensional space information of each piece of event data and second-dimensional space information of the previous piece of event data, to obtain a third group of differences;
converting, using the preset first function, a negative number in second-dimensional space information of the event data ranking first and the third group of differences into a positive number;
obtaining the third sequence based on converted second-dimensional space information of the event data ranking first and each difference in a converted third group of differences;
performing subtraction between event content information of each piece of event data and event content information of the previous piece of event data, to obtain a fourth group of differences;
converting, using the preset first function, a negative number in event content information of the event data ranking first and the fourth group of differences into a positive number; and
obtaining the fourth sequence based on converted event content information of the event data ranking first and each difference in a converted fourth group of differences.
4 . The data storage method according to claim 2 , wherein the plurality of sequences further comprise a fifth sequence and a sixth sequence, the fifth sequence and the sixth sequence comprise the space information of the event data in the event chunk, and the fifth sequence further comprises the event content information of the event data in the event chunk; and
the processing each event chunk, to obtain a plurality of sequences further comprises:
processing event content information of each piece of event data and first-dimensional space information or second-dimensional space information of the event data using a preset second function, to obtain synthetic information;
performing subtraction between synthetic information of each piece of event data and synthetic information of a previous piece of event data, to obtain a fifth group of differences;
converting, using a preset first function, a negative number in synthetic information of the event data ranking first and the fifth group of differences into a positive number;
obtaining the fifth sequence based on converted synthetic information of the event data ranking first and each difference in a converted fifth group of differences;
performing subtraction between second-dimensional space information of each piece of event data and second-dimensional space information of the previous piece of event data, or performing subtraction between first-dimensional space information of each piece of event data and first-dimensional space information of the previous piece of event data, to obtain a sixth group of differences;
converting, using the preset first function, a negative number in second-dimensional information or first-dimensional information of the event data ranking first and the sixth group of differences into a positive number; and
obtaining the sixth sequence based on converted second-dimensional information or first-dimensional information of the event data ranking first and each difference in a converted sixth group of differences.
5 . The data storage method according to claim 1 , wherein based on values of all elements in the sequence being the same and bit widths of all the elements being all b, the preset compression manner comprises a first compression manner; and
in the first compression manner, the values of all the elements in the sequence are recorded using a b-bit binary number, and a value of b is recorded using a B-bit binary number, wherein both b and B are positive integers.
6 . The data storage method according to claim 1 , wherein based on a largest bit width of elements in the sequence being b, the preset compression manner comprises a second compression manner; and
in the second compression manner, a value of each element in the sequence is recorded using a b-bit binary number, and a value of b is recorded using a B-bit binary number, wherein both b and B are positive integers.
7 . The data storage method according to claim 1 , wherein based on bit widths of at least two elements in the sequence being different, and a largest bit width of elements in the sequence being b, the preset compression manner comprises a third compression manner; and
in the third compression manner, a less than b is set, values of b and a each are recorded using a B-bit binary number, low-a-bit data in a binary form of each element in the sequence is recorded using a binary number having a bit width equal to a, high-(b-a)-bit data in the binary form of the element in the sequence is recorded using a binary number having a bit width equal to b-a, and whether each element comprises high-(b-a)-bit data is recorded using a binary number having a bit width equal to a quantity of elements in the sequence, wherein b, B, and a are all positive integers.
8 . A data reading method, wherein the method comprises:
obtaining a plurality of metadata blocks from a second storage area, wherein the plurality of metadata blocks correspond to a plurality of data blocks, each data block of the plurality of data blocks comprises a part of processed event data, and each metadata block of the plurality of metadata blocks comprises a start time, an end time, a quantity, and a compression manner of event data in a data block corresponding to the metadata block, and a storage address of the data block;
finding, in the plurality of metadata blocks, at least one metadata block that comprises a start time and an end time between which any time point is in a target time interval; and
obtaining, from a first storage area based on a storage address in the at least one found metadata block, at least one data block corresponding to the at least one found metadata block, and processing the data block based on a compression manner in a found metadata block corresponding to each data block, to obtain to-be-read event data,
wherein the processing the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain to-be-read event data comprises:
performing decompression processing on the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain a plurality of decompressed sequences, wherein the plurality of sequences comprise time information, space information, and event content information of event data in an event chunk corresponding to the data block, a quantity of elements in each sequence is equal to a quantity of pieces of event data in the event chunk, an element value in each sequence is an integer greater than or equal to 0, and a quantity of elements having element values greater than a first threshold value is greater than a quantity of elements having element values less than or equal to the first threshold value;
processing the plurality of decompressed sequences obtained through decompression processing performed on each data block, to obtain the event chunk corresponding to each data block; and
using, as the read event data, event data that is in event data comprised in all event chunks and whose time information is in the target time interval.
9 . The data reading method according to claim 8 , wherein the plurality of sequences comprise a first sequence, and the first sequence comprises the time information of the event data in the event chunk; and
the processing the plurality of decompressed sequences obtained through decompression processing performed on each data block, to obtain the event chunk corresponding to each data block comprises:
for any first sequence, performing summation on each element in the first sequence and all elements before the element, to obtain a first group of sum values; and
using a first element in the first sequence and all sum values in the first group of sum values respectively as the time information of the event data in the event chunk corresponding to the first sequence.
10 . The data reading method according to claim 9 , wherein the plurality of sequences further comprise a second sequence, a third sequence, and a fourth sequence, the second sequence comprises first-dimensional space information of the event data in the event chunk, the third sequence comprises second-dimensional space information of the event data in the event chunk, and the fourth sequence comprises the event content information of the event data in the event chunk; and
the processing the plurality of decompressed sequences obtained through decompression processing performed on each data block, to obtain the event chunk corresponding to each data block further comprises:
for any second sequence, third sequence, or fourth sequence, converting each element in the second sequence, the third sequence, or the fourth sequence by using an inverse function of a preset first function, to obtain a converted element;
performing summation on each converted element and all elements before the element, to obtain a second group of sum values, a third group of sum values, or a fourth group of sum values; and
using a converted first element and each sum value in a converted second group of sum values as the first-dimensional space information of the event data in the event chunk corresponding to the second sequence, or using a converted first element and each sum value in a converted third group of sum values as the second-dimensional space information of the event data in the event chunk corresponding to the third sequence, or using a converted first element and each sum value in a converted fourth group of sum values as the event content information of the event data in the event chunk corresponding to the fourth sequence.
11 . The data reading method according to claim 9 , wherein the plurality of sequences further comprise a fifth sequence and a sixth sequence, the fifth sequence and the sixth sequence comprise the space information of the event data in the event chunk, and the fifth sequence further comprises the event content information of the event data in the event chunk; and
the processing the plurality of decompressed sequences obtained through decompression processing performed on each data block, to obtain the event chunk corresponding to each data block further comprises:
for any fifth sequence, converting each element in the fifth sequence by using an inverse function of a preset first function, to obtain a converted element;
performing summation on each converted element and all elements before the element, to obtain a fifth group of sum values;
processing a converted first element and each sum value in the fifth group of sum values by using an inverse function of a preset second function, to obtain event content information of each piece of event data and first-dimensional space information or second-dimensional space information of the event data;
for any sixth sequence, converting each element in the sixth sequence by using the inverse function of the preset first function, to obtain a converted element;
performing summation on each converted element and all elements before the element, to obtain a sixth group of sum values; and
using a converted first element and each sum value in a converted sixth group of sum values as second-dimensional space information or first-dimensional space information of the event data in the event chunk corresponding to the sixth sequence.
12 . The data reading method according to claim 8 , wherein the data block comprises a plurality of pieces of compressed data, and the performing decompression processing on the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain a plurality of decompressed sequences comprises:
based on the compression manner corresponding to the metadata block indicating that a compression manner of any compressed data is a first compression manner, reading a B-bit binary number from the compressed data;
determining a value of b based on a decimal form of the B-bit binary number, and further reading a b-bit binary number from the compressed data; and
using a decimal form of the b-bit binary number as a value of each element in the sequence, to obtain a decompressed sequence, wherein both b and B are positive integers.
13 . The data reading method according to claim 8 , wherein the data block comprises a plurality of pieces of compressed data, and the performing decompression processing on the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain a plurality of decompressed sequences comprises:
based on the compression manner corresponding to the metadata block indicating that a compression manner of any compressed data is a second compression manner, reading a B-bit binary number from the compressed data;
determining a value of b based on a decimal form of the B-bit binary number;
further sequentially reading a b-bit binary number from the compressed data until a quantity of read b-bit binary numbers is equal to a quantity that is of pieces of event data and that is indicated by the metadata block; and
using a decimal form of the sequentially read b-bit binary number as a value of each element in the sequence, to obtain a decompressed sequence, wherein both b and B are positive integers.
14 . The data reading method according to claim 8 , wherein the data block comprises a plurality of pieces of compressed data, and the performing decompression processing on the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain a plurality of decompressed sequences comprises:
based on the compression manner corresponding to the metadata block indicating that a compression manner of any compressed data is a third compression manner, sequentially reading two B-bit binary numbers from the compressed data;
determining a value of b and a value of a respectively based on decimal forms of the two B-bit binary numbers;
further sequentially reading a-bit binary number from the compressed data until a quantity of read a-bit binary numbers is equal to a quantity that is of pieces of event data and that is indicated by the metadata block;
using the sequentially read a-bit binary number as low a bits of each element in the sequence;
further reading, from the compressed data, a binary number whose bit width is equal to the quantity that is of pieces of event data and that is indicated by the metadata block, and determining, based on each bit of the binary number, whether each element in the sequence comprises high-(b-a)-bit data;
further sequentially reading a (b-a)-bit binary number from the compressed data until a quantity of read (b-a)-bit binary numbers is equal to a quantity of determined elements that each comprise high-(b-a)-bit data; and
using the read (b-a)-bit binary number as high b-a bits of an element that comprises high-(b-a)-bit data in the sequence, wherein b, B, and a are all positive integers.
15 . A data storage apparatus, comprising:
a processor; and
a memory, configured to store instructions executable by the processor, wherein
execution of the instructions by the processor causes the processor to:
obtain event data;
process the event data, to obtain a plurality of data blocks and a plurality of corresponding metadata blocks, wherein each data block of the plurality of data blocks comprises a part of processed event data, and each metadata block of the plurality of metadata blocks comprises a start time, an end time, a quantity, and a compression manner of event data in a data block corresponding to the metadata block, and a storage address of the data block; and
store the plurality of data blocks into a first storage area, and store the plurality of metadata blocks into a second storage area,
wherein execution of the instructions to process the event data, to obtain the plurality of data blocks and the plurality of corresponding metadata blocks causes the processor to:
perform chunking on the event data based on time information and/or space information in the event data, to obtain a plurality of event chunks;
process each event chunk of the plurality of event chunks, to obtain a plurality of sequences, wherein the plurality of sequences comprise time information, space information, and event content information of event data in the event chunk, a quantity of elements in each sequence is equal to a quantity of pieces of event data in the event chunk, an element value in each sequence of the plurality of sequences is an integer greater than or equal to 0, and a quantity of elements having element values greater than a first threshold value is greater than a quantity of elements having element values less than or equal to the first threshold value;
for each sequence, of the plurality of sequences select, from a preset compression manner, a compression manner matching the sequence, and compress the sequence based on the selected compression manner, to obtain compressed data;
obtain one data block based on a plurality of pieces of compressed data obtained based on the plurality of sequences corresponding to each event chunk; and
obtain, based on time information in event data comprised in each event chunk, a quantity of pieces of event data, a selected compression manner, and a storage address of the data block, to obtain one metadata block corresponding to the data block.
16 . A data reading apparatus, comprising:
a processor; and
a memory, configured to store instructions executable by the processor, wherein
execution of the instructions by the processor causes the processor to:
obtain a plurality of metadata blocks from a second storage area, wherein the plurality of metadata blocks correspond to a plurality of data blocks, each data block of the plurality of data blocks comprises a part of processed event data, and each metadata block of the plurality of metadata blocks comprises a start time, an end time, a quantity, and a compression manner of event data in a data block corresponding to the metadata block, and a storage address of the data block;
find, in the plurality of metadata blocks, at least one metadata block that comprises a start time and an end time between which any time point is in a target time interval; and
obtain, from a first storage area based on a storage address in the at least one found metadata block, at least one data block corresponding to the at least one found metadata block, and process the data block based on a compression manner in a found metadata block corresponding to each data block, to obtain to-be-read event data,
wherein execution of the instructions to process the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain to-be-read event data causes the processor to:
perform decompression processing on the data block based on the compression manner in the found metadata block corresponding to each data block, to obtain a plurality of decompressed sequences, wherein the plurality of sequences comprise time information, space information, and event content information of event data in an event chunk corresponding to the data block, a quantity of elements in each sequence is equal to a quantity of pieces of event data in the event chunk, an element value in each sequence is an integer greater than or equal to 0, and a quantity of elements having element values greater than a first threshold value is greater than a quantity of elements having element values less than or equal to the first threshold value;
process the plurality of decompressed sequences obtained through decompression processing performed on each data block, to obtain the event chunk corresponding to each data block; and
use, as the read event data, event data that is in event data comprised in all event chunks and whose time information is in the target time interval.