Data classification using spatial data
Some examples relate generally to computer architecture software data classification and information security and, in some more particular aspects, to verifying information or events in a file system using spatial data.
1. A data management system, comprising
a first storage device configured to store a base file associated with a first version of a virtual machine;
a second storage device configured to store one or more forward incremental files associated with one or more versions of the virtual machine;
a processor-implemented text content verifier; and
one or more processors in communication with the first storage device and the second storage device, the one or more processors configured to perform operations including:
identifying a spreadsheet stored in a file in a monitored computer system, the spreadsheet including formatted text and row and column dimensions;
generating spreadsheet spatial metadata by parsing the formatted text and row and column dimensions, the spreadsheet spatial metadata comprising at least some of the formatted row and column dimensions;
incorporating at least some of the spreadsheet spatial metadata in a machine-readable key to verify data included within an audit event in a series of audit events including at least a create event, a write event, a read event, or a cleanup event; and
supplying the spreadsheet spatial metadata and the machine-readable key to the text content verifier for verifying the data included in the audit event.
2. The data management system of claim 1 , wherein the one or more processors are further configured to:
identify a data structure in the spreadsheet, the data structure including a table or a header; and
incorporate the identified data structure in the spreadsheet spatial metadata.
3. The data management system of claim 2 , wherein the text content verifier is included in a tiered array of text content verifiers and wherein the spreadsheet spatial metadata is supplied to the tiered array of text content verifiers.
4. The data management system of claim 3 , wherein a first text content verifier in the tiered array of text content verifiers is supplied first spreadsheet spatial metadata comprising keyword data sourced from a first ranked region of the spreadsheet, and a second text content verifier in the tiered array of text content verifiers is supplied second spreadsheet spatial metadata comprising aspects of the data structure sourced from a second ranked region of the spreadsheet.
5. The data management system of claim 4 , wherein the keyword data is sourced from a proximity-based keyword search based on a selection of a row dimension or a column dimension of the spreadsheet.
6. A computer-implemented method by a data management system, the method including operations comprising, at least:
identifying a spreadsheet stored in a file in a monitored computer system, the spreadsheet including formatted text and row and column dimensions;
generating spreadsheet spatial metadata by parsing the formatted text and row and column dimensions, the spreadsheet spatial metadata comprising at least some of the formatted row and column dimensions;
incorporating at least some of the spreadsheet spatial metadata in a machine-readable key to verify data included within an audit event in a series of audit events including at least a create event, a write event, a read event, or a cleanup event; and
supplying the spreadsheet spatial metadata and the machine-readable key to a text content verifier for verifying the data included in the audit event.
7. The method of claim 6 , wherein the operations further comprise:
identify a data structure in the spreadsheet, the data structure including a table or a header; and
incorporate the identified data structure in the spreadsheet spatial metadata.
8. The method of claim 7 , wherein the text content verifier is included in a tiered array of text content verifiers and wherein the spreadsheet spatial metadata is supplied to the tiered array of text content verifiers.
9. The method of claim 8 , wherein a first text content verifier in the tiered array of text content verifiers is supplied first spreadsheet spatial metadata comprising keyword data sourced from a first ranked region of the spreadsheet, and a second text content verifier in the tiered array of text content verifiers is supplied second spreadsheet spatial metadata comprising aspects of the data structure sourced from a second ranked region of the spreadsheet.
10. The method of claim 9 , wherein the keyword data is sourced from a proximity-based keyword search based on a selection of a row or column dimension of the spreadsheet.
11. A machine-storage medium storing instructions which, when read by a machine, cause the machine to perform operations comprising, at least:
identifying a spreadsheet stored in a file in a monitored computer system, he spreadsheet including formatted text and row and column dimensions;
generating spreadsheet spatial metadata by parsing the formatted text and row and column dimensions, the spreadsheet spatial metadata comprising at least some of the formatted row and column dimensions;
incorporating at least some of the spreadsheet spatial metadata in a machine-readable key to verify data included within an audit event in a series of audit events including at least a create event, a write event, a read event, or a cleanup event; and
supplying the spreadsheet spatial metadata and the machine-readable key to a text content verifier for verifying the data included in the audit event.
12. The medium of claim 11 , wherein the operations further comprise:
identify a data structure in the spreadsheet, the data structure including a table or a header; and
incorporate the identified data structure in the spreadsheet spatial metadata.
13. The medium of claim 12 , wherein the text content verifier is included in a tiered array of text content verifiers and wherein the spreadsheet spatial metadata is supplied to the tiered array of text content verifiers.
14. The medium of claim 13 , wherein a first text content verifier in the tiered array of text content verifiers is supplied first spreadsheet spatial metadata comprising keyword data sourced from a first ranked region of the spreadsheet, and a second text content verifier in the tiered array of text content verifiers is supplied second spreadsheet spatial metadata comprising aspects of the data structure sourced from a second ranked region of the spreadsheet.
15. The medium of claim 14 , wherein the keyword data is sourced from a proximity-based keyword search based on a selection of a row or column dimension of the spreadsheet.