IP Library › Granted Patent US 10,872,075
Granted Patent B2
US 10,872,075 · App. 15/655,667 · Granted Dec 22, 2020

Indexing flexible multi-representation storages for time series data

Inventors: Gordon Gaumnitz (Walldorf, DE); Lars Dannecker (Dresden, DE)
Assignee: SAP SE
G06F16/2365G06F16/2272G06F16/2477G06F16/24568G06F16/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,872,075
App. No.
15/655,667
Granted
Dec 22, 2020
Kind
B2
Abstract

Time series data may be represented with multiple representations, optionally using a variety of storage approaches, and the plurality of representations may be indexed using a representation index, which includes a start row identifier, a representation identifier, and an offset within the representation for each segment of one or more rows in the time series data column.

Claims (48)

1. A computer-implemented method comprising:

representing raw time series data in a time series data column with a plurality of representations of the raw time series data, wherein each of the plurality of representations refers to a storage approach of the raw time series data, wherein the plurality of representations comprise at least two storage approaches, wherein the at least two storage approaches comprise at least two different compression formats in which the raw time series data is stored;

indexing the plurality of representations using a representation index, the representation index comprising, for each segment of the plurality of representations, a start row identifier, a representation identifier corresponding to a representation of the plurality of representations, an end row identifier, and an offset value, wherein the offset value indicates a relevant position within the representation of the plurality of representations where the segment begins, wherein each segment comprises one or more rows in the time series data column associated with the representation of the plurality of representations; and

accessing the representation index instead of the time series data column to perform a data operation on the raw time series data.

2. The computer-implemented method of claim 1 , wherein the accessing of the representation index comprises fetching, from the representation index, the start row identifier, the representation identifier, and the offset for each of one or more representations spanning a set of rows to be operated on, and accessing the one or more representations based on the start row identifier, the representation identifier, and the offset.

3. The computer-implemented method of claim 1 , wherein the at least two storage approaches differ in two or more of compression, storage type, and approximation error.

4. The computer-implemented method of claim 1 , wherein the data operation comprises an update and/or an insert of a value in the time series data column, and the method further comprises:

creating a copy of the representation index; and

adding one or more new lines to the copy of the representation index to reflect a new segment of one or more rows in the time series data column.

5. The computer-implemented method of claim 1 , wherein the data operation comprises a deletion of a value in the time series data column and the method further comprises:

creating a copy of the representation index; and

deleting one or more existing lines from the copy of the representation index to reflect deletion of an existing segment one or more existing rows.

6. The computer-implemented method of claim 1 , wherein the data operation comprises two concurrent data modification transactions, and the method further comprises:

creating a first copy of the representation index for a first transaction of the two concurrent data modification transactions;

adding at least one first new line to the representation index to reflect a first new segment of one or more rows in the time series data column and/or deleting at least one first existing line from the first copy of the representation index to reflect deletion of an existing first segment comprising one or more existing rows;

creating a second copy of the representation index for a second transaction of the two concurrent data modification transactions;

adding at least one second new line to the representation index to reflect a second new segment of one or more rows in the time series data column and/or deleting at least one second existing line from the second copy of the representation index to reflect deletion of an existing second segment comprising one or more existing rows.

7. The computer-implemented method of claim 6 , further comprising merging the first copy of the representation index and the second copy of the representation index.

8. The computer-implemented method of claim 6 , further comprising aborting the second transaction when an attempt to merge the first copy of the representation index and the second copy of the representation index results in a conflict.

9. The computer-implemented method of claim 1 , further comprising creating at least one new representation to replace two or more representations referenced by the representation index when the representation index exceeds a threshold number of lines.

10. A computer program product comprising a non-transitory machine-readable medium storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to perform operations comprising:

representing raw time series data in a time series data column with a plurality of representations of the raw time series data, wherein each of the plurality of representations refers to a storage approach of the raw time series data, wherein the plurality of representations comprise at least two storage approaches, wherein the at least two storage approaches comprise at least two different compression formats in which the raw time series data is stored;

indexing the plurality of representations using a representation index, the representation index comprising, for each segment of the plurality of representations, a start row identifier, a representation identifier corresponding to a representation of the plurality of representations, an end row identifier, and an offset value, wherein the offset value indicates a relevant position within the representation of the plurality of representations where the segment begins, wherein each segment comprises one or more rows in the time series data column associated with the representation of the plurality of representations; and

accessing the representation index instead of the time series data column to perform a data operation on the raw time series data.

11. The computer program product of claim 10 , wherein the accessing of the representation index comprises fetching, from the representation index, the start row identifier, the representation identifier, and the offset for each of one or more representations spanning a set of rows to be operated on, and accessing the one or more representations based on the start row identifier, the representation identifier, and the offset.

12. The computer program product of claim 10 , wherein the at least two storage approaches differ in two or more of compression, storage type, and approximation error.

13. The computer program product of claim 10 , wherein the data operation comprises an update and/or an insert of a value in the time series data column, and the method further comprises:

creating a copy of the representation index; and

adding one or more new lines to the copy of the representation index to reflect a new segment of one or more rows in the time series data column.

14. The computer program product of claim 10 , wherein the data operation comprises a deletion of a value in the time series data column and the method further comprises:

creating a copy of the representation index; and

deleting one or more existing lines from the copy of the representation index to reflect deletion of an existing segment one or more existing rows.

15. The computer program product of claim 10 , wherein the data operation comprises two concurrent data modification transactions, and the method further comprises:

creating a first copy of the representation index for a first transaction of the two concurrent data modification transactions;

adding at least one first new line to the representation index to reflect a first new segment of one or more rows in the time series data column and/or deleting at least one first existing line from the first copy of the representation index to reflect deletion of an existing first segment comprising one or more existing rows;

creating a second copy of the representation index for a second transaction of the two concurrent data modification transactions;

adding at least one second new line to the representation index to reflect a second new segment of one or more rows in the time series data column and/or deleting at least one second existing line from the second copy of the representation index to reflect deletion of an existing second segment comprising one or more existing rows.

16. The computer program product of claim 15 , wherein the operations further comprise merging the first copy of the representation index and the second copy of the representation index.

17. The computer program product of claim 15 , wherein the operations further comprise aborting the second transaction when an attempt to merge the first copy of the representation index and the second copy of the representation index results in a conflict.

18. The computer program product of claim 10 , wherein the operations further comprise creating at least one new representation to replace two or more representations referenced by the representation index when the representation index exceeds a threshold number of lines.

19. A system comprising:

computer hardware configured to perform operations comprising:

representing raw time series data in a time series data column with a plurality of representations of the raw time series data, wherein each of the plurality of representations refers to a storage approach of the raw time series data, wherein the plurality of representations comprise at least two storage approaches, wherein the at least two storage approaches comprise at least two different compression formats in which the raw time series data is stored;

indexing the plurality of representations using a representation index, the representation index comprising, for each segment of the plurality of representations, a start row identifier, a representation identifier corresponding to a representation of the plurality of representations, an end row identifier, and an offset value, wherein the offset value indicates a relevant position within the representation of the plurality of representations where the segment begins, wherein each segment comprises one or more rows in the time series data column associated with the representation of the plurality of representations; and

accessing the representation index instead of the time series data column to perform a data operation on the raw time series data.

20. A system as in claim 19 , wherein the computer hardware comprises

a programmable processor; and

a machine-readable medium storing instructions that, when executed by the processor, cause the at least one programmable processor to perform at least some of the operations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2017
From: GAUMNITZ, GORDON; DANNECKER, LARS
To: SAP SE
Reel/Frame 043089/0567 →
Continuity (1)
Related Publication 20190026329A1 · Jan 24, 2019