IP Library › Granted Patent US 10,073,897
Granted Patent B2
US 10,073,897 · App. 15/363,005 · Granted Sep 11, 2018

Transforming timeseries and non-relational data to relational for complex and analytical query processing

Inventors: Kevin Brown (San Rafael, CA); Frederick C. Ho (Campbell, CA); Raghupathi K. Murthy (Union City, CA); Karl Ostner (St. Wolfgang, DE)
Assignee: International Business Machines Corporation
G06F17/30563G06F17/30442G06F17/30486G06F17/30592G06F17/30312
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,073,897
App. No.
15/363,005
Granted
Sep 11, 2018
Kind
B2
Abstract

A system for transforming time series data into data that is accessible by a data warehouse identifies a data table comprising the time series data. The system creates a virtual view of the data table where the time series data is represented as at least one standard relational table in the virtual view, where the virtual view is presented as a virtual table. The system partitions the virtual table into a plurality of virtual partitions according to a time interval. The virtual table is partitioned across a data time range, where the data time range comprises at least one time interval, and where each of the plurality of virtual partitions has a respective partition time range that spans the time interval. The virtual partitions are created to optimize loading of the data into the data warehouse by incrementally refreshing the data according to the respective partition time range.

Claims (78)

1. A method for transforming data to be accessible by a data warehouse, implemented by a computing processor, the method comprising:

identifying, by the computing processor, a non-relational time series database table comprising time series data, wherein the non-relational time series database table comprises a column containing an array of the time series data collected from at least one source, wherein as new time series data is collected from the at least one source, the new time series data is added to the array instead of adding a new row for the new time series data;

creating, by the computing processor, a virtual view of the time series data in the non-relational time series database table by representing the time series data in the non-relational time series database table as a virtual relational database table, wherein the virtual view is stored as an in-memory storage structure without any intermediate storage, wherein the time series data is stored in the non-relational time series database table and not stored in a relational database table;

receiving a request with a user defined time interval to view the time series data, wherein the request is an SQL request; and

presenting a snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into a plurality of virtual partitions according to the user defined time interval across a data time range, wherein each of the plurality of virtual partitions has a respective partition time range that spans the user defined time interval, the plurality of virtual partitions created to optimize loading of the time series data in the virtual view into the snapshot by incrementally refreshing corresponding time series data for a corresponding virtual partition within the data time range.

2. The method of claim 1 comprising:

providing, by the computing processor, the plurality of virtual partitions to the snapshot for analysis of the time series data via a data accelerator, wherein selection of the plurality of virtual partitions, based on the data time range, optimizes analysis of the time series data.

3. The method of claim 1 wherein presenting the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval comprises:

incrementally refreshing the corresponding time series data by extracting new time series data from the non-relational time series database table, according to the user defined time interval;

creating a new virtual partition that spans the user defined time interval, the new virtual partition having a new respective partition time range; and

adding the new virtual partition to the plurality of virtual partitions.

4. The method of claim 1 wherein presenting the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval comprises:

incrementally refreshing the corresponding time series data by identifying the data time range as representing a chosen view of the time series data, wherein the chosen view comprises a subset of the plurality of virtual partitions, wherein each of the subset of the plurality of virtual partitions has a second respective partition time range that is within the data time range.

5. The method of claim 4 comprising:

detecting a first virtual partition, within the subset of the plurality of virtual partitions, having a first partition time range outside of the data time range; and

removing the first virtual partition from the subset of the plurality of virtual partitions, wherein the first virtual partition is no longer represented within the chosen view.

6. The method of claim 4 comprising:

detecting a second virtual partition having a second partition time range within the data time range, wherein the second virtual partition is not in the subset of the plurality of virtual partitions; and

adding the second virtual partition to the subset of the plurality of virtual partitions, wherein the second virtual partition is now represented within the chosen view.

7. The method of claim 1 wherein presenting the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval comprises:

creating a partitioning calendar;

associating the user defined time interval to the partitioning calendar;

assigning the partitioning calendar to the virtual relational database table; and

partitioning the virtual relational database table, using the partitioning calendar, wherein each partition spans the user defined time interval.

8. The method of claim 1 wherein presenting the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval comprises:

defining a time window selected to optimize an amount of relevant data that is loaded into the snapshot;

associating the user defined time interval to the time window; and

applying the time window to the virtual relational database table to partition the virtual table relational database into the plurality of virtual partitions according to the user defined time interval.

9. The method of claim 1 , wherein the time series data in the non-relational time series database table is not accessible using relational database queries, wherein the method further comprises:

providing access to the time series data in the non-relational time series database table through the request issued for the virtual relational database table.

10. A computer program product for transforming data to be accessible by a data warehouse, the computer program product comprising:

a computer readable storage medium having computer readable program code embodied therewith, the program code executable by a processor to:

identify, by the computing processor, a non-relational time series database table comprising time series data, wherein the non-relational time series database table comprises a column containing an array of the time series data collected from at least one source, wherein as new time series data is collected from the at least one source, the new time series data is added to the array instead of adding a new row for the new time series data;

create, by the computing processor, a virtual view of the time series data in the non-relational time series database table by representing the time series data in the non-relational time series database table as a virtual relational database table, wherein the virtual view is stored as an in-memory storage structure without any intermediate storage, wherein the time series data is stored in the non-relational time series database table and not stored in a relational database table;

receive a request with a user defined time interval to view the time series data, wherein the request is an SQL request; and

present a snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into a plurality of virtual partitions according to the user defined time interval across a data time range, wherein each of the plurality of virtual partitions has a respective partition time range that spans the user defined time interval, the plurality of virtual partitions created to optimize loading of the time series data in the virtual view into the snapshot by incrementally refreshing corresponding time series data for a corresponding virtual partition within the data time range.

11. The computer program product of claim 10 wherein the computer readable program code is further configured to:

provide, by the computing processor, the plurality of virtual partitions to the snapshot for analysis of the time series data via a data accelerator, wherein selection of the plurality of virtual partitions, based on the data time range, optimizes analysis of the time series data.

12. The computer program product of claim 10 wherein the computer readable program code configured to present the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval is further configured to:

incrementally refresh the corresponding time series data by extracting new time series data from the non-relational time series database table, according to the user defined time interval;

create a new virtual partition that spans the user defined time interval, the new virtual partition having a new respective partition time range; and

add the new virtual partition to the plurality of virtual partitions.

13. The computer program product of claim 10 wherein the computer readable program code configured to present the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval is further configured to:

incrementally refresh the corresponding time series data by identifying the data time range as representing a chosen view of the time series data, wherein the chosen view comprises a subset of the plurality of virtual partitions, wherein each of the subset of the plurality of virtual partitions has a second respective partition time range that is within the data time range.

14. The computer program product of claim 13 wherein the computer readable program code is further configured to:

detect a first virtual partition, within the subset of the plurality of virtual partitions, having a first partition time range outside of the data time range; and

remove the first virtual partition from the subset of the plurality of virtual partitions, wherein the first virtual partition is no longer represented within the chosen view.

15. The computer program product of claim 13 wherein the computer readable program code is further configured to:

detect a second virtual partition having a second partition time range within the data time range, wherein the second virtual partition is not in the subset of the plurality of virtual partitions; and

add the second virtual partition to the subset of the plurality of virtual partitions, wherein the second virtual partition is now represented within the chosen view.

16. The computer program product of claim 10 wherein the computer readable program code configured to present the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the time interval is further configured to:

create a partitioning calendar;

associate the user defined time interval to the partitioning calendar;

assign the partitioning calendar to the virtual table; and

partition the virtual relational database table, using the partitioning calendar, wherein each partition spans the user defined time interval.

17. The computer program product of claim 10 wherein the computer readable program code configured to present the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval is further configured to:

define a time window selected to optimize an amount of relevant data that is loaded into the snapshot;

associate the time interval to the time window; and

apply the time window to the virtual relational database table to partition the virtual relational database table into the plurality of virtual partitions according to the user defined time interval.

18. The computer program product of claim 10 , wherein the time series data in the non-relational time series database table is not accessible using relational database queries, wherein the computer readable program code is further configured to:

provide access to the time series data in the non-relational time series database table through the request issued for the virtual relational database table.

19. A system comprising:

a processor; and

a computer readable storage medium operationally coupled to the processor, the computer readable storage medium having computer readable program code embodied therewith to be executed by the processor, the computer readable program code configured to:

identify, by the computing processor, a non-relational time series database table comprising time series data, wherein the non-relational time series database table comprises a column containing an array of the time series data collected from at least one source, wherein as new time series data is collected from the at least one source, the new time series data is added to the array instead of adding anew row for the new time series data;

create, by the computing processor, a virtual view of the time series data in the non-relational time series database table by representing the time series data in the non-relational time series database table as a virtual relational database table, wherein the virtual view is stored as an in-memory storage structure without any intermediate storage, wherein the time series data is stored in the non-relational time series database table and not stored in a relational database table;

receive a request with a user defined time interval to view the time series data, wherein the request is an SQL request; and

present a snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into a plurality of virtual partitions according to the user defined time interval across a data time range, wherein each of the plurality of virtual partitions has a respective partition time range that spans the user defined time interval, the plurality of virtual partitions created to optimize loading of the time series data in the virtual view into the snapshot by incrementally refreshing corresponding time series data for a corresponding virtual partition within the data time range.

20. The system of claim 19 wherein the computer readable program code is further configured to:

provide, by the computing processor, the plurality of virtual partitions to the snapshot for analysis of the time series data via a data accelerator, wherein selection of the plurality of virtual partitions, based on the data time range, optimizes analysis of the time series data.

21. The system of claim 19 wherein the computer readable program code configured to present the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval is further configured to:

incrementally refresh the corresponding time series data by extracting new time series data from the non-relational time series database table, according to the user defined time interval;

create a new virtual partition that spans the time interval, the new virtual partition having a new respective partition time range; and

add the new virtual partition to the plurality of virtual partitions.

22. The system of claim 19 wherein the computer readable program code configured to present the snapshot of the time series data, in response to the request, by partitioning, by the computing processor, the virtual relational database table into the plurality of virtual partitions according to the user defined time interval is further configured to:

incrementally refresh the corresponding time series data by identifying the data time range as representing a chosen view of the time series data, wherein the chosen view comprises a subset of the plurality of virtual partitions, wherein each of the subset of the plurality of virtual partitions has a second respective partition time range that is within the data time range.

23. The system of claim 19 , wherein the time series data in the non-relational time series database table is not accessible using relational database queries, wherein the computer readable program code is further configured to:

provide access to the time series data in the non-relational time series database table through the request issued for the virtual relational database table.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE FOURTH ASSIGNOR'S EXECUTION DATE OF ASSIGNMENT DOCUMENT PREVIOUSLY RECORDED ON REEL 040448 FRAME 0931. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 13, 2016
From: BROWN, KEVIN; HO, FREDERICK C.; MURTHY, RAGHUPATHI K.; OSTNER, KARL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 040898/0294 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2016
From: BROWN, KEVIN; HO, FREDERICK C.; MURTHY, RAGHUPATHI K.; OSTNER, KARL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 040448/0931 →
Continuity (2)
Continuation 14153904 · Jan 13, 2014
Related Publication 20170075968A1 · Mar 16, 2017
Cited By (1)
US 12,688,135