IP Library › Granted Patent US 9,940,406
Granted Patent B2
US 9,940,406 · App. 14/670,208 · Granted Apr 10, 2018

Managing database

Inventors: Li Li (Beijing, CN); Liang Liu (Beijing, CN); Junmei Qu (Beijing, CN); Wen Jun Yin (Beijing, CN); Wei Zhuang (Beijing, CN)
Assignee: International Business Machine Corporation
G06F17/30917G06F17/30297G06F17/30315G06F17/30353G06F17/30486G06F17/30548
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,940,406
App. No.
14/670,208
Granted
Apr 10, 2018
Kind
B2
Abstract

The present invention discloses a method and system for managing a database. According to embodiments of the present invention, there is provided a method for managing a database, each item of data in the database being associated with a timestamp and a data point, the timestamps being used as row keys for rows of a table in the database, the method comprising: obtaining a behavior characteristic of a user based on a previous data access to the database by the user; partitioning columns in the table into column families based on the obtained behavior characteristic and system configuration of the database; and causing data in the database to be stored in respective column families at least in part based on the associated data point. There is further disclosed a corresponding system.

Claims (31)

1. A system for managing a database, each item of data in the database being associated with a timestamp and a data point, the timestamps being used as row keys for rows of a table in the database, the rows defined on one of two dimensions of the table, the system comprising:

a hardware processor in communication with said database, said hardware processor configured to perform a method comprising:

obtaining a behavior characteristic of a user based on a previous data access to the database by the user, said behavior characteristic indicating a time span of the previous data access;

partitioning columns in the table into column families based on the obtained behavior characteristic and a system configuration of the database, the system configuration indicating a size of blocks in a file system associated with the database, the columns defined on the other of the two dimensions of the table; and

storing data in the database in respective column families at least in part based on the associated data points.

2. The system according to claim 1 ,

wherein the hardware processor is further configured to:

generate at least one logical table from the table in the database, the number of rows in the at least one logical table determined based on the time span; and

determine the number of columns included in each of the column families at least in part based on the number of rows in the at least one logical table and the size of the blocks, such that the number of the blocks occupied by data in the column family is minimized.

3. The system according to claim 2 , wherein the behavior characteristic further indicates a data type of the previous data access, and wherein the number of columns included in each of the column families is further determined based on the data type.

4. The system according to claim 2 , wherein data in the table of the database is partitioned into a plurality of regions to store, the hardware processor being further configured to:

determine the number of the column families included in the at least one logical table based on a size of a region and the size of the blocks, such that the number of the regions occupied by data in the at least one logical table is minimized.

5. The system according to claim 2 , wherein the hardware processor is further configured to:

maintain an index for the table in the database, the index mapping identifications of the data points to the at least one logical table and the column families.

6. The system according to claim 1 , wherein the hardware processor is further configured to cause data for correlated data points to be stored in a same column family.

7. The system according to claim 1 , wherein the database is a Hadoop database (HBase).

8. A computer program product for managing a database, each item of data in the database being associated with a timestamp and a data point, the timestamps being used as row keys for rows of a table in the database, the rows defined on one of two dimensions of the table, the computer program product comprising a non-transitory storage medium readable by a processing circuit and storing instructions run by the processing circuit for performing a method, the method comprising:

obtaining a behavior characteristic of a user based on a previous data access to the database by the user, said behavior characteristic indicating a time span of the previous data access;

partitioning columns in the table into column families based on the obtained behavior characteristic and a system configuration of the database, the system configuration indicating a size of blocks in a file system associated with the database, the columns defined on the other of the two dimensions of the table; and

storing data in the database in respective column families at least in part based on the associated data points.

9. The computer program product according to claim 8 ,

wherein the partitioning columns in the table into column families comprises:

generating at least one logical table from the table, the number of rows in the at least one logical table determined based on the time span; and

determining the number of columns included in each of the column families at least in part based on the number of rows in the at least one logical table and the size of the blocks, such that the number of the blocks occupied by data in the column family is minimized.

10. The computer program product according to claim 9 , wherein the behavior characteristic further indicates a data type of the previous data access, and wherein the number of columns included in each of the column families is further determined based on the data type.

11. The computer program product according to claim 9 , wherein data in the table of the database is partitioned into a plurality of regions to store, the method further comprising:

determining the number of the column families included in the at least one logical table based on a size of a region and the size of the blocks, such that the number of the regions occupied by data in the at least one logical table is minimized.

12. The computer program product according to claim 9 , further comprising:

maintaining an index for the table in the database, the index mapping identifications of the data points to the at least one logical table and the column families.

13. The computer program product according to claim 8 , wherein causing data in the database to be stored in respective column families at least in part based on the associated data points comprises:

causing data for correlated data points to be stored in a same column family.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2015
From: LI, LI; LIU, LIANG; QU, JUNMEI; YIN, WEN JUN; ZHUANG, WEI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 035268/0482 →
Priority Claims (1)
CN 2014 1 0119830 · Mar 27, 2014 · national
Continuity (1)
Related Publication 20150278394A1 · Oct 1, 2015