IP Library › Granted Patent US 7,912,812
Granted Patent B2
US 7,912,812 · App. 11/970,337 · Granted Mar 22, 2011

Smart data caching using data mining

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,912,812
App. No.
11/970,337
Granted
Mar 22, 2011
Kind
B2
Abstract

Methods and apparatus, including computer program products, implementing and using techniques for populating a data cache on a server. Data requests received by the server are collected in a repository. A data mining algorithm is applied to the collected data requests to predict a set of data that is likely to be requested during an upcoming time period. It is determined whether the complete set of predicted data exists in the data cache. If the complete set of predicted data does not exist in the data cache, the missing data is retrieved from a database and added to the data cache.

Claims (38)

1. A computer-implemented method for populating a data cache on a server the method comprising:

collecting new data requests received by the server from a plurality of users in a repository containing historical data requests received by the server from the plurality of users;

applying a data mining algorithm to the data requests in the repository to discover patterns in the data requests received from the plurality of users and to predict a set of data that is likely to be requested by the plurality of users during an upcoming time period, based on the discovered patterns;

determining whether the complete set of predicted data exists in the data cache; and

in response to determining that the complete set of predicted data does not exist in the data cache, retrieving the missing data from a database and adding the missing data to the data cache.

2. The method of claim 1 , wherein applying the data mining algorithm includes applying the data mining algorithm at a time when no consumers are submitting data requests, wherein the data mining algorithm generates a predicted set of data based on past query patterns for the consumers.

3. The method of claim 1 , wherein applying the data mining algorithm includes applying the data mining algorithm in real time as data requests are received, wherein the data mining algorithm generates a predicted set of data based on identified patterns of the current data requests.

4. The method of claim 1 , wherein the data requests are in one of the following formats: a Structured Query Language request, an Extensible Markup Language request, and a Multi-Dimensional Expression request.

5. The method of claim 1 , wherein the data mining algorithms are selected from the group consisting of: clustering algorithms, associations algorithms and sequences algorithms.

6. The method of claim 1 , further comprising:

serving data from the cache to a consumer in response to a request from a consumer.

7. The method of claim 1 , further comprising:

identifying old data in the data cache that is not expected to be used during the next time period; and

clearing the data cache of the identified old data.

8. The method of claim 1 , wherein applying the data mining algorithm further includes:

parsing the collected data requests; and

storing data request attributes in one or more of: a transactional table layout format and a behavioral format.

9. The method of claim 8 , wherein the data request attributes include one or more of: a username for the user submitting the data request, a date the data request was submitted, a time the data request was submitted, and parsed elements of the requests.

10. The method of claim 1 , wherein the data cache is embodied in one of: a persistent storage medium, a non-persistent storage medium, and a combination persistent and non-persistent storage media.

11. A computer readable storage medium storing program instructions for populating a data cache on a server, the program instructions comprising:

first program instructions to collect new data requests received by the server from a plurality of users in a repository containing historical data requests received by the server from the plurality of users;

second program instructions to apply a data mining algorithm to the data requests in the repository to discover patterns in the data requests received from the plurality of users and to predict a set of data that is likely to be requested by the plurality of users during an upcoming time period, based on the discovered patterns;

third program instructions to determine whether the complete set of predicted data exists in the data cache; and

fourth program instructions to, in response to determining that the complete set of predicted data does not exist in the data cache, retrieve the missing data from a database and adding the missing data to the data cache.

12. The computer program product of claim 11 , wherein the program instructions to apply the data mining algorithm includes program instructions to apply the data mining algorithm at a time when no consumers are submitting data requests, wherein the data mining algorithm generates a predicted set of data based on past query patterns for the consumers.

13. The computer program product of claim 11 , wherein the program instructions to apply the data mining algorithm includes program instructions to apply the data mining algorithm in real time as data requests are received, wherein the data mining algorithm generates a predicted set of data based on identified patterns of the current data requests.

14. The computer program product of claim 11 , wherein the data requests are in one of the following formats: a Structured Query Language request, an Extensible Markup Language request, and a Multi-Dimensional Expression request.

15. The computer program product of claim 11 , wherein the data mining algorithms are selected from the group consisting of: clustering algorithms, associations algorithms and sequences algorithms.

16. The computer program product of claim 11 , wherein the computer readable storage medium further stores:

fifth program instructions to serve data from the cache to a consumer in response to a request from a consumer.

17. The computer program product of claim 11 , wherein the computer readable storage medium further stores:

sixth program instructions to identify old data in the data cache that is not expected to be used during the next time period; and

seventh program instructions to clear the data cache of the identified old data.

18. The computer program product of claim 11 , wherein the program instructions to apply the data mining algorithm include program instructions to:

parse the collected data requests; and

store data request attributes in one or more of: a transactional table layout format and a behavioral format.

19. The computer program product of claim 18 , wherein the data request attributes include one or more of: a username for the user submitting the data request, a date the data request was submitted, a time the data request was submitted, and parsed elements of the requests.

20. The computer program product of claim 11 , wherein the data cache is embodied in one of: a persistent storage medium, a non-persistent storage medium, and a combination persistent and non-persistent storage media.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2008
From: RAMOS, JO ARAO; ROLLINS, JOHN BAXTER; WILHITE, DAVID GIDDENS, JR.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 020327/0518 →
Continuity (1)
Related Publication 20090177667A1 · Jul 9, 2009