IP Library Granted Patent US 9,098,454
Granted Patent B2
US 9,098,454 · App. 14/031,830 · Granted Aug 4, 2015

Speculative recovery using storage snapshot in a clustered database

Inventors: Douglas Griffith (Austin, TX); Angela Astrid Jaehde (Austin, TX); Matthew Ryan Ochs (Austin, TX)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F11/1458G06F11/1412G06F11/1469G06F11/1471G06F11/2028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,098,454
App. No.
14/031,830
Granted
Aug 4, 2015
Kind
B2
Abstract

A method for recovery in a database is provided in the illustrative embodiments. A failure is detected in a first computing node, the first computing node serving the database in a cluster of computing nodes. A snapshot is created of data of the database. A subset of log entries is applied to the snapshot, the applying modifying the snapshot to result in a modified snapshot. An access of the first computing node to the data of the database is preserved. Responsive to receiving a signal of activity from the first computing node during the applying and after a grace period has elapsed, the applying is aborted such that the first computing node can continue serving the database in the cluster.

Claims (27)

1. A method for recovery in a database, the method comprising:

detecting a failure in a first computing node, the first computing node serving the database in a cluster of computing nodes;

creating, using a processor and a memory, a snapshot of data of the database;

applying, using the processor and the memory, a subset of log entries to the snapshot, the applying modifying the snapshot to result in a modified snapshot;

preserving an access of the first computing node to the data of the database; and

aborting, responsive to receiving a signal of activity from the first computing node during the applying and after a grace period has elapsed, the applying such that the first computing node can continue serving the database in the cluster.

2. The method of claim 1 , further comprising:

receiving the signal of activity from the first computing node during the applying and after the grace period for the signal of activity has elapsed;

determining a level of performance of the first computing node;

completing the applying responsive to the level of performance begin below a threshold; and

taking over the serving of the database from the first computing node using the snapshot such that the first computing node cannot serve the database in the cluster.

3. The method of claim 2 , further comprising:

combining the modified snapshot with a modified data of the database, wherein the modified data of the database results from the first computing node continuing to modify the data of the database after the creating of the snapshot.

4. The method of claim 3 , wherein the combining comprises a reverse copy operation.

5. The method of claim 1 , further comprising:

communicating to a cluster management infrastructure a waiting period within which to allow the preserving.

6. The method of claim 5 , wherein the waiting period is an estimate of time needed to apply the subset of log entries, and wherein the waiting period is changeable during the applying.

7. The method of claim 1 , wherein the preserving comprises:

allowing the first computing node to continue manipulating the data of the database during the applying.

8. The method of claim 1 , wherein the creating and applying occur are a second computing node in the cluster of computing nodes, and wherein the second computing node and the first computing node concurrently serve the database from the cluster for a waiting period.

9. The method of claim 1 , the applying comprising:

selecting the subset of log entries, wherein the log entries comprise transaction information processed by the first computing node after a checkpoint operation and one of (i) up to a time of the failure, and (ii) prior to a time of the failure.

10. The method of claim 1 , the creating comprising:

making a copy of the data at a storage subsystem after the detecting, wherein the storage subsystem is used for storing the data of the database, wherein the making the copy comprises a data replication on the storage subsystem such that a change tracking and disambiguation of access data is resolved on a storage layer that includes the storage subsystem.

11. The method of claim 10 , wherein the making comprises performing a FlashCopy operation and wherein the copy uses a copy on write mode.

12. The method of claim 1 , wherein the failure comprises failure to receive the signal of activity from the first computing node within the grace period.

13. The method of claim 12 , wherein the signal of activity is a heartbeat message.

Assignments (6)
RELEASE OF SECURITY INTEREST Recorded May 12, 2021
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: GLOBALFOUNDRIES U.S. INC.
Reel/Frame 056987/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 20, 2020
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: GLOBALFOUNDRIES INC.
Reel/Frame 054636/0001 →
SECURITY AGREEMENT Recorded Nov 29, 2018
From: GLOBALFOUNDRIES INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 049490/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2015
From: GLOBALFOUNDRIES U.S. 2 LLC; GLOBALFOUNDRIES U.S. INC.
To: GLOBALFOUNDRIES INC.
Reel/Frame 036779/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2015
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: GLOBALFOUNDRIES U.S. 2 LLC
Reel/Frame 036550/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2013
From: GRIFFITH, DOUGLAS; JAEHDE, ANGELA ASTRID; OCHS, MATTHEW RYAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 031243/0356 →
Continuity (2)
Continuation 13940013 · Jul 11, 2013
Related Publication 20150019494A1 · Jan 15, 2015