IP Library › Granted Patent US 10,445,197
Granted Patent B1
US 10,445,197 · App. 15/605,833 · Granted Oct 15, 2019

Detecting failover events at secondary nodes

Inventor: Harpreet (Singapore, SG)
Assignee: Amazon Technologies, Inc.
G06F11/2023G06F11/3006G06F11/3409H04L43/0811G06F2201/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,445,197
App. No.
15/605,833
Granted
Oct 15, 2019
Kind
B1
Abstract

Secondary nodes may detect failover operations for applications performing at a primary node. Application state indications may be collected at a primary node and reported to a monitor for the primary node. A secondary node may obtain the state indications from the monitor in order to evaluate whether the application is performing correctly at the primary and if not, trigger a failover operation to switch performance of the application to the secondary node. In some embodiments, the state information may be encoded into a single metric that can be obtained by the secondary node and evaluated to detect failover events.

Claims (45)

1. A system, comprising:

a memory to store program instructions which, if performed by at least one processor, cause the at least one processor to perform a method to at least:

report, by a primary node, one or more state indications of an application performing at the primary node to a monitor for the primary node;

poll, by a secondary node, the one or more state indications at the monitor;

detect, by the secondary node, a communication failure between the primary node and the secondary node;

in response to the detection of the communication failure:

evaluate, by the secondary node, the state indications to identify a failover operation for the primary node; and

perform, by the secondary node, the failover operation to switch performance of the application from the primary node to the secondary node.

2. The system of claim 1 , wherein to report the state indications, the method includes encode the state indications into a single performance metric, wherein the single performance metric is reported to the monitor.

3. The system of claim 1 , wherein to perform the failover operation to switch performance of the application from the primary node to the secondary node, the method includes send a request to halt performance of the application at the primary node.

4. The system of claim 1 , wherein the primary node and the secondary node are implemented as part of a virtual compute service offered by a provider network, wherein the monitor is implemented as part of a monitoring service for the provider network, and wherein the failover operation includes sending a request to a networking service implemented as part of the provider network to redirect requests for the application to the secondary node.

5. A method, comprising:

obtaining, by a first node, one or more state indications of an application from a monitor of a second node performing the application;

detecting, by the first node, a communication failure between the first node and the second node; and

in response to detecting the communication failure, causing, by the first node, performance of the application to switch from the second node to the first node based, at least in part, on an evaluation of the state indications for the application.

6. The method of claim 5 , further comprising:

identifying, by the second node, state indications for the application; and

sending, by the second node, the state indications to the monitor.

7. The method of claim 6 , further comprising encoding, by the second node, the state indications into a single performance metric, wherein the single performance metric is sent to the monitor.

8. The method of claim 7 , further comprising comparing, by the first node, the performance metric with a mapping for the state indications that identifies a failover operation for the application.

9. The method of claim 5 , wherein causing, by the first node, performance of the application to switch from the second node to the first node comprises redirecting work for the application from the second node to the first node.

10. The method of claim 5 , wherein causing, by the first node, performance of the application to switch from the second node to the first node comprises determining one or more failover operations to perform based, at least in part on the state indications.

11. The method of claim 5 , wherein causing, by the first node, performance of the application to switch from the second node to the first node comprises sending a request to perform a failover operation to a migration agent.

12. The method of claim 5 , further comprising:

upon switching performance of the application to the first node:

identifying, by the first node, one or more state indications for the performance of the application at the first node; and

sending, by the first node, the state indications for the performance of the application at the first node to the monitor.

13. The method of claim 5 , wherein the application is a database application that provides access to a database on behalf of database clients, wherein the first node is a secondary node that is preconfigured to access the database, and wherein the second node is a primary node for the database.

14. A non-transitory, computer-readable storage medium, storing program instructions that when executed by one or more computing devices cause the one or more computing devices to implement:

receiving, at a secondary node, one or more state indications of an application collected by a monitor of a primary node performing the application;

detecting, by the secondary node, a communication failure between the primary node and the secondary node;

in response to detecting the communication failure:

evaluating, by the secondary node, the state indications to identify a failover operation for the primary node; and

performing, by the secondary node, the failover operation to switch performance of the application from the primary node to the secondary node.

15. The non-transitory, computer-readable storage medium of claim 14 , wherein the program instructions cause the one or more computing devices to further implement:

identifying, by the primary node, the state indications for the application;

encoding, by the primary node, the state indications into a single performance metric; and

sending, by the primary node, the single performance metric to the monitor, wherein the single performance metric is received at the secondary node.

16. The non-transitory, computer-readable storage medium of claim 15 , wherein, in evaluating, by the secondary node, the state indications to identify the failover operation for the primary node, the program instructions cause the one or more computing devices to implement:

decoding the single performance metric into the state indications; and

comparing the state indications to failover event criteria to identify the failover operation.

17. The non-transitory, computer-readable storage medium of claim 14 , wherein, in receiving the state indications of the application collected by the monitor, the program instructions cause the one or more computing devices to implement sending via a public network a request to the monitor to obtain the state indications for the application, wherein the request is formatted according to an application programming interface (API) for the monitor.

18. The non-transitory, computer-readable storage medium of claim 14 , wherein the failover operation redirects requests for the application from the primary node to the secondary node.

19. The non-transitory, computer-readable storage medium of claim 14 , wherein the failover operation sends a request to halt performance of the application at the primary node.

20. The non-transitory, computer-readable storage medium of claim 14 , wherein the primary node and the secondary node are implemented as part of a virtual compute service offered by a provider network, wherein the monitor is implemented as part of a monitoring service for the provider network, and wherein the failover operation includes sending a request to a networking service implemented as part of the provider network to redirect requests for the application to the secondary node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2017
From: HARPREET, .
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 042574/0393 →
Cited By (4)
US 12,375,542 US 12,493,535 US 12,515,681 US 12,748,664