IP Library › Granted Patent US 12,265,455
Granted Patent B2
US 12,265,455 · App. 17/452,787 · Granted Apr 1, 2025

Task failover

Inventors: Guang Han Sui (Beijing, CN); Wei Ge (Beijing, CN); Lan Zhe Liu (Beijing, CN); Zhang Li Ping (Beijing, CN); Er Tao Zhao (Beijing, CN)
Assignee: International Business Machines Corporation
G06F11/203G06F9/461G06F9/485G06F9/4856G06F9/5038G06F11/0709G06F2209/5021G06F2209/509
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,455
App. No.
17/452,787
Granted
Apr 1, 2025
Kind
B2
Abstract

The present invention relates to a method, system and computer program product for task failover in an unstable environment, wherein the unstable environment includes a plurality of reclaimable nodes. According to the method, it is monitored if any node of the plurality of reclaimable nodes is to be reclaimed. Whether a task on any node of the plurality of reclaimable nodes is recoverable is determined. Responsive to the task being recoverable, data of the recoverable task is stored. Responsive to a node being reclaimed and the task on the reclaimed node being recoverable, at least one associated task of at least one associated node of the reclaimed node is notified to wait.

Claims (31)

1. A computer-implemented method for task failover in a cloud environment, wherein the cloud environment includes a plurality of reclaimable nodes, comprising:

monitoring, by one or more processing units, if any node of the plurality of reclaimable nodes having one or more child nodes is to be reclaimed for executing other tasks, a reclaimable node executing in a spot instance in the cloud environment and being reclaimable at any time without delay for data storing;

determining, by one or more processing units, whether a task of the existing tasks on any node of the reclaimable nodes having one or more child nodes is recoverable, the task defined as recoverable if the task has been executed on the one or more child nodes for a period higher than a threshold and an impact will occur if associated task execution in the one or more child nodes is killed, the impact related to wasted calculations associated with task execution in the one or more child nodes;

storing, by one or more processing units, data of the recoverable task on the reclaimable node remotely in a stable node;

notifying, by one or more processing units, at least one associated task executed by at least one child node associated with a reclaimed node in the cloud environment to wait without abandoning the at least one associated task which results in a waste of calculation resources and delay of response;

connecting, by one or more processing units, the child node of the reclaimed node to the stable node which will not be reclaimed to interrupt task execution; and

continuing execution of the recoverable task and the at least one associated task on the stable node.

2. The method of claim 1 , wherein the step of determining whether a task is recoverable is further based on at least one of the following factors associated with the task on the to-be-reclaimed node: task execution time, a task execution percentage, a task execution result size, a number of child nodes, and a number of times of being re-executed.

3. The method of claim 1 , wherein monitoring if any node of the plurality of reclaimable nodes is to be reclaimed, determining whether a task of the existing tasks on any node of the plurality of reclaimable nodes is recoverable, and storing data of the recoverable task remotely in the stable node occur in parallel.

4. The method of claim 1 , further comprises:

responsive to the at least one node being reclaimed and the task being unrecoverable, exiting, by one or more processing units, the task execution of the unrecoverable task on the reclaimed node and the task execution associated with the unrecoverable task on the child node of the reclaimed node.

5. A computer program product for task failover in a cloud environment, wherein the cloud environment includes a plurality of reclaimable nodes, the computer program product comprising:

a computer readable storage medium having program instructions embodied therewith, the program instructions being executable by a computer to cause the computer to perform a method comprising:

monitoring if any node of the plurality of reclaimable nodes having one or more child nodes is to be reclaimed for executing other tasks, a reclaimable node executing in a spot instance in the cloud environment and being reclaimable at any time without delay for data storing;

determining whether a task of the existing tasks on any node of the reclaimable nodes having one or more child nodes is recoverable, the task defined as recoverable if the task has been executed on the one or more child nodes for a period higher than a threshold and an impact will occur if associated task execution in one or more child nodes is killed, the impact related to wasted calculations associated with task execution in the one or more child nodes;

storing data of the recoverable task on the reclaimable node remotely in a stable node;

notifying at least one associated task executed by at least one child node associated with a reclaimed node in the cloud environment to wait without abandoning the at least one associated task which results in a waste of calculation resources and delay of response;

connecting the child node of the reclaimed node to the stable node which will not be reclaimed to interrupt task execution; and

continuing execution of the recoverable task and the at least one associated task on the stable node.

6. The computer program product of claim 1 , wherein the step of determining whether a task is recoverable is further based on at least one of the following factors associated with the task on the reclaimed node: task execution time, a task execution percentage, a task execution result size, a number of child nodes, and a number of times of being re-executed.

7. A computer system for task failover in a cloud environment, wherein the cloud environment includes a plurality of reclaimable nodes, comprising:

one or more processors;

a memory coupled to at least one of the processors; and

a set of computer program instructions stored in the memory and executed by at least one of the processors to perform a method comprising:

monitoring if any node of the plurality of reclaimable nodes having one or more child nodes is to be reclaimed for executing other tasks, a reclaimable node executing in a spot instance in the cloud environment and being reclaimable at any time without delay for data storing;

determining whether a task of the existing tasks on any node of the reclaimable nodes having one or more child nodes is recoverable, the task defined as recoverable if the task has been executed on the one or more child nodes for a period higher than a threshold and an impact will occur if associated task execution in the one or more child nodes is killed, the impact related to wasted calculations associated with task execution in the one or more child nodes;

storing data of the recoverable task on the reclaimable node remotely in a stable node;

notifying at least one associated task executed by at least one child node associated with a reclaimed node in the cloud environment to wait without abandoning the at least one associated task which results in a waste of calculation resources and delay of response;

connecting, by one or more processing units, the child node of the reclaimed node to the stable node which will not be reclaimed to interrupt task execution; and

continuing execution of the recoverable task and the at least one associated task on the stable node.

8. The computer system of claim 7 , wherein the step of determining whether a task is recoverable is further based on at least one of the following factors associated with the task on the reclaimed node: task execution time, a task execution percentage, a task execution result size, a number of child nodes, and a number of times of being re-executed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2021
From: SUI, GUANG HAN; GE, WEI; LIU, LAN ZHE; PING, ZHANG LI; ZHAO, ER TAO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057957/0513 →
Continuity (1)
Related Publication 20230132831A1 · May 4, 2023
References Cited (47)
US 6122752A · Farah · 2000 [cited by applicant]
US 7752484B2 · Goetz · 2010 [cited by applicant]
US 8266477B2 · Mankovskii · 2012 [cited by applicant]
US 8392750B2 · Kaul · 2013 [cited by applicant]
US 8806490B1 · Pulsipher · 2014 [cited by examiner]
US 9195528B1 · Sarda · 2015 [cited by applicant]
US 9417918B2 · Chin · 2016 [cited by applicant]
US 9658887B2 · Chin · 2017 [cited by applicant]
US 9674302B1 · Khalid · 2017 [cited by examiner]
US 10960304B1 · Pare · 2021 [cited by examiner]
US 20120137164A1 · Uhlig · 2012 [cited by examiner]
US 20120317579A1 · Liu · 2012 [cited by examiner]
US 20140280325A1 · Krishnamurthy · 2014 [cited by examiner]
US 20150154046A1 · Farkas · 2015 [cited by examiner]
US 20170371703A1 · Wagner · 2017 [cited by examiner]
US 20180203728A1 · Yan · 2018 [cited by examiner]
US 20200004580A1 · Sui · 2020 [cited by applicant]
US 20210294651A1 · Misca · 2021 [cited by examiner]
US 20220129322A1 · Schwartz · 2022 [cited by examiner]
US 20220374276A1 · Mitra · 2022 [cited by examiner]
CN 106844083A · 2017 [cited by applicant]
CN 112463437A · 2021 [cited by applicant]
CN 113364603A · 2021 [cited by applicant]
WO 2015195834A1 · 2015 [cited by applicant]
WO 2023072252A1 · 2023 [cited by applicant]
A Cost-Effective Framework for Running Industrial Big Data Analysis Applications in Public Clouds Liduo Lin, Li Pan, and Shijun Liu Date of publication Oct. 22, 2021 (Year: 2021). [cited by examiner]
TR-Spark: Transient Computing for Big Data Analytics Ying Yan, Yanjie Gao, Yang Chen, Zhongxin Guo, Bole Chen, Thomas Moscibroda (Year: 2016). [cited by examiner]
A survey on optimal utilization of preemptible VM instances in cloud computing Ashish Kumar Mishra, Brajesh Kumar Umrao, Dharmendra K. Yadav (Year: 2018). [cited by examiner]
Backup or Not: An Online Cost Optimal Algorithm for Data Analysis Jobs Using Spot Instances Liduo Lin, Li Pan, and Shijun Liu (Year: 2020). [cited by examiner]
Optimizing Bioinformatics Variant Analysis Pipeline for Clinical Use Hatem Mohamed Elshazly Chapters 1 and 4-6 (Year: 2016). [cited by examiner]
Flint: Batch-Interactive Data-Intensive Processing on Transient Servers Prateek Sharma Tian Guo Xin He David Irwin Prashant Shenoy (Year: 2016). [cited by examiner]
A Cost-Effective Framework for Running Industrial Big Data Analysis Applications in Public Clouds (online version) Liduo Lin, Li Pan, Shijun Liu DOI: 10.1109/JIOT.2021.3122196 ieeexplore.ieee.org/abstract/document/95850… [cited by examiner]
SpotCheck: Designing a Derivative IaaS Cloud on the Spot Market Prateek Sharma Stephen Lee Tian Guo David Irwin Prashant Shenoy (Year: 2015). [cited by examiner]
Derivative IaaS Layer Utilizing Low Priority Server Instances for Web Applications With State Persistence (Year: 2018). [cited by examiner]
FaultTolerance in the Borealis Distributed Stream Processing System Magdalena Balazinska, Hari Balakrishnan, Samuel Madden, and Michael Stonebraker (Year: 2005). [cited by examiner]
Beter Safe than Sorry: Grappling with Failures of In-Memory Data Analytics Frameworks Bogdan Ghit and Dick Epema (Year: 2017). [cited by examiner]
Notification of Transmittal of The International Search Report and The Written Opinion of the International Searching Authority, or the declaration, International Application No. PCT/CN2022/128265, International filing … [cited by applicant]
Amazon, “Amazon EC2 Auto Scaling Capacity Rebalancing,” Amazon.com, [accessed Oct. 4, 2021], 8 pgs., Retrieved from the Internet: <https://docs.aws.amazon.com/autoscaling/ec2/userguide/capacity-rebalance.html>. [cited by applicant]
Amazon, “Cost Optimization Pillar,” Amazon Web Service, Jul. 2018, 35 pgs., Retrieved from the Internet: <https://d1.awsstatic.com/whitepapers/architecture/AWS-Cost-Optimization-Pillar.pdf>. [cited by applicant]
Antonopoulos, et al., “Cloud Computing, Principles, Systems and Applications,” Springer, Computer Communications and Networks, 2010, 386 pgs. [cited by applicant]
Disclosed Anonymously, “Efficient Task Execution and Progress Reporting For Complex Jobs,” IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000221077D, Aug. 28, 2012, 15 pgs. [cited by applicant]
Disclosed Anonymously, “Resource Reclaim Optimization for Parent Child Workload,” IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000253984D, May 22, 2018, 3 pgs. [cited by applicant]
IBM, “Backfill Parent Task Resources With Child Tasks,” IBM.com, [accessed Oct. 5, 2021], 4 pgs., Retrieved from the Internet: <https://www.ibm.com/docs/en/spectrum-symphony/7.2.1?topic=applications-backfill-parent-task… [cited by applicant]
IBM, “Parent-child Workload: Advantage IBM Spectrum Symphony,” IBM.com, [accessed Oct. 5, 2021], 7 pgs., Retrieved from the Internet: <https://www.ibm.com/support/pages/parent-child-workload-advantage-ibm-spectrum-symph… [cited by applicant]
Kathalkar, et al., “Study of Checkpoint Restore mechanism for Fault Tolerance in Cloud computing,” International Journal of Advance Research in Science and Engineering, vol. No. 07, Issue No. 04, Apr. 2018, pp. 237-243. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, pp. 1-7. [cited by applicant]
Poojary, “What Are Your Spot Instance Options on AWS, Azure, and Google?” Six Nines, [accessed Oct. 4, 2021], 7 pgs., Retrieved from the Internet: <https://sixninesit.com/what-are-your-spot-instances-options-on-aws-azur… [cited by applicant]