IP Library Granted Patent US 11,089,097
Granted Patent B2
US 11,089,097 · App. 15/920,501 · Granted Aug 10, 2021

Method and system for recovering data in distributed computing system

Inventors: Kaiho Fukuchi (Tokyo, JP); Jun Nemoto (Tokyo, JP); Masakuni Agetsuma (Tokyo, JP)
Assignee: HITACHI, LTD.
H04L67/1095G06F11/1469G06F11/2025G06F11/2028H04L43/0811H04L67/42G06F11/203G06F11/2094
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,089,097
App. No.
15/920,501
Granted
Aug 10, 2021
Kind
B2
Abstract

The time required for recovery in a distributed computing system can be reduced. At least one node (for example a server) or a different computer (for example a management server) are provided in the distributed computing system which includes a plurality of nodes existing at a plurality of sites. One or more sites at which one or more nodes that hold one or more datasets identical to one or more datasets held by a node to be recovered are identified. For the recovery, it is determined, on the basis of the one or more identified sites, a restore destination site that is a site of a node to which the one or more identical datasets are to be restored from among the plurality of sites.

Claims (35)

1. A non-transitory computer readable medium in a distributed computing system including a plurality of computers existing at a plurality of sites, wherein in the distributed computing system, datasets are distributed and held in the plurality of computers existing at the plurality of sites, the non-transitory computer readable medium comprising a program having instructions that when executed by at least one processor, when a failure occurs in one site of the plurality of sites, perform the steps of:

identifying a dataset of one computer at the one site that needs to be restored;

identifying other sites at which other computers hold datasets corresponding to the dataset at the one site that needs to be restored;

computing a transfer time based on a communication speed between the sites and a data size of the dataset to be restored so as to determine a restore destination site; and

selecting the restore destination site from the other sites based on the computed transfer time,

wherein for each of the datasets, the sites that store a corresponding dataset have computers that must belong to a predetermined data sharing range with a computer that stores the respective dataset,

wherein, for the identified data set, the identified other sites are sites that must be within the data sharing range of the one computer,

wherein the selected restore destination site is selected from the other sites that have the computers that are within the data sharing range of the one computer based on the computed transfer time,

wherein the selected restore destination site is a site that has a shortest total data transfer time among the other sites including the selected restore destination site, and

wherein for each of the other sites including the selected restore destination site, the total data transfer time is a sum of one or more transfer times that correspond to the dataset.

2. The non-transitory computer readable medium according to claim 1 ,

wherein the selected restore destination site is a site that has a fastest communication speed among the other sites including the selected restore destination site.

3. At least one computer in a distributed computing system including a plurality of computers existing at a plurality of sites when a failure occurs in one site of the plurality of sites, wherein in the distributed computing system, datasets are distributed and held in the plurality of computers existing at the plurality of sites, comprising:

an interface unit including one or more interface devices for communication with one or more of the plurality of computers; and

a processor unit including one or more processors coupled to the interface unit

wherein the processor unit is configured to:

identify a dataset at the one site that needs to be restored;

identify other sites at which other computers hold datasets corresponding to the dataset at the one site that needs to be restored;

compute a transfer time based on a communication speed between the sites and a data size of the dataset to be restored so as to determine a restore destination site; and

select the restore destination site from the other sites based on the computed transfer time,

wherein for each of the datasets, the sites that store a corresponding dataset have computers that must belong to a predetermined data sharing range with a computer that stores the respective dataset,

wherein, for the identified data set, the identified other sites are sites that must be within the data sharing range of the one computer,

wherein the selected restore destination site is selected from the other sites that have the computers that are within the data sharing range of the one computer based on the computed transfer time,

wherein the selected restore destination site is a site that has a shortest total data transfer time among the other sites including the selected restore destination site, and

wherein for each of the other sites including the selected restore destination site, the total data transfer time is a sum of one or more transfer times that correspond to the dataset.

4. A method for recovering a computer to be recovered in a distributed computing system including a plurality of computers existing at a plurality of sites when a failure occurs in one site of the plurality of sites, wherein in the distributed computing system, datasets are distributed and held in the plurality of computers existing at the plurality of sites, the method comprising:

identifying a dataset at the one site that needs to be restored;

identifying other sites at which other computers hold datasets corresponding to the dataset at the one site that needs to be restored;

computing a transfer time based on a communication speed between the sites and a data size of the dataset to be restored so as to determine a restore destination site; and

selecting the restore destination site from the other sites based on the computed transfer time,

wherein for each of the datasets, the sites that store a corresponding dataset have computers that must belong to a predetermined data sharing range with a computer that stores the respective dataset,

wherein, for the identified data set, the identified other sites are sites that must be within the data sharing range of the one computer,

wherein the selected restore destination site is selected from the other sites that have the computers that are within the data sharing range of the one computer based on the computed transfer time,

wherein the selected restore destination site is a site that has a shortest total data transfer time among the other sites including the selected restore destination site, and

wherein for each of the other sites including the selected restore destination site, the total data transfer time is a sum of one or more transfer times that correspond to the dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2018
From: FUKUCHI, KAIHO; NEMOTO, JUN; AGETSUMA, MASAKUNI
To: HITACHI, LTD.
Reel/Frame 045199/0098 →
Priority Claims (1)
JP JP2017-136209 · Jul 12, 2017 · national
Continuity (1)
Related Publication 20190020716A1 · Jan 17, 2019