IP Library › Granted Patent US 12,585,673
Granted Patent B1
US 12,585,673 · App. 18/910,421 · Granted Mar 24, 2026

Synchronous block level replication across availability zones

Inventors: Andrey Arkharov (Kirkland, WA); Andrei Burago (Kirkland, WA); Jonathan Forbes (Bellevue, WA); Anton Sukhanov (Bellevue, WA); Fabricio Voznika (Kenmore, WA)
Assignee: Google LLC
G06F16/275G06F16/178
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,673
App. No.
18/910,421
Granted
Mar 24, 2026
Kind
B1
Abstract

A replicated block storage service provides durable and high performance network-attached storage replicated in two or more zones of a single region, and remains available despite a single zone failure. A probe file is generated to determine a health state of a replicated disk. When a disk is degraded, a lease is created indicating which replica is trusted and providing visibility to backend jobs to facilitate reconciliation of data between the first replica and the second replica. Moreover, degraded file markers are generated for use by the backend jobs in quickly identifying the data to be copied.

Claims (49)

1 . A method for maintaining a replicated disk in a distributed storage system, comprising:

maintaining, in one or more memories in a first zone of the distributed storage system, a first replica including a first copy of disk data including a plurality of files;

maintaining, in the one or more memories in a second zone of the distributed storage system, a second replica including a second copy of the disk data including the plurality of files;

attempting, by a first virtual machine attached to the replicated disk, a write command to a specified file in both the first replica and the second replica;

determining, based on a result of the attempted write command, a health state of the replicated disk,

wherein the health state of the disk is variable between a fully replicated state, a degraded state, and a partially replicated state;

creating a replication lease for the degraded state, wherein the replication lease is allowed to expire once the disk enters the partially replicated state, wherein expiration of the replication lease indicates to one or more backend processors to begin copying the data from the second replica to the first replica, and wherein the one or more backend processors delete the expired replication lease when copying is complete.

2 . The method of claim 1 , further comprising copying, by one or more backend processors, data from the second replica to the first replica when the disk is in the degraded state or the partially replicated state.

3 . The method of claim 2 , further comprising scanning, by the one or more backend processors, the data in the second replica for files marked as degraded, wherein the copying from the second replica to the first replica comprises copying data from the files marked as degraded identified during scanning.

4 . The method of claim 1 , wherein the attempted write command is directed to a probe file generated for testing the health state of the disk.

5 . The method of claim 1 , further comprising:

attaching the replicated disk to a second virtual machine while the replicated disk is accessible to the first virtual machine; and

preventing the first virtual machine from creating new read-write files.

6 . The method of claim 1 , wherein the attempted write command comprises a write command generated by a client device.

7 . The method of claim 1 , wherein when the determined health state of the replicated disk indicates that the first replica is unhealthy, closing only the specified file in both replicas while continuing to write to other files of the plurality of files on each of the first and second replicas, creating a new file corresponding to the specified file in the second replica, and creating a degraded file corresponding to the specified file in the second replica, wherein the degraded file is marked as degraded.

8 . The method of claim 1 , wherein the replication lease further indicates which replica is degraded and which replica is trusted.

9 . The method of claim 1 , wherein:

in the fully replicated state both replicas are healthy;

in the degraded state the first replica is unhealthy and the second replica is trusted; and

in the partially replicated state, the first replica has been restored to health from the degraded state, but is missing data as compared to the second replica.

10 . A system for maintaining a replicated disk in a distributed storage system, comprising:

one or more memories in a first zone of the distributed storage system, the one or more memories in the first zone storing a first replica including a first copy of disk data including a plurality of files;

one or more memories in a second zone of the distributed storage system, the one or more memories in the second zone storing a second replica including a second copy of the disk data including the plurality of files;

one or more processors in communication with at least one of the first replica or the second replica, the one or more processors configured to:

attempt a write command to a specified file in both the first replica and the second replica; and

determine, based on a result of the attempted write command, a health state of the replicated disk,

wherein the health state of the disk is variable between a fully replicated state, a degraded state, and a partially replicated state; and

create a replication lease for the degraded state, wherein the replication lease is allowed to expire once the disk enters the partially replicated state.

11 . The system of claim 10 , further comprising one or more backend processors configured to copy data from the second replica to the first replica when the disk is in the degraded state or the partially replicated state.

12 . The system of claim 11 , wherein the one or more backend processors are further configured to scan the data in the second replica for files marked as degraded, wherein the copying from the second replica to the first replica comprises copying data from the files marked as degraded identified during scanning.

13 . The system of claim 10 , wherein the attempted write command is directed to a probe file generated for testing the health state of the disk.

14 . The system of claim 13 , wherein the attempted write command is to a log file.

15 . The system of claim 10 , wherein when the determined health state of the replicated disk indicates that the first replica is unhealthy:

close only the specified file in both replicas while continuing to write to other files of the plurality of files on each of the first and second replicas,

create a new file corresponding to the specified file in the second replica, and

create a degraded file corresponding to the specified file in the second replica, wherein the degraded file is marked as degraded.

16 . The system of claim 10 , wherein expiration of the replication lease indicates to one or more backend processors to begin copying the data from the second replica to the first replica, and wherein the one or more backend processors delete the expired replication lease when copying is complete.

17 . The system of claim 10 , wherein the replication lease further indicates which replica is degraded and which replica is trusted.

18 . The system of claim 10 , wherein:

in the fully replicated state both replicas are healthy;

in the degraded state the first replica is unhealthy and the second replica is trusted; and

in the partially replicated state, the first replica has been restored to health from the degraded state, but is missing data as compared to the second replica.

19 . A non-transitory computer-readable medium storing instructions executable by one or more processor for performing a method of maintaining a replicated disk in a distributed storage system, the method comprising:

maintaining a first replica including a first copy of disk data including a plurality of files;

maintaining a second replica including a second copy of the disk data including the plurality of files;

attempting a write command to a specified file in both the first replica and the second replica;

determining, based on a result of the attempted write command, a health state of the replicated disk,

wherein the health state of the disk is variable between a fully replicated state, a degraded state, and a partially replicated state; and

creating a replication lease for the degraded state, wherein the replication lease is allowed to expire once the disk enters the partially replicated state.

Continuity (3)
Continuation 18521240 · Nov 28, 2023
Continuation 17551914 · Dec 15, 2021
Continuation 15893262 · Feb 9, 2018
References Cited (18)
US 6662198B2 · Satyanarayanan et al. · 2003 [cited by applicant]
US 8041818B2 · Gupta et al. · 2011 [cited by applicant]
US 8074107B2 · Sivasubramanian et al. · 2011 [cited by applicant]
US 8838539B1 · Ashcraft et al. · 2014 [cited by applicant]
US 10360057B1 · Vashishtha et al. · 2019 [cited by applicant]
US 20060253504A1 · Lee et al. · 2006 [cited by applicant]
US 20120254116A1 · Thereska et al. · 2012 [cited by applicant]
US 20140222878A1 · Avati · 2014 [cited by examiner]
US 20150113324A1 · Factor et al. · 2015 [cited by applicant]
US 20170011061A1 · Hansen · 2017 [cited by examiner]
US 20170161160A1 · Helmick et al. · 2017 [cited by applicant]
US 20170249334A1 · Karampuri et al. · 2017 [cited by applicant]
Zhuan Chen et al., Replication-based Highly Available Metadata Management for Cluster File Systems, ICT, Cluster 2010, Heraklion, Creece, Sep. 23, 2010, 34 pages. [cited by applicant]
Mike Burrows, The Chubby lock service for loosely-coupled distributed systems, OSDI '06 Paper, Sep. 6, 2006, 22 pages. [cited by applicant]
Yang Wang et al., Gnothi: Separating Data and Metadata for Efficient andAvailable Storage Replication, The University of Texas at Austin, Usenix ATC 2012, 12 pages. [cited by applicant]
Proceedings of the 13th USENIX Conference on File and Storage Technologies (FAST '15), Feb. 16-19, 2015, Santa Clara, CA, 397 pages. [cited by applicant]
Sanjay Ghemawat et al., The Google File System, SOSP'03, Oct. 19-22, 2003, Bolton Landing, New York, pp. 29-43. [cited by applicant]
DbVisit, “Physical Or Logical Replication? | The Smart Alternative—Dbvisit” http://www.dbvisit.com/physical-vs-logical-replication/ (2015) 2 pgs. [cited by applicant]