IP Library › Granted Patent US 12,135,685
Granted Patent B2
US 12,135,685 · App. 17/322,113 · Granted Nov 5, 2024

Verifying data has been correctly replicated to a replication target

Inventors: David Grunwald (San Francisco, CA); Luke Paulsen (Mountain View, CA); Ronald Karr (Palo Alto, CA); Thomas Gill (Bury St Edmunds, GB); Yao-Cheng Tien (Cupertino, CA)
Assignee: PURE STORAGE, INC.
G06F16/128G06F3/0619G06F3/0646G06F3/067G06F16/27G06F21/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,135,685
App. No.
17/322,113
Granted
Nov 5, 2024
Kind
B2
Abstract

Verifying that data has been correctly replicated to a replication target, including: replicating a dataset stored at a first computing system to a second computing system; and determining, based at least on a comparison of a first hash and a second hash, validity of the dataset stored at the second computing system, wherein the first hash is generated by applying a hash function to a copy of the dataset that is stored at the first computing system and the second hash is generated by applying the hash function to a copy of the dataset that is stored at the second computing system.

Claims (32)

1. A method comprising:

replicating a dataset stored at a first computing system to a second computing system, wherein the dataset is asynchronously replicated to the second computing system at a first time;

determining, by the first computing system comparing a first hash and a second hash, that the asynchronously replicated dataset stored at the second computing system is valid, wherein the first hash is generated by applying a hash function to the dataset stored at the first computing system and the second hash is generated by applying the hash function to the asynchronously replicated dataset stored at the second computing system and generated from the dataset replicated from the first computing system, and wherein the second hash is received by the first computing system from the second computing system;

replicating the dataset at the first computing system to the second computing system at a second time, the dataset comprising an updated portion of data that is different than the dataset at the first time and a same portion; and

determining whether the dataset is valid based on comparing a third hash generated by applying the hash function on only the updated portion of the dataset and a fourth hash generated by applying the hash function on only a replicated updated portion of the replicated dataset.

2. The method of claim 1 , wherein the first hash is generated by the first computing system and the second hash is generated by the second computing system after receiving the dataset.

3. The method of claim 1 , wherein determining validity of the asynchronously replicated dataset stored at the second computing system is carried out by a hash validation engine.

4. The method of claim 3 , wherein the hash validation engine is executed on one or more of the computing systems.

5. The method of claim 3 , wherein the hash validation engine is executed externally to the computing systems.

6. The method of claim 3 , wherein periodically or aperiodically after determining validity of the dataset stored at the second computing system, the hash validation engine continues to determine validity of the dataset stored at the second computing system by comparing an updated first hash and an updated second hash, wherein the updated first hash is generated by applying a hash function to a copy of the dataset that is stored at the first computing system and the updated second hash is generated by applying the hash function to a copy of the dataset that is stored at the second computing system.

7. The method of claim 1 , wherein the asynchronously replicated dataset that is stored at the second computing system is a replicated snapshot of the dataset that is stored at the first computing system.

8. An apparatus comprising:

a memory; and

a processing device, operatively coupled to the memory, configured to:

replicate a dataset stored at a first computing system to a second computing system, wherein the dataset is asynchronously replicated to the second computing system at a first time;

determine, by the first computing system comparing a first hash and a second hash, that the asynchronously replicated dataset stored at the second computing system is valid, wherein the first hash is generated by applying a hash function to the dataset stored at the first computing system and the second hash is generated by applying the hash function to the asynchronously replicated dataset stored at the second computing system and generated from the dataset replicated from the first computing system, and wherein the second hash is received by the first computing system from the second computing system;

replicate the dataset at the first computing system to the second computing system at a second time, the dataset comprising an updated portion of data that is different than the dataset at the first time and a same portion; and

determine whether the dataset is valid based on comparing a third hash generated by applying the hash function on only the updated portion of the dataset and a fourth hash generated by applying the hash function on only a replicated updated portion of the replicated dataset.

9. The apparatus of claim 8 , wherein the first hash is generated by the first computing system and the second hash is generated by the second computing system after receiving the dataset.

10. The apparatus of claim 8 , wherein determining validity of the asynchronously replicated dataset stored at the second computing system is carried out by a hash validation engine.

11. The apparatus of claim 8 , wherein periodically or aperiodically after determining validity of the dataset stored at the second computing system, the processing device continues to determine validity of the dataset stored at the second computing system by comparing an updated first hash and an updated second hash, wherein the updated first hash is generated by applying a hash function to a copy of the dataset that is stored at the first computing system and the updated second hash is generated by applying the hash function to a copy of the dataset that is stored at the second computing system.

12. The apparatus of claim 8 , wherein the asynchronously replicated dataset that is stored at the second computing system is a replicated snapshot of the dataset that is stored at the first computing system.

13. A non-transitory computer readable storage medium storing instructions which, when executed, cause a processing device to:

replicate a dataset stored at a first computing system to a second computing system, wherein the dataset is asynchronously replicated to the second computing system at a first time;

determine, by the first computing system comparing a first hash and a second hash, that the asynchronously replicated dataset stored at the second computing system is valid, wherein the first hash is generated by applying a hash function to the dataset stored at the first computing system and the second hash is generated by applying the hash function to the asynchronously replicated dataset stored at the second computing system and generated from the dataset replicated from the first computing system, and wherein the second hash is received by the first computing system from the second computing system;

replicate the dataset at the first computing system to the second computing system at a second time, the dataset comprising an updated portion of data that is different than the dataset at the first time and a same portion; and

determine whether the dataset is valid based on comparing a third hash generated by applying the hash function on only the updated portion of the dataset and a fourth hash generated by applying the hash function on only a replicated updated portion of the replicated dataset.

14. The non-transitory computer readable storage medium of claim 13 , wherein the first hash is generated by the first computing system and the second hash is generated by the second computing system after receiving the asynchronously replicated dataset.

15. The non-transitory computer readable storage medium of claim 14 , wherein determining validity of the dataset stored at the second computing system is carried out by a hash validation engine.

16. The non-transitory computer readable storage medium of claim 15 , wherein the hash validation engine is executed on one or more of the computing systems.

17. The non-transitory computer readable storage medium of claim 15 , wherein the hash validation engine is executed in a cloud environment.

18. The non-transitory computer readable storage medium of claim 13 , wherein periodically or aperiodically after determining validity of the dataset stored at the second computing system, the processing device continues to determine validity of the dataset stored at the second computing system by comparing an updated first hash and an updated second hash, wherein the updated first hash is generated by applying a hash function to a copy of the dataset that is stored at the first computing system and the updated second hash is generated by applying the hash function to a copy of the dataset that is stored at the second computing system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: GRUNWALD, DAVID; PAULSEN, LUKE; KARR, RONALD; GILL, THOMAS; TIEN, YAO-CHENG
To: PURE STORAGE, INC.
Reel/Frame 056261/0135 →
Continuity (4)
Continuation 16174498 · Oct 30, 2018
Provisional Application 62722247 · Aug 24, 2018
Provisional Application 62598989 · Dec 14, 2017
Related Publication 20210279204A1 · Sep 9, 2021