IP Library Granted Patent US 11,036,677
Granted Patent B1
US 11,036,677 · App. 16/174,498 · Granted Jun 15, 2021

Replicated data integrity

Inventors: David Grunwald (San Francisco, CA); Luke Paulsen (Mountain View, CA); Ronald Karr (Palo Alto, CA); Thomas Gill (Mountain View, CA); Yao-Cheng Tien (Milpitas, CA)
Assignee: Pure Storage, Inc.
G06F16/128G06F3/067G06F3/0619G06F3/0646G06F21/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,036,677
App. No.
16/174,498
Granted
Jun 15, 2021
Kind
B1
Abstract

Performing replicated data integrity, including: generating, at a first computer system, a local hash of a local dataset; replicating the local dataset; receiving, at the first computer system from a second computer system, a remote hash of a remote dataset generated from the local dataset replicated from the first computer system; and determining, based at least on a comparison of the local hash of the local dataset with the remote hash of the remote dataset, validity of the remote dataset generated from the local dataset replicated from the first computer system.

Claims (45)

1. A method comprising:

replicating the local dataset from a first computer system to a second computer system;

receiving, at the first computer system from the second computer system, a remote hash of a remote dataset generated from the local dataset replicated from the first computer system; and

determining, by the first computer system, based at least on a comparison of a local hash of the local dataset with the remote hash of the remote dataset, validity of the remote dataset generated from the local dataset replicated from the first computer system,

wherein the local dataset is a subset of a volume of data selected based at least upon a hashing policy, and wherein the hashing policy specifies random portions of the volume of data for hashing.

2. The method of claim 1 , wherein the first computer system is a first storage system, wherein the second computer system is a second storage system, and wherein the dataset is synchronously replicated between the first storage system and the second storage system.

3. The method of claim 2 , wherein determining the validity of the remote dataset generated from the local dataset replicated from the first computer system to the second computer system is performed as part of synchronizing the dataset among the first storage system and the second storage system.

4. The method of claim 2 , wherein the dataset is asynchronously replicated between the first computer system and the second computer system.

5. The method of claim 1 , wherein the local dataset is a first snapshot of a volume of data at a first point in time, and wherein the method further comprises:

generating a second snapshot of the volume of data at a second point in time, wherein the second snapshot comprises differences between the volume of data at the first point in time and the volume of data at the second point in time;

generating an incremental local hash of the second snapshot of the volume of data at the second point in time;

replicating the second snapshot;

receiving, at the first computer system, an incremental remote hash of a remote second snapshot generated from the local second snapshot replicated from the first computer system; and

determining, based at least on a comparison of the incremental local hash with the incremental remote hash of the remote second snapshot generated from the local second snapshot replicated from the first computer system, validity of the remote second snapshot.

6. The method of claim 1 , wherein, periodically or aperiodically after determining validity of the remote dataset, the first computer system continues to determine validity of the remote dataset by generating a current local hash of the dataset and comparing the current local hash to a requested current remote hash of the dataset from the second computer system.

7. An apparatus comprising a computer processor, a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions that, when executed by the computer processor, cause the apparatus to carry out:

replicating the local dataset from a first computer system to a second computer system;

receiving, at the first computer system from the second computer system, a remote hash of a remote dataset generated from the local dataset replicated from the first computer system; and

determining, by the first computer system, based at least on a comparison of a local hash of the local dataset with the remote hash of the remote dataset, validity of the remote dataset generated from the local dataset replicated from the first computer system,

wherein the local dataset is a subset of a volume of data selected based at least upon a hashing policy, and wherein the hashing policy specifies random portions of the volume of data for hashing.

8. The apparatus of claim 7 , wherein the first computer system is a first storage system, wherein the second computer system is a second storage system, and wherein the dataset is synchronously replicated between the first storage system and the second storage system.

9. The apparatus of claim 8 , wherein determining the validity of the remote dataset generated from the local dataset replicated from the first computer system to the second computer system is performed as part of synchronizing the dataset among the first storage system and the second storage system.

10. The apparatus of claim 8 , wherein the dataset is asynchronously replicated between the first computer system and the second computer system.

11. The apparatus of claim 10 , wherein the local dataset is a first snapshot of a volume of data at a first point in time, and wherein the computer program instructions further cause the apparatus to carry out the steps of:

generating a second snapshot of the volume of data at a second point in time, wherein the second snapshot comprises differences between the volume of data at the first point in time and the volume of data at the second point in time;

generating an incremental local hash of the second snapshot of the volume of data at the second point in time;

replicating, from the first computer system to a second computer system, the second snapshot;

receiving, at the first computer system from the second computer system, an incremental remote hash of a remote second snapshot generated from the local second snapshot replicated from the first computer system to the second computer system; and

determining, based at least on a comparison of the incremental local hash with the incremental remote hash of the remote second snapshot generated from the local second snapshot replicated from the first computer system to the second computer system, validity of the remote second snapshot.

12. The apparatus of claim 7 , wherein, periodically or aperiodically after determining validity of the remote dataset, the first computer system continues to determine validity of the remote dataset by generating a current local hash of the dataset and comparing the current local hash to a requested current remote hash of the dataset from the second computer system.

13. A computer program product disposed upon a computer readable medium, the computer program product comprising computer program instructions that, when executed, cause a computer to carry out:

replicating the local dataset from a first computer system to a second computer system;

receiving, at the first computer system from the second computer system, a remote hash of a remote dataset generated from the local dataset replicated from the first computer system; and

determining, by the first computer system, based at least on a comparison of a local hash of the local dataset with the remote hash of the remote dataset, validity of the remote dataset generated from the local dataset replicated from the first computer system,

wherein the local dataset is a subset of a volume of data selected based at least upon a hashing policy, and wherein the hashing policy specifies random portions of the volume of data for hashing.

14. The computer program product of claim 13 , wherein the first computer system is a first storage system, wherein the second computer system is a second storage system, and wherein the dataset is synchronously replicated between the first storage system and the second storage system.

15. The computer program product of claim 14 , wherein determining the validity of the remote dataset generated from the local dataset replicated from the first computer system to the second computer system is performed as part of synchronizing the dataset among the first storage system and the second storage system.

16. The computer program product of claim 14 , wherein the dataset is asynchronously replicated between the first computer system and the second computer system.

17. The computer program product of claim 13 , wherein the local dataset is a first snapshot of a volume of data at a first point in time, and wherein the computer program instructions, when executed, further cause the computer to carry out the steps of:

generating a second snapshot of the volume of data at a second point in time, wherein the second snapshot comprises differences between the volume of data at the first point in time and the volume of data at the second point in time;

generating an incremental local hash of the second snapshot of the volume of data at the second point in time;

replicating, from the first computer system to a second computer system, the second snapshot;

receiving, at the first computer system from the second computer system, an incremental remote hash of a remote second snapshot generated from the local second snapshot replicated from the first computer system to the second computer system; and

determining, based at least on a comparison of the incremental local hash with the incremental remote hash of the remote second snapshot generated from the local second snapshot replicated from the first computer system to the second computer system, validity of the remote second snapshot.

18. The computer program product of claim 13 , wherein, periodically or aperiodically after determining validity of the remote dataset, the first computer system continues to determine validity of the remote dataset by generating a current local hash of the dataset and comparing the current local hash to a requested current remote hash of the dataset from the second computer system.

Assignments (3)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 11, 2025
From: BARCLAYS BANK PLC, AS ADMINISTRATIVE AGENT
To: PURE STORAGE, INC.
Reel/Frame 071558/0523 →
SECURITY INTEREST Recorded Aug 26, 2020
From: PURE STORAGE, INC.
To: BARCLAYS BANK PLC AS ADMINISTRATIVE AGENT
Reel/Frame 053867/0581 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2018
From: GRUNWALD, DAVID; PAULSEN, LUKE; KARR, RONALD; GILL, THOMAS; TIEN, YAO-CHENG
To: PURE STORAGE, INC.
Reel/Frame 047668/0302 →
Cited By (29)
US 12,197,762 US 12,299,087 US 12,306,801 US 12,306,802 US 12,306,804 US 12,309,271 US 12,341,887 US 12,368,588 US 12,386,950 US 12,393,485 US 12,406,056 US 12,417,038 US 12,443,587 US 12,443,684 US 12,445,283 US 12,455,861 US 12,481,703 US 12,487,972 US 12,493,417 US 12,493,431 US 12,530,262 US 12,569,722 US 12,572,513 US 12,579,109 US 12,608,401 US 12,664,237 US 12,681,953 US 12,693,993 US 12,717,755