IP Library Granted Patent US 11,561,860
Granted Patent B2
US 11,561,860 · App. 16/122,447 · Granted Jan 24, 2023

Methods and systems for power failure resistance for a distributed storage system

Inventors: Maor Ben Dayan (Tel Aviv, IL); Omri Palmon (Tel Aviv, IL); Liran Zvibel (Tel Aviv, IL); Kanael Arditti (Tel Aviv, IL)
G06F11/142G06F3/067G06F3/0619G06F3/0632G06F3/0652
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,561,860
App. No.
16/122,447
Granted
Jan 24, 2023
Kind
B2
Abstract

A plurality of computing devices are communicatively coupled to each other via a network, and each of the plurality of computing devices is operably coupled to one or more of a plurality of storage devices. One or more of the computing devices and/or the storage devices may be used to rebuild data that may be lost due to a power failure.

Claims (36)

1. A method for data recovery in a storage system with a file system that is integrated with a protection layer, comprising:

detecting power up after a non-scheduled power down;

determining active regions of associated storage devices, wherein the active regions of the associated storage devices are part of an address space owned by a same bucket of a plurality of buckets, wherein each bucket of the plurality of buckets is operable to write to the associated storage devices without any need to coordinate with other buckets of the plurality of buckets;

searching the active regions of the associated storage devices for a journal that corresponds to at least a portion of the associated storage devices; and

rebuilding data in the at least a portion of the associated storage devices based on information in the journal,

wherein the active regions comprise new stripes that are capable of receiving and storing new data, where the new data are data that are not committed to the associated storage devices, and wherein each bucket of the plurality of buckets is operable to store into a group of storage devices, and wherein no two groups, of storage devices, are identical.

2. The method of claim 1 , comprising, after rebuilding the data, scrubbing memory to identify an abnormal data block.

3. The method of claim 2 , wherein scrubbing memory comprises freeing the abnormal data block.

4. The method of claim 2 , wherein scrubbing memory comprises fixing an error in the abnormal data block.

5. The method of claim 2 , wherein scrubbing memory occurs in the background.

6. The method of claim 2 , wherein scrubbing memory occurs continuously.

7. The method of claim 2 , wherein scrubbing memory occurs on demand.

8. The method of claim 1 , wherein information regarding the active regions are stored in memory.

9. A system for data recovery in a storage system with a file system that is integrated with a protection layer, comprising:

a processor configured to:

detect power up after a non-scheduled power down;

determine active regions of associated storage devices, wherein the active regions of the associated storage devices are part of an address space owned by a same bucket of a plurality of buckets, wherein each bucket of the plurality of buckets is operable to write to the associated storage devices without any need to coordinate with other buckets of the plurality of buckets;

search the active regions of the associated storage devices for a journal that corresponds to at least a portion of the associated storage devices; and

rebuild data in the at least a portion of the associated storage devices based on information in the journal,

wherein the processor is configured to write new data to the active regions, where the new data are data that are not committed to the associated storage devices, and wherein each bucket of the plurality of buckets is operable to store into a group of storage devices, and wherein no two groups, of storage devices, are identical.

10. The system of claim 9 , wherein the processor is configured to, after rebuilding the data, scrub memory to identify an abnormal data block.

11. The system of claim 10 , wherein scrubbing memory comprises freeing the abnormal data block.

12. The system of claim 10 , wherein scrubbing memory comprises fixing an error in the abnormal data block.

13. The system of claim 10 , wherein scrubbing memory occurs in the background.

14. The system of claim 10 , wherein scrubbing memory occurs continuously.

15. The system of claim 10 , wherein scrubbing memory occurs on demand.

16. The system of claim 9 , wherein the processor is configured to read information regarding the active regions in memory.

17. A non-transitory machine-readable storage having stored thereon, a computer program having at least one code section for data recovery in a storage system with a file system that is integrated with a protection layer, the at least one code section comprising machine executable instructions for causing the machine to perform steps comprising:

detecting power up after a non-scheduled power down;

determining active regions of associated storage devices, wherein the active regions of the associated storage devices are part of an address space owned by a same bucket of a plurality of buckets, wherein each bucket of the plurality of buckets is operable to write to the associated storage devices without any need to coordinate with other buckets of the plurality of buckets;

searching the active regions of the associated storage devices for a journal that corresponds to at least a portion of the associated storage devices; and

rebuilding data in the at least a portion of the associated storage devices based on information in the journal,

wherein the active regions comprise new stripes that are capable of receiving and storing new data, where the new data are data that are not committed to the associated storage devices, and wherein each bucket of the plurality of buckets is operable to store into a group of storage devices, and wherein no two groups, of storage devices, are identical.

18. The non-transitory machine-readable storage of claim 17 , comprising machine executable instructions for, after rebuilding the data, scrubbing memory to identify an abnormal data block to perform one of: freeing the abnormal data block and fixing an error in the abnormal data block.

19. The non-transitory machine-readable storage of claim 18 , wherein scrubbing memory comprises one or more of: executing in the background, executing continuously, and executing on demand.

20. The non-transitory machine-readable storage of 17 , wherein information regarding the active regions are stored in memory.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Jun 20, 2024
From: BANK LEUMI LE-ISRAEL B.M.
To: WEKAIO LTD.
Reel/Frame 067783/0962 →
SECURITY INTEREST Recorded Mar 29, 2020
From: WEKAIO LTD.
To: BANK LEUMI LE-ISRAEL B.M.
Reel/Frame 052253/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2018
From: DAYAN, MAOR BEN; PALMON, OMRI; ZVIBEL, LIRAN; ARDITTI, KANAEL
To: WEKA.IO LTD.
Reel/Frame 047101/0168 →
Continuity (2)
Provisional Application 62585186 · Nov 13, 2017
Related Publication 20190146879A1 · May 16, 2019
Cited By (1)
US 12,346,203