IP Library Granted Patent US 9,811,529
Granted Patent B1
US 9,811,529 · App. 13/760,933 · Granted Nov 7, 2017

Automatically redistributing data of multiple file systems in a distributed storage system

Inventors: Silvius V. Rus (Orina, CA); Thileepan Subramaniam (Mountain View, CA)
Assignee: Quantcast Corporation
G06F17/30194G06F3/067G06F17/30584
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,811,529
App. No.
13/760,933
Granted
Nov 7, 2017
Kind
B1
Abstract

A distributed storage system maintains multiple logically independent file systems. Each file system includes a data set stored by a distributed storage of the distributed storage system. During operation, access pattern levels for the multiple logically independent file systems are determined. Thereafter, the data sets included in the multiple logically independent file systems are redistributed across multiple storage devices of the distributed storage. In one aspect, redistribution of a particular data set is based at least in part on the particular file system including the particular data set and on the determined access pattern levels for the multiple logically independent file systems. In one implementation, redistribution is performed according to a uniform redistribution scheme. In another implementation, redistribution is performed according to a proportional distribution scheme.

Claims (67)

1. A computer-implemented method for redistributing data stored in a distributed storage, the method comprising:

maintaining a plurality of logically independent file systems, wherein each file system includes a data set stored by the distributed storage and metadata including a unique identifier and organizational structure information;

observing accesses to data files in the data set included in each of the file systems to determine access pattern levels;

determining a respective access pattern level for each of the plurality of logically independent file systems;

determining that a first file system from the plurality of logically independent file systems has an access pattern level specifying a higher probability of future access than a probability of future access specified by an access pattern level of a second file system from the plurality of logically independent file systems;

determining storage requirements for a plurality of file systems;

obtaining device characteristic information for the plurality of storage devices of the distributed storage;

redistributing the data set of the first file system having the higher probability of future access before redistributing the data set of the second file system across a plurality of storage devices of the distributed storage; and

redistributing data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage;

wherein redistributing the data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage comprises:

determining a respective performance level for each of the plurality of storage devices based on the device characteristic information;

computing an aggregate performance level for the distributed storage by summing the determined respective performance levels for the plurality of storage devices of the distributed storage;

computing a proportional performance level for a particular storage device of the plurality of storage devices by dividing the determined performance level for the particular storage device by the aggregate performance level;

computing a target amount of storage space assigned to the first file system from the plurality of logically independent file systems for the particular storage device by multiplying a determined storage requirement for the first file system and the proportional performance level for the particular storage device; and

redistributing the data set of the first file system based at least in part on the computed target amount of storage space assigned to the first file system for the particular storage device.

2. The computer-implemented method of claim 1 , wherein

redistributing the data set of the first file system comprises redistributing the data set of the first file system uniformly across the plurality of storage devices of the distributed storage.

3. The computer-implemented method of claim 1 , wherein redistributing the data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage comprises: redistributing the data set of the first file system from the plurality of logically independent file systems across the plurality of storage devices in proportion to device characteristic information.

4. The computer-implemented method of claim 1 , wherein redistributing the data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage comprises: calculating a target amount of storage space assigned to a first file system for each of the plurality of storage devices by dividing a storage requirement for the first file system by a number of the plurality of storage devices; and redistributing the data set of the first file system across the plurality of storage devices based on the calculated target amounts of storage space.

5. The method of claim 1 , wherein:

the data set of the first file system and the data set of the second file system each comprises a plurality of data files.

6. The method of claim 1 , wherein:

the first file system and the second file system are each associated with a respective type of data selected from the list of: work data, temporary data, log data, and backup data; and

the first file system and the second file system are not associated with the same type of data.

7. A non-transitory computer readable storage medium executing computer program instructions for redistributing data stored in a distributed storage, the computer program instructions comprising instructions for:

maintaining a plurality of logically independent file systems, wherein each file system includes a data set stored by the distributed storage and metadata including a unique identifier and organizational structure information;

observing accesses to data files in the data set included in each of the file systems to determine access pattern levels; determining a respective access pattern level for each of the plurality of logically independent file systems;

determining that a first file system from the plurality of logically independent file systems has an access pattern level specifying a higher probability of future access than a probability of future access specified by an access pattern level of a second file system from the plurality of logically independent file systems;

determining storage requirements for a plurality of file systems;

obtaining device characteristic information for the plurality of storage devices of the distributed storage;

redistributing the data set of the first file system having the higher probability of future access before redistributing the data set of the second file system across a plurality of storage devices of the distributed storage; and

redistributing data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage;

wherein redistributing the data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage comprises:

determining a respective performance level for each of the plurality of storage devices based on the device characteristic information;

computing an aggregate performance level for the distributed storage by summing the determined respective performance levels for the plurality of storage devices of the distributed storage;

computing a proportional performance level for a particular storage device of the plurality of storage devices by dividing the determined performance level for the particular storage device by the aggregate performance level;

computing a target amount of storage space assigned to the first file system from the plurality of logically independent file systems for the particular storage device by multiplying a determined storage requirement for the first file system and the proportional performance level for the particular storage device; and

redistributing the data set of the first file system based at least in part on the computed target amount of storage space assigned to the first file system for the particular storage device.

8. The medium of claim 7 , wherein the instructions for redistributing the data set of the first file system comprises redistributing the data set of the first file system uniformly across the plurality of storage devices of the distributed storage.

9. The medium of claim 7 , wherein:

the data set of the first file system and the data set of the second file system each comprises a plurality of data files.

10. The medium of claim 7 , wherein:

the first file system and the second file system are each associated with a respective type of data selected from the list of: work data, temporary data, log data, and backup data; and

the first file system and the second file system are not associated with the same type of data.

11. A system comprising:

a non-transitory computer readable storage medium storing processor-executable computer program instructions for redistributing data stored in a distributed storage, the instructions comprising instructions for:

maintaining a plurality of logically independent file systems, wherein each file system includes a data set stored by the distributed storage and metadata including a unique identifier and organizational structure information;

observing accesses to data files in the data set included in each of the file systems to determine access pattern levels;

determining a respective access pattern level for each of the plurality of logically independent file systems;

determining that a first file system from the plurality of logically independent file systems has an access pattern level specifying a higher probability of future access than a probability of future access specified by an access pattern level of a second file system from the plurality of logically independent file systems;

determining storage requirements for a plurality of file systems;

obtaining device characteristic information for the plurality of storage devices of the distributed storage;

redistributing the data set of the first file system having the higher probability of future access before redistributing the data set of the second file system across a plurality of storage devices of the distributed storage; and

redistributing data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage;

wherein redistributing the data sets across the plurality of storage devices based on the storage requirements for the plurality of file systems and the obtained device characteristic information for the plurality of storage devices of the distributed storage comprises:

determining a respective performance level for each of the plurality of storage devices based on the device characteristic information;

computing an aggregate performance level for the distributed storage by summing the determined respective performance levels for the plurality of storage devices of the distributed storage;

computing a proportional performance level for a particular storage device of the plurality of storage devices by dividing the determined performance level for the particular storage device by the aggregate performance level;

computing a target amount of storage space assigned to the first file system from the plurality of logically independent file systems for the particular storage device by multiplying a determined storage requirement for the first file system and the proportional performance level for the particular storage device; and

redistributing the data set of the first file system based at least in part on the computed target amount of storage space assigned to the first file system for the particular storage device; and

a computer processor for executing the computer program instructions.

12. The system of claim 11 , wherein redistributing the data set of the first file system comprises redistributing the data set of the first file system uniformly across the plurality of storage devices of the distributed storage.

13. The system of claim 11 , wherein:

the data set of the first file system and the data set of the second file system each comprises a plurality of data files.

14. The system of claim 11 , wherein:

the first file system and the second file system are each associated with a respective type of data selected from the list of: work data, temporary data, log data, and backup data; and

the first file system and the second file system are not associated with the same type of data.

Assignments (13)
RELEASE OF SECURITY INTEREST Recorded Jun 21, 2024
From: BANK OF AMERICA, N.A.
To: QUANTCAST CORPORATION
Reel/Frame 067807/0017 →
SECURITY INTEREST Recorded Jun 18, 2024
From: QUANTCAST CORPORATION
To: CRYSTAL FINANCIAL LLC D/B/A SLR CREDIT SOLUTIONS
Reel/Frame 067777/0613 →
SECURITY INTEREST Recorded Dec 5, 2022
From: QUANTCAST CORPORATION
To: VENTURE LENDING & LEASING IX, INC.; WTI FUND X, INC.
Reel/Frame 062066/0265 →
SECURITY INTEREST Recorded Sep 30, 2021
From: QUANTCAST CORPORATION
To: BANK OF AMERICA, N.A., AS AGENT
Reel/Frame 057677/0297 →
RELEASE OF SECURITY INTEREST Recorded Sep 30, 2021
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: QUANTCST CORPORATION
Reel/Frame 057678/0832 →
RELEASE OF SECURITY INTEREST Recorded May 6, 2021
From: VENTURE LENDING & LEASING VI, INC.; VENTURE LENDING & LEASING VII, INC.
To: QUANTCAST CORPORATION
Reel/Frame 056159/0702 →
RELEASE OF SECURITY INTEREST Recorded Mar 15, 2021
From: TRIPLEPOINT VENTURE GROWTH BDC CORP.
To: QUANTCAST CORPORATION
Reel/Frame 055599/0282 →
SECURITY INTEREST Recorded Aug 7, 2018
From: QUANTCAST CORPORATION
To: TRIPLEPOINT VENTURE GROWTH BDC CORP.
Reel/Frame 046733/0305 →
PATENT SECURITY AGREEMENT Recorded Jun 26, 2015
From: QUANTCAST CORPORATION
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 036020/0721 →
SECURITY AGREEMENT Recorded Oct 18, 2013
From: QUANTCAST CORPORATION
To: VENTURE LENDING & LEASING VI, INC.; VENTURE LENDING & LEASING VII, INC.
Reel/Frame 031438/0474 →
SECURITY AGREEMENT Recorded Jul 10, 2013
From: QUANTCAST CORPORATION
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 030772/0488 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2013
From: RUS, SILVIUS V.; SUBRAMANIAM, THILEEPAN
To: QUANTCAST CORPORATION
Reel/Frame 030333/0475 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2013
From: RUS, SILVIUS V.; SUBRAMANIAM, THILEEPAN
To: QUANTCAST CORP.
Reel/Frame 030330/0576 →