IP Library › Granted Patent US 12,566,568
Granted Patent B2
US 12,566,568 · App. 18/598,554 · Granted Mar 3, 2026

Migrating erasure-coded data fragments within a growing distributed data storage system with at least data plus parity nodes

Inventors: Anand Vishwanath Vastrad (San Jose, CA); Suhani Gupta (San Jose, CA)
Assignee: Commvault Systems, Inc.
G06F3/0655G06F3/0604G06F3/0619G06F3/064G06F3/065G06F3/0652G06F3/0664G06F3/067
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,568
App. No.
18/598,554
Filed
Mar 7, 2024
Granted
Mar 3, 2026
Kind
B2
Art Unit
2135
USPC
711/154
Abstract

A distributed data storage system that employs erasure coding grows from fewer than data-plus-parity (D+P) storage service nodes to at least D+P nodes. The system detects an increase in the number of available storage service nodes, i.e., at least D+P; analyzes how storage for each virtual disk is distributed among storage containers in the existing (pre-growth) nodes; identifies containers that are co-hosted on the same node; and, on a node-by-node basis, migrates data fragments from a co-hosted storage container to a corresponding new container that is configured on another node. For each virtual disk in the illustrative system, the migration causes containers, including the erasure-coded data fragments they host, to be re-distributed so that the containers for a virtual disk are NOT doubled up or co-hosted on the same node. The disclosed computer-implemented process for re-distributing erasure-coded data fragments operates organically, without requiring the system to restart or reboot.

Claims (52)

1 . A system comprising:

a first plurality of storage service nodes, wherein physical data storage resources are configured among the first plurality of storage service nodes; and

wherein a first storage service node among the first plurality of storage service nodes is configured to:

receive a write request comprising a first data block in unfragmented form, wherein the write request indicates that the first data block is to be written to a virtual disk configured in the system;

determine that the first data block is to be stored according to an erasure-coding scheme, which generates a count of D data fragments and a count of P parity fragments that total N erasure-coded fragments of the first data block;

based on determining that the first plurality of storage service nodes is fewer than N, cause the N erasure-coded fragments to be written at one or more storage service nodes among the first plurality of storage service nodes, wherein only one instance of each of the N erasure-coded fragments is stored among the first plurality of storage service nodes in the system, and wherein the first storage service node comprises a first one of the N erasure-coded fragments and further comprises a second one of the N erasure-coded fragments;

after the N erasure-coded fragments have been stored among the first plurality of storage service nodes that is fewer than N, detect that more storage service nodes are currently operational in the system;

based on determining that at least N storage service nodes are currently operational in the system, initiate a migration that moves the second one of the N erasure-coded fragments from the first storage service node to a second storage service node, wherein after the migration is complete, each of the N erasure-coded fragments is stored on a distinct storage service node in the system.

2 . The system of claim 1 , wherein the migration re-distributes storage containers of the virtual disk, including erasure-coded data fragments hosted therein, such that the storage containers of the virtual disk are not co-hosted on a same storage service node after the migration is complete.

3 . The system of claim 1 , wherein at least one of the first plurality of storage service nodes is configured to, responsive to the first storage service node initiating the migration, identify, among the more storage service nodes, the second storage service node that is suitable for the migration.

4 . The system of claim 1 , wherein a data storage subsystem that executes at the first storage service node is configured to: based on determining that the first storage service node comprises two storage containers of the virtual disk, request a metadata subsystem to identify another storage service node that is suitable to host one of the two storage containers;

wherein the metadata subsystem, which executes at one of the first plurality of storage service nodes, is configured to identify, among the more storage service nodes, a second storage service node that does not comprise any storage containers of the virtual disk and is further configured to indicate the second storage service node to the data storage subsystem for the migration.

5 . The system of claim 1 , wherein the first storage service node is further configured to:

after the migration is complete, receive a second write request comprising a second data block in unfragmented form, wherein the second write request indicates that the second data block is to be written to the virtual disk;

based on a configuration of the virtual disk programmed into the system, determine that the second data block is to be stored according to the erasure-coding scheme, which generates N erasure-coded fragments of the second data block; and

based on determining that at least N storage service nodes are currently operational in the system, cause the N erasure-coded fragments of the second data block to be written at N distinct storage service nodes among the at least N storage service nodes that are currently operational in the system, wherein only one instance of each of the N erasure-coded fragments of the second data block is stored at any one storage service node.

6 . The system of claim 1 , wherein the first storage service node is further configured to:

receive a read request for the first data block;

determine that the first data block is stored in the system as the N erasure-coded fragments of the first data block, which consist of D data fragments and P parity fragments;

obtain D erasure-coded fragments from among the N erasure-coded fragments of the first data block, wherein the count of D is sufficient to reconstruct the first data block according to the erasure-coding scheme;

reconstruct the first data block in unfragmented form from the D erasure-coded fragments; and

transmit the first data block in unfragmented form in response to the read request.

7 . The system of claim 1 , wherein the virtual disk comprises storage containers that are distributed among the first plurality of storage service nodes, and wherein the migration causes a first storage container configured at the first storage service node to move to the second storage service node, wherein after the migration is complete, none of the storage containers of the virtual disk are co-located at a same storage service node.

8 . The system of claim 1 , wherein a data storage subsystem that executes at the first storage service node is configured to initiate the migration.

9 . The system of claim 8 , wherein a metadata subsystem that executes at one of the first plurality of storage service nodes is configured to identify the second storage service node as a destination for the migration.

10 . The system of claim 1 , wherein the migration does not comprise a reboot of the system.

11 . The system of claim 1 , wherein the migration does not comprise a restart of the system.

12 . A system comprising:

a first plurality of storage service nodes, wherein physical data storage resources are configured among the first plurality of storage service nodes; and

wherein a first storage service node among the first plurality of storage service nodes is configured to:

receive a write request comprising a first data block in unfragmented form, wherein the write request indicates that the first data block is to be written to a virtual disk configured in the system, wherein the virtual disk comprises storage containers that are distributed among the first plurality of storage service nodes;

determine that the first data block is to be stored according to an erasure-coding scheme, which generates a count of N erasure-coded fragments of the first data block, wherein N consists of D data fragments and P parity fragments;

based on determining that the first plurality of storage service nodes is fewer than N, cause the N erasure-coded fragments to be written at one or more storage service nodes among the first plurality of storage service nodes, wherein only one instance of each of the N erasure-coded fragments is stored among the first plurality of storage service nodes in the system, and wherein a first storage container of the virtual disk configured at the first storage service node comprises a first one of the N erasure-coded fragments, and wherein a second storage container of the virtual disk also configured at the first storage service node comprises a second one of the N erasure-coded fragments;

after the N erasure-coded fragments have been stored among the first plurality of storage service nodes that is fewer than N, based on determining that at least N storage service nodes are currently operational in the system, initiate a migration that moves the second storage container of the virtual disk, including the second one of the N erasure-coded fragments, to a second storage service node that lacks any storage container of the virtual disk; and

wherein after the migration is complete, each of the N erasure-coded fragments is stored on a distinct storage service node in the system.

13 . The system of claim 12 , wherein after the migration is complete, none of the storage containers of the virtual disk are co-located at a same storage service node.

14 . The system of claim 12 , wherein the migration re-distributes the storage containers of the virtual disk, including erasure-coded data fragments hosted therein, such that the storage containers of the virtual disk are not co-hosted on a same storage service node after the migration is complete.

15 . The system of claim 12 , wherein at least one of the first plurality of storage service nodes is configured to, responsive to the first storage service node initiating the migration, identify the second storage service node that is suitable for the migration.

16 . The system of claim 12 , wherein a data storage subsystem that executes at the first storage service node is configured to: based on determining that the first storage service node comprises two storage containers of the virtual disk, request a metadata subsystem to identify another storage service node that is suitable to host one of the two storage containers;

wherein the metadata subsystem, which executes at one of the first plurality of storage service nodes, is configured to identify the second storage service node that does not comprise any storage containers of the virtual disk, and is further configured to indicate the second storage service node to the data storage subsystem for the migration.

17 . The system of claim 12 , wherein the first storage service node is further configured to:

after the migration is complete, receive a second write request comprising a second data block in unfragmented form, wherein the second write request indicates that the second data block is to be written to the virtual disk;

based on a configuration of the virtual disk programmed into the system, determine that the second data block is to be stored according to the erasure-coding scheme, which generates N erasure-coded fragments of the second data block; and

based on determining that at least N storage service nodes are currently operational in the system, cause the N erasure-coded fragments of the second data block to be written at N distinct storage service nodes among the at least N storage service nodes that are currently operational in the system, wherein only one instance of each of the N erasure-coded fragments of the second data block is stored at any one storage service node.

18 . The system of claim 12 , wherein the first storage service node is further configured to:

receive a read request for the first data block;

determine that the first data block is stored in the system as the N erasure-coded fragments of the first data block, which consist of D data fragments and P parity fragments;

obtain D erasure-coded fragments from among the N erasure-coded fragments of the first data block, wherein the count of D is sufficient to reconstruct the first data block according to the erasure-coding scheme;

reconstruct the first data block in unfragmented form from the D erasure-coded fragments; and

transmit the first data block in unfragmented form in response to the read request.

19 . The system of claim 12 , wherein a data storage subsystem that executes at the first storage service node is configured to initiate the migration, based on determining that more than one storage container of the virtual disk is configured at the first storage service node.

20 . The system of claim 19 , wherein a metadata subsystem that executes at one of the first plurality of storage service nodes is configured to identify the second storage service node as a destination for the migration.

Assignments (2)
SUPPLEMENTAL CONFIRMATORY GRANT OF SECURITY INTEREST IN UNITED STATES PATENTS Recorded Apr 16, 2025
From: COMMVAULT SYSTEMS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 070864/0344 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2024
From: VASTRAD, ANAND VISHWANATH; GUPTA, SUHANI
To: COMMVAULT SYSTEMS, INC.
Reel/Frame 066927/0166 →
Continuity (6)
Continuation In Part 18112078 · Feb 21, 2023
Continuation 17336103 · Jun 1, 2021
Provisional Application 63454441 · Mar 24, 2023
Provisional Application 63065722 · Aug 14, 2020
Provisional Application 63053414 · Jul 17, 2020
Related Publication 20240211167A1 · Jun 27, 2024
References Cited (150)
US 4084231A · Capozzi et al. · 1978 [cited by applicant]
US 4267568A · Dechant et al. · 1981 [cited by applicant]
US 4283787A · Chambers · 1981 [cited by applicant]
US 4417321A · Chang et al. · 1983 [cited by applicant]
US 4641274A · Swank · 1987 [cited by applicant]
US 4654819A · Stiffler et al. · 1987 [cited by applicant]
US 4686620A · Ng · 1987 [cited by applicant]
US 4912637A · Sheedy et al. · 1990 [cited by applicant]
US 4995035A · Cole · 1991 [cited by applicant]
US 5005122A · Griffin · 1991 [cited by applicant]
US 5093912A · Dong et al. · 1992 [cited by applicant]
US 5133065A · Cheffetz et al. · 1992 [cited by applicant]
US 5193154A · Kitajima et al. · 1993 [cited by applicant]
US 5212772A · Masters · 1993 [cited by applicant]
US 5226157A · Nakano et al. · 1993 [cited by applicant]
US 5239647A · Anglin et al. · 1993 [cited by applicant]
US 5241668A · Eastridge et al. · 1993 [cited by applicant]
US 5241670A · Eastridge et al. · 1993 [cited by applicant]
US 5276860A · Fortier et al. · 1994 [cited by applicant]
US 5276867A · Kenley et al. · 1994 [cited by applicant]
US 5287500A · Stoppani, Jr. · 1994 [cited by applicant]
US 5301286A · Rajani · 1994 [cited by applicant]
US 5321816A · Rogan et al. · 1994 [cited by applicant]
US 5347653A · Flynn et al. · 1994 [cited by applicant]
US 5410700A · Fecteau et al. · 1995 [cited by applicant]
US 5420996A · Aoyagi · 1995 [cited by applicant]
US 5454099A · Myers et al. · 1995 [cited by applicant]
US 5559991A · Kanfi · 1996 [cited by applicant]
US 5642496A · Kanfi · 1997 [cited by applicant]
US 6418478B1 · Ignatius et al. · 2002 [cited by applicant]
US 6542972B2 · Ignatius et al. · 2003 [cited by applicant]
US 6658436B2 · Oshinsky et al. · 2003 [cited by applicant]
US 6721767B2 · DeMeno et al. · 2004 [cited by applicant]
US 6760723B2 · Oshinsky et al. · 2004 [cited by applicant]
US 7003641B2 · Prahlad · 2006 [cited by applicant]
US 7035880B1 · Crescenti · 2006 [cited by applicant]
US 7107298B2 · Prahlad · 2006 [cited by applicant]
US 7130970B2 · Devassy · 2006 [cited by applicant]
US 7162496B2 · Amarendran et al. · 2007 [cited by applicant]
US 7174433B2 · Kottomtharayil et al. · 2007 [cited by applicant]
US 7246207B2 · Kottomtharayil · 2007 [cited by applicant]
US 7315923B2 · Retnamma · 2008 [cited by applicant]
US 7343453B2 · Prahlad · 2008 [cited by applicant]
US 7389311B1 · Crescenti et al. · 2008 [cited by applicant]
US 7395282B1 · Crescenti · 2008 [cited by applicant]
US 7440982B2 · Lu · 2008 [cited by applicant]
US 7454569B2 · Kavuri · 2008 [cited by applicant]
US 7490207B2 · Amarendran et al. · 2009 [cited by applicant]
US 7500053B1 · Kavuri · 2009 [cited by applicant]
US 7529782B2 · Prahlad · 2009 [cited by applicant]
US 7536291B1 · Vijayan Retnamma et al. · 2009 [cited by applicant]
US 7543125B2 · Gokhale · 2009 [cited by applicant]
US 7546324B2 · Prahlad et al. · 2009 [cited by applicant]
US 7603386B2 · Amarendran et al. · 2009 [cited by applicant]
US 7606844B2 · Kottomtharayil · 2009 [cited by applicant]
US 7613752B2 · Prahlad · 2009 [cited by applicant]
US 7617253B2 · Prahlad et al. · 2009 [cited by applicant]
US 7617262B2 · Prahlad · 2009 [cited by applicant]
US 7620710B2 · Kottomtharayil · 2009 [cited by applicant]
US 7636743B2 · Erofeev · 2009 [cited by applicant]
US 7651593B2 · Prahlad · 2010 [cited by applicant]
US 7657550B2 · Prahlad · 2010 [cited by applicant]
US 7660807B2 · Prahlad · 2010 [cited by applicant]
US 7661028B2 · Erofeev · 2010 [cited by applicant]
US 7734669B2 · Kottomtharayil · 2010 [cited by applicant]
US 7747579B2 · Prahlad · 2010 [cited by applicant]
US 7801864B2 · Prahlad · 2010 [cited by applicant]
US 7809914B2 · Kottomtharayil · 2010 [cited by applicant]
US 8156086B2 · Lu et al. · 2012 [cited by applicant]
US 8170995B2 · Prahlad · 2012 [cited by applicant]
US 8229954B2 · Kottomtharayil · 2012 [cited by applicant]
US 8230195B2 · Amarendran · 2012 [cited by applicant]
US 8285681B2 · Prahlad · 2012 [cited by applicant]
US 8307177B2 · Prahlad · 2012 [cited by applicant]
US 8364652B2 · Vijayan · 2013 [cited by applicant]
US 8370542B2 · Lu et al. · 2013 [cited by applicant]
US 8578120B2 · Attarde · 2013 [cited by applicant]
US 8954446B2 · Vijayan Retnamma et al. · 2015 [cited by applicant]
US 9020900B2 · Vijayan Retnamma et al. · 2015 [cited by applicant]
US 9098495B2 · Gokhale · 2015 [cited by applicant]
US 9239687B2 · Vijayan · 2016 [cited by applicant]
US 9411534B2 · Lakshman · 2016 [cited by applicant]
US 9424151B2 · Lakshman · 2016 [cited by applicant]
US 9483205B2 · Lakshman · 2016 [cited by applicant]
US 9558085B2 · Lakshman · 2017 [cited by applicant]
US 9633033B2 · Vijayan · 2017 [cited by applicant]
US 9639274B2 · Maranna · 2017 [cited by applicant]
US 9798489B2 · Lakshman · 2017 [cited by applicant]
US 9864530B2 · Lakshman · 2018 [cited by applicant]
US 9875063B2 · Lakshman · 2018 [cited by applicant]
US 10007587B2 · Richardson · 2018 [cited by examiner]
US 10067722B2 · Lakshman · 2018 [cited by applicant]
US 10248174B2 · Lakshman et al. · 2019 [cited by applicant]
US 10310953B2 · Vijayan · 2019 [cited by applicant]
US 10691187B2 · Lakshman · 2020 [cited by applicant]
US 10740300B1 · Lakshman et al. · 2020 [cited by applicant]
US 10795577B2 · Lakshman et al. · 2020 [cited by applicant]
US 10846024B2 · Lakshman · 2020 [cited by applicant]
US 10848468B1 · Lakshman et al. · 2020 [cited by applicant]
US 11016696B2 · Ankireddypalle et al. · 2021 [cited by applicant]
US 11314687B2 · Kavaipatti Anantharamakrishnan et al. · 2022 [cited by applicant]
US 11347590B1 · Resch et al. · 2022 [cited by applicant]
US 11487468B2 · Gupta et al. · 2022 [cited by applicant]
US 11614883B2 · Vastrad et al. · 2023 [cited by applicant]
US 11647075B2 · Camargos · 2023 [cited by applicant]
US 20060224846A1 · Amarendran · 2006 [cited by applicant]
US 20090319534A1 · Gokhale · 2009 [cited by applicant]
US 20120150818A1 · Vijayan Retnamma et al. · 2012 [cited by applicant]
US 20160188218A1 · Gray et al. · 2016 [cited by applicant]
US 20160350302A1 · Lakshman · 2016 [cited by applicant]
US 20160350391A1 · Vijayan et al. · 2016 [cited by applicant]
US 20170168903A1 · Dornemann et al. · 2017 [cited by applicant]
US 20170185330A1 · Danilov · 2017 [cited by applicant]
US 20170185488A1 · Kumarasamy et al. · 2017 [cited by applicant]
US 20170193003A1 · Vijayan et al. · 2017 [cited by applicant]
US 20170235647A1 · Kilaru et al. · 2017 [cited by applicant]
US 20170242871A1 · Kilaru et al. · 2017 [cited by applicant]
US 20180150351A1 · Li · 2018 [cited by examiner]
US 20180349043A1 · Eda et al. · 2018 [cited by applicant]
US 20190114094A1 · Ki · 2019 [cited by examiner]
US 20210208782A1 · Zhu · 2021 [cited by applicant]
US 20220066669A1 · Naik et al. · 2022 [cited by applicant]
US 20220100618A1 · Jain et al. · 2022 [cited by applicant]
US 20220103622A1 · Camargos et al. · 2022 [cited by applicant]
US 20230214152A1 · Vastrad · 2023 [cited by applicant]
EP 0259912 · 1988 [cited by applicant]
EP 0405926 · 1991 [cited by applicant]
EP 0467546 · 1992 [cited by applicant]
EP 0541281 · 1993 [cited by applicant]
EP 0774715 · 1997 [cited by applicant]
EP 0809184 · 1997 [cited by applicant]
EP 0899662 · 1999 [cited by applicant]
EP 0981090 · 2000 [cited by applicant]
WO 9513580 · 1995 [cited by applicant]
WO 9912098 · 1999 [cited by applicant]
WO 2006052872 · 2006 [cited by applicant]
WO 2016004120 · 2016 [cited by applicant]
Arneson, “Mass Storage Archiving in Network Environments,” Digest of Papers, Ninth IEEE Symposium on Mass Storage Systems, Oct. 31, 1988Nov. 3, 1988, pp. 45-50, Monterey, CA. [cited by applicant]
Arneson, David A., “Development of Omniserver,” Control Data Corporation, Tenth IEEE Symposium on Mass Storage Systems, May 1990, ‘Crisis in Mass Storage’ Digest of Papers, pp. 88-93, Monterey, CA. [cited by applicant]
Cabrera et al., “ADSM: A Multi-Platform, Scalable, Backup and Archive Mass Storage System,” Digest of Papers, Compcon '95, Proceedings of the 40th IEEE Computer Society International Conference, Mar. 5, 1995-Mar. 9, 199… [cited by applicant]
Commvault HyperScale™ Appliance 3300—Technical Specifications, dated Jul. 6, 2020, in 12 pages (https://www.commvault.com/resources/commvault-hyperscale-appliance-3300-technical-specifications). [cited by applicant]
Commvault Launches HyperScale X, Marking First Portfolio Integration Of Hedvig Technology, dated Jul. 21, 2020, in 10 pages (https://www.commvault.com/news/commvault-launches-hyperscale-x-marking-first-portfolio-integra… [cited by applicant]
Eitel, “Backup and Storage Management in Distributed Heterogeneous Environments,” IEEE, Jun. 12-16, 1994, pp. 124-126. [cited by applicant]
Hedvig distributed storage platform technical and architectural overview white paper—Executive Overview, dated Jul. 6, 2020, in 40 pages (https://www.commvault.com/resources/hedvig-distributed-storage-platform-technical… [cited by applicant]
Hedvig Erasure Coding Architectural Overview, copyright date listed on document is 2019, in 13 pages. [cited by applicant]
Huff, KL, “Data Set Usage Sequence Number,” IBM Technical Disclosure Bulletin, vol. 24, No. 5, Oct. 1981 New York, US, pp. 2404-2406. [cited by applicant]
Lakshman et al., “Cassandra—A Decentralized Structured Storage System”, https://doi.org/10.1145/1773912.1773922, ACM SIGOPS Operating Systems Review, vol. 44, Issue 2, Apr. 2010, pp. 35-40. [cited by applicant]
Rosenblum et al., “The Design and Implementation of a Log-Structure File System,” Operating Systems Review SIGOPS, vol. 25, No. 5, May 1991, New York, US, pp. 1-15. [cited by applicant]
SwiftStack, Inc., The OpenStack Object Storage System, Feb. 2012, pp. 1-29. [cited by applicant]
Wikipedia , “Erasure Code”, accessed on http://en.wikipedia.org/wiki/Erasure code, May 20, 2021, available on https://web.archive.org/web/20210413201654/https://en.wikipedia.org/wiki/Erasure_code, Jan. 8, 2024, 07 Pages. [cited by applicant]