IP Library Granted Patent US 12,487,970
Granted Patent B2
US 12,487,970 · App. 18/957,212 · Granted Dec 2, 2025

Container-based erasure coding

Inventors: Apurv Gupta (Bengaluru, IN); Akshat Agarwal (Delhi, IN); Manvendra Singh Tomar (Bengaluru, IN); Donthula Akshith Reddy (Bheemaram, IN); Kushal Singh (Bengaluru, IN); Tarun Kumar Yadav (Rajasthan, IN); Mandar Suresh Naik (Pune, IN)
Assignee: Cohesity, Inc.
G06F16/1748
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,487,970
App. No.
18/957,212
Granted
Dec 2, 2025
Kind
B2
Abstract

A repository of replicated chunk files is analyzed to identify chunk files that meet at least a portion of combination criteria. Selected chunk files are associated together under a data protection grouping container. Erasure coding is applied to the data protection grouping container including by utilizing the selected chunk files as different data stripes of the erasure coding and generating one or more parity stripes based on the different data stripes.

Claims (47)

1 . A method, comprising:

selecting one or more replicated chunk files stored on a first storage device associated with an original chunk file stored on a second storage device different than the first storage device;

updating a metadata table to indicate the one or more replicated chunk files selected are associated with a data protection grouping container without updating metadata within the one or more replicated chunk files associated with the data protection grouping container; and

applying erasure coding to the data protection grouping container, including utilizing the one or more replicated chunk files associated with the data protection grouping container as different data stripes of the erasure coding and generating one or more parity stripes based on the different data stripes.

2 . The method of claim 1 , further comprising:

selecting the one or more replicated chunk files based on combination criteria including at least the one or more replicated chunk files stored on the first storage device different from the second storage device to which the original chunk file is stored.

3 . The method of claim 2 , further comprising:

selecting the one or more replicated chunk files based on the combination criteria further including on one or more of:

an age of one of the one or more replicated chunk files selected, a size of the one of the one or more replicated chunk files selected, whether the one of the one or more replicated chunk files selected includes non-deduplicated data chunks, a storage node that includes the first storage device or the second storage device storing the one of the one or more replicated chunk files selected, a chassis including the storage node that includes the first storage device or the second storage device storing the one of the one or more replicated chunk files selected, and/or a rack including the chassis including the storage node that includes the first storage device or the second storage device storing the one of the one or more replicated chunk files selected.

4 . The method of claim 1 , further comprising:

associating the one or more replicated chunk files selected in the data protection grouping container by updating the metadata table to indicate the one or more replicated chunk files are associated with the data protection grouping container without writing the one or more replicated chunk files associated with the data protection grouping container to a new chunk file.

5 . The method of claim 4 , further comprising:

updating the metadata table to include an entry, corresponding to the data protection grouping container, that associates the one or more replicated chunk files selected with a corresponding storage device of a storage system.

6 . The method of claim 1 , further comprising:

analyzing a repository of replicated chunk files to select the one or more replicated chunk files; and

wherein the analyzing includes selecting and including a first replicated chunk file in the data protection grouping container and selecting a second replicated chunk file based on the second replicated chunk file satisfying combination criteria.

7 . The method of claim 6 , further comprising:

determining whether the second replicated chunk file satisfies the combination criteria.

8 . The method of claim 7 , further comprising:

associating the second replicated chunk file with the data protection grouping container based on determining the second replicated chunk file satisfies the combination criteria.

9 . The method of claim 8 , further comprising determining whether the data protection grouping container is full based on an erasure coding configuration.

10 . The method of claim 9 , wherein the erasure coding configuration indicates a number of data stripes and a number of parity stripes to include in the data protection grouping container.

11 . The method of claim 1 , wherein the one or more replicated chunk files selected and one or more parity stripes generated are stored on different storage devices of a storage system.

12 . The method of claim 1 , further comprising:

monitoring the data protection grouping container; and

performing a garbage collection process.

13 . The method of claim 12 , wherein the garbage collection process includes determining whether a measure of unreferenced data chunks associated with the data protection grouping container is greater than a threshold measure of unreferenced data chunks.

14 . The method of claim 13 , further comprising:

generating a new data protection grouping container based on whether the measure of unreferenced data chunks associated with the data protection grouping container is greater than the threshold measure of unreferenced data chunks.

15 . The method of claim 14 , further comprising:

updating the data protection grouping container based on whether the measure of unreferenced data chunks associated with the data protection grouping container is greater than the threshold measure of unreferenced data chunks.

16 . Non-transitory computer-readable storage media encoded with instructions that, when executed, cause one or more processors to:

select one or more replicated chunk files stored on a first storage device associated with an original chunk file stored on a second storage device different than the first storage device;

update a metadata table to indicate the one or more replicated chunk files selected are associated with a data protection grouping container without updating metadata within the one or more replicated chunk files associated with the data protection grouping container; and

apply erasure coding to the data protection grouping container, including utilizing the one or more replicated chunk files associated with the data protection grouping container as different data stripes of the erasure coding and generating one or more parity stripes based on the different data stripes.

17 . The non-transitory computer-readable storage media of claim 16 , wherein the instructions, when executed, further cause the one or more processors to:

select the one or more replicated chunk files based on combination criteria including at least the one or more replicated chunk files stored on the first storage device different from the second storage device to which the original chunk file is stored.

18 . The non-transitory computer-readable storage media of claim 16 , wherein the instructions, when executed, further cause the one or more processors to:

associate the one or more replicated chunk files selected in the data protection grouping container by updating the metadata table to indicate the one or more replicated chunk files are associated with the data protection grouping container without writing the one or more replicated chunk files associated with the data protection grouping container to a new chunk file.

19 . A system, comprising:

one or more processors; and

memory, coupled to the one or more processors, the memory storing instructions that when executed cause the one or more processors to:

select one or more replicated chunk files stored on a first storage device associated with an original chunk file stored on a second storage device different than the first storage device;

update a metadata table to indicate the one or more replicated chunk files selected are associated with a data protection grouping container without updating metadata within the one or more replicated chunk files associated with the data protection grouping container; and

apply erasure coding to the data protection grouping container, including utilizing the one or more replicated chunk files associated with the data protection grouping container as different data stripes of the erasure coding and generating one or more parity stripes based on the different data stripes.

20 . The system of claim 19 , wherein the instructions, when executed, further cause the one or more processors to:

associate the one or more replicated chunk files selected in the data protection grouping container by updating the metadata table to indicate the one or more replicated chunk files are associated with the data protection grouping container without writing the one or more replicated chunk files associated with the data protection grouping container to a new chunk file.

Assignments (1)
PATENT SECURITY AGREEMENT SUPPLEMENT Recorded Aug 6, 2025
From: COHESITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 072373/0649 →
Continuity (2)
Continuation 17582763 · Jan 24, 2022
Related Publication 20250086145A1 · Mar 13, 2025
References Cited (15)
US 11681586B2 · Terei · 2023 [cited by examiner]
US 11687286B2 · Aguilera · 2023 [cited by examiner]
US 20130339818A1 · Baker et al. · 2013 [cited by applicant]
US 20150220398A1 · Schirripa · 2015 [cited by examiner]
US 20160378612A1 · Hipsh · 2016 [cited by examiner]
US 20170177473A1 · Danilov · 2017 [cited by examiner]
US 20170277630A1 · Wideman et al. · 2017 [cited by applicant]
US 20190370170A1 · Oltean et al. · 2019 [cited by applicant]
US 20230079486A1 · Yarlagadda et al. · 2023 [cited by applicant]
US 20230237020A1 · Gupta et al. · 2023 [cited by applicant]
US 20230237029A1 · Tal · 2023 [cited by examiner]
Haddock et al., “High Performance Erasure Coding for Very Large Stripe Sizes”, 2019 Spring Simulation Conference (SpringSim), IEEE, Apr. 29, 2019, 12 pp. [cited by applicant]
Hu et al., Exploiting Combined Locality for Wide-Stripe Erasure Coding in Distributed Storage, Proceedings of the 19th USENIX Conference on File and Storage Technologies, Feb. 2021, pp. 233-248. [cited by applicant]
Huang et al., “Erasure Coding in Windows Azure Storage”, 2012 USENIX Annual Technical Conference (USENIX ATC 12), 2012, pp. 15-26, (Applicant points out, in accordance with MPEP 609.04(a), that the year of publication, … [cited by applicant]
Prosecution History from U.S. Appl. No. 17/582,763, dated Feb. 21, 2023 through Nov. 8, 2024, 95 pp. [cited by applicant]