IP Library Granted Patent US 12,373,312
Granted Patent B1
US 12,373,312 · App. 18/421,744 · Granted Jul 29, 2025

Failure recovery of solid state disk (SSD) storage for cluster file system serviceability

Inventors: Vishal Tiwary (Sunnyvale, CA); Philip Shilane (Newtown, PA)
Assignee: Dell Products L.P.
G06F11/1469G06F3/0616G06F3/065G06F3/0679G06F11/1453
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,312
App. No.
18/421,744
Granted
Jul 29, 2025
Kind
B1
Abstract

A process of recovering from a failure of a solid-state device (SSD) storing a persistent volume (PV) for logging in a cluster network. The PV is recreated on a different SSD that has spare capacity and IOPS resources, and log information is redirected from an application to the relocated PV on the different SSD. Upon replacement of the original failed SSD, the PV can be copied from the different SSD to the new replacement SSD, or the replacement SSD can be left empty, and the relocated PV on the different SSD can continue to be used. This process provides a high degree of redundancy with regard to logging so that a failure of single SSD does not bring down the node or the entire cluster network.

Claims (30)

1. A method of maintaining solid state disk (SSD) devices for storage of persistent volumes (PVs) for support bundle processing in a cluster system operated having a plurality of nodes executing containerized applications, comprising:

recreating, upon encountering a failure of an original SSD storing an original PV, a new PV for logging on a different SSD that has spare capacity and IOPS resources;

redirecting an application pod generating log information for the logging to the new PV on the different SSD;

replacing the original SSD with a replacement SSD;

keeping the replacement SSD empty for storage capacity to accommodate PVs from future failing SSD devices in the cluster system; and

maintaining the new PV on the different SSD for the logging, wherein the logging comprises collecting logs for each node of a plurality of nodes in the cluster system that record transactions of the nodes during execution of the applications, generating, for the entire cluster system, log files from the logs for each node, and storing the log files for a corresponding node in the original PV, wherein the applications comprise non-critical and critical applications requiring node affinity in the cluster system, and further wherein the original PV comprises storage for critical applications writing directly to a persistent location dedicated to each critical application.

2. The method of claim 1 wherein the original PV comprises a central PV maintained for writing log files for non-critical applications to a local file using one of a standard output function and a log forwarder utility, or a sidecar container that tracks log files in a pod and that transfers the log files to other services.

3. The method of claim 1 wherein the logs are collected for events including system changes, authorization activities, privileged access events, and audit related activities.

4. The method of claim 3 wherein the system changes comprise component failures and changes to availability, configuration, or source code; the authorization activities comprise login or access failures and access provisioning; the privileged access events comprise using a superuser status or running an administrator console; and the audit-related activities comprise changing event logging configurations, deleting event logs, or audit log failures.

5. The method of claim 1 wherein the cluster system comprises a Santorini network processing containerized data utilizing a Kubernetes-based framework, and wherein the cluster system comprises part of a Data Domain deduplication backup system performing backup and restore operations for the nodes, and wherein the containerized applications comprise at least one of a Data Domain container running deduplication and compression processes, a cloud-native data protection manager, and a scalable object storage manager.

6. A method of maintaining solid state disk (SSD) devices for storage of persistent volumes (PVs) for support bundle processing in a cluster system operated having a plurality of nodes executing containerized applications, comprising:

recreating, upon encountering a failure of an original SSD storing an original PV, a new PV for logging on a different SSD that has spare capacity and IOPS resources;

first redirecting an application pod generating log information for the logging to the new PV on the different SSD;

replacing the original SSD with a replacement SSD;

copying the new PV from the different SSD to a replacement PV on the replacement SSD;

second redirecting an application pod generating log information for the logging to the replacement PV on the replacement SSD; and

deleting the new PV on the different SSD, wherein the logging comprises collecting logs for each node of a plurality of nodes in the cluster system that record transactions of the nodes during execution of the applications, generating, for the entire cluster system, log files from the logs for each node, and storing the log files for a corresponding node in the original PV, wherein the applications comprise non-critical and critical applications requiring node affinity in the cluster system, and further wherein the original PV comprises storage for critical applications writing directly to a persistent location dedicated to each critical application.

7. The method of claim 6 wherein the original PV comprises a central PV maintained for writing log files for non-critical applications to a local file using one of a standard output function and a log forwarder utility, or a sidecar container that tracks log files in a pod and that transfers the log files to other services.

8. The method of claim 6 wherein the logs are collected for events including system changes, authorization activities, privileged access events, and audit related activities.

9. The method of claim 8 wherein the system changes comprise component failures and changes to availability, configuration, or source code; the authorization activities comprise login or access failures and access provisioning; the privileged access events comprise using a superuser status or running an administrator console; and the audit-related activities comprise changing event logging configurations, deleting event logs, or audit log failures.

10. The method of claim 6 wherein the cluster system comprises a Santorini network processing containerized data utilizing a Kubernetes-based framework, and wherein the cluster system comprises part of a Data Domain deduplication backup system performing backup and restore operations for the nodes, and wherein the containerized applications comprise at least one of a Data Domain container running deduplication and compression processes, a cloud-native data protection manager, and a scalable object storage manager.

11. A method of recovering from solid state device (SSD) failure while logging activities in a cluster system operated by a user and having a plurality of nodes executing containerized applications, the method comprising:

collecting logs for each node of a plurality of nodes in the cluster system, wherein the logs record transactions of the nodes during execution of the applications;

storing, on an original SSD, log files for a node in an original persistent volume (PV);

recreating, upon encountering a failure of the original SSD, a new PV on a different SSD for continued logging in the node;

redirecting an application generating log information for the logging to the new PV on the different SSD;

first replacing the original SSD with a replacement SSD; and

keeping the replacement SSD empty for storage capacity to accommodate PVs from future failing SSD devices in the cluster system while maintaining the new PV on the different SSD for the logging, or copying the new PV from the different SSD to a replacement PV on the replacement SSD and second redirecting the application to the replacement PV on the replacement SSD, wherein the cluster system comprises a Santorini network processing containerized data utilizing a Kubernetes-based framework, and wherein the cluster network comprises part of a Data Domain deduplication backup system performing backup and restore operations for the nodes, and wherein the containerized applications comprise at least one of a Data Domain container running deduplication and compression processes.

12. The method of claim 11 wherein the original PV comprises a central PV maintained for writing log files for non-critical applications to a local file using one of a standard output function and a log forwarder utility, or a sidecar container that tracks log files in a pod and that transfers the log files to other services.

13. The method of claim 11 wherein the original PV comprises storage for critical applications writing directly to a persistent location dedicated to each critical application, wherein a critical application is an application that requires node affinity in the cluster system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2024
From: SHILANE, PHILIP; TIWARY, VISHAL
To: DELL PRODUCTS L.P.
Reel/Frame 066236/0202 →
References Cited (27)
US 6460141B1 · Olden · 2002 [cited by applicant]
US 10592418B2 · Kelly · 2020 [cited by applicant]
US 11379421B1 · Ciubotariu · 2022 [cited by applicant]
US 11561868B1 · Poornachandran · 2023 [cited by examiner]
US 11709760B1 · Abogado · 2023 [cited by applicant]
US 12045198B2 · Mathew · 2024 [cited by applicant]
US 12166811B2 · Morgan · 2024 [cited by applicant]
US 20080306977A1 · Karuppiah · 2008 [cited by applicant]
US 20110246826A1 · Hsieh · 2011 [cited by applicant]
US 20140006881A1 · Loimuneva · 2014 [cited by applicant]
US 20140304401A1 · Jagadish · 2014 [cited by applicant]
US 20150112935A1 · French · 2015 [cited by applicant]
US 20170310537A1 · Henry · 2017 [cited by applicant]
US 20180027115A1 · Kamboh · 2018 [cited by applicant]
US 20180307545A1 · Ganesan · 2018 [cited by applicant]
US 20200201699A1 · Yu · 2020 [cited by applicant]
US 20200356292A1 · Ippatapu · 2020 [cited by applicant]
US 20210157710A1 · Alexander · 2021 [cited by applicant]
US 20220060398A1 · Shishir · 2022 [cited by applicant]
US 20230050796A1 · Ghergu · 2023 [cited by applicant]
US 20230236953A1 · Kumar · 2023 [cited by applicant]
US 20230342222A1 · Mathew · 2023 [cited by applicant]
US 20240037069A1 · Mathew · 2024 [cited by applicant]
US 20240111557A1 · Kulkarni · 2024 [cited by applicant]
US 20240320012A1 · Zhang · 2024 [cited by applicant]
US 20240419684A1 · Achi Vasudevan · 2024 [cited by examiner]
Kubernetes cluster-wide logging using sidecars, Dec. 12, 2023 https://web.archive.org/web/20231212102624/https://kubernetes.io/docs/concepts/cluster-administration/logging/ (Year: 2023). [cited by examiner]