IP Library › Granted Patent US 12,381,787
Granted Patent B1
US 12,381,787 · App. 18/430,109 · Granted Aug 5, 2025

Data center monitoring and management operation for data unavailable/data loss (DU/DL) predictions, including service level agreement (SLA) failure prediction

Inventors: Elie A Jreij (Pflugerville, TX); Amihai Savir (Newton, MA); Romulo Teixeira De Abreu Pinho (Niteroi, BR); Peter Ziesmann (Spicewood, TX); Rachel Shalom (Tel Aviv, IL); Roberto Stelling (Rio de Janeiro, BR); Aviv Zeilig (Rosh Haain, IL)
Assignee: Dell Products L.P.
H04L41/147H04L41/16H04L43/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,381,787
App. No.
18/430,109
Granted
Aug 5, 2025
Kind
B1
Abstract

A system, method, and computer-readable medium for performing a data center management and monitoring operation to predict event failures, comprising taking time stamped telemetry metrics of data asset clusters; sampling the time stamped telemetry metrics; separating the sampled time stamped telemetry metrics into groups based on metric functionality; providing data of the groups into a residual neural network to learn patterns of the three groups; and concatenating output of the residual neural network of the groups, wherein concatenated output is feed back to the residual neural network to learn relationships of the three groups.

Claims (39)

1. A computer-implementable method for performing a data center management and monitoring operation to predict event failures, comprising:

taking time stamped telemetry metrics of data asset clusters;

sampling the time stamped telemetry metrics;

separating the sampled time stamped telemetry metrics into groups based on metric functionality;

providing data of the groups into a residual neural network to learn patterns of three groups; and

concatenating output of the residual neural network of the groups, wherein the concatenated output is feed back to the residual neural network to learn relationships of the three groups.

2. The method of claim 1 , further comprising handling based on service level agreements (SLA) factors in predicting event failures.

3. The method of claim 2 wherein an SLA policy is implemented in predicting event failures if handling based on SLA factors.

4. The method of claim 1 , further comprising preprocessing the sampled time stamped telemetry metrics before the providing data of the groups into the residual neural network.

5. The method of claim 1 wherein holes in data are designated with a value of minus one (−1).

6. The method of claim 1 , wherein the three groups are metrics as to nodes on a cluster, metrics as to capacity on clusters, and metrics as to performance of clusters.

7. The method of claim 1 , wherein metrics of the three groups are sliced per periodic windows and labeled as healthy or not healthy.

8. A system comprising:

a processor;

a data bus coupled to the processor;

a data center asset client module; and,

a non-transitory, computer-readable storage medium embodying computer program code, the non-transitory, computer-readable storage medium being coupled to the data bus, the computer program code interacting with a plurality of computer operations for performing a data center management and monitoring operation to predict event failures and comprising instructions executable by the processor and configured for:

taking time stamped telemetry metrics of data asset clusters;

sampling the time stamped telemetry metrics;

separating the sampled time stamped telemetry metrics into groups based on metric functionality;

providing data of the groups into a residual neural network to learn patterns of three groups; and

concatenating output of the residual neural network of the groups, wherein the concatenated output is feed back to the residual neural network to learn relationships of the three groups.

9. The system of claim 8 , further comprising handling based on service level agreements (SLA) factors in predicting event failures.

10. The system of claim 9 wherein an SLA policy is implemented in predicting event failures if handling based on SLA factors.

11. The system of claim 8 , further comprising preprocessing the sampled time stamped telemetry metrics before the providing data of the groups into the residual neural network.

12. The system of claim 11 wherein holes in data are designated with a value of minus one (−1).

13. The system of claim 8 wherein the three groups are metrics as to nodes on a cluster, metrics as to capacity on clusters, and metrics as to performance of clusters.

14. The system of claim 8 wherein metrics of the three groups are sliced per periodic windows and labeled as healthy or not healthy.

15. A non-transitory, computer-readable storage medium embodying computer program code for performing a data center management and monitoring operation to predict event failures, the computer program code comprising computer executable instructions configured for:

taking time stamped telemetry metrics of data asset clusters;

sampling the time stamped telemetry metrics;

separating the sampled time stamped telemetry metrics into groups based on metric functionality;

providing data of the groups into a residual neural network to learn patterns of three groups; and

concatenating output of the residual neural network of the groups, wherein the concatenated output is feed back to the residual neural network to learn relationships of the three groups.

16. The non-transitory, computer-readable storage medium of claim 15 further comprising handling based on service level agreements (SLA) factors in predicting event failures.

17. The non-transitory, computer-readable storage medium of claim 16 wherein an SLA policy is implemented in predicting event failures if handling based on SLA factors.

18. The non-transitory, computer-readable storage medium of claim 15 , further comprising preprocessing the sampled time stamped telemetry metrics before the providing data of the groups into the residual neural network.

19. The non-transitory, computer-readable storage medium of claim 15 wherein holes in data are designated with a value of minus one (−1).

20. The non-transitory, computer-readable storage medium of claim 15 wherein the three groups are metrics as to nodes on a cluster, metrics as to capacity on clusters, and metrics as to performance of clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: JREIJ, ELIE A.; SAVIR, AMIHAI; TEIXEIRA DE ABREU PINHO, ROMULO; ZIESMANN, PETER; SHALOM, RACHEL; STELLING, ROBERTO; ZEILIG, AVIV
To: DELL PRODUCTS L.P.,
Reel/Frame 066329/0738 →
References Cited (7)
US 10776721B1 · Shi · 2020 [cited by examiner]
US 12164503B1 · Khan · 2024 [cited by examiner]
US 20200387565A1 · Caglar · 2020 [cited by examiner]
US 20200387797A1 · Ryan · 2020 [cited by examiner]
US 20230229675A1 · Hautyunyan · 2023 [cited by examiner]
US 20240005777A1 · Doshi · 2024 [cited by examiner]
US 20240193165A1 · Spannhake, II · 2024 [cited by examiner]