IP Library Granted Patent US 11,249,835
Granted Patent B2
US 11,249,835 · App. 16/879,157 · Granted Feb 15, 2022

Automatic repair of computing devices in a data center

Inventors: Ian Ferreira (Issaquah, WA); Ganesh Balakrishnan (Sammamish, WA); Evan Adams (Bainbridge Island, WA); Carla Cortez (Playa Del Rey, CA); Eric Hullander (Edmonds, WA)
Assignee: Core Scientific, Inc.
G06F11/0793G06F11/0721
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,249,835
App. No.
16/879,157
Granted
Feb 15, 2022
Kind
B2
Abstract

A management device for managing a plurality of computing devices in a data center may comprise a network interface, a first module that periodically sends health status queries to the computing devices via the network interface, a second module configured to receive responses to the health status queries and collect and store health status data for the computing devices, a third module configured to create support tickets, and/or a fourth module configured to (i) create and periodically update a Cox proportional hazards (CPH) model based on the health status data; (ii) apply a deep neural network (DNN) to the input of the CPH model; (iii) determine a probability of failure for each computing device; (iv) compare each probability of failure with a threshold; and (v) cause the third module to generate a pre-failure support ticket for each computing device having determined probabilities of failure above the threshold.

Claims (67)

1. A management device for managing a plurality of computing devices in a data center, wherein the management device comprises:

a network interface for communicating with the plurality of computing devices,

a first module that periodically sends health status queries to each of the computing devices via the network interface,

a second module configured to receive responses to the health status queries and collect and store health status data for each of the computing devices,

a third module configured to create support tickets, and

a fourth module configured to:

(i) create and periodically update a Cox proportional hazards (CPH) model based on the collected health status data;

(ii) apply a deep neural network (DNN) to an input of the CPH model;

(iii) determine a probability of failure for each of the plurality of computing devices;

(iv) compare each determined probability of failure with a predetermined threshold; and

(v) cause the third module to generate a pre-failure support ticket for each of the plurality of computing devices having determined probabilities of failure that exceed the predetermined threshold.

2. The management device of claim 1 , wherein in response to not receiving an acceptable response to a particular health status query within a first predetermined time, the third module is configured to:

(i) send a first repair instruction,

(ii) wait at least enough time for the first repair instruction to complete,

(iii) send a second health status query, and

(iv) in response to not receiving an acceptable response to the second health status query within a second predetermined time:

(a) send a second repair instruction,

(b) wait at least enough time for the second repair instruction to complete,

(c) send a third health status query, and

(d) in response to not receiving an acceptable response to the third health status query within a third predetermined time, cause the third module to create a repair ticket.

3. The management device of claim 1 , wherein the health status data comprises hash rate.

4. The management device of claim 1 , wherein the health status data comprises computing device temperature.

5. The management device of claim 1 , wherein the second module is further configured to collect temperature and humidity data for the data center.

6. A method for managing a plurality of computing devices in a data center, the method comprising:

collecting health status data from each of the plurality of computing devices by periodically transmitting health status queries to each of the plurality of computing devices;

creating and periodically updating a Cox proportional hazards (CPH) model based on the collected health status data;

applying a deep neural network (DNN) to an input of the CPH model;

determining a probability of failure for each of the plurality of computing devices;

comparing each determined probability of failure with a predetermined threshold; and

generating a support ticket for each of the plurality of computing devices that have a determined probability of failure that exceed the predetermined threshold.

7. The method of claim 6 , wherein the collected health status data comprises telemetry data.

8. The method of claim 7 , wherein the collected health status data comprises temperature and humidity readings from within the data center.

9. The method of claim 6 , wherein the collected health status data comprises device fan speed.

10. The method of claim 6 , wherein the collected health status data comprises device hash rate.

11. The method of claim 6 , wherein the collected health status data comprises device temperature.

12. The method of claim 6 , wherein the predetermined threshold is 80%.

13. The method of claim 6 , further comprising:

in response to not receiving an acceptable response to a first health status query within a first predetermined time:

(i) issuing a first repair instruction,

(ii) waiting at least enough time for the first repair instruction to complete,

(iii) issuing a second health status query, and

(iv) in response to not receiving an acceptable response to the second health status query within a second predetermined time:

(a) issuing a second repair instruction,

(b) waiting at least enough time for the second repair instruction to complete,

(c) issuing a third health status query, and

(d) in response to not receiving an acceptable response to the third health status query within a third predetermined time, issuing a repair ticket.

14. A non-transitory, computer-readable storage medium storing instructions executable by a processor of a computational device, which when executed cause the computational device to:

collect health status data from each of a plurality of computing devices in a data center by periodically transmitting health status queries to each of the computing devices;

create and periodically update a Cox proportional hazards (CPH) model based on the collected health status data;

apply a deep neural network (DNN) to an input of the CPH model;

determine a probability of failure for each of the plurality of computing devices;

compare each determined probability of failure with a predetermined threshold; and

generate a predictive support ticket for each of the plurality of computing devices that have a determined probability of failure that exceed the predetermined threshold.

15. The storage medium of claim 14 , wherein in response to not receiving an acceptable response to a first health status query within a first predetermined time:

(i) issuing a first repair instruction,

(ii) waiting at least enough time for the first repair instruction to complete,

(iii) issuing a second health status query, and

(iv) in response to not receiving an acceptable response to the second health status query within a second predetermined time:

(a) issuing a second repair instruction,

(b) waiting at least enough time for the second repair instruction to complete,

(c) issuing a third health status query, and

(d) in response to not receiving an acceptable response to the third health status query within a third predetermined time, issuing a repair ticket.

16. The storage medium of claim 15 , wherein the first health status query, the second health status query, and the third health status query are queries for a hash rate, wherein the first repair instruction is to restart a mining application, and wherein the second repair instruction is a device restart instruction.

17. The storage medium of claim 14 , wherein the health status data comprises device fan speed.

18. The storage medium of claim 14 , wherein the health status data comprises device hash rate.

19. The storage medium of claim 14 , wherein the health status data comprises device temperature.

20. The storage medium of claim 19 , wherein the health status data further comprises temperature and humidity readings from within the data center.

Assignments (18)
RELEASE OF SECURITY INTEREST Recorded May 6, 2026
From: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
To: CORE SCIENTIFIC, INC.
Reel/Frame 075520/0577 →
SECURITY INTEREST Recorded Mar 10, 2026
From: CORE SCIENTIFIC, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 075099/0814 →
RELEASE OF SECURITY INTEREST Recorded Sep 13, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: CORE SCIENTIFIC, INC.
Reel/Frame 068968/0373 →
RELEASE OF SECURITY INTEREST Recorded Aug 28, 2024
From: B. RILEY COMMERCIAL CAPITAL, LLC
To: CORE SCIENTIFIC, INC.; CORE SCIENTIFIC OPERATING COMPANY
Reel/Frame 068803/0146 →
RELEASE OF SECURITY INTEREST Recorded Aug 20, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: CORE SCIENTIFIC, INC.
Reel/Frame 068693/0052 →
RELEASE OF SECURITY INTEREST Recorded Aug 20, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: CORE SCIENTIFIC, INC.
Reel/Frame 068719/0600 →
SECURITY INTEREST Recorded Jul 29, 2024
From: CORE SCIENTIFIC, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 068178/0585 →
SECURITY INTEREST Recorded Jul 29, 2024
From: CORE SCIENTIFIC, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 068178/0677 →
SECURITY INTEREST Recorded Jul 29, 2024
From: CORE SCIENTIFIC, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 068178/0700 →
MERGER Recorded Feb 6, 2024
From: CORE SCIENTIFIC OPERATING COMPANY
To: CORE SCIENTIFIC, INC.
Reel/Frame 066507/0675 →
RELEASE OF SECURITY INTEREST Recorded Jan 26, 2024
From: U.S. BANK NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: CORE SCIENTIFIC, INC.
Reel/Frame 066375/0365 →
RELEASE OF SECURITY INTEREST Recorded Jan 26, 2024
From: U.S. BANK NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: CORE SCIENTIFIC OPERATING COMPANY; CORE SCIENTIFIC ACQUIRED MINING LLC
Reel/Frame 066375/0324 →
SECURITY INTEREST Recorded Mar 1, 2023
From: CORE SCIENTIFIC, INC.; CORE SCIENTIFIC OPERATING COMPANY
To: B. RILEY COMMERCIAL CAPITAL, LLC
Reel/Frame 062899/0741 →
RELEASE OF SECURITY INTEREST Recorded Feb 3, 2023
From: WILMINGTON SAVINGS FUND SOCIETY, FSB
To: CORE SCIENTIFIC INC.; CORE SCIENTIFIC OPERATING COMPANY
Reel/Frame 063272/0450 →
SECURITY INTEREST Recorded Dec 23, 2022
From: CORE SCIENTIFIC OPERATING COMPANY; CORE SCIENTIFIC INC.
To: WILMINGTON SAVINGS FUND SOCIETY, FSB
Reel/Frame 062218/0713 →
SECURITY INTEREST Recorded Feb 10, 2022
From: CORE SCIENTIFIC OPERATING COMPANY; CORE SCIENTIFIC ACQUIRED MINING LLC
To: U.S. BANK NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 059004/0831 →
SECURITY INTEREST Recorded Apr 21, 2021
From: CORE SCIENTIFIC, INC.
To: U.S. BANK NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 055996/0839 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: FERREIRA, IAN; BALAKRISHNAN, GANESH; ADAMS, EVAN; CORTEZ, CARLA; HULLANDER, ERIC
To: CORE SCIENTIFIC, INC.
Reel/Frame 052720/0861 →