IP Library Granted Patent US 10,031,797
Granted Patent B2
US 10,031,797 · App. 15/054,948 · Granted Jul 24, 2018

Method and apparatus for predicting GPU malfunctions

Inventor: Fei Hui (Beijing, CN)
Assignee: Alibaba Group Holding Limited
G06F11/079G06F11/008G06F11/076G06F11/0721G06T1/20G06T2200/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,031,797
App. No.
15/054,948
Granted
Jul 24, 2018
Kind
B2
Abstract

A method of predicting GPU malfunctions includes installing a daemon program at a GPU node, the daemon program periodically collecting GPU status parameters corresponding to the GPU node at a pre-determined time period. The method also includes obtaining the GPU status parameters from the GPU node and comparing the obtained GPU status parameters with mean status fault parameters to determine whether the GPU is to malfunction, where the mean status fault parameters are obtained by use of a pre-configured statistical model. Prior to a GPU enters a malfunction state, the GPU can be replaced, or the programs executing on the GPU can be migrated to other GPUs for execution, without affecting the normal business operations.

Claims (56)

1. A method of predicting GPU malfunctions in a cluster of GPUs, the method comprising:

collecting a plurality of measurements of a condition of a first GPU in the cluster of GPUs during a pre-determined time period;

determining a GPU count that represents how many of the measurements of the condition exceeded a threshold value during the pre-determined time period, and obtaining a mean fault count;

determining a GPU standard deviation based on the plurality of measurements of the condition collected during the pre-determined time and a plurality of measurements of the condition previously collected from the first GPU;

detecting when the GPU count is greater than the mean fault count, and the GPU standard deviation is less than a fault standard deviation threshold; and

migrating an application program executing on the first GPU to a second GPU in the cluster of GPUs in response to detecting when the GPU count exceeds the mean fault count and the GPU standard deviation falls below the fault standard deviation threshold.

2. The method of claim 1 , wherein:

the condition is temperature;

the plurality of measurements collected from the first GPU during the pre-determined time period include a plurality of currently-collected temperature measurements, and the plurality of measurements of the condition previously collected from the first GPU include a plurality of previously-collected temperature measurements;

the GPU count represents how many temperature measurements exceeded a temperature threshold; and

the GPU standard deviation is based on the plurality of currently-collected temperature measurements and the plurality of previously-collected temperature measurements.

3. The method of claim 2 , wherein the fault count is a mean count based on the plurality of currently-collected temperature measurements and the plurality of previously-collected temperature measurements.

4. The method of claim 3 , wherein the fault standard deviation threshold is an assigned value based on experience.

5. The method of claim 1 , wherein:

the condition is power consumption;

the plurality of measurements collected from the first GPU during the pre-determined time period include a plurality of currently-collected power consumption measurements, and the plurality of measurements of the condition previously collected from the first GPU include a plurality of previously-collected power consumption measurements;

the GPU count represents how many power consumption measurements exceeded a power consumption threshold; and

the GPU standard deviation is based on the plurality of currently-collected power consumption measurements, and the plurality of previously-collected power consumption measurements.

6. The method of claim 5 , wherein the fault count is a mean count based on the plurality of currently-collected power consumption measurements collected during the pre-determined time and the plurality of previously-collected power consumption measurements.

7. The method of claim 6 , wherein the fault standard deviation threshold is an assigned value based on experience.

8. A method of predicting GPU malfunctions in a cluster of GPUs, the method comprising:

collecting a GPU usage duration of a first GPU in the cluster of GPUs during a pre-determined time period;

detecting when the GPU usage duration is greater than a fault usage duration; and

migrating an application program executing on the first GPU to a second GPU in the cluster of GPUs in response to detecting when the GPU usage duration is greater than the fault usage duration.

9. The method of claim 8 , wherein the GPU usage duration is a mean duration based on the GPU usage duration collected during the pre-determined time and a plurality of GPU usage durations previously collected from the first GPU.

10. The method of claim 9 , wherein the plurality of GPU usage durations previously collected from the first GPU are stored in an information storage space as a plurality of stored measurements that correspond to the GPU.

11. An apparatus for predicting GPU malfunctions in a cluster of GPUs, the apparatus comprising:

a processor; and

a non-transitory computer-readable medium coupled to the processor, the non-transitory computer-readable medium having computer-readable instructions stored thereon to be executed when accessed by the processor, the instructions comprising:

collecting a plurality of measurements of a condition of a first GPU in the cluster of GPUs during a pre-determined time period;

determining a GPU count that represents how many of the measurements of the condition exceeded a threshold value during the pre-determined time period;

determining a GPU standard deviation based on the plurality of measurements of the condition collected during the pre-determined time and a plurality of measurements of the condition previously collected from the first GPU;

detecting when the GPU count is greater than a fault count, and the GPU standard deviation is less than a fault standard deviation threshold; and

migrating an application program executing on the first GPU to a second GPU in the cluster of GPUs in response to detecting when the GPU count exceeds the fault count and the GPU standard deviation falls below the fault standard deviation threshold.

12. The apparatus of claim 11 , wherein:

the condition is temperature;

the plurality of measurements collected from the first GPU during the pre-determined time period include a plurality of currently-collected temperature measurements, and the plurality of measurements of the condition previously collected from the first GPU include a plurality of previously-collected temperature measurements;

the GPU count represents how many temperature measurements exceeded a temperature threshold; and

the GPU standard deviation is based on the plurality of currently-collected temperature measurements and the plurality of previously-collected temperature measurements.

13. The apparatus of claim 12 , wherein the fault count is a mean count based on the plurality of currently-collected temperature measurements and the plurality of previously-collected temperature measurements.

14. The apparatus of claim 13 , wherein the fault standard deviation threshold is an assigned value based on experience.

15. The apparatus of claim 11 , wherein:

the condition is power consumption;

the plurality of measurements collected from the first GPU during the pre-determined time period include a plurality of currently-collected power consumption measurements, and the plurality of measurements of the condition previously collected from the first GPU include a plurality of previously-collected power consumption measurements;

the GPU count represents how many power consumption measurements exceeded a power consumption threshold; and

the GPU standard deviation is based on the plurality of currently-collected power consumption measurements, and the plurality of previously-collected power consumption measurements.

16. The apparatus of claim 15 , wherein the fault count is a mean count based on the plurality of currently-collected power consumption measurements collected during the pre-determined time and the plurality of previously-collected power consumption measurements.

17. The apparatus of claim 16 , wherein the fault standard deviation threshold is an assigned value based on experience.

18. An apparatus for predicting GPU malfunctions in a cluster of GPUs, the apparatus comprising:

a processor; and

a non-transitory computer-readable medium coupled to the processor, the non-transitory computer-readable medium having computer-readable instructions stored thereon to be executed when accessed by the processor, the instructions comprising:

collecting a GPU usage duration of a first GPU in the cluster of GPUs during a pre-determined time period;

detecting when the GPU usage duration is greater than a fault usage duration; and

migrating an application program executing on the first GPU to a second GPU in the cluster of GPUs in response to detecting when the GPU usage duration is greater than the fault usage duration.

19. The apparatus of claim 18 , wherein the GPU usage duration is a mean duration based on the GPU usage duration collected during the pre-determined time and a plurality of GPU usage durations previously collected from the first GPU.

20. The apparatus of claim 19 , wherein the plurality of GPU usage durations previously collected from the first GPU are stored in an information storage space as a plurality of stored measurements that correspond to the first GPU.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2026
From: ALIBABA GROUP HOLDING LIMITED
To: CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PRIVATE LIMITED
Reel/Frame 075478/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2016
From: HUI, FEI
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 037842/0070 →
Priority Claims (1)
CN 2015 1 0088768 · Feb 26, 2015 · national
Continuity (1)
Related Publication 20160253230A1 · Sep 1, 2016