IP Library Granted Patent US 11,775,344
Granted Patent B1
US 11,775,344 · App. 18/036,864 · Granted Oct 3, 2023

Training task queuing cause analysis method and system, device and medium

Inventor: Wenxiao Wang (Shandong, CN)
Assignee: INSPUR SUZHOU INTELLIGENT TECHNOLOGY CO., LTD.
G06F9/4881G06F9/3891
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,344
App. No.
18/036,864
Granted
Oct 3, 2023
Kind
B1
Abstract

The present application discloses a system, readable storage medium, and method for analyzing a cause of training-task queuing, wherein the method includes the steps of: obtaining a required resource that is required by a training task inputted by a user and a remaining resource of a cluster; in responding to that the remaining resource does not satisfy the required resource, obtaining a plurality of cluster center data that are pre-generated in a cluster model; regarding the required resource and the remaining resource as sample data, and calculating distances between the sample data and each of the cluster center data; and feeding back a cause corresponding to the cluster center datum that has a minimum distance with the sample data.

Claims (22)

1. A method for analyzing a cause of training-task queuing, comprising:

obtaining a required resource that is required by a training task inputted by a user and a remaining resource of a cluster;

in responding to that the remaining resource does not satisfy the required resource, obtaining a plurality of cluster center data that are pre-generated in a cluster model;

regarding the required resource and the remaining resource as sample data, and

calculating distances between the sample data and each of the cluster center data;

feeding back a cause corresponding to the cluster center datum that has a minimum distance with the sample data;

saving the sample data;

in responding to that a quantity of the saved sample data reaches a threshold, by using the saved sample data, updating the cluster model,

wherein the step of, by using the saved sample data, updating the cluster model comprises:

randomly generating the plurality of cluster center data;

calculating distances between the saved sample data and each of the current cluster center data and distances between original sample data in the cluster model and each of the current cluster center data, to divide the saved sample data and the original sample data in the cluster model into the corresponding cluster center data;

by using the sample data in each of the cluster center data, recalculating the corresponding cluster center data; and

in responding to that the cluster center data obtained by the calculation are different from the cluster center data that are currently used for the sample-data division, performing division of the sample data again by using the cluster center data obtained by the calculation to perform iterative training, till the cluster center data obtained by the calculation and the cluster center data that are currently used for the sample-data division are the same.

2. The method according to claim 1 , wherein the step of regarding the required resource and the remaining resource as the sample data comprises:

performing quantization processing to the sample data.

3. A computer device, comprising: at least one processor; and a memory, the memory storing a computer program that is executable in the processor, wherein the processor, when executing the computer program, implements the steps of the method according to claim 1 .

4. A non-transitory computer-readable storage medium, the computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according claim 1 .

5. The method according to claim 1 , wherein in responding to that the remaining resource does not satisfy the required resource, obtaining the plurality of cluster center data that are pre-generated in the cluster model comprises:

starting up a real-time monitoring thread to monitor the training-task queuing.

6. The method according to claim 1 , wherein calculating distances between the sample data and each of the cluster center data comprises:

calculating Euclidean distances between the sample data and each of the cluster center data.

7. The method according to claim 1 , wherein the sample data comprises configuration parameters related to the training task, and the configuration parameters comprises: a quantity of used CPUs, a quantity of used GPUs, resource groups, GPU types and specified scheduling nodes.

Assignments (2)
LICENSE Recorded Jun 30, 2026
From: IEIT SYSTEMS CO., LTD
To: AIVRES SYSTEMS INC.
Reel/Frame 075857/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: WANG, WENXIAO
To: INSPUR SUZHOU INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 063631/0512 →
Priority Claims (1)
CN 202011402706.2 · Dec 4, 2020 · national