IP Library Granted Patent US 11,656,919
Granted Patent B2
US 11,656,919 · App. 17/734,041 · Granted May 23, 2023

Real-time simulation of compute accelerator workloads for distributed resource scheduling

Inventor: Matthew D. McClure (Alameda, CA)
Assignee: VMWARE, INC.
G06F9/5088G06F9/4856G06F2209/501
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,656,919
App. No.
17/734,041
Granted
May 23, 2023
Kind
B2
Abstract

Disclosed are various embodiments of real-time simulation of the performance of a compute accelerator workload for distributed resource scheduling. A compute kernel of a compute accelerator workload is augmented to include instructions that increment an execution counter at artificial halting points. Execution of the compute accelerator workload is suspended at an artificial halting point. The compute accelerator workload is executed on a plurality of candidate hosts and a performance counter is incremented during the execution of the compute accelerator workload on the various hosts. The compute accelerator workload is migrated to a destination host selected using an efficiency metric that is identified using the performance counter.

Claims (43)

1. A system, comprising:

at least one computing device comprising at least one processor and at least one memory; and

machine-readable instructions accessible to the at least one computing device, wherein the instructions, when executed by the at least one processor, cause the at least one computing device to at least:

suspend execution of a compute accelerator workload at a halting point indicated in a compute kernel of the compute accelerator workload;

clone the compute accelerator workload and a working set of the compute accelerator workload to a plurality of candidate hosts;

execute the compute accelerator workload on the plurality of candidate hosts, wherein at least one performance counter for a respective candidate host is incremented during the execution of the compute accelerator workload on the respective candidate host;

determine a plurality of efficiency metrics corresponding to plurality of candidate hosts based at least in part on the at least one performance counter for the respective candidate host, a respective efficiency metric comprising at least one of: a page reference velocity, a page dirty velocity, and an execution velocity; and

migrate the compute accelerator workload to a destination host selected from the plurality of candidate hosts based at least in part on an efficiency metric for the destination host.

2. The system of claim 1 , wherein the instructions, when executed by the at least one processor, further cause the at least one computing device to at least:

augment the compute kernel of the compute accelerator workload to include instructions to increment a performance counter for a memory load instruction or a memory store instruction.

3. The system of claim 1 , wherein the instructions, when executed by the at least one processor, further cause the at least one computing device to at least:

augment the compute kernel of the compute accelerator workload to include the halting point at a loop iteration or of the compute kernel; and

augment the compute kernel of the compute accelerator workload to increment an execution counter at the halting point.

4. The system of claim 1 , wherein the respective candidate host comprises a hardware compute accelerator.

5. The system of claim 1 , wherein the compute accelerator workload is executed on the respective candidate host from the halting point, and wherein the working set comprises intermediate data and a program counter that enables the compute accelerator workload to be resumed from the halting point.

6. The system of claim 1 , wherein the at least one performance counter is provided by a hardware performance counter.

7. The system of claim 1 , wherein the instructions, when executed by the at least one processor, cause the at least one computing device to at least:

query the performance counter in real time to determine the respective efficiency metric.

8. A non-transitory, computer-readable medium comprising machine-readable instructions that, when executed by at least one processor, cause at least one computing device to at least:

augment a compute kernel of a compute accelerator workload to be an augmented compute kernel comprising instructions that increment an execution counter at a plurality of artificial halting points;

suspend execution of the compute accelerator workload at an artificial halting point of the augmented compute kernel of the compute accelerator workload;

execute the compute accelerator workload on a plurality of candidate hosts, wherein at least one performance counter for a respective candidate host is incremented during the execution of the compute accelerator workload on the respective candidate host; and

migrate the compute accelerator workload to a destination host selected from the plurality of candidate hosts based at least in part on an efficiency metric that is identified using the at least one performance counter.

9. The non-transitory, computer-readable medium of claim 8 , wherein the instructions, when executed by the at least one processor, cause the at least one computing device to at least:

determine a plurality of efficiency metrics corresponding to the plurality of candidate hosts based at least in part on the at least one performance counter, wherein the efficiency metric corresponds to the destination host that is selected.

10. The non-transitory, computer-readable medium of claim 9 , wherein a respective efficiency metric comprises at least one of: a non-local page reference velocity, a non-local page dirty velocity, a device-local page reference velocity, a device-local page dirty velocity, and an execution velocity.

11. The non-transitory, computer-readable medium of claim 10 , wherein the instructions, when executed by the at least one processor, cause the at least one computing device to at least:

query the at least one performance counter in real time to determine the respective efficiency metric.

12. The non-transitory, computer-readable medium of claim 8 , wherein the at least one performance counter comprises a hardware performance counter.

13. The non-transitory, computer-readable medium of claim 8 , wherein the instructions, when executed by the at least one processor, cause the at least one computing device to at least:

insert, into a compute kernel of the compute accelerator workload, instructions to increment the at least one performance counter at a memory load instruction, a memory store instruction, or a halting point of the compute kernel.

14. The non-transitory, computer-readable medium of claim 8 , wherein the working set comprises at least one of: initialization data provided to the compute accelerator workload to begin execution, and intermediate data identified at a halting point of the compute accelerator workload.

15. A method, comprising:

augmenting a compute kernel of a compute accelerator workload to be an augmented compute kernel comprising instructions that increment an execution velocity counter at a plurality of artificial halting points;

suspending execution of the compute accelerator workload at an artificial halting point of the augmented compute kernel of the compute accelerator workload;

executing the compute accelerator workload on a plurality of candidate hosts, wherein at least one performance counter for a respective candidate host is incremented during the execution of the compute accelerator workload on the respective candidate host; and

migrating the compute accelerator workload to a destination host selected from the plurality of candidate hosts based at least in part on an efficiency metric that identified using the at least one performance counter.

16. The method of claim 15 , wherein the efficiency metric is based at least in part on the at least one performance counter and the execution velocity counter.

17. The method of claim 16 , wherein a scheduling service utilizes the efficiency metric as part of a load balancing algorithm.

18. The method of claim 17 , wherein the load balancing algorithm utilizes hardware resource utilization information and the efficiency metric to select the destination host.

19. The method of claim 15 , further comprising:

removing the compute accelerator workloads from a subset of the plurality of candidate hosts that excludes the destination host.

20. The method of claim 15 , wherein the at least one performance counter comprises at least one of: a non-local page reference velocity, a non-local page dirty velocity, a device-local page reference velocity, a device-local page dirty velocity, and the execution velocity counter.

Assignments (1)
CHANGE OF NAME Recorded Apr 15, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067102/0395 →
Continuity (2)
Continuation 16923137 · Jul 8, 2020
Related Publication 20220261296A1 · Aug 18, 2022