IP Library › Granted Patent US 11,429,450
Granted Patent B2
US 11,429,450 · App. 16/394,063 · Granted Aug 30, 2022

Aggregated virtualized compute accelerators for assignment of compute kernels

Inventor: Matthew D. McClure (Alameda, CA)
Assignee: VMWARE, INC.
G06F9/505G06F8/75G06F9/5077
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,450
App. No.
16/394,063
Granted
Aug 30, 2022
Kind
B2
Abstract

Disclosed are various embodiments for assigning compute kernels to compute accelerators that form an aggregated virtualized compute accelerator. A directed, acyclic graph (DAG) representing a workload assigned to a virtualized compute accelerator is generated. The workload can include a plurality of compute kernels and the DAG comprising a plurality of nodes and a plurality of edges, each of the nodes representing a respective compute kernel, each edge representing a dependency between a respective pair of the compute kernels, and the virtualized compute accelerator representing a logical interface for a plurality of compute accelerators. The DAG can be analyzed to identify sets of dependent compute kernels, each set of dependent compute kernels being independent of the other sets of dependent compute kernels and execution of at least one compute kernel in a set of dependent compute kernels depending on a previous execution of another computer kernel in the set of dependent compute kernels. Then, each set of dependent compute kernels can be assigned to a respective one of the plurality of compute accelerators.

Claims (76)

1. A system, comprising:

a computing device comprising a processor and a memory; and

machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least:

generate a directed acyclic graph (DAG) representing a workload assigned to a virtualized compute accelerator that is executed by the processor of the computing device, wherein:

the workload comprises a plurality of compute kernels and the DAG comprising a plurality of nodes and a plurality of edges,

each of the nodes represents a respective compute kernel,

each of the edges represents a dependency between a respective pair of the compute kernels, and

the virtualized compute accelerator represents a logical interface for a plurality of compute accelerators;

analyze, by a management layer of the virtualized compute accelerator, the DAG to identify sets of dependent compute kernels, a respective set of dependent compute kernels being independent of other sets of dependent compute kernels and execution of at least one compute kernel in the respective set of dependent compute kernels depending on a previous execution of another compute kernel in the respective set of dependent compute kernels;

determine, by the management layer, a set-specific computing resource profile for a particular set of dependent compute kernels by an analysis of resource requirements specified by a set of nodes corresponding to the particular set of dependent compute kernels, the set-specific computing resource profile comprising a total memory and a total processing power determined for the particular set of dependent compute kernels as a whole;

identify, by a static analysis engine executed by the processor of the computing device, a host device that captures image data used for an image processing operation of at least one compute kernel of the particular set of dependent compute kernels;

assign, by the management layer, the particular set of dependent compute kernels to a particular compute accelerator of the plurality of compute accelerators that is identified to be installed to the host device that captures the image data, and is identified to include sufficient available resources based at least in part on the total memory and the total processing power for the particular set of dependent compute kernels for execution; and

execute the particular set of dependent compute kernels using the particular compute accelerator.

2. The system of claim 1 , wherein the machine-readable instructions that cause the computing device to generate the DAG representing the workload further cause the computing device to at least:

perform static analysis on an object code or a source code representation of the workload to identify the plurality of compute kernels; and

perform static analysis on the object code or the source code representation of the workload to identify dependencies between pairs of the plurality of compute kernels.

3. The system of claim 1 , wherein the machine-readable instructions that cause the computing device to assign the respective set of dependent compute kernels to a respective one of the compute accelerators further cause the computing device to at least:

determine that the respective one of the compute accelerators complies with a predefined criterion;

select the respective one of the compute accelerators from the plurality of compute accelerators based on a determination that the respective one of the compute accelerators complies with the predefined criterion; and

send the respective set of dependent compute kernels to the respective one of the compute accelerators.

4. The system of claim 3 , wherein the machine-readable instructions further cause the computing device to encrypt the respective set of dependent compute kernels sent to the respective one of the compute accelerators.

5. The system of claim 3 , wherein the predefined criterion comprises the respective one of the compute accelerators being configured to use a remote direct memory access (RDMA) protocol to access a single copy of a working set.

6. The system of claim 1 , wherein the machine-readable instructions that cause the computing device to assign the respective set of dependent compute kernels to a respective one of the compute accelerators further cause the computing device to at least:

determine that a set of the dependent compute kernels is performing a predefined computation;

select the respective one of the compute accelerators from the plurality of compute accelerators based on a determination that the set of the dependent compute kernels is performing the predefined computation; and

send the set of the dependent compute kernels to the respective one of the compute accelerators.

7. The system of claim 6 , wherein the predefined computation involves a modification to a predefined resource.

8. A method, comprising:

generating, by a computing device, a directed acyclic graph (DAG) representing a workload assigned to a virtualized compute accelerator that is executed by a processor of the computing device, wherein:

the workload comprises a plurality of compute kernels and the DAG comprising a plurality of nodes and a plurality of edges,

each of the nodes represents a respective compute kernel,

each of the edges represents a dependency between a respective pair of the compute kernels, and

the virtualized compute accelerator represents a logical interface for a plurality of compute accelerators;

analyzing, by a management layer of the virtualized compute accelerator, the DAG to identify sets of dependent compute kernels, a respective set of dependent compute kernels being independent of other sets of dependent compute kernels and execution of at least one compute kernel in the respective set of dependent compute kernels depending on a previous execution of another compute kernel in the respective set of dependent compute kernels;

determining, by the management layer, a set-specific computing resource profile for a particular set of dependent compute kernels by an analysis of resource requirements specified by a set of nodes corresponding to the particular set of dependent compute kernels, the set-specific computing resource profile comprising a total memory and a total processing power determined for the particular set of dependent compute kernels as a whole;

identifying, by a static analysis engine executed by the processor of the computing device, a host device that originates image data used for an image processing operation of at least one compute kernel of the particular set of dependent compute kernels;

assigning, by the management layer, the particular set of dependent compute kernels to a particular compute accelerator of the plurality of compute accelerators that is identified to be installed to the host device that originates the image data used for the image processing operation, and is identified to include available resources based at least in part on the total memory and the total processing power for the particular set of dependent compute kernels for execution; and

executing the particular set of dependent compute kernels using the particular compute accelerator.

9. The method of claim 8 , wherein generating the DAG representing the workload further comprises:

performing static analysis on an object code or a source code representation of the workload to identify the plurality of compute kernels; and

performing static analysis on the object code or the source code representation of the workload to identify dependencies between pairs of the plurality of compute kernels.

10. The method of claim 8 , wherein assigning the respective set of dependent compute kernels to a respective one of the compute accelerators further comprises:

determining that the respective one of the compute accelerators complies with a predefined criterion;

selecting the respective one of the compute accelerators from the plurality of compute accelerators based on a determination that the respective one of the compute accelerators complies with the predefined criterion; and

sending the respective set of dependent compute kernels to the respective one of the compute accelerators.

11. The method of claim 10 , further comprising encrypting the respective set of dependent compute kernels sent to the respective one of the compute accelerators.

12. The method of claim 8 , wherein assigning the respective set of dependent compute kernels to a respective one of the compute accelerators further comprises:

determine that a set of the dependent compute kernels is performing a predefined computation;

select the respective one of the compute accelerators from the plurality of compute accelerators based on a determination that the set of the dependent compute kernels is performing the predefined computation; and

send the set of the dependent compute kernels to the respective one of the compute accelerators.

13. The method of claim 12 , wherein the predefined computation involves a modification to a predefined resource.

14. The method of claim 8 , wherein a respective one of the compute accelerators uses a remote direct memory access (RDMA) protocol to access a single copy of a working set.

15. A non-transitory, computer-readable medium comprising machine-readable instruction that, when executed by a processor, cause a computing device to at least:

generate a directed acyclic graph (DAG) representing a workload assigned to a virtualized compute accelerator that is executed by the processor of the computing device, wherein:

the workload comprises a plurality of compute kernels and the DAG comprising a plurality of nodes and a plurality of edges,

each of the nodes represents a respective compute kernel,

each of the edges represents a dependency between a respective pair of the compute kernels, and

the virtualized compute accelerator represents a logical interface for a plurality of compute accelerators;

analyze, by a management layer of the virtualized compute accelerator, the DAG to identify sets of dependent compute kernels, a respective set of dependent compute kernels being independent of other sets of dependent compute kernels and execution of at least one compute kernel in the respective set of dependent compute kernels depending on a previous execution of another compute kernel in the respective set of dependent compute kernels;

determine, by the management layer, a set-specific computing resource profile for a particular set of dependent compute kernels by an analysis of resource requirements specified by a set of nodes corresponding to the particular set of dependent compute kernels, the set-specific computing resource profile comprising a total memory and a total processing power determined for the particular set of dependent compute kernels as a whole;

identify, by a static analysis engine executed by the processor of the computing device, a host device that captures image data used for an image processing operation of at least one compute kernel of the particular set of dependent compute kernels;

assign, by the management layer, the particular set of dependent compute kernels to a particular compute accelerator of the plurality of compute accelerators that is identified to be installed to the host device that captures the image data, and is identified to include available resources based at least in part on the total memory and the total processing power for the particular set of dependent compute kernels for execution; and

execute the particular set of dependent compute kernels using the particular compute accelerator.

16. The non-transitory, computer-readable medium of claim 15 , wherein the machine-readable instructions that cause the computing device to generate the DAG representing the workload further cause the computing device to at least:

perform static analysis on an object code or a source code representation of the workload to identify the plurality of compute kernels; and

perform static analysis on the object code or the source code representation of the workload to identify dependencies between pairs of the plurality of compute kernels.

17. The non-transitory, computer-readable medium of claim 15 , wherein the machine-readable instructions that cause the computing device to assign the respective set of dependent compute kernels to a respective one of the compute accelerators further cause the computing device to at least:

determine that the respective one of the compute accelerators complies with a predefined criterion;

select the respective one of the compute accelerators from the plurality of compute accelerators based on a determination that the respective one of the compute accelerators complies with the predefined criterion; and

send the respective set of dependent compute kernels to the respective one of the compute accelerators.

18. The non-transitory, computer-readable medium of claim 17 , wherein the machine-readable instructions further cause the computing device to encrypt the respective set of dependent compute kernels sent to the respective one of the compute accelerators.

19. The non-transitory, computer-readable medium of claim 15 , wherein the machine-readable instructions that cause the computing device to assign the respective set of dependent compute kernels to a respective one of the compute accelerators further cause the computing device to at least:

determine that a set of the dependent compute kernels is performing a predefined computation;

select the respective one of the compute accelerators from the plurality of compute accelerators based on a determination that the set of the dependent compute kernels is performing the predefined computation; and

send the set of the dependent compute kernels to the respective one of the compute accelerators.

20. The non-transitory, computer-readable medium of claim 15 , wherein a respective one of the compute accelerators uses a remote direct memory access (RDMA) protocol to access a single copy of a working set.

Assignments (2)
CHANGE OF NAME Recorded Apr 15, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067102/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2019
From: MCCLURE, MATTHEW D.
To: VMWARE, INC.
Reel/Frame 049868/0578 →
Continuity (1)
Related Publication 20200341812A1 · Oct 29, 2020