Federated distribution of computation and operations using networked processing units
Various approaches for deploying and controlling distributed compute operations with the use of infrastructure processing units (IPUs) and similar network-addressable processing units are disclosed. A device for orchestrating functions in a network compute mesh is configured to receive, at a network-addressable processing unit of a network-addressable processing unit mesh from a requestor device, a computation request to execute a workflow with a set of objectives; query at least one other network-addressable processing units of the network-addressable processing unit mesh using the set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow; transmit a list of recommended resources available to execute the workflow to the requestor device, the list of recommended resources being ranked based on at least one dimension of the resources; obtain a compute chain from the requestor device, the compute chain describing resource control transitions and data flow provided from the recommended resources and data in the network-addressable processing unit mesh; and schedule the execution of the workflow at one or more network-addressable processing units in the network-addressable processing unit mesh in accordance with the compute chain.
1 . A device for orchestrating functions in a network compute mesh, comprising:
a memory device configured to store instructions; and
a processor subsystem, which when configured by the instructions, is operable to:
receive, at a network-addressable processing unit of a network-addressable processing unit mesh from a requestor device, a computation request to execute a workflow with a set of objectives;
query at least one other network-addressable processing units of the network- addressable processing unit mesh using the set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow;
transmit a list of recommended resources available to execute the workflow to the requestor device, the list of recommended resources being ranked based on at least one dimension of the available resources;
obtain a compute chain from the requestor device, the compute chain describing resource control transitions and data flow provided from resources in the list of recommended resources and data in the network-addressable processing unit mesh; and
schedule the execution of the workflow at one or more network-addressable processing units in the network-addressable processing unit mesh in accordance with the compute chain.
2 . The device of claim 1 , wherein the set of objectives are expressed as service level objectives.
3 . The device of claim 1 , wherein the set of objectives are expressed as a multi-objective function.
4 . The device of claim 1 , wherein the set of objectives are default objectives.
5 . The device of claim 1 , wherein the aspects of available resources is provided by a second network-addressable processing unit of the network-addressable processing unit mesh in a resource map.
6 . The device of claim 1 , wherein the aspects of available resources include a percentage of compute, a number of cycles of compute, an amount of memory, an amount of storage, or network resources of a second network-addressable processing unit or a host managed by the network-addressable processing unit.
7 . The device of claim 1 , wherein the aspects of available data is provided by a second network-addressable processing unit of the network-addressable processing unit mesh in a data map.
8 . The device of claim 1 , wherein the aspects of available data include a location, a version, a type, or an amount or data.
9 . The device of claim 1 , wherein the processor subsystem is to:
receive a revised set of objectives from the requestor device;
query at least one other network-addressable processing units of the network-addressable processing unit mesh using the revised set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow; and
transmit revised recommended resources available to execute the workflow to the requestor device, the revised recommended resources including a revised ranked list of resources based on at least one dimension of the available resources.
10 . The device of claim 1 , wherein the list of recommended resources includes a top N of resources based on the at least one dimension of the available resources.
11 . The device of claim 1 , wherein to schedule the execution of the workflow across the network-addressable processing unit mesh in accordance with the compute chain, the processor subsystem is to:
transmit the compute chain to each network-addressable processing unit in the network-addressable processing unit mesh that is assigned to a resource used in the compute chain, wherein the respective network-addressable processing units associated with the respective resources used in the compute chain cooperatively coordinate resource scheduling and data movements to execute the compute chain.
12 . The device of claim 1 , wherein intermediate results of the execution of the compute chain are stored in a logging database.
13 . The device of claim 1 , wherein the execution of the compute chain produces a result, which is stored in a logging database.
14 . A method for orchestrating functions in a network compute mesh, comprising:
receiving, at a network-addressable processing unit of a network-addressable processing unit mesh from a requestor device, a computation request to execute a workflow with a set of objectives;
querying at least one other network-addressable processing units of the network-addressable processing unit mesh using the set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow;
transmitting a list of recommended resources available to execute the workflow to the requestor device, the list of recommended resources being ranked based on at least one dimension of the available resources;
obtaining a compute chain from the requestor device, the compute chain describing resource control transitions and data flow provided from resources in the list of recommended resources and data in the network-addressable processing unit mesh; and
scheduling the execution of the workflow at one or more network-addressable processing units in the network-addressable processing unit mesh in accordance with the compute chain.
15 . The method of claim 14 , wherein the set of objectives are expressed in a service level agreement.
16 . The method of claim 14 , comprising:
receiving a revised set of objectives from the requestor device;
querying at least one other network-addressable processing units of the network-addressable processing unit mesh using the revised set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow; and
transmitting revised recommended resources available to execute the workflow to the requestor device, the revised recommended resources including a revised ranked list of resources based on at least one dimension of the available resources.
17 . The method of claim 14 , wherein the list of recommended resources includes a top N of resources based on the at least one dimension of the available resources.
18 . The method of claim 14 , wherein scheduling the execution of the workflow across the network-addressable processing unit mesh in accordance with the compute chain comprises:
transmitting the compute chain to each network-addressable processing unit in the network-addressable processing unit mesh that is assigned to a resource used in the compute chain, wherein the respective network-addressable processing units associated with the respective resources used in the compute chain cooperatively coordinate resource scheduling and data movements to execute the compute chain.
19 . The method of claim 14 , wherein intermediate results of the execution of the compute chain are stored in a logging database.
20 . The method of claim 14 , wherein the execution of the compute chain produces a result, which is stored in a logging database.
21 . At least one non-transitory machine-readable medium including instructions for orchestrating functions in a network compute mesh, which when executed by a machine, cause the machine to:
receive, at a network-addressable processing unit of a network-addressable processing unit mesh from a requestor device, a computation request to execute a workflow with a set of objectives;
query at least one other network-addressable processing units of the network-addressable processing unit mesh using the set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow;
transmit a list of recommended resources available to execute the workflow to the requestor device, the list of recommended resources being ranked based on at least one dimension of the available resources;
obtain a compute chain from the requestor device, the compute chain describing resource control transitions and data flow provided from resources in the list of recommended resources and data in the network-addressable processing unit mesh; and
schedule the execution of the workflow at one or more network-addressable processing units in the network-addressable processing unit mesh in accordance with the compute chain.
22 . The at least one non-transitory machine-readable medium of claim 21 , comprising instructions to:
receive a revised set of objectives from the requestor device;
query at least one other network-addressable processing units of the network-addressable processing unit mesh using the revised set of objectives, to determine aspects of available resources and data in the network-addressable processing unit mesh to apply to the workflow; and
transmit revised recommended resources available to execute the workflow to the requestor device, the revised recommended resources including a revised ranked list of resources based on at least one dimension of the available resources.
23 . The at least one non-transitory machine-readable medium of claim 21 , wherein the list of recommended resources includes a top N of resources based on the at least one dimension of the available resources.
24 . The at least one non-transitory machine-readable medium of claim 21 , wherein the instructions to schedule the execution of the workflow across the network-addressable processing unit mesh in accordance with the compute chain comprise instructions to:
transmit the compute chain to each network-addressable processing unit in the network-addressable processing unit mesh that is assigned to a resource used in the compute chain, wherein the respective network-addressable processing units associated with the respective resources used in the compute chain cooperatively coordinate resource scheduling and data movements to execute the compute chain.
25 . The at least one non-transitory machine-readable medium of claim 21 , wherein intermediate results of the execution of the compute chain are stored in a logging database.