Deploying instances of artificial intelligence (AI) models having different levels of computational complexity in a heterogeneous computing platform
Systems and methods for deploying instances of Artificial Intelligence (AI) models having different levels of complexity in a heterogenous computing platform are described. In an embodiment, an Information Handling System (IHS), may include a heterogeneous computing platform comprising a plurality of devices and a memory having a plurality of sets of firmware instructions that, upon execution by a respective device, enable the respective device to provide a corresponding service, and where at least one of the devices operates as an orchestrator configured to: select an instance of an AI model among a plurality of instances of the AI model based, at least in part, upon context or telemetry data received from at least a subset of the plurality of devices, where each of the plurality of instances of the AI model has a different level of complexity, and instruct a device to execute the selected instance of the AI model.
1 . An Information Handling System (IHS), comprising:
a heterogeneous computing platform comprising a plurality of devices; and
a memory coupled to the heterogeneous computing platform, wherein the memory comprises a plurality of sets of firmware instructions, wherein each of the sets of firmware instructions, upon execution by a respective device among the plurality of devices, enables the respective device to provide a respective firmware service, and wherein at least one of the plurality of devices operates as an orchestrator configured to:
select an instance of an Artificial Intelligence (AI) model among a plurality of instances of the AI model based, at least in part, upon context data received from at least a subset of the plurality of devices, wherein the context data comprises a metric indicative of at least one of: a user's presence, an IHS posture, or an application in execution by the IHS, wherein each of the plurality of instances of the AI model has a different level of complexity, and wherein, to receive the context data, the orchestrator is configured to send one or more messages to one or more firmware services executed by the subset of the plurality of devices via one or more Application Programming Interfaces (APIs) without any involvement by any host Operating System (OS) to collect the context data and provide the collected context data in a format wherein each individual piece or set of the context data includes a common clock time stamp; and
instruct a device among the plurality of devices to execute the selected instance of the AI model.
2 . The IHS of claim 1 , wherein the heterogeneous computing platform comprises: a System-On-Chip (SoC), a Field-Programmable Gate Array (FPGA), or an Application-Specific Integrated Circuit (ASIC).
3 . The IHS of claim 1 , wherein the orchestrator comprises at least one of: a sensing hub, an Embedded Controller (EC), or a Baseboard Management Controller (BMC).
4 . The IHS of claim 1 , wherein the orchestrator is further configured to select the subset of the plurality of devices from which to receive the context data based upon a policy, wherein the policy is configured to identify at least one of: devices to collect the context data from, and one or more context data collection parameters comprising at least one of a collection frequency or sampling rate, collection start and end times, a duration of the collection, or a maximum amount of context data to be collected, and wherein the subset of the plurality of devices is dynamically chosen by the orchestrator based upon previously collected context data.
5 . The IHS of claim 1 , wherein the selected instance of the AI model has a degree of quantization different than at least one other instance of the AI model among the plurality of instances of the AI model.
6 . The IHS of claim 1 , wherein the selected instance of the AI model has a degree of pruning different than at least one other instance of the AI model among the plurality of instances of the AI model.
7 . The IHS of claim 1 , wherein the selected instance of the AI model has a degree of weight sharing different than at least one other instance of the AI model among the plurality of instances of the AI model.
8 . The IHS of claim 1 , wherein the orchestrator is further configured to receive a policy from an Information Technology Decision Maker (ITDM) or Original Equipment Manufacturer (OEM), wherein the policy comprises at least one of commands, program instructions, routines, or rules that conform to the one or more APIs, and wherein, based upon the policy, the orchestrator is configured to enable, disable, or modify at least one firmware service provided by at least one other device among the plurality of devices without any involvement by the host OS.
9 . The IHS of claim 8 , wherein the policy is configured to identify at least one of: the AI model, the plurality of instance of the AI model, the context data, the subset of the plurality of devices, the level of complexity, or a target performance metric, and wherein the policy is further configured to identify at least one of: context data collection routines, scripts, or algorithms to process and/or or produce the context data, context data collection parameters, or whether each individual piece or set of the context data includes the common clock time stamp.
10 . The IHS of claim 9 , wherein the policy comprises one or more rules, wherein each rule associates: (a) a difference between a current performance metric, calculated based upon the context data, and the target performance metric, with (b) the selected instance of the AI model, and wherein the instance of the AI model is selected based upon the rule.
11 . The IHS of claim 10 , wherein the orchestrator is configured to select another instance of the plurality of instances of the AI model based, at least in part, upon a change in the context data.
12 . The IHS of claim 11 , wherein the other instance of the AI model has a lower level of complexity than the selected instance of the AI model in response to the difference between the current performance metric and the target performance metric being greater than or equal to a threshold value.
13 . The IHS of claim 11 , wherein the other instance of the AI model has a higher level of complexity than the selected instance of the AI model in response to a determination the difference between the current performance metric and the target performance metric is smaller than or equal to a threshold value.
14 . The IHS of claim 11 , wherein the device comprises a Central Processing Unit (CPU), a Graphical Processing Unit (GPU), a Video Processing Unit (VPU), an Image Signal Processor (ISP), a Neural Processing Unit (NPU), a Tensor Processing Unit (TSU), a Neural Network Processor (NNP), or an Intelligence Processing Unit (IPU).
15 . The IHS of claim 14 , wherein to instruct the selected device, the orchestrator is configured to send a message to one or more firmware services executed by the selected device via an Application Programming Interface (API) without any involvement by any host Operating System (OS) to execute the selected instance of the AI model.
16 . The IHS of claim 8 , wherein the orchestrator is further configured to instruct a selected one of the plurality of devices to execute the selected instance of the AI model according to the policy.
17 . A memory coupled to a heterogeneous computing platform of an Information Handling System (IHS), wherein the heterogeneous computing platform comprises a plurality of devices, wherein the memory is configured to receive a plurality of sets of firmware instructions, wherein each set of firmware instructions, upon execution by a respective device among the plurality of devices, enables the respective device to provide a corresponding respective firmware service without any involvement by any host Operating System (OS), and wherein at least one of the plurality of devices operates as an orchestrator configured to: instruct a device among the plurality of devices to execute an instance of an Artificial Intelligence (AI) model selected among a plurality of instances of the AI model based, at least in part, upon context or telemetry data, wherein the context or telemetry data comprises a metric indicative of at least one of: a user's presence, an IHS posture, or an application in execution by the IHS, wherein, to receive the context or telemetry data, the orchestrator is configured to send one or more messages to one or more firmware services executed by one or more of the plurality of devices via one or more Application Programming Interfaces (APIs) without any involvement by the host OS to collect the context or telemetry data and provide the collected context or telemetry data in a format wherein each individual piece or set of the context or telemetry data includes a common clock time stamp, and wherein the one or more messages include an instruction configured to identify at least one of: (i) one or more other devices among the plurality of devices to deliver the collected context or telemetry data to, (ii) one or more acceptable data format(s), (iii) one or more protocol(s), or (iv) a manner or frequency of data delivery; and in response to a change in the context or telemetry data, instruct another device among the plurality of devices to execute another instance of the AI model.
18 . A method, comprising:
selecting a policy; and
transmitting the policy to an Information Handling System (IHS) over a network, wherein the IHS comprises a heterogeneous computing platform having a plurality of devices, and wherein an orchestrator among the plurality of devices is configured to: send one or more messages to one or more firmware services executed by one or more of the plurality of devices via one or more Application Programming Interfaces (APIs) without any involvement by any host Operating System (OS) to collect context or telemetry data according to the policy and provide the collected context or telemetry data with a common clock time stamp for each individual piece or set of the context or telemetry data, and wherein the one or more messages include an instruction identifying at least one of: (i) one or more other devices among the plurality of devices to deliver the collected context or telemetry data to, (ii) one or more acceptable data format(s), (iii) one or more protocol(s), or (iv) a manner or frequency of data delivery;
instruct a device among the plurality of devices to execute an instance of an Artificial Intelligence (AI) model selected among a plurality of instances of the AI model based, at least in part, upon the policy; and
in response to a change in a performance level of the IHS, instruct another device among the plurality of devices to execute another instance of the AI model.