Self-contained artificial intelligence (AI) exectuable to automate model hosting, deployment, and AI life cycle
An artificial intelligence (AI) system generates a self-contained AI container packaged as an executable that can be installed on client devices or computing infrastructure. The AI container envelopes a set of components capable of performing one or more processes related to machine learning operations, including data collection, data processing, feature engineering, data labelling, model design, model training and performance optimization for underlying hardware, model deployment, and model feedback monitoring, and the like in a local environment of a client device or cloud environment of a computing infrastructure. The AI system packages the self-contained AI container as an executable that can be installed on client devices of any hardware configuration or operating system.
1 . A method, comprising:
generating an artificial intelligence (AI) executable, the AI executable when installed on a computing resource configured to generate a self-contained environment on the computing resource including a set of components, the set of components including at least one or more machine-learning (ML) models for performing a set of inference tasks, the set of components, when installed on the computing resource, further performing:
receiving a request from an application on the computing resource to perform one or more inference tasks,
generating predictions using the one or more ML models and provide the predictions as a response to the request to the application,
receiving feedback information from the application, the feedback information indicative of feedback from a user of the application on the generated predictions by a ML model for an inference task, and wherein the feedback information is generated by user interaction with the application,
normalizing, for each of one or more feedback instances for the inference task, the feedback instance to indicate a difference between a generated prediction for the feedback instance and a desired prediction of the user,
obtaining preferences of the user for the inference task, applying a reinforcement learning (RL) model to update parameters of the ML model to reflect the obtained preferences of the user,
performing machine learning operations for updating the one or more ML models using at least data generated from the client device,
storing, in data storage, information related to the request and the generated predictions for the request as a data instance,
receiving a second request from the application to perform the one or more inference tasks,
generating second predictions and provide the second predictions as a response to the second request to the application, and
determining, responsive to determining that a similarity metric between the second request and the data instance is above a threshold, not to store data related to the second request in the data storage; and
providing the AI executable for download to the computing resource.
2 . The method of claim 1 , wherein the computing resource is a client device including a laptop computer, a mobile device, a desktop computer, and wherein the application is installed on a local environment of the client device.
3 . The method of claim 1 , wherein the computing resource is a computing infrastructure including a public cloud infrastructure or an on-premise infrastructure, and wherein the application is deployed on the computing infrastructure.
4 . The method of claim 1 , wherein at least one ML model is configured to process data of two or more data modalities.
5 . A non-transitory computer readable medium comprising stored instructions, the stored instructions when executed by at least one processor of one or more computing devices, cause the one or more computing devices to:
generate an artificial intelligence (AI) executable, the AI executable when installed on a computing resource configured to generate a self-contained environment on the computing resource including a set of components, the set of components including at least one or more machine-learning (ML) models for performing a set of inference tasks, the set of components, when installed on the computing resource, further configured to:
receive a request from an application on the computing resource to perform one or more inference tasks,
generate predictions using the one or more ML models and provide the predictions as a response to the request to the application,
receive feedback information from the application, the feedback formation indicative of feedback from a user of the application on the generated predictions by a ML model for an inference task, and wherein the feedback information is generated by user interaction with the application,
for each of one or more feedback instances for the inference task, normalize the feedback instance to indicate a difference between a generated prediction for the feedback instance and a desired prediction of the user,
obtain preferences of the user for the inference task,
apply a reinforcement learning (RL) model to update parameters of the ML model to reflect the obtained preferences of the user,
perform machine learning operations for updating the one or more ML models using at least data generated from the client device,
store, in data storage, information related to the request and the generated predictions for the request as a data instance,
receive a second request from the application to perform the one or more inference tasks,
generate second predictions and provide the second predictions as a response to the second request to the application, and
responsive to determining that a similarity metric between the second request and the data instance is above a threshold, determining not to store data related to the second request in the data storage; and
provide the AI executable for download to the computing resource.
6 . The non-transitory computer readable medium of claim 5 , wherein the computing resource is a client device including a laptop computer, a mobile device, a desktop computer, and wherein the application is installed on a local environment of the client device.
7 . The non-transitory computer readable medium of claim 5 , wherein the computing resource is a computing infrastructure including a public cloud infrastructure or an on-premise infrastructure, and wherein the application is deployed on the computing infrastructure.
8 . The non-transitory computer readable medium of claim 5 , wherein at least one ML model is configured to process data of two or more data modalities.
9 . A computer system comprising:
one or more computer processors; and
one or more computer readable mediums storing instructions that, when executed by the one or more computer processors, cause the computer system to:
generate an artificial intelligence (AI) executable, the AI executable when installed on a computing resource configured to generate a self-contained environment on the computing resource including a set of components, the set of components including at least one or more machine-learning (ML) models for performing a set of inference tasks, the set of components, when installed on the computing resource, further configured to:
receive a request from an application on the computing resource to perform one or more inference tasks,
generate predictions using the one or more ML models and provide the predictions as a response to the request to the application,
receive feedback information from the application, the feedback information indicative of feedback from a user of the application on the generated predictions by a ML model for an inference task, and wherein the feedback information is generated by user interaction with the application,
for each of one or more feedback instances for the inference task, normalize the feedback instance to indicate a difference between a generated prediction for the feedback instance and a desired prediction of the user,
obtain preferences of the user for the inference task,
apply a reinforcement learning (RL) model to update parameters of the ML model to reflect the obtained preferences of the user,
perform machine learning operations for updating the one or more ML models using at least data generated from the client device,
store, in data storage, information related to the request and the generated predictions for the request as a data instance,
receive a second request from the application to perform the one or more inference tasks,
generate second predictions and provide the second predictions as a response to the second request to the application, and
responsive to determining that a similarity metric between the second request and the data instance is above a threshold, determining not to store data related to the second request in the data storage; and
provide the AI executable for download to the computing resource.
10 . The computer system of claim 9 , wherein the computing resource is a client device including a laptop computer, a mobile device, a desktop computer, and wherein the application is installed on a local environment of the client device.
11 . The computer system of claim 9 , wherein the computing resource is a computing infrastructure including a public cloud infrastructure or an on-premise infrastructure, and wherein the application is deployed on the computing infrastructure.
12 . The computer system of claim 9 , wherein at least one ML model is configured to process data of two or more data modalities.