Using container information to select containers for executing models
Using container information to select containers for executing models is described. A system receives a request from an application and identifies a version of a machine-learning model associated with the request. The system identifies a set of each serving container corresponding to the machine-learning model from a cluster of available serving containers associated with the version of the machine-learning model. The system selects a serving container from the set of each serving container corresponding to the machine-learning model. If the machine-learning model is not loaded in the serving container, the system loads the machine-learning model in the serving container. If the machine-learning model is loaded in the serving container, the system executes, in the serving container, the machine-learning model on behalf of the request. The system responds to the request based on executing the machine-learning model on behalf of the request.
1 . A system for using container information to select containers for executing models, the system comprising:
one or more processors; and
a non-transitory computer readable medium storing a plurality of instructions, which when executed, cause the one or more processors to:
watch, by a watcher associated with a routing container of a cluster of routing containers, for changes in available serving containers in a cluster of available serving containers associated with the routing container, wherein the changes include addition of one or more new serving containers to the cluster of available serving containers;
provide, by the watcher, information about the changes in the available serving containers to the routing container to update a mapping of the available serving containers based on the information;
identify a version of a machine-learning model associated with a request, in response to receiving the request from an application;
identify a set of available serving containers corresponding to the version of the machine-learning model from a cluster of available serving containers associated with the version of the machine-learning model, wherein the set of available serving containers is identified based, at least in part, on executing a hashing function applied to identifiers of the set of available serving containers, an identifier of any corresponding machine-learning model, and a replication factor for the version of the machine-learning model, wherein the replication factor identifies a number of serving containers of the set of available serving containers from the cluster of available serving containers for loading the version of the machine-learning model;
select an existing serving container and a new serving container from the set of available serving containers corresponding to the version of the machine-learning model;
load the version of the machine-learning model in the existing serving container, in response to a determination that the version of the machine-learning model is not loaded in the existing serving container;
execute, in the existing serving container, the machine-learning model on behalf of the request, in response to a determination that the machine-learning model is loaded in the existing serving container;
when the existing serving container is executing the version of the machine-learning model by another request, load another copy of the version of the machine-learning model into the new serving container and execute, in the new serving container, the version of the machine-learning model; and
respond to the request based on executing the version of the machine-learning model on behalf of the request.
2 . The system of claim 1 , wherein selecting from any cluster of available serving containers is based on updating a data structure comprising container information associated with serving containers in any corresponding cluster of available serving containers.
3 . The system of claim 1 , comprising further instructions, which when executed, cause the one or more processors to:
identify an other version of an other machine-learning model associated with the request;
identify an other set of available serving containers corresponding to the other machine-learning model from an other cluster of available serving containers associated with the other version of the other machine-learning model;
select an other serving container from the other set of available serving containers corresponding to the other machine-learning model;
load the other machine-learning model in the other serving container, in response to a determination that the other machine-learning model is not loaded in the other serving container; and
execute, in the other serving container, the other machine-learning model on behalf of the request, in response to a determination that the other machine-learning model is loaded in the other serving container;
wherein responding to the request is further based on executing the other machine-learning model on behalf of the request.
4 . The system of claim 1 , comprising further instructions, which when executed, cause the one or more processors to:
identify the version of the machine-learning model associated with an additional request, in response to receiving the additional request from the application;
identify the set of available serving containers corresponding to the machine-learning model from the cluster of available serving containers associated with the version of the machine-learning model;
select an additional serving container from the set of available serving containers corresponding to the machine-learning model;
load a copy of the machine-learning model in the additional serving container, in response to a determination that the copy of the machine-learning model is not loaded in the additional serving container;
execute, in the additional serving container, the copy of the machine-learning model on behalf of the additional request, in response to a determination that the copy of the machine-learning model is loaded in the additional serving container; and
respond to the additional request based on executing the copy of the machine-learning model on behalf of the additional request.
5 . The system of claim 1 , comprising further instructions, which when executed, cause the one or more processors to:
identify an extra version of an extra machine-learning model associated with an extra request, in response to receiving the extra request from an extra application;
identify an extra set of available serving containers corresponding to the extra machine-learning model from the cluster of available serving containers which is associated with both the extra version of the extra machine-learning model and the version of the machine-learning model;
select an extra serving container from the extra set of available each serving containers container corresponding to the extra machine-learning model;
load the extra machine-learning model in the extra serving container, in response to a determination that the extra machine-learning model is not loaded in the extra serving container;
execute, in the extra serving container, the extra machine-learning model on behalf of the extra request, in response to a determination that the extra machine-learning model is loaded in the extra serving container; and
respond to the extra request based on executing the extra machine-learning model on behalf of the extra request.
6 . The system of claim 5 , wherein the application is associated with a first tenant and the extra application is associated with a second tenant.
7 . The system of claim 1 , wherein any set of available serving containers corresponding to any machine-learning model is identified based on executing a consistent hashing function applied to identifiers of each serving container associated with any version of any corresponding machine-learning model and an identifier of any corresponding machine-learning model.
8 . A computer program product comprising computer-readable program code to be executed by one or more processors when retrieved from a non-transitory computer-readable medium, the computer-readable program code including instructions to:
watch, by a watcher associated with a routing container of a cluster of routing containers, for changes in available serving containers in a cluster of available serving containers associated with the routing container, the changes include addition of one or more new serving containers to the cluster of available serving containers;
provide, by the watcher, information about the changes in the available serving containers to the routing container to update a mapping of the available serving containers based on the information;
identify a version of a machine-learning model associated with a request, in response to receiving the request from an application;
identify a set of available serving containers corresponding to the version of the machine-learning model from a cluster of available serving containers associated with the version of the machine-learning model, wherein the set of available serving containers is identified based, at least in part, on executing a hashing function applied to identifiers of the available serving containers, an identifier of any corresponding machine-learning model, and a replication factor for the version of the machine-learning model, wherein the replication factor identifies a number of serving containers of the set of available serving containers from the cluster of available serving containers for loading the version of the machine-learning model;
select an existing serving container and a new serving container from the set of available serving containers corresponding to the version of the machine-learning model;
load the version of the machine-learning model in the existing serving container, in response to a determination that the version of the machine-learning model is not loaded in the existing serving container;
execute, in the existing serving container, the machine-learning model on behalf of the request, in response to a determination that the machine-learning model is loaded in the existing serving container;
when the existing serving container is executing the version of the machine-learning model by another request, load another copy of the version of the machine-learning model into the new serving container and execute, in the new serving container, the version of the machine-learning model; and
respond to the request based on executing version of the machine-learning model on behalf of the request.
9 . The computer program product of claim 8 , wherein selecting from any cluster of available serving containers is based on updating a data structure comprising container information associated with serving containers in any corresponding cluster of available serving containers.
10 . The computer program product of claim 8 , wherein the computer-readable program code comprises further instructions to:
identify an other version of an other machine-learning model associated with the request;
identify an other set of available serving containers corresponding to the other machine-learning model from an other cluster of available serving containers associated with the other version of the other machine-learning model;
select an other serving container from the other set of available serving containers corresponding to the other machine-learning model;
load the other machine-learning model in the other serving container, in response to a determination that the other machine-learning model is not loaded in the other serving container; and
execute, in the other serving container, the other machine-learning model on behalf of the request, in response to a determination that the other machine-learning model is loaded in the other serving container;
wherein responding to the request is further based on executing the other machine-learning model on behalf of the request.
11 . The computer program product of claim 8 , wherein the computer-readable program code comprises further instructions to:
identify the version of the machine-learning model associated with an additional request, in response to receiving the additional request from the application;
identify the set of available serving containers corresponding to the machine-learning model from the cluster of available serving containers associated with the version of the machine-learning model;
select an additional serving container from the set of available serving containers corresponding to the machine-learning model;
load a copy of the machine-learning model in the additional serving container, in response to a determination that the copy of the machine-learning model is not loaded in the additional serving container;
execute, in the additional serving container, the copy of the machine-learning model on behalf of the additional request, in response to a determination that the copy of the machine-learning model is loaded in the additional serving container; and
respond to the additional request based on executing the copy of the machine-learning model on behalf of the additional request.
12 . The computer program product of claim 8 , wherein the computer-readable program code comprises further instructions to:
identify an extra version of an extra machine-learning model associated with an extra request, in response to receiving the extra request from an extra application, wherein the application is associated with a first tenant and the extra application is associated with a second tenant;
identify an extra set of available serving containers corresponding to the extra machine-learning model from the cluster of available serving containers which is associated with both the extra version of the extra machine-learning model and the version of the machine-learning model;
select an extra serving container from the extra set of available serving containers corresponding to the extra machine-learning model;
load the extra machine-learning model in the extra serving container, in response to a determination that the extra machine-learning model is not loaded in the extra serving container;
execute, in the extra serving container, the extra machine-learning model on behalf of the extra request, in response to a determination that the extra machine-learning model is loaded in the extra serving container; and
respond to the extra request based on executing the extra machine-learning model on behalf of the extra request.
13 . The computer program product of claim 8 , wherein any set of available serving containers corresponding to any machine-learning model is identified based on executing a consistent hashing function applied to identifiers of each serving container associated with any version of any corresponding machine-learning model and an identifier of any corresponding machine-learning model.
14 . A computer-implemented method for using container information to select containers for executing models, the computer-implemented method comprising:
watching, by a watcher associated with a routing container of a cluster of routing containers, for changes in available serving containers in a cluster of available serving containers associated with the routing container, wherein the changes include addition of one or more new serving containers to the cluster of available serving containers;
providing, by the watcher, information about the changes in the available serving containers to the routing container to update a mapping of the available serving containers based on the information;
identifying a version of a machine-learning model associated with a request, in response to receiving the request from an application;
identifying a set of available serving containers corresponding to the version of the machine-learning model from a cluster of available serving containers associated with the version of the machine-learning model, wherein the set of available serving containers is identified based, at least in part, on executing a hashing function applied to identifiers of the available serving containers and an identifier of any corresponding machine-learning model, and a replication factor for the version of the machine-learning model, wherein the replication factor identifies a number of serving containers of the set of available serving containers from the cluster of available serving containers for loading the version of the machine-learning model;
selecting an existing serving container and a new serving container from the set of available serving containers corresponding to the version of the machine-learning model;
loading the version of the machine-learning model in the existing serving container, in response to a determination that the version of the machine-learning model is not loaded in the existing serving container;
executing, in the existing serving container, the machine-learning model on behalf of the request, in response to a determination that the machine-learning model is loaded in the existing serving container;
when the existing serving container is executing the version of the machine-learning model by another request, loading another copy of the version of the machine-learning model into the new serving container and executing, in the new serving container, the version of the machine-learning model; and
responding to the request based on executing the version of the machine-learning model on behalf of the request.
15 . The computer-implemented method of claim 14 , wherein selecting from any cluster of available serving containers is based on updating a data structure comprising container information associated with available serving containers in any corresponding cluster of available serving containers.
16 . The computer-implemented method of claim 14 , the computer-implemented method further comprising:
identifying an other version of an other machine-learning model associated with the request;
identifying an other set of available serving containers corresponding to the other machine-learning model from an other cluster of available serving containers associated with the other version of the other machine-learning model;
selecting an other serving container from the other set of available serving containers corresponding to the other machine-learning model;
loading the other machine-learning model in the other serving container, in response to a determination that the other machine-learning model is not loaded in the other serving container; and
executing, in the other serving container, the other machine-learning model on behalf of the request, in response to a determination that the other machine-learning model is loaded in the other serving container;
wherein responding to the request is further based on executing the other machine-learning model on behalf of the request.
17 . The computer-implemented method of claim 14 , the computer-implemented method further comprising:
identifying the version of the machine-learning model associated with an additional request, in response to receiving the additional request from the application;
identifying the set of available serving containers corresponding to the machine-learning model from the cluster of available serving containers associated with the version of the machine-learning model;
selecting an additional serving container from the set of available serving containers corresponding to the machine-learning model;
loading a copy of the machine-learning model in the additional serving container, in response to a determination that the copy of the machine-learning model is not loaded in the additional serving container;
executing, in the additional serving container, the copy of the machine-learning model on behalf of the additional request, in response to a determination that the copy of the machine-learning model is loaded in the additional serving container; and
responding to the additional request based on executing the copy of the machine-learning model on behalf of the additional request.
18 . The computer-implemented method of claim 14 , the computer-implemented method further comprising:
identifying an extra version of an extra machine-learning model associated with an extra request, in response to receiving the extra request from an extra application;
identifying an extra set of available serving containers corresponding to the extra machine-learning model from the cluster of available serving containers which is associated with both the extra version of the extra machine-learning model and the version of the machine-learning model;
selecting an extra serving container from the extra set of available serving containers corresponding to the extra machine-learning model;
loading the extra machine-learning model in the extra serving container, in response to a determination that the extra machine-learning model is not loaded in the extra serving container;
executing, in the extra serving container, the extra machine-learning model on behalf of the extra request, in response to a determination that the extra machine-learning model is loaded in the extra serving container; and
responding to the extra request based on executing the extra machine-learning model on behalf of the extra request.
19 . The computer-implemented method of claim 18 , wherein the application is associated with a first tenant and the extra application is associated with a second tenant.
20 . The computer-implemented method of claim 14 , wherein any set of available serving containers corresponding to any machine-learning model is identified based on executing a consistent hashing function applied to identifiers of each serving container associated with any version of any corresponding machine-learning model and an identifier of any corresponding machine-learning model.