System and method for communication load balancing in unseen traffic scenarios
Several policies are trained for determining communication parameters used by mobile devices in selecting a cell of a first communication network to operate on. The several policies form a policy bank. By adjusting the communication parameters, load balancing among cells of the first communication network is achieved. A policy selector is trained so that a target communication network, different than the first communication network, can be load balanced. The policy selector selects a policy from the policy bank for the target communication network. The target communication network applies the policy and the load is balanced on the target communication network. Improved load balancing leads to a reduction of the number of base stations needed in the target communication network.
1 . A method comprising:
selecting a first policy from a policy bank by a policy selector, the policy bank comprising a plurality of policies pre-trained for each of a plurality of traffic profiles, based on a traffic state of a target network;
determining a plurality of load balancing parameters based on the first policy;
balancing, using the plurality of load balancing parameters, a load of the target network for a first period of time;
obtaining a plurality of traffic profiles, wherein a first traffic profile of the plurality of traffic profiles comprises a first time series of traffic demand values and the first traffic profile is associated with a first base station situated at a first geographic location, wherein a second traffic profile of the plurality of traffic profiles comprises a second time series of traffic demand values and the second traffic profile is associated with a second base station situated at a second geographic location different from the first geographic location;
obtaining, by clustering over the plurality of traffic profiles, a vector, each element of the vector corresponding to a plurality of representative traffic profiles;
obtaining, for each plurality of representative traffic profiles, one policy of the plurality of policies,
and
the balancing the load of the target network using the plurality of load balancing parameters comprises performing load balancing of a third base station and a fourth base station using the first policy of the plurality of policies, wherein the target network includes the third base station and the fourth base station.
2 . The method of claim 1 , wherein the policy selector comprises a feed-forward neural network classifier with three hidden layers and one output layer.
3 . The method of claim 2 , wherein the method further comprises obtaining the plurality of policies, wherein the first policy of the plurality of policies indicates a first reward which will result when moving from a first state to a second state based on a first action.
4 . The method of claim 3 , wherein the method further comprises obtaining a plurality of traffic profiles, wherein a first traffic profile of the plurality of traffic profiles is a time series of traffic demand values and the first traffic profile is associated with a first channel of a first base station situated at a first geographic location.
5 . The method of claim 4 , wherein the method further comprises:
running the plurality of policies on the plurality of traffic profiles and obtaining a plurality of state representations, wherein each state representation includes a plurality of state vectors corresponding to a traffic profile and a policy;
retaining the plurality of state representations as a training set; and
training the policy selector based on the training set.
6 . The method of claim 5 further comprising:
deploying the plurality of policies and the policy selector to the target network, wherein the target network exhibits traffic profiles not identical to the plurality of traffic profiles; and
balancing the load of the target network using the plurality of policies.
7 . A server comprising:
one or more processors; and
one or more memories, the one or more memories storing a program, wherein execution of the program by the one or more processors is configured to cause the server to at least:
implement a policy selector for selecting a first policy from a policy bank, the policy bank comprising a plurality of policies;
determine a plurality of load balancing parameters based on the first policy and a current traffic state of a target network;
balance, using the plurality of load balancing parameters, a load of the target network for a first period of time;
obtain a plurality of traffic profiles, wherein a first traffic profile of the plurality of traffic profiles comprises a first time series of traffic demand values and the first traffic profile is associated with a first base station situated at a first geographic location, wherein a second traffic profile of the plurality of traffic profiles comprises a second time series of traffic demand values and the second traffic profile is associated with a second base station situated at a second geographic location different from the first geographic location;
obtain, by clustering over the plurality of traffic profiles, a vector, each element of the vector corresponding to a plurality of representative traffic profiles;
obtain, for each plurality of representative traffic profiles, one policy of the plurality of policies, and
perform load balancing of a third base station and a fourth base station using the first policy of the plurality of policies, wherein the target network comprises the third base station and the fourth base station.
8 . The server of claim 7 , wherein the policy selector comprises a feed-forward neural network classifier with three hidden layers and one output layer.
9 . The server of claim 8 , wherein execution of the program by the one or more processors is further configured to cause the server to obtain the plurality of policies, wherein the first policy of the plurality of policies indicates a first reward which will result when moving from a first state to a second state based on a first action.
10 . The server of claim 7 , wherein execution of the program by the one or more processors is further configured to cause the server to obtain a plurality of traffic profiles, wherein a first traffic profile of the plurality of traffic profiles is a time series of traffic demand values and the first traffic profile is associated with a first channel of a first base station situated at a first geographic location.
11 . The server of claim 10 , wherein execution of the program by the one or more processors is further configured to cause the server to at least:
run the plurality of policies on the plurality of traffic profiles and obtaining a plurality of state representations, wherein each state representation includes a plurality of state vectors corresponding to a traffic profile and a policy;
retain the plurality of state representations as a training set; and
train the policy selector based on the training set.
12 . The server of claim 11 wherein execution of the program by the one or more processors is further configured to cause the server to at least:
deploy the plurality of policies and the policy selector to the target network, wherein the target network exhibits traffic profiles not identical to the plurality of traffic profiles; and
balance the load of the target network using the plurality of policies.
13 . A non-transitory computer readable medium configured to store a program, wherein execution of the program by one or more processors of a server is configured to cause the server to at least:
implement a policy selector for selecting a first policy from a policy bank, the policy bank comprising a plurality of policies;
determine a plurality of load balancing parameters based on the first policy and a current traffic state of a target network;
balance, using the plurality of load balancing parameters, a load of the target network for a first period of time;
obtain a plurality of traffic profiles, wherein a first traffic profile of the plurality of traffic profiles comprises a first time series of traffic demand values and the first traffic profile is associated with a first base station situated at a first geographic location, wherein a second traffic profile of the plurality of traffic profiles comprises a second time series of traffic demand values and the second traffic profile is associated with a second base station situated at a second geographic location different from the first geographic location;
obtain, by clustering over the plurality of traffic profiles, a vector, each element of the vector corresponding to a plurality of representative traffic profiles;
obtain, for each plurality of representative traffic profiles, one policy of the plurality of policies, and
perform load balancing of a third base station and a fourth base station using the first policy of the plurality of policies, wherein the target network comprises the third base station and the fourth base station.
14 . The non-transitory computer readable medium of claim 13 , wherein the policy selector comprises a feed-forward neural network classifier with three hidden layers and one output layer.
15 . The non-transitory computer readable medium of claim 14 , wherein execution of the program by the one or more processors is further configured to cause the server to obtain the plurality of policies, wherein the first policy of the plurality of policies indicates a first reward which will result when moving from a first state to a second state based on a first action.
16 . The non-transitory computer readable medium of claim 15 , wherein execution of the program by the one or more processors is further configured to cause the server to obtain a plurality of traffic profiles, wherein a first traffic profile of the plurality of traffic profiles is a time series of traffic demand values and the first traffic profile is associated with a first channel of a first base station situated at a first geographic location.