Optimal AP connection method and system using reinforcement learning to improve energy efficiency and latency of IoT devices
Disclosed is an optimal AP connection method including transmitting, by an IoT device, a probe request message to a plurality of iAPs; transmitting, by each of the plurality of iAPs that receives the probe request message, the probe request message to an iAP controller and transmitting local information to each iAP to the iAP controller; performing, by the iAP controller, reinforcement learning for IoT device using global information that is updated from the local information; selecting, by the iAP controller, an optimal iAP based on RLreinforcement learning and transmitting recommended Tx power value information on the IoT device and a probe response message to the selected corresponding iAP; and transmitting, by the corresponding iAP that receives its selection as the optimal iAP in response to the probe request from the iAP controller, the probe response message and the recommended Tx power value information to the IoT device.
1 . An access point (AP) connection method using reinforcement learning, the method comprising:
transmitting, by an Internet of things (IoT) device present within a multiple intelligent access point (iAP) coverage area, a probe request message for iAP connection to a plurality of iAPs;
transmitting, by each of the plurality of iAPs that receives the probe request message, a received signal strength indication (RSSI) value and the probe request message to an iAP controller and periodically transmitting local information that includes the number of IoT devices connected to each iAP to the iAP controller;
performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs, the iAP controller including an energy and latency reinforcement learning (EL-RL) model, a location estimation model, and a recommended Tx power model to perform reinforcement learning;
selecting, by the iAP controller, an optimal iAP based on reinforcement learning and transmitting recommended Tx power value information on the IoT device and a probe response message to the selected corresponding iAP;
transmitting, by the corresponding iAP that receives its selection as the optimal iAP in response to the probe request from the iAP controller, the probe response message and the recommended Tx power value information to the IoT device; and
transmitting, by the IoT device that receives the probe response message from the optimal iAP, IoT data at recommended Tx power through a connection process with the optimal iAP.
2 . The method of claim 1 , wherein the performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs comprises numerically analyzing average energy consumption and average latency of the IoT devices according to the number of uplink transmission attempts of the IoT device and a successful transmission probability according to each transmission attempt through the EL-RL model of the iAP controller and performing reinforcement learning with a policy that minimizes an objective function configured with a weighted sum of the average energy consumption and the average latency of the IoT devices according to analysis results.
3 . The method of claim 2 , wherein the performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs comprises transmitting a state to an EL-RL agent of the EL-RL model for optimal iAP selection in an environment of a simulator for performing reinforcement learning through the EL-RL model, and
the state is set based on an RSSI value between an IoT device to be connected and a candidate iAP and the number of IoT devices connected to the iAP, a reward is calculated based on a distance from the connected IoT device according to an action that represents an iAP to be connected between candidate iAPs with the IoT device to be connected through numerical analysis of the simulator and the average energy consumption and the average latency, and minimizing the average energy consumption and the average latency of all connected IoT devices is set as the objective function of the EL-RL model.
4 . The method of claim 1 , wherein the performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs comprises performing pretraining using a fingerprinting method of estimating a location of the IoT device by comparing RSSI values input to a fingerprinting map including reference point values prestored in a database in a data collection of an offline process through the location estimation model of the iAP controller.
5 . The method of claim 4 , wherein the performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs comprises collecting RSSI in real time through the IoT device and estimating a location of the IoT device using a model pretrained through the fingerprint method in an online process through the location estimation model of the iAP controller.
6 . The method of claim 1 , wherein the performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs comprises calculating a distance from a candidate iAP according to an estimated location value of the IoT device estimated through the location estimation model of the iAP controller and then calculating a recommended Tx power value of the IoT device through the recommended Tx power model of the iAPcontroller.
7 . The method of claim 1 , wherein the selecting, by the iAP controller, the optimal iAP based on reinforcement learning and the transmitting the recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP comprises:
selecting the optimal iAP through the EL-RL model of the iAP controller using an estimated location value of the IoT device estimated through the location estimation model of the iAP controller and a recommended Tx power value calculated through the recommended Tx power model of the iAP controller; and
transmitting the recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP to reduce energy consumption in an uplink transmission of the IoT device.
8 . The method of claim 1 , wherein the transmitting, by the IoT device that receives the probe response message from the optimal iAP, the IoT data at recommended Tx power through the connection process with the optimal iAP comprises numerically analyzing average energy consumption and average latency of the IoT devices using the number of transmission attempts including a retransmission of the IoT device due to a packet collision and a successful transmission probability according to each transmission attempt and transmitting the IoT data at the recommended Tx power according to analysis results, to reduce IoT device energy consumption due to the packet collision that occurs when the IoT device and another IoT device simultaneously transmit a packet during an uplink transmission of the IoT device.
9 . The method of claim 8 , wherein the transmitting, by the IoT device that receives the probe response message from the optimal iAP, the IoT data at the recommended Tx power through the connection process with the optimal iAP comprises:
calculating the average energy consumption of the IoT devices with a sum of product of probability of all transmission attempts, probability of successful transmission without a packet collision, and an energy consumption value of each transmission attempt; and
calculating the average latency of the IoT devices with a sum of product of probability of all transmission attempts, probability of successful transmission without a packet collision, and a latency value of each transmission attempt, and product of a packet collision probability and a time consumed for the packet collision.
10 . An access point (AP) connection system using reinforcement learning, the AP connection system comprising:
a plurality of intelligent access points (iAPs) configured to receive a probe request message for iAP connection from an Internet of things (IoT) device present within a multiple iAP coverage area, each of the plurality of iAPs that receives the probe request message transmitting a received signal strength indication (RSSI) value and the probe request message to an iAP controller and periodically transmitting local information that includes the number of IoT devices connected to each iAP to the iAP controller;
the iAP controller configured to perform reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs, to select an optimal iAP based on reinforcement learning, and to transmit recommended Tx power value information on the IoT device and a probe response message to the selected corresponding iAP, the iAP controller including an energy and latency reinforcement learning (EL-RL) model, a location estimation model, and a recommended Tx power model to perform reinforcement learning; and
the IoT device configured to receive the probe response message and the recommended Tx power value information from the corresponding iAP that receives its selection as the optimal iAP in response to the probe request, and to transmit IoT data at recommended Tx power through a connection process with the optimal iAP.
11 . The AP connection system of claim 10 , wherein, to perform reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs, the iAP controller is configured to numerically analyze average energy consumption and average latency of the IoT devices according to the number of uplink transmission attempts of the IoT device and a successful transmission probability according to each transmission attempt through the EL-RL model of the iAP controller and to perform reinforcement learning with a policy that minimizes an objective function configured with a weighted sum of the average energy consumption and the average latency of the IoT devices according to analysis results.
12 . The AP connection system of claim 11 , wherein the iAP controller is configured to transmit a state to an EL-RL agent of the EL-RL model for optimal iAP selection in an environment of a simulator for performing reinforcement learning through the EL-RL model, and
the state is set based on an RSSI value between an IoT device to be connected and a candidate iAP and the number of IoT devices connected to the iAP, a reward is calculated based on a distance from the connected IoT device according to an action that represents an iAP to be connected between candidate iAPs with the IoT device to be connected through numerical analysis of the simulator and the average energy consumption and the average latency, and minimizing the average energy consumption and the average latency of all connected IoT devices is set as the objective function of the EL-RL model.
13 . The AP connection system of claim 10 , wherein, to perform reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs, the iAP controller is configured to perform pretraining using a fingerprinting method of estimating a location of the IoT device by comparing RSSI values input to a fingerprinting map including reference point values prestored in a database in a data collection of an offline process through the location estimation model of the iAP controller.
14 . The AP connection system of claim 13 , wherein the iAP controller is configured to collect RSSI in real time through the IoT device and to estimate a location of the IoT device using a model pretrained through the fingerprint method in an online process through the location estimation model of the iAP controller.
15 . The AP connection system of claim 10 , wherein, to perform reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs, the iAP controller is configured to calculate a distance from a candidate iAP according to an estimated location value of the IoT device estimated through the location estimation model of the iAP controller and then calculate a recommended Tx power value of the IoT device through the recommended Tx power model of the iAP controller.
16 . The AP connection system of claim 10 , wherein, to select the optimal iAP based on reinforcement learning and to transmit the recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP, the iAP controller is configured to,
select the optimal iAP through the EL-RL model of the iAP controller using an estimated location value of the IoT device estimated through the location estimation model of the iAP controller and a recommended Tx power value calculated through the recommended Tx power model of the iAP controller, and
transmit the recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP to reduce energy consumption in an uplink transmission of the IoT device.
17 . The AP connection system of claim 10 , wherein, to transmit the IoT data at the recommended Tx power through the connection process with the optimal iAP when the probe response message is received from the optimal iAP, the IoT device is configured to numerically analyze average energy consumption and average latency of the IoT devices using the number of transmission attempts including a retransmission of the IoT device due to a packet collision and a successful transmission probability according to each transmission attempt and to transmit the IoT data at the recommended Tx power according to analysis results, to reduce IoT device energy consumption due to the packet collision that occurs when the IoT device and another IoT device simultaneously transmit a packet during an uplink transmission of the IoT device.
18 . The AP connection system of claim 17 , wherein the IoT device is configured to, calculate the average energy consumption of the IoT devices with a sum of product of probability of all transmission attempts, probability of successful transmission without a packet collision, and an energy consumption value of each transmission attempt, and
calculate the average latency of the IoT devices with a sum of product of probability of all transmission attempts, probability of successful transmission without a packet collision, and a latency value of each transmission attempt, and product of a packet collision probability and a time consumed for the packet collision.
19 . A non-transitory computer-readable recording medium to perform an optimal access point (AP) connection method using reinforcement learning to improve energy efficiency and latency of Internet of things (IoT) devices, the method comprising:
transmitting, by an IoT device present within a multiple intelligent access point (iAP) coverage area, a probe request message for iAP connection to a plurality of iAPs;
transmitting, by each of the plurality of iAPs that receives the probe request message, a received signal strength indication (RSSI) value and the probe request message to an iAP controller and periodically transmitting local information that includes the number of IoT devices connected to each iAP to the iAP controller;
performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs, the iAP controller including an energy and latency reinforcement learning (EL-RL) model, a location estimation model, and a recommended Tx power model to perform reinforcement learning;
selecting, by the iAP controller, an optimal iAP based on reinforcement learning and transmitting recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP;
transmitting, by the corresponding iAP that receives its selection as the optimal iAP in response to the probe request from the iAP controller, the probe response message and the recommended Tx power value information to the IoT device; and
transmitting, by the IoT device that receives the probe response message from the optimal iAP, IoT data at recommended Tx power through a connection process with the optimal iAP.
20 . The non-transitory computer-readable recording medium of claim 19 , wherein the performing, by the iAP controller, reinforcement learning for IoT device energy efficiency and latency using global information that is updated from the local information obtained from the plurality of iAPs comprises:
numerically analyzing average energy consumption and average latency of the IoT devices according to the number of uplink transmission attempts of the IoT device and a successful transmission probability according to each transmission attempt through the EL-RL model of the iAP controller and performing reinforcement learning with a policy that minimizes an objective function configured with a weighted sum of the average energy consumption and the average latency of the IoT devices according to analysis results;
performing pretraining using a fingerprinting method of estimating a location of the IoT device by comparing RSSI values input to a fingerprinting map including reference point values prestored in a database in a data collection of an offline process through the location estimation model of the iAP controller; and
calculating a distance from a candidate iAP according to an estimated location value of the IoT device estimated through the location estimation model of the iAP controller and then calculating a recommended Tx power value of the IoT device through the recommended Tx power model of the iAP controller, and
the selecting, by the iAP controller, the optimal iAP based on reinforcement learning and the transmitting the recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP comprises:
selecting the optimal iAP through the EL-RL model of the iAP controller using an estimated location value of the IoT device estimated through the location estimation model of the iAP controller and a recommended Tx power value calculated through the recommended Tx power model of the iAP controller; and
transmitting the recommended Tx power value information on the IoT device and the probe response message to the selected corresponding iAP to reduce energy consumption in an uplink transmission of the IoT device.