Method of supporting reinforcement learning in mobile communication system and devices for performing the same
View Patent ↗A method of supporting reinforcement learning (RL) in a mobile communication system and devices for performing the same are disclosed. The method of supporting RL in a mobile communication system includes receiving a subscription request to train an RL model from a consumer network function (NF), training the RL model by collecting data in response to the subscription request, and when training of the RL model is completed based on a training completion condition, transmitting a notification including RL model information to the consumer NF.
1 . A method of operating a device including a network data analytics function (NWDAF) that supports reinforcement learning (RL) in a mobile communication system, the method comprising:
receiving a subscription request to train an RL model from a consumer network function (NF);
training the RL model by collecting data in response to the subscription request; and
when training of the RL model is completed based on a training completion condition, transmitting a notification comprising RL model information to the consumer NF, and
wherein the subscription request comprises parameters that the consumer NF is able to provide, and
wherein the parameters comprise an expected RL training completion time, an expected RL training completion condition other than time, an expected RL model confidence level, an expected reward, a candidate action space, an RL model update period, an RL preparation flag, an RL correlation identifier (ID), and an available data requirement.
2 . The method of claim 1 , wherein the RL preparation flag is for identifying whether the subscription request is for preparing RL or executing RL, and
the RL correlation ID is for identifying an RL procedure for RL model training.
3 . The method of claim 1 , wherein the training completion condition comprises at least one of the number of epochs, an RL training completion time, a case in which an indicator, such as a reward value or a loss function, shows a difference within or greater than or equal to a predetermined level, and an RL model confidence level.
4 . The method of claim 1 , further comprising:
registering, to a network repository function (NRF), RL capability information comprising information indicating whether RL training is supported at the NWDAF as an RL agent.
5 . The method of claim 1 , wherein the NWDAF that trains the RL model comprises a model training logical function (MTLF).
6 . The method of claim 5 , wherein the NWDAF is selected based on criteria by the consumer NF from a plurality of NWDAFs registered as an RL agent in the NRF.
7 . The method of claim 6 , wherein the criteria comprise RL capability information, a candidate action space, a state, an analytics ID of a RL model, whether an NWDAF, which is a selected RL agent, executes an RL procedure, a time period of interest, and a service area.
8 . A device including a network data analytics function (NWDAF) that supports reinforcement learning (RL) in a mobile communication system, the device comprising:
a processor; and
a memory electrically connected to the processor and configured to store instructions executable by the processor,
wherein the processor performs a plurality of operations when the instructions are executed by the processor, and
the plurality of operations comprises:
receiving a subscription request to train an RL model from a consumer network function (NF);
training the RL model by collecting data in response to the subscription request; and
when training of the RL model is completed based on a training completion condition, transmitting a notification comprising RL model information to the consumer NF, and
wherein the subscription request comprises parameters that the consumer NF is able to provide, and
wherein the parameters comprise an expected RL training completion time, an expected RL training completion condition other than time, an expected RL model confidence level, an expected reward, a candidate action space, an RL model update period, an RL preparation flag, an RL correlation identifier (ID), and an available data requirement.
9 . The device of claim 8 , wherein the RL preparation flag is for identifying whether the subscription request is for preparing RL or executing RL, and
the RL correlation ID is for identifying an RL procedure for RL model training.
10 . The device of claim 8 , wherein the training completion condition comprises at least one of a number of epochs, an RL training completion time, a case in which an indicator, such as a reward value or a loss function, shows a difference within or greater than or equal to a predetermined level, and an RL model confidence level.
11 . The device of claim 8 , wherein the plurality of operations comprises:
registering, to a network repository function (NRF), RL capability information comprising information indicating whether RL training is supported at the NWDAF as an RL agent.
12 . The device of claim 8 , wherein the NWDAF that trains the RL model comprises a model training logical function (MTLF).
13 . The device of claim 12 , wherein the NWDAF is selected based on criteria by the consumer NF from a plurality of NWDAFs registered as an RL agent in the NRF.
14 . The device of claim 13 , wherein the criteria comprise RL capability information, a candidate action space, a state, an analytics ID of a RL model, whether an NWDAF, which is a selected RL agent, executes an RL procedure, a time period of interest, and a service area.