QoS aware reinforcement learning prevention intrusion system
Methods, systems, and apparatuses are disclosed. A network node configured for performing network routing associated with a plurality of wireless devices, WDs, in a communication system is described. The network node includes processing circuit configured to collect, from a control plane, a plurality of graph states associated with a plurality of graphs. Each graph of the plurality of graphs has at least one graph node associated with one WD of the plurality of WDs. At least one action is determined, using self-learning, to update at least one route in at least one graph of the plurality of graphs based on the collected plurality of graphs states. The at least one action is transmitted to a controller for instructing at least one WD to update at least one network route based on the at least one action.
1 . A network node configured for performing network routing associated with a plurality of wireless devices, WDs, in a communication system, the network node comprising processing circuitry configured to:
collect, from a control plane, a plurality of graph states associated with a plurality of graphs, each graph of the plurality of graphs having at least one graph node associated with one WD of the plurality of WDs;
determine, using self-learning, at least one action to update at least one route in at least one graph of the plurality of graphs based on the collected plurality of graphs states, the self-learning comprising:
entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and
entering a production phase including:
exploiting the explored plurality of graph states of the warm-up phase; and
determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and
cause the network node to transmit the at least one action to a controller for instructing at least one WD to update at least one network route based on the at least one action.
2 . The network node of claim 1 , wherein the self-learning is based at least in part on a quality of service parameter.
3 . The network node of claim 1 , wherein the plurality of graph states includes flowtables and metrics, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.
4 . The network node of claim 1 , wherein the self-learning further includes any one of:
monitoring a topology of at least one graph of the plurality of graphs;
when at least one WD has been one of removed from and added to the at least one graph, one of enter and continue with the warm-up phase; and
when at least one WD has not been one of removed from and added to the at least one graph, one of enter and continue with the production phase.
5 . The network node of claim 1 , wherein the self-learning further includes:
selecting a random state depicting a graph snapshot and a random plurality of actions for each graph node; and
evaluating the selected random state and the random plurality of actions using a probabilistic policy based on a derived quality value.
6 . The network node of claim 5 , wherein the self-learning further includes:
determining a reward based on a cost of the random plurality of actions, the selected random state, an overall transmission delay, a queue delay, and overall packet loss rate;
determining a future state and a future action;
evaluating an action selection policy based on the derived quality value;
learning another quality value by evaluating an impact of the reward and how the future state and the future action compare with the selected random state and random plurality of actions;
updating a current state with the future state and a current action with the future action; and
capturing an overall quality value determine a convergence.
7 . The network node of claim 6 , wherein updating the current state and the current action includes selecting a graph node that is a parent to another graph node, the selecting being based at least on the derived quality value.
8 . The network node of claim 1 , wherein the controller is in the control plane, the at least one WD is in a data plane, and transmitting the at least one action triggers the WD to update the at least one network route.
9 . The network node of claim 1 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.
10 . The network node of claim 1 , wherein the communication system includes a wireless sensor network, the network routing is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.
11 . A method implemented in a network node configured for performing network routing associated with a plurality of wireless devices, WDs, in a communication system, the method comprising:
collecting, from a control plane, a plurality of graph states associated with a plurality of graphs, each graph of the plurality of graphs having at least one graph node associated with one WD of the plurality of WDs;
determining, using self-learning, at least one action to update at least one route in at least one graph of the plurality of graphs based on the collected plurality of graphs states, the self-learning comprising:
entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and
entering a production phase including:
exploiting the explored plurality of graph states of the warm-up phase; and
determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and
transmitting the at least one action to a controller for instructing at least one WD to update at least one network route based on the at least one action.
12 . The method of claim 11 , wherein the self-learning is based at least in part on a quality of service parameter.
13 . The method of claim 11 , wherein the plurality of graph states includes flowtables and metrics, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.
14 . The method of claim 11 , wherein the self-learning further includes any one of:
monitoring a topology of at least one graph of the plurality of graphs;
when at least one WD has been one of removed from and added to the at least one graph, one of enter and continue with the warm-up phase; and
when at least one WD has not been one of removed from and added to the at least one graph, one of enter and continue with the production phase.
15 . The method of claim 11 , wherein the self-learning further includes:
selecting a random state depicting a graph snapshot and a random plurality of actions for each graph node; and
evaluating the selected random state and the random plurality of actions using a probabilistic policy based on a derived quality value.
16 . The method of claim 15 , wherein the self-learning further includes:
determining a reward based on a cost of the random plurality of actions, the selected random state, an overall transmission delay, a queue delay, and overall packet loss rate;
determining a future state and a future action;
evaluating an action selection policy based on the derived quality value;
learning another quality value by evaluating an impact of the reward and how the future state and the future action compare with the selected random state and random plurality of actions;
updating a current state with the future state and a current action with the future action; and
capturing an overall quality value determine a convergence.
17 . The method of claim 16 , wherein updating the current state and the current action includes selecting a graph node that is a parent to another graph node, the selecting being based at least on the derived quality value.
18 . The method of claim 11 , wherein the controller is in the control plane, the at least one WD is in a data plane, and transmitting the at least one action triggers the WD to update the at least one network route.
19 . The method of claim 11 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.
20 . The method of claim 11 , wherein the communication system includes a wireless sensor network, the network routing is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.
21 . A wireless device, WD, configured to communicate with a network node in a communication system, the WD comprising processing circuitry and a radio interface in communication with the processing circuitry, the radio interface being configured to:
receive at least one action for instructing the WD to update at least one network route, the at least one action being determined using self-learning, the at least one network route being in at least one graph of a plurality of graphs, the self-learning comprising:
entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and
entering a production phase including:
exploiting the explored plurality of graph states of the warm-up phase; and
determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and
the processing circuitry being configured to:
update the at least one network route based on the received at least one action.
22 . The WD of claim 21 , wherein the self-learning is based at least in part on a quality of service parameter.
23 . The WD of claim 21 , wherein the received at least one action is further determined based on flowtables and metrics associated with the at least one graph, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.
24 . The WD of claim 21 , wherein the instructing is in a data plane.
25 . The WD of claim 21 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.
26 . The WD of claim 21 , wherein the communication system includes a wireless sensor network, the at least one network route is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.
27 . A method implemented in a wireless device, WD, configured to communicate with a network node in a communication system, the method including:
receiving a at least one action for instructing the WD to update at least one network route, the at least one action being determined using self-learning, the at least one network route being in at least one graph of a plurality of graphs, the self-learning comprising:
entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and
entering a production phase including:
exploiting the explored plurality of graph states of the warm-up phase; and
determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and
updating the at least one network route based on the received at least one action.
28 . The method of claim 27 , wherein the self-learning is based at least in part on a quality of service parameter.
29 . The method of claim 27 , wherein the received at least one action is further determined based on flowtables and metrics associated with the at least one graph, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.
30 . The method of claim 27 , wherein the instructing is in a data plane.
31 . The method of claim 27 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.
32 . The method of claim 27 , wherein the communication system includes a wireless sensor network, the at least one network route is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.