IP Library › Granted Patent US 12,621,233
Granted Patent B2
US 12,621,233 · App. 18/250,120 · Granted May 5, 2026

QoS aware reinforcement learning prevention intrusion system

Inventors: Amine Boukhtouta (Laval, CA); Hyame Alameddine (Montreal, CA); Taous Madi (Thuwal, SA); Christian Miranda Moreira (Montreal, CA); Georges Kaddoum (Laval, CA)
Assignee: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
H04L45/02H04L45/00H04L45/08H04W28/0268H04W40/02H04W84/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,233
App. No.
18/250,120
Granted
May 5, 2026
Kind
B2
Abstract

Methods, systems, and apparatuses are disclosed. A network node configured for performing network routing associated with a plurality of wireless devices, WDs, in a communication system is described. The network node includes processing circuit configured to collect, from a control plane, a plurality of graph states associated with a plurality of graphs. Each graph of the plurality of graphs has at least one graph node associated with one WD of the plurality of WDs. At least one action is determined, using self-learning, to update at least one route in at least one graph of the plurality of graphs based on the collected plurality of graphs states. The at least one action is transmitted to a controller for instructing at least one WD to update at least one network route based on the at least one action.

Claims (81)

1 . A network node configured for performing network routing associated with a plurality of wireless devices, WDs, in a communication system, the network node comprising processing circuitry configured to:

collect, from a control plane, a plurality of graph states associated with a plurality of graphs, each graph of the plurality of graphs having at least one graph node associated with one WD of the plurality of WDs;

determine, using self-learning, at least one action to update at least one route in at least one graph of the plurality of graphs based on the collected plurality of graphs states, the self-learning comprising:

entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and

entering a production phase including:

exploiting the explored plurality of graph states of the warm-up phase; and

determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and

cause the network node to transmit the at least one action to a controller for instructing at least one WD to update at least one network route based on the at least one action.

2 . The network node of claim 1 , wherein the self-learning is based at least in part on a quality of service parameter.

3 . The network node of claim 1 , wherein the plurality of graph states includes flowtables and metrics, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.

4 . The network node of claim 1 , wherein the self-learning further includes any one of:

monitoring a topology of at least one graph of the plurality of graphs;

when at least one WD has been one of removed from and added to the at least one graph, one of enter and continue with the warm-up phase; and

when at least one WD has not been one of removed from and added to the at least one graph, one of enter and continue with the production phase.

5 . The network node of claim 1 , wherein the self-learning further includes:

selecting a random state depicting a graph snapshot and a random plurality of actions for each graph node; and

evaluating the selected random state and the random plurality of actions using a probabilistic policy based on a derived quality value.

6 . The network node of claim 5 , wherein the self-learning further includes:

determining a reward based on a cost of the random plurality of actions, the selected random state, an overall transmission delay, a queue delay, and overall packet loss rate;

determining a future state and a future action;

evaluating an action selection policy based on the derived quality value;

learning another quality value by evaluating an impact of the reward and how the future state and the future action compare with the selected random state and random plurality of actions;

updating a current state with the future state and a current action with the future action; and

capturing an overall quality value determine a convergence.

7 . The network node of claim 6 , wherein updating the current state and the current action includes selecting a graph node that is a parent to another graph node, the selecting being based at least on the derived quality value.

8 . The network node of claim 1 , wherein the controller is in the control plane, the at least one WD is in a data plane, and transmitting the at least one action triggers the WD to update the at least one network route.

9 . The network node of claim 1 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.

10 . The network node of claim 1 , wherein the communication system includes a wireless sensor network, the network routing is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.

11 . A method implemented in a network node configured for performing network routing associated with a plurality of wireless devices, WDs, in a communication system, the method comprising:

collecting, from a control plane, a plurality of graph states associated with a plurality of graphs, each graph of the plurality of graphs having at least one graph node associated with one WD of the plurality of WDs;

determining, using self-learning, at least one action to update at least one route in at least one graph of the plurality of graphs based on the collected plurality of graphs states, the self-learning comprising:

entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and

entering a production phase including:

exploiting the explored plurality of graph states of the warm-up phase; and

determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and

transmitting the at least one action to a controller for instructing at least one WD to update at least one network route based on the at least one action.

12 . The method of claim 11 , wherein the self-learning is based at least in part on a quality of service parameter.

13 . The method of claim 11 , wherein the plurality of graph states includes flowtables and metrics, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.

14 . The method of claim 11 , wherein the self-learning further includes any one of:

monitoring a topology of at least one graph of the plurality of graphs;

when at least one WD has been one of removed from and added to the at least one graph, one of enter and continue with the warm-up phase; and

when at least one WD has not been one of removed from and added to the at least one graph, one of enter and continue with the production phase.

15 . The method of claim 11 , wherein the self-learning further includes:

selecting a random state depicting a graph snapshot and a random plurality of actions for each graph node; and

evaluating the selected random state and the random plurality of actions using a probabilistic policy based on a derived quality value.

16 . The method of claim 15 , wherein the self-learning further includes:

determining a reward based on a cost of the random plurality of actions, the selected random state, an overall transmission delay, a queue delay, and overall packet loss rate;

determining a future state and a future action;

evaluating an action selection policy based on the derived quality value;

learning another quality value by evaluating an impact of the reward and how the future state and the future action compare with the selected random state and random plurality of actions;

updating a current state with the future state and a current action with the future action; and

capturing an overall quality value determine a convergence.

17 . The method of claim 16 , wherein updating the current state and the current action includes selecting a graph node that is a parent to another graph node, the selecting being based at least on the derived quality value.

18 . The method of claim 11 , wherein the controller is in the control plane, the at least one WD is in a data plane, and transmitting the at least one action triggers the WD to update the at least one network route.

19 . The method of claim 11 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.

20 . The method of claim 11 , wherein the communication system includes a wireless sensor network, the network routing is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.

21 . A wireless device, WD, configured to communicate with a network node in a communication system, the WD comprising processing circuitry and a radio interface in communication with the processing circuitry, the radio interface being configured to:

receive at least one action for instructing the WD to update at least one network route, the at least one action being determined using self-learning, the at least one network route being in at least one graph of a plurality of graphs, the self-learning comprising:

entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and

entering a production phase including:

exploiting the explored plurality of graph states of the warm-up phase; and

determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and

the processing circuitry being configured to:

update the at least one network route based on the received at least one action.

22 . The WD of claim 21 , wherein the self-learning is based at least in part on a quality of service parameter.

23 . The WD of claim 21 , wherein the received at least one action is further determined based on flowtables and metrics associated with the at least one graph, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.

24 . The WD of claim 21 , wherein the instructing is in a data plane.

25 . The WD of claim 21 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.

26 . The WD of claim 21 , wherein the communication system includes a wireless sensor network, the at least one network route is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.

27 . A method implemented in a wireless device, WD, configured to communicate with a network node in a communication system, the method including:

receiving a at least one action for instructing the WD to update at least one network route, the at least one action being determined using self-learning, the at least one network route being in at least one graph of a plurality of graphs, the self-learning comprising:

entering a warm-up phase comprising exploring the plurality of graph states as a ground truth to self-learn optimizing network routes; and

entering a production phase including:

exploiting the explored plurality of graph states of the warm-up phase; and

determining a plurality of actions including the at least one action to update the at least on route in at least one graph; and

updating the at least one network route based on the received at least one action.

28 . The method of claim 27 , wherein the self-learning is based at least in part on a quality of service parameter.

29 . The method of claim 27 , wherein the received at least one action is further determined based on flowtables and metrics associated with the at least one graph, the metrics including at least one of a transmission delay, a packet loss rate, and a queue delay.

30 . The method of claim 27 , wherein the instructing is in a data plane.

31 . The method of claim 27 , wherein the plurality of graphs is a plurality of Destination Oriented Directed Acyclic Graphs, DODAGs.

32 . The method of claim 27 , wherein the communication system includes a wireless sensor network, the at least one network route is a Quality of Service, QoS, awareness-based routing that is performed in Routing Protocol for low Power and Lossy networks, RPL, in the wireless sensor network, the network node is a border router to the wireless sensor network, and the wireless sensor network is a IPv6 low power wireless personal area network, SD6LowPAN, network.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2023
From: KADDOUM, GEORGES; MOREIRA, CHRISTIAN MIRANDA
To: ECOLE DE TECHNOLOGIE SUPERIEURE
Reel/Frame 065913/0278 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2023
From: ECOLE DE TECHNOLOGIE SUPERIEURE
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 065913/0533 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2023
From: BOUKHTOUTA, AMINE; ALAMEDDINE, HYAME; MADI, TAOUS
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 065885/0854 →
Continuity (2)
Provisional Application 63104607 · Oct 23, 2020
Related Publication 20230403600A1 · Dec 14, 2023
References Cited (40)
US 10230419B2 · Bharadia et al. · 2019 [cited by applicant]
US 20120155475A1 · Vasseur et al. · 2012 [cited by applicant]
US 20130159550A1 · Vasseur · 2013 [cited by applicant]
US 20130223218A1 · Vasseur et al. · 2013 [cited by applicant]
US 20140029432A1 · Vasseur et al. · 2014 [cited by applicant]
US 20200259736A1 · She · 2020 [cited by examiner]
US 20210099378A1 · Alaettinoglu · 2021 [cited by examiner]
US 20210119882A1 · Orhan · 2021 [cited by examiner]
US 20220103462A1 · Sidebottom · 2022 [cited by examiner]
US 20220376984A1 · Wen · 2022 [cited by examiner]
CN 110049527A · 2019 [cited by applicant]
CN 112822109A · 2021 [cited by applicant]
Zhu et al., “Network planning with deep reinforcement learning”, Aug. 9, 2021, In Proceedings of the 2021 ACM SIGCOMM 2021 Conference (SIGCOMM '21). Association for Computing Machinery, https://doi.org/10.1145/3452296.3… [cited by examiner]
International Search Report and Written Opinion dated Jan. 26, 2022 issued in PCT Application No. PCT/IB2021/059777 filed Oct. 22, 2021, consisting of 11 pages. [cited by applicant]
Mustafa Kocakulak et al., An Overview of Wireless Sensor Networks Towards Internet of Things; Computing and Communication Workshop and Conference, 2017, consisting of 6 pages. [cited by applicant]
Al-Kashoash, Congestion Control for 6LoWPAN Wireless Sensor Networks: Toward the Internet of Things; Springer, Sep. 2017, consisting of 221 pages. [cited by applicant]
M.L.F. Miguel, et al., SDN Architecture for 6LoWPAN Wireless Sensor Networks; Sensors, vol. No. 18, No. 11, 2018, consisting of 23 pages. [cited by applicant]
Hlabishi I. Kobo, et al., A Survey on Software-Defined Wireless Sensor Networks: Challenges and Design Requirements; IEEE Access; vol. 5, 2017, consisting of 28 pages. [cited by applicant]
Ivan Nedyalkov, Studying of a Modeled IP—Based Network Using Different Dynamic Routing Protocols; IEEE National Conference with International Participation Conference “Electronica 2019”, May 16-17, 2019, Sofia, Bulgaria… [cited by applicant]
Nicolas Tsiftes et al., Poster Abstract: Low-Power Wireless IPV6 Routing with ContikiRPL; IEEE International Conference Information Processing in Sensor Networks; 2010 Stockholm, Sweden, consisting of 2 pages. [cited by applicant]
P. Levis et al., The Trickle Algorithm; Internet Engineering Task Force (IETF), RFC6206; 2011, consisting of 13 pages. [cited by applicant]
Zibuyisile Magubane et al., RPL-Based on Load Balancing Routing Objective Functions for IoTs in Distributed Networks; in IEEE Int. Mult. Inf. Technol. Eng. Conf., 2019, consisting of 6 pages. [cited by applicant]
Muneer Bani Yassein et al., Performance Evaluation of RPL in High Density Networks for Internet of Things (IoT); In Int. Conf. Soft. and Inf. Eng., 2019, consisting of 5 pages. [cited by applicant]
Walid Khallef et al., Multiple Constrained QoS Routing with RPL; IEEE ICC 2017 Ad-Hoc and Sensor Networking Symposium, consisting of 7 pages. [cited by applicant]
P. Thubert et al., Objective Function Zero for the Routing Protocol for Low-Power and Lossy Networks (RPL); Cisco Systems, Mar. 2012, consisting of 16 pages. [cited by applicant]
Nurrahmat Pradeska et al., Performance Analysis of Objective Function MRHOF and OF0 in Routing Protocol RPL IPV6 Over Low Power Wireless Personal Area Networks (6LoWPAN); 2016 8th International Conference on Information… [cited by applicant]
Nabil Djedjig et al., New Trust Metric for RPL Routing Protocol; IEEE Int. Conf. Inf. and Commun. Sys., 2017, consisting 8 pages. [cited by applicant]
Muhammad Ali Lodhi et al., Rank Attack Using Objective Function in RPL for Low Power and Lossy Networks; IEEE Int. Conf. Ind. Inf. and Comp. Syst., 2016, consisting of 6 pages. [cited by applicant]
P. Thubert, An Architecture for IPV6 over the TSCH mode of IEEE 802.15.4 draft-ietf-6tisch-architecture-15; Jan. 2017, work in Progress. [Online], Internet Draft, Cisco; consisting of 51 pages. [cited by applicant]
Michael Baddeley et al., Evolving SDN for Low-Power IoT Networks; in IEEE Conference Network Software and Workshops, 2018; consisting of 9 pages. [cited by applicant]
Michael Baddeley A Low-Overhead SDN Stack and Embedded SDN Controller for Contiki; Jun. 2018, consisting of 8 pages. [cited by applicant]
Haoyu Song, Protocol-Oblivious Forwarding: Unleash the Power of SDN through a Future-Proof Forwarding Plane; Proceedings of the Second ACM SIGCOMM Workshop on Hot Topics in Software Defined Net-working, ser. HotSDN 13. … [cited by applicant]
Tie Luo et al., Enhancing Responsiveness and Scalability for OpenFlow Networks via Control-Message Quenching; 2012 International Conference on ICT Convergence (ICTC), Oct. 2012, consisting of 6 pages. [cited by applicant]
Saloua Chettibi et al., An Adaptive Energy-Aware Routing Protocol for MANETs Using The SARSA Reinforcement Learning Algorithm; IEEE Conf. Evolv. and Adapt. Intel. System, 2012, consisting of 6 pages. [cited by applicant]
Kazunori Iwata, Extending the Peak Bandwidth of Parameters for Softmax Selection in Reinforcement Learning; IEEE Transactions on Neural Networks and Learning Systems, vol. 28, No. 8, Aug. 2017, consisting of 13 pages. [cited by applicant]
T. Winter et al., IPv6 Routing Protocol for Low-Power and Lossy Networks; Internet Engineering Task Force (IETF); Mar. 2012, consisting of 157 pages. [cited by applicant]
Communication Under Rule 71(3) EPC dated Jul. 23, 2025, issued in corresponding European Patent Application No. 21 801 640.0, consisting of 8 pages. [cited by applicant]
Chinese Office Action and English translation of the Chinese Office Action dated Jan. 26, 2026 issued in corresponding Chinese Application No. 202180071930.4, consisting of 18 pages. [cited by applicant]
Boudouaia et al., “Security Against Rank Attach in RPL Protocol”; IEEE Network, July/Aug. 2020, consisting of 7 pages. [cited by applicant]
Yu et al., “Computer Applications and Software”; Research Progress on Routing Algorithms for Smart Grid Neighborhood Area Networks; vol. 34, No. 1; Jan. 2017, consisting of 8 pages. [cited by applicant]