IP Library › Granted Patent US 12,192,820
Granted Patent B2
US 12,192,820 · App. 17/484,743 · Granted Jan 7, 2025

Reinforcement learning for multi-access traffic management

Inventors: Shu-ping Yeh (Campbell, CA); Jingwen Bai (San Jose, CA); Shilpa Talwar (Cupertino, CA)
Assignee: Intel Corporation
H04W28/0268G06N3/08H04W28/0231
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,192,820
App. No.
17/484,743
Filed
Sep 24, 2021
Granted
Jan 7, 2025
Kind
B2
Art Unit
2414
USPC
370/235
Abstract

The present disclosure is related to multi-access traffic management in edge computing environments, and in particular, artificial intelligence (AI) and/or machine learning (ML) techniques for multi-access traffic management. A scalable AI/ML architecture for multi-access traffic management is provided. Reinforcement learning (RL) and/or Deep RL (DRL) approaches that learn policies and/or parameters for traffic management and/or for distributing multi-access traffic through interacting with the environment are also provided. Deep contextual bandit RL techniques for intelligent traffic management for edge networks are also provided. Other embodiments may be described and/or claimed.

Claims (69)

1. An apparatus comprising:

interface circuitry to collect observation data from a plurality of multi-access user equipment (UE) within an environment;

a non-transitory machine-readable medium including machine-readable instructions; and

at least one processor circuit to be programmed based on the machine-readable instructions to:

determine a state of the environment based on the collected observation data;

operate a Reinforcement Learning Model (RLM) to determine one or more actions based on the determined state, the one or more actions to cause one or more of the multi-access UEs to perform one or more operations according to one or more traffic management strategies;

obtain action data associated with one or more other actions taken by one or more of the multi-access UEs;

train the RLM based on the observation data and the action data; and

deploy the RLM to the one or more of the multi-access UEs, the deployed RLM to identify one or more subsequent actions to be taken by at least one of the multi-access UEs.

2. The apparatus of claim 1 , wherein the one or more traffic management strategies include one or more of a traffic steering strategy or a traffic splitting strategy.

3. The apparatus of claim 1 , wherein the interface circuitry is to communicate the one or more actions to one or more of the multi-access UEs.

4. The apparatus of claim 1 , wherein one or more of the multi-access UEs are to use the deployed RLM to independently predict one or more of the traffic management strategies according to their individually collected observations from interacting with the environment.

5. The apparatus of claim 4 , wherein one or more of the at least one processor circuit is to:

calculate a reward value based on the collected observation data;

update the RLM based on the calculated reward value; and

distribute the updated RLM to the one or more multi-access UEs when the updated RLM passes a verification process.

6. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

trigger collection of the observation data when a trigger condition is met, the trigger condition including one or more of expiration of a timer or a measurement value meeting a threshold.

7. The apparatus of claim 1 , wherein the interface circuitry is to collect additional observation data from one or more network access nodes (NANs).

8. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

calculate, based on a reward function, a reward value based on the collected observation data.

9. The apparatus of claim 8 , wherein the reward function is a utility function of a network quality of service (QOS) target.

10. The apparatus of claim 9 , wherein inputs to the utility function include one or more QoS parameters, the one or more QOS parameters including one or more of packet delay, packet loss rate, packet drop rate, physical (PHY) rate, goodput, UE throughput, cell throughput, jitter, alpha fairness, channel quality indicator (CQI) related measurements, Modulation Coding Scheme (MCS) related measurements, physical resource block (PRB) utilization, radio utilization level per NAN, or data volume.

11. The apparatus of claim 1 , wherein the observation data includes one or more of packet delay, packet loss rate, packet drop rate, PHY rate, goodput, UE throughput, cell throughput, jitter, alpha fairness, CQI related measurements, MCS related measurements, PRB utilization, radio utilization level per NAN, or data volume.

12. The apparatus of claim 1 , wherein the apparatus corresponds to a Multi-Access Edge Computing (MEC) server/host or an Open RAN Alliance (O-RAN) Radio Access Network (RAN) intelligent controller (RIC).

13. One or more non-transitory computer readable media (NTCRM) comprising instructions to cause at least one processor circuit of a computing device to at least:

operate a representation network of a Reinforcement Learning Model (RLM) to generate a representation of a multi-access communication environment based on observation data collected from the multi-access communication environment;

operate an actor network of the RLM to determine one or more actions based on the generated representation and feedback from a critic network of the RLM, the one or more actions including one or more traffic management strategies for one or more of a plurality of multi-access user equipment (UE), and the one or more actions are to cause one or more of the multi-access UEs to perform one or more operations according to the one or more traffic management strategies;

operate the critic network to determine the feedback for the one or more actions;

communicate the one or more actions to one or more of the multi-access UEs;

obtain action data associated with one or more other actions taken by one or more of the multi-access UEs;

train the RLM based on the observation data and the action data; and

deploy the RLM to one or more of the multi-access UEs, execution of the deployed RLM to derive subsequent actions by one or more of the multi-access UEs.

14. The one or more NTCRM of claim 13 , wherein the feedback is (i) a state-value based on a state-value function or (ii) a quality value (Q value) based on a Q value function.

15. The one or more NTCRM of claim 13 , wherein the feedback includes a probability distribution over the one or more actions.

16. The one or more NTCRM of claim 13 , wherein the representation network includes a recurrent neural network (RNN) to learn the representation of the multi-access communication environment based on inputs having variable sizes.

17. The one or more NTCRM of claim 13 , wherein the critic network includes an RNN that is to learn a measurement sequence correlation to determine the feedback.

18. The one or more NTCRM of claim 13 , wherein the actor network includes an RNN that is to learn a measurement sequence correlation to determine the one or more actions.

19. The one or more NTCRM of claim 13 , wherein the representation network, the actor network, and the critic network include respective Long Short-Term Memory (LSTM) networks.

20. The one or more NTCRM of claim 13 , wherein the instructions are to cause one or more of the at least one processor circuit of the computing device to:

randomly initialize the critic network, the actor network, and the representation network with respective weight parameters;

initialize a replay buffer;

initialize a random process to perform action exploration; and

operate the representation network to obtain a representation of the observation data for a target UE to whom an action recommendation is to be provided, the target UE is among the multi-access UEs.

21. The one or more NTCRM of claim 20 , wherein the instructions are to cause one or more of the at least one processor circuit of the computing device to:

operate the actor network to select an action according to a current policy and exploration noise, and deploy the selected action to the target UE;

collect a reward based on performance of the selected action; and

cause storage of the selected action, a state, and the collected reward in the replay buffer.

22. The one or more NTCRM of claim 21 , wherein, to select the action, the instructions are to cause one or more of the at least one processor circuit of the computing device to:

randomly sample an action space for the action based on a probability; or

select the action based on a heuristic algorithm.

23. The one or more NTCRM of claim 21 , wherein the instructions are to cause one or more of the at least one processor circuit of the computing device to:

sample a minibatch of experience data from the replay buffer;

determine a temporal difference (TD) target for a TD error computation based on the collected reward; and

train the critic network to minimize a loss function that includes the TD target.

24. The one or more NTCRM of claim 23 , wherein the instructions are to cause one or more of the at least one processor circuit of the computing device to:

train the actor network based on a policy gradient to maximize a reward function, the policy gradient produced by the critic network.

25. An apparatus comprising:

interface circuitry to collect observation data from a multi-access communication environment;

a non-transitory machine-readable medium including machine-readable instructions; and

at least one processor circuit to be programmed based on the machine-readable instructions to:

determine a state of the multi-access communication environment based on the collected observation data; and

operate a Reinforcement Learning (RL) agent to determine one or more actions based on the determined state and an enforce safety action space (ESAS) procedure, the determined one or more actions including one or more traffic management strategies for one or more of a plurality of multi-access user equipment (UE), the one or more actions to cause one or more of the multi-access UEs to perform one or more operations according to the one or more traffic management strategies, and the ESAS procedure based on an acceptable range of actions or a set of constraints described by one or more linear or non-linear functions of the one or more actions.

26. The apparatus of claim 25 , wherein one or more of the at least one processor circuit is to operate the RL agent to determine the one or more actions based on the determined state, the ESAS procedure and a guided exploration procedure, the guided exploration procedure including at least one of one or more rule-based algorithms, one or more model-based heuristic algorithms, or one or more pre-trained machine learning algorithms.

27. The apparatus of claim 25 , wherein one or more of the at least one processor circuit is to operate the RL agent to determine the one or more actions based on the determined state, the ESAS procedure and an early warning procedure, the early warning procedure to:

trigger implementation of a back-up model when one or more performance metrics meet or exceed a corresponding threshold, the one or more performance metrics including one or more of an average value of one way delay for one or more data flows, an average value of end-to-end (e2e) delay for one or more data flows, a delay variation trend, buffer accumulation, a channel quality indicator (CQI) variation trend, or a modulation coding scheme (MCS) distribution trend.

28. The apparatus of claim 25 , wherein one or more of the at least one processor circuit is to operate the RL agent to determine the one or more actions based on the determined state, the ESAS procedure and an opportunistic exploration control (OEC) procedure, the OEC procedure to:

identify one or more of the one or more actions as high risk actions; and

apply the one or more high risk actions to test flows or flows with less stringent QoS targets than other flows.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2021
From: YEH, SHU-PING; BAI, JINGWEN; TALWAR, SHILPA
To: INTEL CORPORATION
Reel/Frame 057721/0339 →
Continuity (4)
Provisional Application 63165011 · Mar 23, 2021
Provisional Application 63165015 · Mar 23, 2021
Provisional Application 63164440 · Mar 22, 2021
Related Publication 20220014963A1 · Jan 13, 2022
References Cited (65)
US 20190371299A1 · Jiang · 2019 [cited by examiner]
US 20200359256A1 · Liu · 2020 [cited by examiner]
US 20210012227A1 · Fang · 2021 [cited by examiner]
US 20210119881A1 · Shirazipour · 2021 [cited by examiner]
US 20210383212A1 · Ramachandran · 2021 [cited by examiner]
US 20220121941A1 · Goldberg · 2022 [cited by examiner]
3GPP, “3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NG-RAN; Architecture description (Release 16)”, Jul. 2021, 79 pages, 3GPP TS 38.401 V16.6.0, Valbonne, France. [cited by applicant]
O-Ran Alliance, “O-RAN Operations and Maintenance Interface Specification”, Nov. 2020, 66 pages, O-RAN. WG1.O1-Interface.0-v04.00, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Operations and Maintenance Architecture”, Nov. 2020, 54 pages, O-RAN.WG1.OAM-Architecture-v04.00 2020, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Architecture Description”, Mar. 2021, 33 pages, O-RAN.WG1.O-RAN-Architecture-Description-v04.00, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Working Group 3, Near-Real-time RAN Intelligent Controller, E2 Application Protocol (E2AP)”, Jul. 2020, 84 pages, O-RAN.WG3.E2AP-v01.01, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Working Group 3 Near-Real-time RAN Intelligent Controller Architecture & E2 General Aspects and Principles”, Jul. 2020, 36 pages, O-RAN.WG3.E2GAP-v01.01, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Fronthaul Working Group Control, User and Synchronization Plane Specification”, Mar. 2021, 298 pages, O-RAN.WG4.CUS.0-v06.00, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “Cloud Architecture and Deployment Scenarios for O-RAN Virtualized RAN”, Jul. 2020, 52 pages, O-RAN.WG6.CAD-v02.01, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Working Group 3 Near-Real-time RAN Intelligent Controller E2 Service Model (E2SM) KPM”, Feb. 2020, 44 pages, ORAN-WG3.E2SM-KPM-v01.00.00, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Working Group 3 Near-Real-time RAN Intelligent Controller E2 Service Model (E2SM), RAN Function Network Interface (NI)”, Feb. 2020, 44 pages, ORAN-WG3.E2SM-NI-v01.00.00, Alfter, Germany. [cited by applicant]
O-Ran Alliance, “O-Ran Working Group 3 Near-Real-time RAN Intelligent Controller”, Feb. 2020, 27 pages, ORAN-WG3.E2SM-v01.00.00, Alfter, Germany. [cited by applicant]
Chia-Yu Chang, “Cloudification and Slicing in 5G Radio Access Network”, Networking and Internet Architecture [cs.NI], Sorbonne Université, 2018, NNT: 2018SORUS293, HAL id: tel-02501244, version 1, 227 pages (Mar. 6, 202… [cited by applicant]
Coronado et al., “Zero Touch Management: A Survey of Network Automation Solutions for 5G and 6G Networks”, IEEE Communications Survey s & Tutorials, vol. 24, No. 4, Fourth Quarter, pp. 2535-2578, 44 pages, (2022). [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; System architecture for the 5G System (5GS); Stage 2 (Release 17)”, 3GPP TS 23.501 v17.3.0, 559 pages (Dec. 23, 2021). [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects;Management and orchestration; Concepts, use cases and requirements (Release 17)” 3GPP TS 28.530 v17.2.0, 37 pages (Dec. 23, … [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Management and orchestration; Performance assurance (Release 16)”, 3GPP TS 28.550 v16.8.0, 85 pages (Sep. 23, 2021). [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Telecommunication management; Performance Management (PM); Concept and requirements (Release 16)”, 3GPP TS 32.401 v16.0.0, … [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Telecommunication management; Subscriber and equipment trace; Trace concepts and requirements (Release 17)”, 3GPP TS 32.421… [cited by applicant]
“Zero-touch network and Service Management (ZSM); Reference Architecture”, ETSI GS ZSM 002 V1.1.1, 80 pages (Aug. 2019). [cited by applicant]
X. Geng et al., “5G End-to-end Network Slice Mapping from the view of Transport Network”, IETF, draft-geng-teas-network-slice-mapping-04, 19 pages (Oct. 25, 2021), https://www.ietf.org/archive/id/draft-geng-teas-network… [cited by applicant]
A. Farrel et al., “Framework for IETF Network Slices”, IETF, draft-ietf-teas-ietf-network-slices-05, 40 pages (Oct. 25, 2021), https://www.ietf.org/archive/id/draft-ietf-teas-ietf-network-slices-05.txt. [cited by applicant]
Navarro, “Architecture”, ONAP Developer Wiki, 4 pages (Jul. 26, 2019; retrieved on Dec. 23, 2021), https://wiki.onap.org/display/DW/Architecture. [cited by applicant]
Zhao et al., “Improving Worst-Case Delay Analysis for Traffic of Additional Stream Reservation Class in Ethernet-AVB Network”, Sensors 18, No. 11: 3849, 18 pages (Nov. 9, 2018), https://www.mdpi.com/1424-8220/18/11/3849… [cited by applicant]
Hunt et al., “Enhanced Utilization Telemetry for Polling Workloads with collectd and the Data Plane Development Kit (DPDK) User Guide”, Intel Corp., 14 pages (last updated: Aug. 19, 2020) https://networkbuilders.intel.c… [cited by applicant]
“In-band Network Telemetry (INT) Dataplane Specification Version 2.1”, The P4.org Applications Working Group, 56 pages (Nov. 22, 2020), https://p4.org/p4-spec/docs/INT_v2_1.pdf. [cited by applicant]
Johnson et al., “NexRAN: Closed-loop RAN slicing in Powder—A top-to-bottom open-source open-RAN use case”, The 15th ACM Workshop on Wireless Network Testbeds, Experimental evaluation & CHaracterization (WiNTECH), pp. 17… [cited by applicant]
Milić et al., “New Concepts of Asynchronous Circuits Worst-Case Delay and Yield Estimation”, Micro Electronic and Mechanical Systems, Ch. 25, Kenichi Takahata (ed.), IntechOpen, 24 pages (Dec. 1, 2009), https://www.inte… [cited by applicant]
Yaguang Yang, “A Flow Network Model for Software Reliability Assessment”, Sixth American Nuclear Society International Topical Meeting on Nuclear Plant Instrumentation, Control, and Human-Machine Interface Technologies … [cited by applicant]
Michael J. Neely, “Opportunistic Scheduling with Worst Case Delay Guarantees in Single and Multi-Hop Networks”, Proceedings IEEE INFOCOM, 2011, pp. 1728-1736 (Apr. 10, 2011), https://ee.usc.edu/stochastic-nets/docs/wc-d… [cited by applicant]
Henry D. Pfister, “A Short Introduction to Channel Coding”, Supplemental Material for Graphical Models and Inference, 21 pages (Sep. 15, 2015), http://pfister.ee.duke.edu/courses/ece590_gmi/coding_intro.pdf. [cited by applicant]
Prof. John A. Stankovic et al., “Admission Control, Reservation, and Reflection in Operating Sytsems”, Appeared in IEEE Bulletin of the Technical Committee on Operating Systems and Application Environments (TCOS), vol. … [cited by applicant]
Kim et al., “Barefoot Networks Advanced Data-Plane Telemetry”, One Connect, 23 pages (Dec. 2018). [cited by applicant]
Chromy et al., “Admission Control Methods in IP Networks”, Advances in Multimedia, vol. 2013, Article ID 918930, 7 pages (2013), https://doi.org/10.1155/2013/918930. [cited by applicant]
Shu Fan et al., “Cross-Layer Control with Worst Case Delay Guarantees in Multihop Wireless Networks”, Journal of Electrical and Computer Engineering, vol. 2016, Article ID 5762851, 11 pages (Oct. 10, 2016), https://www.… [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; System architecture for the 5G System (5GS); Stage 2 (Release 17)”, 3GPP TS 23.501 v17.2.0, 542 pages (Sep. 24, 2021). [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Architecture for enabling Edge Applications; (Release 17)”, 3GPP TS 23.558 v17.1.0, 162 pages (Sep. 24, 2021). [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Radio Access Network; Evolved Universal Terrestrial Radio Access (E-UTRA); LTE/WLAN Radio Level Integration Using IPsec Tunnel (LWIP) encapsulation; Pro… [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Radio Access Network; Evolved Universal Terrestrial Radio Access (E-UTRA) and NR; Multi-connectivity; Stage 2 (Release 16)”, 3GPP TS 37.340 v16.7.0, 89 … [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NG-RAN; Architecture description (Release 16)”, 3GPP TS 38.401 v16.7.0, 79 pages (Oct. 1, 2021). [cited by applicant]
Kok-Lim Alvin Yau et al., “Application of Reinforcement Learning in Cognitive Radio Networks: Models and Algorithms”, 2014, 24 pages, The Scientific World Journal, vol. 2014, Article ID 209810. [cited by applicant]
Volodymyr Mnih et al, “Human-level control through deep reinforcement learning”, Feb. 26, 2015, 13 pages, vol. 518. [cited by applicant]
Julien Perez et al, “Utility-based Reinforcement Learning for Reactive Grids”, May 2008, 10 pages, The 5th IEEE International Conference on Autonomic Computing, Chicago, USA. [cited by applicant]
Miguel Perez-Enciso et al, “A Guide for Using Deep Learning for Complex Trait Genomic Prediction”, Jul. 20, 2019, 19 pages. [cited by applicant]
Max Pumperla et a., “Deep Learning and the Game of Go”, 2019, 22 pages. [cited by applicant]
David Silver et al., “Deterministic Policy Gradient Algorithms”, 2014, 9 pages, Proceedings of the 31st International Conference on Machine Learning, Beijing, China, 2014. JMLR: W&CP vol. 32. [cited by applicant]
Richard S. Sutton et al., “Reinforcement Learning: An Introduction”, Nov. 5, 2017, 445 pages, London, England. [cited by applicant]
Csaba Szepesvari, “Algorithms for Reinforcement Learning”, Jun. 9, 2009, 98 pages. [cited by applicant]
Yunhao Tang, “Introduction to Deep Learning with Tensorflow”, Feb. 25, 2019, 153 pages, Department of IEOR, Columbia University. [cited by applicant]
Huasen Wu et al., “Adaptive Exploration-Exploitation Tradeoff for Opportunistic Bandits”, Nov. 30, 2018, 18 pages, arXiv:1709.04004v2 [cs.LG]. [cited by applicant]
Timothy P. Lillicrap et al., “Continuous Control With Deep Reinforcement Learning”, Jul. 5, 2019, 14 pages, arXiv:1509.02971v6 [cs.LG], London, U.K. [cited by applicant]
John Langford et al., “The Epoch-Greedy Algorithm for Multi-armed Bandits with Side Information”, 2007, 8 pages, Advances in Neural Information Processing Systems (NIPS), vol. 20. [cited by applicant]
Vijay R. Konda et al., “Actor-Critic Algorithms”, 2000, pp. 1008-1014, Advances in Neural Information Processing Systems. [cited by applicant]
Vijaymohan Konda, “Actor-Critic Algorithms”, Jun. 2002, 147 pages. [cited by applicant]
Rodrigo Toro Icarte et al., “Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning”, Oct. 6, 2020, 31 pages, arXiv:2010.03950v1 [cs.LG]. [cited by applicant]
Gal Dalal et al., “Safe Exploration in Continuous Action Spaces”, Jan. 26, 2018, 9 pages, arXiv:1801.08757v1. [cited by applicant]
Christian Wirth et al., “Model-Free Preference-Based Reinforcement Learning”, 2016, 7 pages, Association for the Advancement of Artificial Intelligence, Germany. [cited by applicant]
Shpra Agrawal, “Reinforcement Learning: Lecture Notes, Spring 2019”, 2019, 104 pages, Columbia University, IEOR. [cited by applicant]
Yoshua Bengio et al., “Learning Long-Term Dependencies with Gradient Descent is Difficult”, Mar. 1994, pp. 157-166, IEEE Transactions on Neural Networks, vol. 5, No. 2. [cited by applicant]
Yoshua Bengio et al., “Representation Learning: A Review and New Perspectives”, Apr. 23, 2014, 30 pages, arXiv:1206.5538v3 [cs.LG]. [cited by applicant]
Cited By (13)
US 12,356,319 US 12,388,719 US 12,477,315 US 12,513,041 US 12,526,617 US 12,526,653 US 12,543,158 US 12,556,943 US 12,596,594 US 12,603,837 US 12,701,436 US 12,706,810 US 12,707,338