Hierarchical management policies for data queues
A method for improving performance of a communications network comprising: using RL to develop a hierarchical policy framework to dynamically configure local flow-control and traffic-shaping policies in a multi-queue system, each queue having QoS and a traffic dynamic that differ from other queues; establishing first and second goals with a Queue System Management (QSM) policy for a BAM policy and a LFM policy respectively, wherein The BAM and LFM policies have a collection of BAM and LFM worker policies available to choose from; choosing, with the BAM policy, one of the BAM worker policies that best enable the BAM policy to achieve the first goal; choosing, with the LFM policy, one of the LFM worker policies for each queue that best enable the LFM policy to achieve the second goal; altering the first and second goals with the QSM policy dynamically according to a surrounding environment.
1 . A method for improving performance of an intermittent, lossy communications network comprising:
using reinforcement learning to develop a hierarchical policy framework to dynamically configure local flow-control and traffic-shaping policies in a multi-queue system, wherein the multi-queue system comprises a plurality of queues, wherein each queue is configured to handle a corresponding data flow that has a quality of service (QoS) requirement and a traffic dynamic that differ from other queues in the multi-queue system;
establishing first and second goals with a Queue System Management (QSM) policy for a Bandwidth Allocation Management (BAM) policy and a Local Flow Management (LFM) policy respectively, wherein the BAM and LFM policies have a collection of predefined traffic shaping and congestion-control policies (hereinafter referred to as BAM and LFM worker policies) available to choose from, and wherein only policies at a bottom of the hierarchical policy framework interact directly with the QSM policy;
choosing, with the BAM policy, one of the BAM worker policies that best enable the BAM policy to achieve the first goal;
choosing, with the LFM policy, one of the LFM worker policies for each queue that best enable the LFM policy to achieve the second goal;
altering the first and second goals with the QSM policy dynamically according to a current state of the plurality of queues and a surrounding environment; and
wherein policies contained within the hierarchical policy framework operate at different levels of temporal resolution, with lower level policies operating at finer temporal resolutions than higher-level policies which operate at coarser temporal resolutions such that a given higher-level policy observes a coarse view of the current state that aligns with a temporal resolution at which the higher-level policy is expected to operate.
2 . The method of claim 1 , wherein individual instances of the LFM and the collection of LFM worker policies are deployed with each queue.
3 . The method of claim 2 , wherein the current state and the first and second goals are represented by respective vectors.
4 . The method of claim 3 , wherein the first goal is represented by a vector ğ:=ŭ⊙g, where ŭ is a vector that contains an average packet loss and a largest packet sojourn-time across the multi-queue system, ⊙ denotes a Hadamard product, and g defines a relative change in an average packet loss rate and the largest sojourn time that the QSM policy seeks to achieve.
5 . The method of claim 4 , wherein the second goal is represented by a vector G:=U⊙g1′ C , where C is a total number of queues in the multi-queue system, and U:=[u 1 , . . . , u C ]∈ , where is a set of real numbers.
6 . The method of claim 5 , wherein the BAM policy and the LFM policy have a reduced view of the current state when choosing the BAM worker policy and the LFM worker policy to achieve the first and second goals respectively.
7 . The method of claim 6 , wherein the BAM and LFM worker policies chosen during the choosing steps generate primitive actions that execute in a single time step.
8 . The method of claim 7 , wherein the step of using reinforcement learning to develop the hierarchical policy framework comprises;
using a hierarchical reinforcement learning (HRL) agent to observe the current state, wherein the HRL agent receives a reward or penalty depending on the current state;
updating the first and second goals based on the reward or penalty; and
observing the current state after the first and second goals have been updated.
9 . The method of claim 8 , wherein the step of using reinforcement learning comprises using a Markov Decision Process (MDP) model to describe an evolution of the BAM, LFM, BAM worker, and LFM worker policies in response to the current state.