IP Library › Granted Patent US 12,461,994
Granted Patent B2
US 12,461,994 · App. 17/636,990 · Granted Nov 4, 2025

User plane selection using reinforcement learning

Inventors: Dinand Roeland (Sollentuna, SE); Andreas Yokobori Sävö (Täby, SE); Jaeseong Jeong (Solna, SE)
Assignee: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
G06F18/217G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,461,994
App. No.
17/636,990
Granted
Nov 4, 2025
Kind
B2
Abstract

A method of reinforcement learning is used for placement of a plurality of service functions at nodes of a telecommunications network. The state of the system is defined by an allocation matrix, wherein each first vector of the allocation matrix corresponds to a respective one of the nodes of the telecommunications network, each second vector of the allocation matrix corresponds to a respective one of the plurality of service functions. Moreover, each cell of the allocation matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective second vector is placed on the one of the nodes of the telecommunications network corresponding to the respective first vector, and otherwise contains a value 0.

Claims (53)

1 . A method of reinforcement learning for placement of a plurality of service functions at nodes of a telecommunications network, the method comprising:

defining a state of a system by means of an allocation matrix, wherein the state of the system includes an indication of service functions supported in the system, wherein:

each first vector of the allocation matrix corresponds to a respective one of the nodes of the telecommunications network,

each second vector of the allocation matrix corresponds to a respective one of the plurality of service functions, and

each cell of the allocation matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective second vector is placed on the one of the nodes of the telecommunications network corresponding to the respective first vector, and otherwise contains a value 0; and

utilizing the allocation matrix as input to train a reinforcement learning (RL) agent to function as a placement algorithm for placing the plurality of service functions among the nodes of the telecommunication network.

2 . The method according to claim 1 , comprising further defining the state of the system by at least one second matrix, wherein one or more of at least one second matrix contains information specific to a family of service sets that share requirements.

3 . The method according to claim 2 , comprising further defining the state of the system by a service function type matrix, wherein:

each first vector of the service function type matrix corresponds to a respective type of service function,

each second vector of the service function type matrix corresponds to a respective one of the plurality of service functions, and

each cell of the service function type matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective second vector comprises a service function of the type corresponding to the respective first vector, and otherwise contains a value 0.

4 . The method according to claim 2 , comprising further defining the state of the system by a key performance indicator matrix, wherein:

each first vector of the key performance indicator matrix corresponds to a respective type of key performance indicator,

each second vector of the key performance indicator matrix corresponds to a respective one of the nodes of the telecommunications network, and

each cell of the key performance indicator matrix contains a value indicating a value of the corresponding key performance indicator for the corresponding one of the nodes of the telecommunications network.

5 . The method according to claim 2 , comprising further defining the state of the system by an ordering matrix, wherein:

each first vector of the ordering matrix corresponds to a respective one of the plurality of service functions,

each second vector of the ordering matrix also corresponds to a respective one of the plurality of service functions, and

each cell of the ordering matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective first vector should be traversed by data passing through the plurality of service functions before the one of the plurality of service functions corresponding to the respective second vector, and otherwise contains a value 0.

6 . The method according to claim 2 , comprising further defining the state of the system by a latency constraint matrix, wherein:

each first vector of the latency goal matrix corresponds to a respective one of the plurality of service functions,

each second vector of the latency goal matrix also corresponds to a respective latency value, and

each cell of the latency goal matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective first vector has a latency requirement corresponding to the latency value of the respective second vector, and otherwise contains a value 0.

7 . The method according to claim 2 , comprising further defining the state of the system by at least one goal matrix, wherein the at least one goal matrix contains information specific to a subset of a family of service sets.

8 . The method according to claim 7 , comprising further defining the state of the system by a latency goal matrix, wherein:

each first vector of the latency goal matrix corresponds to a respective one of the plurality of service functions,

each second vector of the latency goal matrix also corresponds to a respective latency value, and

each cell of the latency goal matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective first vector has a latency requirement corresponding to the latency value of the respective second vector, and otherwise contains a value 0.

9 . The method according to claim 7 , comprising further defining the state of the system by a co-location goal matrix, wherein:

each first vector of the co-location goal matrix corresponds to a respective one of the plurality of service functions,

each second vector of the co-location goal matrix also corresponds to a respective one of the plurality of service functions, and

each cell of the co-location goal matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective row should be co-located with the one of the plurality of service functions corresponding to the respective second vector, and otherwise contains a value 0.

10 . A method of reinforcement learning for placement of a plurality of service functions at nodes of a telecommunications network, the method comprising:

determining a plurality of possible goal matrices;

running a reward calculator for each of the possible goal matrices, to calculate a respective reward value for each of the possible goal matrices;

selecting one of the calculated reward values; and

outputting the selected one of the calculated reward values and the corresponding one of the possible goal matrices as a virtual reward and a virtual goal matrix, wherein the virtual reward and virtual goal matrix are used as input to train a reinforcement learning (RL) agent to function as a placement algorithm for placing the plurality of service functions among the nodes of the telecommunication network.

11 . The method according to claim 10 , wherein the step of selecting one of the calculated reward values comprises selecting a largest reward value of the calculated reward values.

12 . The method according to claim 10 , comprising performing the steps of running the reward calculator for each of the possible goal matrices, selecting one of the calculated reward values, and outputting the selected one of the calculated reward values as a virtual reward and a virtual goal matrix only in response to determining that a cost of running the reward calculator is below a threshold and/or that a dimension of each goal matrix is below a threshold.

13 . A computer program comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out a method according to claim 1 .

14 . A carrier containing a computer program according to claim 13 , wherein the carrier comprises one of an electronic signal, optical signal, radio signal or computer readable storage medium.

15 . A computer program product comprising non transitory computer readable media having stored thereon a computer program according to claim 13 .

16 . Apparatus for performing a method of reinforcement learning for placement of a plurality of service functions at nodes of a telecommunications network, the apparatus comprising a processor and a memory, the memory containing instructions executable by the processor such that the apparatus is operable to:

define a state of a system by an allocation matrix, wherein the state of the system includes an indication of service functions supported in the system, wherein:

each first vector of the allocation matrix corresponds to a respective one of the nodes of the telecommunications network,

each second vector of the allocation matrix corresponds to a respective one of the plurality of service functions, and

each cell of the allocation matrix contains a value 1 if the one of the plurality of service functions corresponding to the respective second vector is placed on the one of the nodes of the telecommunications network corresponding to the respective first vector, and otherwise contains a value 0; and

utilizing the allocation matrix as input to train a reinforcement learning (RL) agent to function as a placement algorithm for placing the plurality of service functions among the nodes of the telecommunication network.

17 . Apparatus for performing a method of reinforcement learning for placement of a plurality of service functions at nodes of a telecommunications network, the apparatus comprising a processor and a memory, the memory containing instructions executable by the processor such that the apparatus is operable to:

determine a plurality of possible goal matrices;

run a reward calculator for each of the possible goal matrices, to calculate a respective reward value for each of the possible goal matrices;

select one of the calculated reward values; and

output the selected one of the calculated reward values and the corresponding one of the possible goal matrices as a virtual reward and a virtual goal matrix, wherein the virtual reward and virtual goal matrix are used as input to train a reinforcement learning (RL) agent to function as a placement algorithm for placing the plurality of service functions among the nodes of the telecommunication network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2022
From: JEONG, JAESEONG; ROELAND, DINAND; YOKOBORI SÄVÖ, ANDREAS
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 059059/0001 →
Continuity (1)
Related Publication 20220358335A1 · Nov 10, 2022
References Cited (13)
US 20140200036A1 · Egner et al. · 2014 [cited by applicant]
CN 109120457A · 2019 [cited by applicant]
CN 105636062B · 2019 [cited by applicant]
CN 109743735A · 2019 [cited by applicant]
CN 109768940A · 2019 [cited by applicant]
WO 2019139510A1 · 2019 [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority for PCT International Application No. PCT/SE2019/050813 dated Jun. 1, 2020. [cited by applicant]
Tavakoli-Someh et al., “Multi-objective virtual network function placement using NSGA-II meta-heuristic approach,” The Journal of Supercomputing, (2019) 75: pp. 6451-6487. [cited by applicant]
Gai et al., “Fusion of Cognitive Wireless Networks and Edge Computing,” Technical report of Wireless Networking Group, IEEE Wireless Communications Jun. 2019, vol. 26, Issue: 3: pp. 69-75. [cited by applicant]
Florensa et al., “Automatic Goal Generation for Reinforcement Learning Agents,” Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018, 14 pages. [cited by applicant]
Communication pursuant to Rule 164(1) EPC regarding the partial supplementary European search report and provisional opinion for European Patent Application No. 19943405.1 dated Apr. 12, 2023. [cited by applicant]
Zhang et al., “Q-Placement: Reinforcement-Learning-Based Service Placement in Software-Defined Networks”, 2018 IEEE 38th International Conference on Distributed Computing Systems (ICDCS), IEEE, Jul. 2, 2018 (Jul. 2, 201… [cited by applicant]
Reynolds et al., “Provisioning Norm: An Asymmetric Quality Measure for SaaS Resource Allocation”, 2011 IEEE International Conference On Services Computing (SCC), IEEE, Jul. 4, 2011 (Jul. 4, 2011), pp. 112-119. [cited by applicant]