IP Library Granted Patent US 12,732,864
Granted Patent B2
US 12,732,864 · App. 17/965,294 · Granted Sep 8, 2026

Coordinated load balancing in mobile edge computing network

Inventors: Di Wu (Saint-Laurent, CA); Manyou Ma (Vancouver, CA); Yi Tian Xu (Mount Royal, CA); Jimmy Li (Longueuil, CA); Seowoo Jang (Seoul, KR); Xue Liu (Montreal, CA); Gregory Lewis Dudek (West Mount, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H04W28/0925H04W28/0226
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,732,864
App. No.
17/965,294
Granted
Sep 8, 2026
Kind
B2
Abstract

A method includes obtaining at least one policy parameter of a neural network corresponding to a load balancing policy, receiving trajectories for each mobile device in a plurality of mobile devices of the wireless network, each trajectory corresponding to a sequence of states of a respective mobile device, wherein the sequence of states is generated based on a continuous interaction of an existing policy of the respective mobile device with the wireless network, estimating advantage functions for each mobile device in the plurality of mobile devices based on the trajectories for each respective mobile device, and updating the at least one policy parameter based on the estimated advantage functions such that the load balancing policy is determined based on states of each mobile device in the plurality of mobile devices.

Claims (47)

1 . A method comprising:

obtaining at least one policy parameter of a neural network corresponding to a load balancing policy;

receiving trajectories for each mobile device in a plurality of mobile devices of the wireless network, each trajectory corresponding to a sequence of states, actions and rewards for the actions of a respective mobile device, wherein the sequence of states is generated based on a continuous interaction of an existing policy of the respective mobile device with the wireless network, the actions indicating a base station to which the mobile device should connect for a request and the rewards being received after the actions guide a learning process;

estimating advantage functions for each mobile device in the plurality of mobile devices based on the trajectories for each respective mobile device; and

updating the at least one policy parameter based on the estimated advantage functions such that the load balancing policy is determined based on states of each mobile device in the plurality of mobile devices,

wherein the advantage functions are determined based on a difference between a cost-to-go function and a value function, and

wherein the sequence of states of each trajectory corresponds to states over a predetermined number of time steps for each mobile device of the plurality of mobile devices.

2 . The method of claim 1 , further comprising:

obtaining at least one value parameter of the neural network corresponding to the load balancing policy; and

updating the at least one value parameter based on the estimated advantage functions.

3 . The method of claim 1 , further comprising deploying the neural network corresponding to the load balancing policy to each mobile device of the plurality of mobile devices in the wireless network.

4 . The method of claim 1 , further comprising:

receiving, as a first input to the neural network corresponding to the load balancing policy, statuses of queues of each base station of a plurality of base stations in the wireless network; and

receiving, as a second input to the neural network corresponding to the load balancing policy, a task request from a first mobile device of the plurality of mobile devices.

5 . The method of claim 4 , further comprising determining a base station of the plurality of base stations for performing the requested task based on the first input and the second input, and

performing a handover operation connecting the first mobile device to the determined base station for performing the requested task.

6 . The method of claim 1 , wherein the wireless network comprising a mobile edge computing (MEC) network.

7 . A system comprising:

a memory storing instructions; and

a processor configured to execute the instructions to:

obtain at least one policy parameter of a neural network corresponding to a load balancing policy;

receive trajectories for each mobile device in a plurality of mobile devices of a mobile edge computing (MEC) network, each trajectory corresponding to a sequence of states, actions and rewards for the actions of a respective mobile device, the actions indicating a base station to which the mobile device should connect for a request and the rewards being received after the actions guide a learning process, wherein the sequence of states is generated based on a continuous interaction of an existing policy of the respective mobile device with the MEC network;

estimate advantage functions for each mobile device in the plurality of mobile devices based on the trajectories for each respective mobile device; and

update the at least one policy parameter based on the estimated advantage functions such that the load balancing policy is determined based on states of each mobile device in the plurality of mobile devices,

wherein the advantage functions are determined based on a difference between a cost-to-go function and a value function, and

wherein the sequence of states of each trajectory corresponds to states over a predetermined number of time steps for each mobile device of the plurality of mobile devices.

8 . The system of claim 7 , wherein the processor is further configured to execute the instructions to:

obtain at least one value parameter of the neural network corresponding to the load balancing policy; and

update the at least one value parameter based on the estimated advantage functions.

9 . The system of claim 7 , wherein the processor is further configured to execute the instructions to deploy the neural network corresponding to the load balancing policy to each mobile device of the plurality of mobile devices in the MEC network.

10 . The system of claim 7 , wherein the processor is further configured to execute the instructions to:

receive, as a first input to the neural network corresponding to the load balancing policy, statuses of queues of each base station of a plurality of base stations in the MEC network; and

receive, as a second input to the neural network corresponding to the load balancing policy, a task request from a first mobile device of the plurality of mobile devices.

11 . The system of claim 10 , wherein the processor is further configured to execute the instructions to determine a base station of the plurality of base stations for performing the requested task based on the first input and the second input, and

perform a handover operation connecting the first mobile device to the determined base station for performing the requested task.

12 . The system of claim 11 , wherein the base station for performing the requested task with the first mobile device is determined at the first mobile device.

13 . A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to:

obtain at least one policy parameter of a neural network corresponding to a load balancing policy;

receive trajectories for each mobile device in a plurality of mobile devices of a mobile edge computing (MEC) network, each trajectory corresponding to a sequence of states, actions and rewards for the actions of a respective mobile device, the actions indicating a base station to which the mobile device should connect for a request and the rewards being received after the actions guide a learning process, wherein the sequence of states is generated based on a continuous interaction of an existing policy of the respective mobile device with the MEC network;

estimate advantage functions for each mobile device in the plurality of mobile devices based on the trajectories for each respective mobile device; and

update the at least one policy parameter based on the estimated advantage functions such that the load balancing policy is determined based on states of each mobile device in the plurality of mobile devices,

wherein the advantage functions are determined based on a difference between a cost-to-go function and a value function, and

wherein the sequence of states of each trajectory corresponds to states over a predetermined number of time steps for each mobile device of the plurality of mobile devices.

14 . The storage medium of claim 13 , wherein the instructions, when executed, further cause the at least one processor to:

obtain at least one value parameter of the neural network corresponding to the load balancing policy; and

update the at least one value parameter based on the estimated advantage functions.

15 . The storage medium of claim 13 , wherein the instructions, when executed, further cause the at least one processor to deploy the neural network corresponding to the load balancing policy to each mobile device of the plurality of mobile devices in the MEC network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: WU, DI; MA, MANYOU; XU, YI TIAN; LI, JIMMY; JANG, SEOWOO; LIU, XUE; DUDEK, GREGORY LEWIS
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061415/0244 →
Continuity (2)
Provisional Application 63278984 · Nov 12, 2021
Related Publication 20230156520A1 · May 18, 2023
References Cited (24)
US 5950134A · Agrawal · 1999 [cited by examiner]
US 10440096B2 · Sabella et al. · 2019 [cited by applicant]
US 10846594B2 · Choque · 2020 [cited by examiner]
US 11558276B2 · Sivaraj · 2023 [cited by examiner]
US 11703853B2 · Hong · 2023 [cited by examiner]
US 20030017831A1 · Lee · 2003 [cited by examiner]
US 20060268783A1 · Julian · 2006 [cited by examiner]
US 20120236712A1 · Park · 2012 [cited by examiner]
US 20130343281A1 · Bakker · 2013 [cited by examiner]
US 20150156122A1 · Singh · 2015 [cited by examiner]
US 20170086121A1 · Kaushik · 2017 [cited by examiner]
US 20190138934A1 · Prakash · 2019 [cited by examiner]
US 20210117249A1 · Doshi · 2021 [cited by examiner]
US 20210282222A1 · Kapoor · 2021 [cited by examiner]
US 20220116811A1 · Li · 2022 [cited by examiner]
US 20220167236A1 · Melodia · 2022 [cited by examiner]
US 20220400532A1 · Kalkunte · 2022 [cited by examiner]
US 20230068386A1 · Akdeniz · 2023 [cited by examiner]
US 20240064105A1 · Sharma · 2024 [cited by examiner]
US 20240306049A1 · Yan · 2024 [cited by examiner]
EP 3869847A1 · 2021 [cited by examiner]
WO WO2013091341A1 · 2013 [cited by examiner]
WO WO2024256902A1 · 2024 [cited by examiner]
WO WO2025099302A1 · 2025 [cited by examiner]