IP Library › Granted Patent US 12,262,330
Granted Patent B2
US 12,262,330 · App. 17/842,211 · Granted Mar 25, 2025

Apparatus and method for controlling transmission power based on reinforcement learning

Inventors: Ukhyeon Shin (Suwon-si, KR); Kwonyeol Park (Jeonju-si, KR)
Assignee: Samsung Electronics Co., Ltd.
H04W52/346H04W52/241H04W52/242
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,262,330
App. No.
17/842,211
Granted
Mar 25, 2025
Kind
B2
Abstract

A method of controlling transmission power for wireless communication includes obtaining detected transmission power; generating a state variable and a reward variable of a reinforcement learning model based on the detected transmission power, threshold transmission power, and a channel state; and training a reinforced learning agent based on the state variable and the reward variable to output an action variable of the reinforcement learning model representing the transmission power.

Claims (58)

1. A method of controlling transmission power for wireless communication, the method comprising:

obtaining detected transmission power;

generating a state variable and a reward variable based on the detected transmission power, a threshold transmission power, and a channel state; and

training a reinforced learning agent based on the state variable and the reward variable to output an action variable representing the transmission power,

wherein the training of the reinforced learning agent comprises generating, by the reinforced learning agent, the action variable based on the state variable and the reward variable, and

wherein the generating of the action variable comprises randomly generating the action variable with a probability ε, and greedily generating the action variable with a probability (1−ε).

2. The method of claim 1 , wherein the generating the state variable and the reward variable comprises:

calculating a transmission power residual rate of a unit period based on the threshold transmission power and the detected transmission power.

3. The method of claim 2 , wherein the generating of the state variable and the reward variable comprises:

obtaining an environment variable based on at least one communication parameter indicating the channel state; and

calculating the state variable based on the transmission power residual rate and the environment variable.

4. The method of claim 2 , wherein the generating the state variable and the reward variable comprises:

calculating the reward variable as a positive value based on the transmission power residual rate and the channel state when the transmission power residual rate is positive.

5. The method of claim 4 , wherein the calculating of the reward variable comprises:

calculating an average error rate during the unit period; and

calculating the reward variable based on the transmission power residual rate and the average error rate.

6. The method of claim 1 , wherein the training of the reinforced learning agent further comprises:

gradually reducing the probability ε.

7. The method of claim 1 , wherein the greedily generating of the action variable comprises:

setting a range of transmission power based on a transmission power of a previous unit period;

calculating a plurality of Q-values of Q-learning respectively corresponding to a plurality of transmission power candidates included in the range of transmission power;

selecting one transmission power candidate from among the transmission power candidates based on the plurality of Q-values; and

generating the action variable and updating a Q-table based on the selected transmission power candidate.

8. The method of claim 7 , wherein the range of transmission power includes the transmission power of the previous unit period.

9. The method of claim 7 , wherein the selecting the transmission power candidate comprises:

applying a weight to at least one of the plurality of transmission power candidates, the weight equal to or less than the threshold transmission power; and

selecting, as the selected transmission power candidate, a transmission power candidate corresponding to the largest sum of a weight and a Q-value from among the transmission power candidates.

10. The method of claim 1 , wherein the threshold transmission power is defined based on a specific absorption rate (SAR).

11. The method of claim 1 , further comprising:

adjusting the transmission power based on the action variable.

12. An apparatus comprising:

a memory configured to store instructions; and

at least one processor configured to communicate with the memory and, by executing the instructions, control transmission power for wireless communication,

wherein, to control the transmission power, the at least one processor is configured to

obtain detected transmission power;

generate a state variable and a reward variable based on the detected transmission power, a threshold transmission power, and a channel state; and

train a reinforced learning agent based on the state variable and the reward variable to output an action variable representing the transmission power,

wherein the at least one processor is configured to calculate a transmission power residual rate of a unit period based on the threshold transmission power and the detected transmission power to generate the state variable and the reward variable.

13. The apparatus of claim 12 , wherein, to train the reinforced learning agent, the at least one processor is further configured to:

set a range of transmission power based on a transmission power of a previous unit period,

calculate a plurality of Q-values of Q-learning respectively corresponding to a plurality of transmission power candidates included in the range of transmission power,

select one transmission power candidate from among the transmission power candidates based on the plurality of Q-values, and

generate the action variable and update a Q-table based on the selected transmission power candidate.

14. A method of controlling transmission power for wireless communication, the method comprising:

obtaining detected transmission power; and

training a reinforced learning agent, based on the detected transmission power, a threshold transmission power, and a channel state, to output an action variable representing the transmission power,

wherein the training of the reinforced learning agent comprises

setting a range of transmission power based on a transmission power of a previous unit period;

calculating a plurality of Q-values of Q-learning respectively corresponding to a plurality of transmission power candidates included in the range of transmission power;

selecting one transmission power candidate from among the plurality of transmission power candidates based on the plurality of Q-values; and

generating the action variable and updating a Q-table based on the selected transmission power candidate.

15. The method of claim 14 , wherein the range of transmission power comprises the transmission power of the previous unit period.

16. The method of claim 15 , wherein the selecting of the transmission power candidate comprises:

applying weights to at least one of the plurality transmission power candidates, the weights equal to or less than the threshold transmission power; and

selecting, as the selected transmission power candidate, a transmission power candidate corresponding to the largest sum of a weight and a Q-value from among the transmission power candidates.

17. The method of claim 14 , wherein the threshold transmission power is defined based on a specific absorption rate (SAR).

18. The method of claim 14 , further comprising:

adjusting the transmission power based on the action variable.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2022
From: SHIN, UKHYEON; PARK, KWONYEOL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060475/0350 →
Priority Claims (1)
KR 10-2021-0081038 · Jun 22, 2021 · national
Continuity (1)
Related Publication 20220408375A1 · Dec 22, 2022
References Cited (16)
US 8798664B2 · Yun · 2014 [cited by applicant]
US 9603105B2 · Yun · 2017 [cited by applicant]
US 10004048B2 · Komulainen et al. · 2018 [cited by applicant]
US 10686538B2 · Seyed et al. · 2020 [cited by applicant]
US 10826550B2 · Choi et al. · 2020 [cited by applicant]
US 10924145B2 · Mercer et al. · 2021 [cited by applicant]
US 20180175944A1 · Seyed · 2018 [cited by examiner]
US 20200142057A1 · Pendse et al. · 2020 [cited by applicant]
US 20200267662A1 · Godala · 2020 [cited by examiner]
US 20200412459A1 · Seyed et al. · 2020 [cited by applicant]
US 20210051465A1 · Koshy et al. · 2021 [cited by applicant]
US 20210055385A1 · Rimini · 2021 [cited by examiner]
US 20210067189A1 · Yu · 2021 [cited by examiner]
US 20220217645A1 · Gupta · 2022 [cited by examiner]
KR 101722237B1 · 2017 [cited by applicant]
WO WO2012122116A1 · 2012 [cited by applicant]