IP Library Granted Patent US 11,234,141
Granted Patent B2
US 11,234,141 · App. 16/998,102 · Granted Jan 25, 2022

Parameter selection for network communication links using reinforcement learning

Inventors: Sharath Ananth (Cupertino, CA); Jin Zhang (Mountain View, CA)
Assignee: SoftBank Corp.
H04W24/02G06N3/08H04W76/10H04W84/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,234,141
App. No.
16/998,102
Granted
Jan 25, 2022
Kind
B2
Abstract

The disclosure provides a method of operating a communication network. The method includes receiving input data related to a state of the communication network and determining an implementation policy for the communication network based on the input data. The implementation policy is a set of features for forming one or more communication links in the communication network over a time interval. The one or more communication links includes at least one communication link between a terrestrial terminal and a high-altitude platform terminal. Determining the implementation policy is based at least in part on utility values of previous policies. The utility values of previous policies are derived using simulation and/or real-world implementation of the previous policies. The communication network is then operated to implement the implementation policy in the time interval.

Claims (53)

1. A method of operating a communication network that includes a plurality of nodes, the method comprising:

receiving, by one or more processors, input data related to a state of the communication network including locations of the plurality of nodes;

determining, by the one or more processors, a first implementation policy for the communication network based on the input data, the first implementation policy being a set of features for forming one or more communication links in the communication network;

detecting, by the one or more processors, one or more performance metrics in the communication network using the first implementation policy;

determining, by the one or more processors, a utility value associated with the first implementation policy as a function of the one or more performance metrics;

determining, by the one or more processors, a second implementation policy for the communication network based at least in part on the utility value associated with the first implementation policy; and

transmitting, by the one or more processors, instructions to the plurality of nodes for implementing the second implementation policy.

2. The method of claim 1 , wherein the plurality of nodes includes one or more high-altitude platforms.

3. The method of claim 1 , wherein the input data also includes data related to operation of the communication network in a geographic area that includes a node on a terrestrial terminal.

4. The method of claim 1 , wherein the set of features includes characteristics for one or more communication beams for transmission or reception at each node of the communication network.

5. The method of claim 1 , wherein the one or more processors form a neural network.

6. The method of claim 5 , wherein further comprising training, by the one or more processors, the neural network by:

receiving training input data related to state information of the communication network;

determining a training policy based on the training input data;

simulating the training policy based on internal and external influences of the communication network; and

determining a training utility value of the training policy according to the simulation.

7. The method of claim 1 , wherein the determining the first implementation policy includes:

identifying a trend in features of policies stored in a database; and

selecting features for the first implementation policy that increase the one or more performance metrics of the communication network relative to other performance metrics in the trend.

8. A method of operating a communication network that includes a plurality of nodes, the method comprising:

receiving, by one or more processors, input data related to a state of the communication network including locations of the plurality of nodes;

determining, by the one or more processors, a training policy based on the input data, the training policy being a set of features for forming a first set of communication links in the communication network;

simulating, by the one or more processors, the training policy based on internal and external influences of the communication network;

determining, by the one or more processors, a utility value of the training policy as a function of one or more performance metrics of the communication network in the simulation;

determining, by the one or more processors, a first implementation policy based at least in part on the utility value associated with the training policy, the first implementation policy being a set of features for forming a second set of communication links in the communication network; and

transmitting, by the one or more processors, instructions to the plurality of nodes for implementing the first implementation policy.

9. The method of claim 8 , further comprising:

determining, by the one or more processors, a second utility value associated with the first implementation policy as a function of one or more second performance metrics of the communication network;

determining, by the one or more processors, a second implementation policy for the communication network based at least in part on the second utility value associated with the first implementation policy; and

transmitting, by the one or more processors, updated instructions to the plurality of nodes for implementing the second implementation policy.

10. The method of claim 8 , wherein the input data also includes data related to operation of the communication network in a geographic area that includes a node on a terrestrial terminal.

11. The method of claim 8 , wherein the set of features includes characteristics for one or more communication beams for transmission or reception at each node of the communication network.

12. The method of claim 8 , wherein the one or more processors form a neural network.

13. The method of claim 8 , wherein determining the training policy includes:

identifying a trend in features of policies stored in a database; and

selecting features that increase the one or more performance metrics of the communication network relative to other performance metrics in the trend.

14. The method of claim 8 , wherein the determining the first implementation policy includes:

identifying a trend in features of policies stored in a database, the policies stored in the database including the training policy; and

selecting features that increase one or more second performance metrics of the communication network relative to other performance metrics in the trend.

15. The method of claim 8 , wherein the plurality of nodes includes one or more balloons.

16. A system comprising:

a memory storing policies for a communication network, each policy being a set of features for forming one or more communication links in the communication network and being associated with a corresponding utility value, each utility value being a function of one or more performance metrics of the communication network;

one or more processors capable of accessing the memory, the one or more processors being configured to:

receive input data related to a state of the communication network;

determine a training policy based on the input data, the training policy being a set of features for forming one or more first communication links in the communication network;

simulate the training policy based on internal and external influences of the communication network;

determine a training utility value of the training policy as a function of one or more training performance metrics of the communication network in the simulation;

determine a first implementation policy based at least in part on the training utility value associated with the training policy, the first implementation policy being a set of features for forming one or more second communication links in the communication network; and

transmit instructions to one or more nodes of the communication network, the instructions being configured to cause the one or more nodes of the communication network to implement the first implementation policy.

17. The system of claim 16 , wherein the training policy is determined based on a trend in features of the policies stored in the memory that maximizes the one or more training performance metrics of the communication network.

18. The system of claim 16 , wherein the one or more processors are further configured to store the training policy in the memory in association with the determined utility value.

19. The system of claim 18 , wherein the first implementation policy is determined based on a trend in features of the policies stored in the memory that maximizes the one or more performance metrics of the communication network.

20. The system of claim 16 , wherein the one or more processors form a neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2021
From: ANANTH, SHARATH; ZHANG, JIN
To: LOON LLC
Reel/Frame 057099/0951 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2021
From: LOON LLC
To: SOFTBANK CORP.
Reel/Frame 056988/0485 →
Cited By (1)
US 12,514,236