IP Library Granted Patent US 10,785,664
Granted Patent B2
US 10,785,664 · App. 16/222,550 · Granted Sep 22, 2020

Parameter selection for network communication links using reinforcement learning

Inventors: Sharath Ananth (Cupertino, CA); Jin Zhang (Mountain View, CA)
Assignee: LOON LLC
H04W24/02G06N3/08H04W76/10H04W84/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,785,664
App. No.
16/222,550
Granted
Sep 22, 2020
Kind
B2
Abstract

The disclosure provides a method of operating a communication network. The method includes receiving input data related to a state of the communication network and determining an implementation policy for the communication network based on the input data. The implementation policy is a set of features for forming one or more communication links in the communication network over a time interval. The one or more communication links includes at least one communication link between a terrestrial terminal and a high-altitude platform terminal. Determining the implementation policy is based at least in part on utility values of previous policies. The utility values of previous policies are derived using simulation and/or real-world implementation of the previous policies. The communication network is then operated to implement the implementation policy in the time interval.

Claims (53)

1. A method of operating a communication network that includes a plurality of nodes, the method comprising:

receiving, by one or more processors, input data related to a state of the communication network including locations of the plurality of nodes and trajectories of the plurality of nodes for a first time interval;

determining, by the one or more processors, a first implementation policy for the communication network based on the input data, the first implementation policy being a set of features for forming one or more communication links in the communication network over the first time interval, the one or more communication links including a given communication link between a terrestrial terminal and a high-altitude platform terminal;

operating, by the one or more processors, the communication network to implement the first implementation policy in the first time interval;

determining, by the one or more processors, a utility value associated with the first implementation policy as a function of a performance metric of the communication network in the first time interval;

determining, by the one or more processors, a second implementation policy for the communication network for a second time interval based at least in part on the utility value associated with the first implementation policy; and

operating, by the one or more processors, the communication network to implement the second implementation policy in the second time interval.

2. The method of claim 1 , wherein the plurality of nodes includes one or more balloons.

3. The method of claim 1 , wherein the input data also includes data related to operation of the communication network in a geographic area that includes the terrestrial terminal.

4. The method of claim 1 , wherein the set of features includes characteristics for one or more communication beams for transmission or reception at each node of the communication network.

5. The method of claim 1 , wherein the one or more processors form a neural network.

6. The method of claim 5 , wherein further comprising training, by the one or more processors, the neural network by:

receiving training input data related to state information of the communication network;

determining a training policy based on the training input data;

simulating the training policy based on internal and external influences of the communication network; and

determining a mining-utility value of the training policy according to the simulation.

7. The method of claim 1 , wherein determining the first implementation policy includes:

identifying a trend in features of policies stored in a database; and

selecting features for the first implementation policy that increase performance metric of the communication network according to the trend.

8. A method of operating a communication network that includes a plurality of nodes, the method comprising:

receiving, by one or more processors, input data related to a state of the communication network including locations of the plurality of nodes and trajectories of the plurality of nodes for a first time interval;

determining, by the one or more processors, a training policy based on the input data, the training policy being a set of features for forming a first set of communication links in the commination network over the first time interval;

simulating by the one or more processors, the training policy over the first time interval based on internal and external influences of the communication network;

determining, by the one or more processors, a utility value of the training policy as a function of a performance metric of the communication network in the simulation;

determining, by the one or more processors, a first implementation policy based at least in part on the utility value associated with the training policy, the first implementation policy being a set of features for forming a second set of communication links in the communication network over a second time interval, the second set of communication links including a given communication link between a terrestrial terminal and a high-altitude platform terminal; and

operating, by the one or more processors the communication network to implement the first implementation policy in the second time interval.

9. The method of claim 8 , further comprising:

determining, by the one or more processors, a second utility value associated with the first implementation policy as a function of a second performance metric of the communication network in the second time interval;

determining, by the one or more processors, a second implementation policy for the communication network for a third time interval based at least in part on the second utility value associated with the first implementation policy; and

operating, by the one or more processors, the communication network to implement the second implementation policy in the third time interval.

10. The method of claim 8 , wherein the input data also includes data related to operation of the communication network in a geographic area that includes the terrestrial terminal.

11. The method of claim 8 , wherein the set of features includes characteristics for one or more communication beams for transmission or reception at each node of the communication network.

12. The method of claim 8 , wherein the one or more processors form a neural network.

13. The method of claim 8 , wherein determining the training policy includes:

identifying a trend in features of policies stored in a database; and

selecting features that increase the performance metric of the communication network according to the trend.

14. The method of claim 8 , wherein determining the first implementation policy includes:

identifying a trend in features of policies stored in a database, the policies stored in the database including the training policy; and

selecting features that increase a second performance metric of the communication network according to the trend.

15. The method of claim 8 , wherein the plurality of nodes includes one or more balloons.

16. A system comprising:

a memory storing policies for a communication network, each policy being a set of features for forming one or more communication links in the communication network over a given time interval and being associated with a corresponding utility value, each utility value being a function of a performance metric of the communication network for the given time interval;

one or more processors capable of accessing the memory, the one or more processors being configured to:

receive input data related to a state of the communication network;

determine a training policy based on the input data, the training policy being a set of features for forming one or more first communication links in the communication network over a first time interval;

simulate the training policy over the first time interval based on internal and external influences of the communication network;

determine a training utility value of the training policy as a function of a training performance metric of the communication network in the simulation;

determine a first implementation policy based at least in part on the training utility value associated with the training policy, the first implementation policy being a set of features for forming one or more second communication links in the communication network over a second time interval, the one or more second communication links including a given communication link between a terrestrial terminal and a high-altitude platform terminal; and

transmit instructions to one or more nodes of the communication network, the instructions being configured to came the one or more nodes of the communication network to implement the first implementation policy in the second time interval.

17. The system of claim 16 , wherein the training policy is determined based on a trend in features of the policies stored in the memory that maximizes the training performance metric of the communication network.

18. The system of claim 16 , wherein the one or more processors are further configured to store the training policy in the memory in association with the determined utility value.

19. The system of claim 18 , wherein the first implementation policy is determined based on a trend in features of the policies stored in the memory that maximizes the performance metric of the communication network.

20. The system of claim 16 , wherein the one or more processors form a neural network.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2021
From: LOON LLC
To: SOFTBANK CORP.
Reel/Frame 056988/0485 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2020
From: ANANTH, SHARATH; ZHANG, JIN
To: LOON LLC
Reel/Frame 053550/0920 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2018
From: ANANTH, SHARATH; ZHANG, JIN
To: LOON LLC
Reel/Frame 047832/0478 →
Cited By (1)
US 12,395,865