IP Library Granted Patent US 12,641,006
Granted Patent B2
US 12,641,006 · App. 17/315,167 · Granted May 26, 2026

Dynamic expansion and contraction of edge clusters for managing access to cloud-based applications

Inventors: Santosh Ghanshyam Pandey (Fremont, CA); Sidhesh Divekar (Milpitas, CA); Linus Aranha (Los Gatos, CA)
Assignee: Palo Alto Networks, Inc.
H04L45/124H04L43/0852H04L45/121H04L45/126H04L45/74H04L63/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,641,006
App. No.
17/315,167
Granted
May 26, 2026
Kind
B2
Abstract

Edge clusters execute in a plurality of regional clouds of a cloud computing platforms, which may include cloud POPs. Edge clusters may be programmed to control access to applications executing in the cloud computing platform. Edge clusters and an intelligent routing module route traffic to applications executing in the cloud computing platform. Cost and latency may be managed by the intelligent routing module by routing requests over the Internet or a cloud backbone network and using or bypassing cloud POPs. The placement of edge clusters may be selected according to measured or estimated latency. Latency may be estimated using speed test servers and the locations of speed test servers may be verified.

Claims (37)

1 . A method comprising:

providing an application instance executing on a cloud computing platform including a plurality of regional clouds each associated with a geographic region and connected to one another by a cloud backbone network and by a wide area network WAN that is external to the cloud computing platform and does not include the cloud backbone network, the application instance executing in either a first regional cloud of the plurality of regional clouds or a computing device connected to the first regional cloud;

providing a first plurality of edge clusters having a first configuration including a first number of the first plurality of edge clusters and first locations of the first plurality of edge clusters in the plurality of regional clouds;

processing, by the first plurality of edge clusters, requests to access the application instance from a plurality of user endpoints in the WAN;

obtaining, by a computer module, a plurality of L1 values corresponding to latency between the plurality of user endpoints and the plurality of regional clouds;

obtaining, by the computer module, a plurality of L2 values corresponding to latency between pairs of regional clouds of the plurality of regional clouds; and

selecting, by the computer module, a second configuration of a second plurality of edge clusters according to the plurality of L1 values, the plurality of L2 values, and quantities of the requests received from across the geographic regions associated with the plurality of regional clouds, the second configuration including a second number of the second plurality of edge clusters and second locations of the first plurality of edge clusters, wherein one or both of (a) the second number is different from the first number and (b) the second locations are not all identical to the first locations;

wherein selecting the second configuration comprises executing an optimization algorithm with respect to the plurality of L1 values, the plurality of L2 values, and the quantities of the requests received from the geographic regions associated with the plurality of regional clouds;

wherein executing the optimization algorithm comprises executing the optimization algorithm subject to bounds to obtain the second configuration; and

wherein the bounds include at least one of a maximum increase of the second number relative to the first number and a maximum decrease of the second number relative to the first number.

2 . The method of claim 1 , further comprising outputting, by the computer module, a recommendation to implement the second configuration.

3 . The method of claim 1 , further comprising executing the optimization algorithm by evaluating each configuration of a plurality of configurations according to a cost function that is a function of the plurality of L1 values, the plurality of L2 values, the quantities of the requests received from the geographic regions associated with the plurality of regional clouds, and monetary costs for network usage and computational usage associated with each configuration.

4 . The method of claim 1 , wherein the bounds include a maximum permitted increase in latency experienced by the plurality of user endpoints relative to the first configuration.

5 . The method of claim 1 , further comprising characterizing cacheability of responses provided by the application instance in response to the requests and scaling down the plurality of L2 values according to the cacheability such that an amount of scaling down increases with increase in the cacheability.

6 . The method of claim 1 , further comprising scaling each L1 value according to one or both of a quantity of the requests received and sizes of the requests received from a geographic region corresponding to each L1 value.

7 . The method of claim 1 , further comprising:

evaluating the second configuration according to one or more validation criteria; and

determining that the second configuration satisfies the one or more validation criteria; and

in response to determining that the second configuration satisfies the one or more validation criteria, one or both of (a) outputting a report of the second configuration and (b) implementing the second configuration.

8 . The method of claim 7 , wherein the one or more validation criteria are a minimum improvement in overall latency of the second configuration relative to the first configuration.

9 . A system comprising:

a cloud computing platform comprising a plurality of regional clouds connected by a cloud backbone network, each regional cloud comprising a plurality of computing devices associated with a geographic region and connected by a regional network, the plurality of regional clouds being further connected to a wide area network (WAN) that does not include the cloud backbone network, an application instance executing in a first regional cloud of the plurality of regional clouds;

a first plurality of edge clusters having a first configuration including a first number of the first plurality of edge clusters and first locations of the first plurality of edge clusters in the plurality of regional clouds, the first plurality of edge clusters being configured to process requests to access the application instance from a plurality of user endpoints in the WAN; and

a computer module programmed to:

obtain a plurality of L1 values corresponding to latency between the plurality of user endpoints and the plurality of regional clouds;

obtain a plurality of L2 values corresponding to latency between pairs of regional clouds of the plurality of regional clouds; and

select a second configuration of a second plurality of edge clusters according to the plurality of L1 values, the plurality of L2 values, and quantities of the requests received from the geographic regions associated with the plurality of regional clouds, the second configuration including a second number of the second plurality of edge clusters and second locations of the first plurality of edge clusters, wherein one or both of (a) the second number is different from the first number and (b) the second locations are not all identical to the first locations;

wherein the computer module is further programmed to select the second configuration by executing an optimization algorithm with respect to the plurality of L1 values, the plurality of L2 values, and the quantities of the requests received from the geographic regions associated with the plurality of regional clouds; and

wherein the computer module is further programmed to execute the optimization algorithm by executing the optimization algorithm subject to a maximum increase of the second number relative to the first number and a maximum decrease of the second number relative to the first number.

10 . The system of claim 9 , wherein the computer module is further programmed to execute the optimization algorithm by evaluating each configuration of a plurality of configurations according to a cost function that is a function of the plurality of L1 values, the plurality of L2 values, the quantities of the requests received from the geographic regions associated with the plurality of regional clouds, and monetary costs for network usage and computational usage associated with each configuration.

11 . The system of claim 9 , wherein the computer module is further programmed to execute the optimization algorithm by executing the optimization algorithm subject to a maximum permitted increase in latency experienced by the plurality of user endpoints relative to the first configuration.

12 . The system of claim 9 , wherein the computer module is further programmed to characterize cacheability of responses provided by the application instance in response to the requests and scale down the plurality of L2 values according to the cacheability such that an amount of scaling down increases with increase in the cacheability.

13 . The system of claim 9 , wherein the computer module is further programmed to scale each L1 value according to a quantity of the requests received from a geographic region corresponding to each L1 value.

14 . The system of claim 9 , wherein the computer module is further programmed to:

evaluate the second configuration according to one or more validation criteria; and

if the second configuration satisfies the one or more validation criteria, one or both of (a) output a report of the second configuration and (b) implement the second configuration;

wherein the one or more validation criteria are a minimum improvement in overall latency of the second configuration relative to the first configuration.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: PROSIMO INC.
To: PALO ALTO NETWORKS, INC.
Reel/Frame 071425/0477 →
RELEASE OF SECURITY INTEREST Recorded Jan 24, 2025
From: FIRST-CITIZENS BANK & TRUST COMPANY
To: PROSIMO INC.
Reel/Frame 069996/0300 →
CORRECTIVE ASSIGNMENT TO CORRECT THE NAME OF THE RECEIVING PARTY PREVIOUSLY RECORDED ON REEL 56176 FRAME 806. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 15, 2025
From: PANDEY, SANTOSH GHANSHYAM; DIVEKAR, SIDHESH; ARANHA, LINUS
To: PROSIMO INC.
Reel/Frame 069933/0487 →
SECURITY INTEREST Recorded Sep 27, 2024
From: PROSIMO INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 068719/0171 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2021
From: PANDEY, SANTOSH GHANSHYAM; DIVEKAR, SIDHESH; ARANHA, LINUS
To: PROSIMO INC
Reel/Frame 056176/0806 →
Continuity (2)
Continuation In Part 17127876 · Dec 18, 2020
Related Publication 20220201673A1 · Jun 23, 2022
References Cited (33)
US 7904541B2 · Swildens et al. · 2011 [cited by applicant]
US 8949459B1 · Scholl · 2015 [cited by examiner]
US 9391856B2 · Kazerani et al. · 2016 [cited by applicant]
US 10218781B2 · Byers et al. · 2019 [cited by applicant]
US 10887276B1 · Parulkar et al. · 2021 [cited by applicant]
US 11070453B2 · Thiagarajan et al. · 2021 [cited by applicant]
US 11095534B1 · Dunsmore et al. · 2021 [cited by applicant]
US 11128597B1 · Johnson et al. · 2021 [cited by applicant]
US 11394636B1 · Walker · 2022 [cited by examiner]
US 20120124194A1 · Shouraboura · 2012 [cited by applicant]
US 20150188823A1 · Williams et al. · 2015 [cited by applicant]
US 20160119279A1 · Maslak et al. · 2016 [cited by applicant]
US 20160191600A1 · Scharber et al. · 2016 [cited by applicant]
US 20170310709A1 · Foxhoven et al. · 2017 [cited by applicant]
US 20170317954A1 · Masurekar et al. · 2017 [cited by applicant]
US 20200076685A1 · Vaidya et al. · 2020 [cited by applicant]
US 20200099659A1 · Cometto et al. · 2020 [cited by applicant]
US 20200162386A1 · Radlein · 2020 [cited by examiner]
US 20210075729A1 · Fedorov · 2021 [cited by examiner]
US 20210314291A1 · Chandrashekhar et al. · 2021 [cited by applicant]
US 20210328893A1 · Cherkas et al. · 2021 [cited by applicant]
US 20220060431A1 · Vadayadiyil Raveendran · 2022 [cited by examiner]
US 20220200957A1 · Prabagaran et al. · 2022 [cited by applicant]
US 20220377131A1 · Szilagyi et al. · 2022 [cited by applicant]
US 20240380654A1 · Zaicenko et al. · 2024 [cited by applicant]
Li et al., “Internet Anycast: Performance, Problems, & Potential”, Aug. 25, 2018, SIGCOMM '18, Aug. 20-25, 2018, Budapest, Hungary, ACM ISBN 978-1-4503-5567-4/18/08, https://doi.org/10.1145/3230543.3230547, pp. 1-15. (Y… [cited by examiner]
Yu et al., A Survey on the Edge Computing for the Internet of Things, Nov. 29, 2017, Digital Object Identifier 10.1109/ACCESS.2017.2778504, vol. 6, 2018, pp. 6900-6919. (Year: 2017). [cited by examiner]
U.S. Appl. No. 17/127,876, Notice of Allowance mailed Oct. 15, 2025, 6 pages. [cited by applicant]
U.S. Appl. No. 17/127,876, Non-Final Office Action, mailed Nov. 3, 2022, 16 pages. [cited by applicant]
U.S. Appl. No. 17/315,175, Non-Final Office Action, mailed Dec. 22, 2022, 25 pages. [cited by applicant]
U.S. Appl. No. 17/315,192, Non-Final Office Action, mailed Dec. 21, 2022, 13 pages. [cited by applicant]
U.S. Appl. No. 18/530,458, Non-Final Office Action mailed Jul. 15, 2025, 10 Pages. [cited by applicant]
U.S. Appl. No. 17/127,876 Final Office Action, mailed May 8, 2023, 16 pages. [cited by applicant]