IP Library › Granted Patent US 12,530,241
Granted Patent B2
US 12,530,241 · App. 18/676,628 · Granted Jan 20, 2026

Cloud instance scaling method and related device thereof

Inventors: Haomin Cai (Hangzhou, CN); Rui Jing (Hangzhou, CN); Zhongkai Lei (Hangzhou, CN); Jingxiao Lu (Hangzhou, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F9/5072H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,241
App. No.
18/676,628
Granted
Jan 20, 2026
Kind
B2
Abstract

This application provides a cloud instance scaling method and a related device thereof, to ensure that a service running on a cloud instance is not interrupted when a resource quota of the cloud instance is increased or decreased. The method in this application includes: A first worker node obtains status information of a plurality of cloud instances. The first worker node determines, based on the status information, a to-be-scaled-up cloud instance from the plurality of cloud instances and a quantity of resources required for scale-up. If a quantity of idle resources of the first worker node is greater than or equal to the quantity of resources required for scale-up, the first worker node increases a resource quota of the to-be-scaled-up cloud instance based on the quantity of resources required for scale-up.

Claims (53)

1 . A cloud instance scale-up method, wherein the method is applied to a cloud service system that comprises a first worker node on which a plurality of cloud instances are deployed, and the method comprises:

obtaining, by the first worker node, status information of the plurality of cloud instances;

determining, by the first worker node based on the status information, a to-be-scaled-up cloud instance from the plurality of cloud instances and a quantity of resources required for scale-up; and

if a quantity of idle resources of the first worker node is greater than or equal to the quantity of resources required for scale-up, increasing, by the first worker node, a resource quota of the to-be-scaled-up cloud instance based on the quantity of resources required for scale-up.

2 . The method according to claim 1 , wherein the determining, by the first worker node based on the status information, a to-be-scaled-up cloud instance from the plurality of cloud instances and a quantity of resources required for scale-up comprises:

determining, by the first worker node, a cloud instance whose status information meets a preset scale-up condition as the to-be-scaled-up cloud instance; and

determining, by the first worker node, the quantity of resources required for scale-up based on status information of the to-be-scaled-up cloud instance.

3 . The method according to claim 2 , wherein the status information comprises at least one of: a resource utilization rate, a load degree, and a service success rate.

4 . The method according to claim 3 , wherein the preset scale-up condition comprises at least one of: the resource utilization rate is greater than or equal to a preset first resource utilization rate, the load degree is greater than or equal to a preset first load degree, and the service success rate is less than a preset first service success rate.

5 . The method according to claim 1 , wherein the cloud service system further comprises a second worker node, and the method further comprises:

if the quantity of idle resources of the first worker node is less than the quantity of resources required for scale-up, determining, by the first worker node, a to-be-migrated cloud instance from the plurality of cloud instances, wherein a priority of a service running on the to-be-migrated cloud instance is lower than a preset priority;

migrating, by the first worker node, the to-be-migrated cloud instance to the second worker node, and correspondingly updating the quantity of idle resources of the first worker node; and

if an updated quantity of idle resources of the first worker node is greater than or equal to the quantity of resources required for scale-up, increasing, by the first worker node, a resource quota of the to-be-scaled-up cloud instance based on the quantity of resources required for scale-up.

6 . The method according to claim 5 , wherein the cloud service system further comprises a third worker node, and the method further comprises:

if the updated quantity of idle resources of the first worker node is less than the quantity of resources required for scale-up, detecting, by the first worker node, a type of a service running on the to-be-scaled-up cloud instance;

if the service running on the to-be-scaled-up cloud instance is a stateless application, creating, by the first worker node, a new cloud instance on the third worker node, and the new cloud instance and the to-be-scaled-up cloud instance are jointly used to run the stateless application; and

if the service running on the to-be-scaled-up cloud instance is a stateful application, migrating, by the first worker node, the to-be-scaled-up cloud instance to the third worker node.

7 . The method according to claim 6 , wherein the migration is a cold migration or a hot migration.

8 . The method according to claim 1 , wherein the cloud service system further comprises a master node, and after the increasing, by the first worker node, a resource quota of the to-be-scaled-up cloud instance, the method further comprises:

sending, by the first worker node, the resource quota of the to-be-scaled-up cloud instance to the master node.

9 . A cloud instance scale-down method, wherein the method is applied to a cloud service system that comprises a first worker node on which a plurality of cloud instances are deployed, and the method comprises:

obtaining, by the first worker node, status information of the plurality of cloud instances;

determining, by the first worker node, a to-be-scaled-down cloud instance from the plurality of cloud instances based on the status information;

releasing, by the first worker node, idle resources of the to-be-scaled-down cloud instance, and determining a quantity of released idle resources of the to-be-scaled-down cloud instance; and

decreasing, by the first worker node, a resource quota of the to-be-scaled-down cloud instance based on the quantity of idle resources of the to-be-scaled-down cloud instance.

10 . The method according to claim 9 , wherein the determining, by the first worker node, a to-be-scaled-down cloud instance from the plurality of cloud instances based on the status information comprises:

determining, by the first worker node, a cloud instance whose status information meets a preset scale-down condition as a to-be-scaled-down cloud instance from the plurality of cloud instances.

11 . The method according to claim 9 , wherein the status information comprises at least one of: a resource utilization rate, a load degree, and a service success rate.

12 . The method according to claim 11 , wherein the preset scale-down condition comprises at least one of: the resource utilization rate is less than a preset second resource utilization rate, the load degree is less than a preset second load degree, and the service success rate is greater than or equal to a preset second service success rate.

13 . The method according to claim 9 , wherein the cloud service system further comprises a master node, and after the decreasing, by the first worker node, a resource quota of the to-be-scaled-down cloud instance based on the quantity of idle resources of the to-be-scaled-down cloud instance, the method further comprises:

sending, by the first worker node, the resource quota of the to-be-scaled-down cloud instance to the master node.

14 . A worker node configured to operate in a cloud service system, the worker node comprising a memory and a processor, wherein the memory stores program code, and the processor is configured to execute the program code to perform operations comprising:

obtaining status information of a plurality of cloud instances deployed on the worker node;

determining, based on the status information, a to-be-scaled-up cloud instance from the plurality of cloud instances and a quantity of resources required for scale-up; and

if a quantity of idle resources of the worker node is greater than or equal to the quantity of resources required for scale-up, increasing a resource quota of the to-be-scaled-up cloud instance based on the quantity of resources required for scale-up.

15 . The worker node according to claim 14 , wherein the determining, based on the status information, a to-be-scaled-up cloud instance from the plurality of cloud instances and a quantity of resources required for scale-up comprises:

determining a cloud instance whose status information meets a preset scale-up condition as a to-be-scaled-up cloud instance from the plurality of cloud instances; and

determining the quantity of resources required for scale-up based on status information of the to-be-scaled-up cloud instance.

16 . The worker node according to claim 14 , wherein the cloud service system further comprises a second worker node, and the processor is configured to perform operations further comprising:

if the quantity of idle resources of the first-worker node is less than the quantity of resources required for scale-up, determining a to-be-migrated cloud instance from the plurality of cloud instances, wherein a priority of a service running on the to-be-migrated cloud instance is lower than a preset priority;

migrating the to-be-migrated cloud instance to the second worker node, to update the quantity of idle resources of the worker node; and

if an updated quantity of idle resources of the worker node is greater than or equal to the quantity of resources required for scale-up, increasing a resource quota of the to-be-scaled-up cloud instance based on the quantity of resources required for scale-up.

17 . The worker node according to claim 16 , wherein the cloud service system further comprises a third worker node, and the operations further comprise:

if the updated quantity of idle resources of the worker node is less than the quantity of resources required for scale-up, detecting a type of a service running on the to-be-scaled-up cloud instance;

if the service running on the to-be-scaled-up cloud instance is a stateless application, creating a new cloud instance on the third worker node, and the new cloud instance and the to-be-scaled-up cloud instance are jointly used to run the stateless application; and

if the service running on the to-be-scaled-up cloud instance is a stateful application, migrating the to-be-scaled-up cloud instance to the third worker node.

18 . A worker node configured to operate in a cloud service system, the worker node comprises a memory and a processor, wherein the memory stores program code, and the processor is configured to execute the program code to perform operations comprising:

obtaining status information of a plurality of cloud instances deployed on the worker node;

determining a to-be-scaled-down cloud instance from the plurality of cloud instances based on the status information;

releasing idle resources of the to-be-scaled-down cloud instance, and determining a quantity of released idle resources of the to-be-scaled-down cloud instance; and

decreasing a resource quota of the to-be-scaled-down cloud instance based on the quantity of idle resources of the to-be-scaled-down cloud instance.

19 . A non-transitory computer storage medium storing one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method according to claim 1 .

20 . A non-transitory computer program product storing instructions that, when executed by a computer, cause the computer to perform the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2025
From: CAI, HAOMIN; JING, RUI; LEI, ZHONGKAI; LU, JINGXIAO
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 069831/0954 →
Priority Claims (1)
CN 202111450334.5 · Nov 30, 2021 · national
Continuity (2)
Continuation PCTCN2022134647 · Nov 28, 2022
Related Publication 20240320055A1 · Sep 26, 2024
References Cited (13)
US 12069128B2 · Einkauf · 2024 [cited by examiner]
US 12223361B2 · Boyapalle · 2025 [cited by examiner]
US 20150106520A1 · Breitgand et al. · 2015 [cited by applicant]
US 20200272526A1 · Bhole · 2020 [cited by examiner]
US 20200364086A1 · Gavali · 2020 [cited by examiner]
CN 111491006A · 2020 [cited by applicant]
CN 112783608A · 2021 [cited by applicant]
CN 113037794A · 2021 [cited by applicant]
CN 113296880A · 2021 [cited by applicant]
KR 102172607B1 · 2020 [cited by applicant]
Exploring Potential for Non-Disruptive Vertical Auto Scaling and Resource Estimation in Kubernetes. Gourav Rattihalli et al, 2019 IEEE 12th International Conference on Cloud Computing (CLOUD), total 8 pages. [cited by applicant]
Aliyun, Challenges and selection considerations encountered by Alibaba Cloud in application expansion and contraction, Dec. 1, 2020 ,https://developer.aliyun.com/article/779281, 23 pages. [cited by applicant]
International Search Report and Written Opinion issued in PCT/CN2022/134647, dated Feb. 2, 2023, 8 pages. [cited by applicant]