IP Library Granted Patent US 12,379,971
Granted Patent B2
US 12,379,971 · App. 17/586,818 · Granted Aug 5, 2025

Reliability-aware resource allocation method and apparatus in disaggregated data centers

Inventors: Chao Guo (Hong Kong, HK); Xinyu Wang (Hong Kong, HK); Moshe Zukerman (Hong Kong, HK)
Assignee: City University of Hong Kong
G06F9/5083
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,379,971
App. No.
17/586,818
Granted
Aug 5, 2025
Kind
B2
Abstract

A method for resource allocation in a disaggregated data center (DDC), comprising: a reliability model to determine an achievable reliability for a service request to the DDC; a integer linear programming (ILP) model to perform a resource allocation for the service request to the DDC such that maximizing total number of service requests received by the DDC accepted for execution is maximized, while the number of the accepted service requests allocated with backup computing resources is minimized; and a heuristic process to perform a resource allocation for the service request to the DDC such that the least reliable node of each needed computing resource type is allocated but still meeting the reliability requirement of the service request.

Claims (92)

1. A disaggregated data center (DDC), comprising:

a plurality of working nodes, each of the working nodes comprises one or more computing resources of only one computing resource type, the computing resource type selected from a group of computing resource types comprising: central processing unit (CPU), graphical processing unit (GPU), transient memory circuitry, and non-transient memory circuitry;

a plurality of backup nodes, each of the backup nodes comprising one or more computing resources of only one computing resource type, the computing resource type selected from a group of computing resource types comprising: a central processing unit (CPU), a graphical processing unit (GPU), transient memory circuitry, and non-transient memory circuitry;

a first processor configured to execute a reliability model to determine an achievable reliability for a service request to the DDC; and

a second processor configured to execute an integer linear programming (ILP) model to perform a resource allocation for the service request to the DDC;

wherein the DDC comprises multiple computing resource types, and the execution of the service request received by the DDC requires performance of computing resource of at least one of the computing resource types; and

wherein nodes of same computing resource type are configured to form a parallel system such that as long as at least one of the nodes in the parallel system is available, the parallel system is available for performance in the execution of the service request received by the DDC;

wherein each service request received by the DDC is executed by one or more working nodes corresponding to one or more necessary computing resource type respectively for an execution of the service request; and if a reliability of the one or more working nodes is lower than a reliability requirement for the service request, the service request is executed by one or more backup nodes corresponding to the one or more necessary computing resource types respectively;

wherein the performance of the ILP model comprises:

maximizing total number of service requests received by the DDC accepted for execution;

minimizing number of the accepted service requests allocated with backup nodes; and

subjecting to one or more of constraints comprising:

a working node and a backup node allocated to the service request do not share a same computing resource;

if the reliability of the one or more working nodes allocated to the service request is equal or higher than the reliability requirement for the service request, the service request is accepted for execution regardless of whether one or more backup nodes are allocated to the service request;

a total resource demand of all service requests allocated with a node and pending for execution is not higher than a resource capacity of the node; and

a reliability of computing resources of a computing resource type in the DDC must equal or higher than the reliability requirement of the service request before its acceptance for execution.

2. The disaggregated data center (DDC) of claim 1 ,

wherein the reliability model determines the achievable reliability by computing a product of reliabilities of all of the parallel systems;

wherein a reliability is a probability of normal working of a system or a node;

wherein each of the parallel systems comprises a working node r W and a backup node r B of computing resource type r;

wherein the working node r W having a reliability r W ;

wherein the backup node r B having a reliability r B ;

wherein the working node r W and the backup node r B are arranged to form the parallel system of computing resource type r; and

wherein a reliability of the parallel system of computing resource type r is obtained by computing:

1−(1− r W )·(1− r B ).

3. A method for autonomously allocating resources in the disaggregated data center (DDC) of claim 1 in order to improve reliability of the disaggregated data center (DDC), comprising:

executing a reliability model in the first processor to determine an achievable reliability for a service request to the DDC, and

executing an integer linear programming (ILP) model in the second processor to perform a resource allocation for the service request to the DDC;

wherein the DDC comprises multiple computing resource types, and the execution of the service requested received by the DDC requires performance of computing resource of at least one of the computing resource types; and

wherein nodes of same computing resource type are configured to form a parallel system such that as long as at least one of the nodes in the parallel system is available, the parallel system is available for performance in the execution of the service requested received by the DDC;

wherein each service request received by the DDC is executed by one or more working nodes corresponding to one or more necessary computing resource type respectively for an execution of the service request; and if a reliability of the one or more working nodes is lower than a reliability requirement for the service request, the service request is executed by one or more backup nodes corresponding to the one or more necessary computing resource types respectively;

wherein the performance of the ILP model comprises:

maximizing total number of service requests received by the DDC accepted for execution;

minimizing number of the accepted service requests allocated with backup nodes; and

subjecting to one or more of constraints comprising:

a working node and a backup node allocated to the service request do not share a same computing resource;

if the reliability of the one or more working nodes allocated to the service request is equal or higher than the reliability requirement for the service request, the service request is accepted for execution regardless of whether one or more backup nodes are allocated to the service request;

a total resource demand of all service requests allocated with a node and pending for execution is not higher than a resource capacity of the node; and

a reliability of computing resources of a computing resource type in the DDC must equal or higher than & the reliability requirement of the service request before its acceptance for execution.

4. The method of claim 3 ,

wherein the reliability model determines the achievable reliability by computing a product of reliabilities of all of the parallel systems;

wherein a reliability is a probability of normal working of a system or a node;

wherein each of the parallel systems comprises a working node r W and a backup node r B of computing resource type r;

wherein the working node r W having a reliability r W ;

wherein the backup node r B having a reliability r B ;

wherein the working node r W and the backup node r B are arranged to form the parallel system of computing resource type r; and

wherein a reliability of the parallel system of computing resource type r is obtained by computing:

1−(1− r W )·(1− r B ).

5. A disaggregated data center (DDC), comprising:

a plurality of working nodes, each of the working nodes comprises one or more computing resources of only one computing resource type, the computing resource type selected from a group of computing resource types comprising: central processing unit (CPU), graphical processing unit (GPU), transient memory circuitry, and non-transient memory circuitry;

a plurality of backup nodes, each of the backup nodes comprising one or more computing resources of only one computing resource type, the computing resource type selected from a group of computing resource types comprising: a central processing unit (CPU), a graphical processing unit (GPU), transient memory circuitry, and non-transient memory circuitry;

a first processor configured to execute a reliability model to determine an achievable reliability for a service request to the DDC; and

a second processor configured to execute a heuristic process to perform the resource allocation for the service request to the DDC;

wherein the DDC comprises multiple computing resource types, and the execution of the service requested received by the DDC requires performance of computing resource of at least one of the computing resource types;

wherein nodes of same computing resource type are configured to form a parallel system such that as long as at least one of the nodes in the parallel system is available, the parallel system is available for performance in the execution of the service requested received by the DDC;

wherein each service request received by the DDC is executed by one or more working nodes corresponding to one or more necessary computing resource type respectively for an execution of the service request; and if a reliability of the one or more working nodes is lower than a reliability requirement for the service request, the service request is executed by one or more backup nodes corresponding to the one or more necessary computing resource types respectively;

wherein the heuristic process comprises:

excluding one or more of the working or backup nodes for being available for execution of a service request received by the DDC, wherein the one or more excluded nodes have insufficient resource capacity to satisfy a resource demand of the service request;

sorting the working and backup nodes of each of the computing resource types necessary for an execution of the service request based on a reliability of each of the working and backup nodes of the computing resource type into lists of nodes of the computing resource types necessary for the execution of the service request;

iterating through each of the lists to find a first working or backup node with a least reliability among the working or backup nodes of each of the computing resource types necessary for the execution of the service request such that a product of the reliabilities of the first nodes of all of the computing resource types necessary for the execution of the service request is equal or higher than a reliability requirement of the service request;

if the first working or backup nodes are found such that the product of the reliabilities of the first working or backup nodes is equal or higher than a reliability requirement of the service request, the first working or backup nodes found execute the service request as its working nodes;

else if the first working or backup nodes are not found such that the product of the reliabilities of the first nodes is equal or higher than a reliability requirement of the service request, then iterating through each of the lists to find a first and second node pair with two least reliabilities among the nodes of each of the computing resource types necessary for the execution of the service request such that a product of the reliabilities of the first and second node pairs each acting as a parallel system of all of the computing resource types necessary for the execution of the service request is equal or higher than a reliability requirement of the service request.

6. A method for autonomously allocating resources in the disaggregated data center (DDC) of claim 5 in order to improve reliability of the disaggregated data center (DDC) comprising:

executing a reliability model in the first processor to determine an achievable reliability for a service request to the DDC; and

executing a heuristic process in the second processor to perform a resource allocation for the service request to the DDC,

wherein the DDC comprises multiple computing resource types, and the execution of the service request received by the DDC requires performance of computing resource of at least one of the computing resource types;

wherein nodes of same computing resource type are configured to form a parallel system such that as long as at least one of the nodes in the parallel system is available, the parallel system is available for performance in the execution of the service request received by the DDC;

wherein each service request received by the DDC is executed by one or more working nodes corresponding to one or more necessary computing resource type respectively for an execution of the service request, and if a reliability of the one or more working nodes is lower than a reliability requirement for the service request, the service request is executed by one or more backup nodes corresponding to the one or more necessary computing resource types respectively;

wherein the heuristic process comprises:

excluding one or more of the working or backup nodes for being available for execution of a service request received by the DDC, wherein the one or more excluded nodes have insufficient resource capacity to satisfy a resource demand of the service request,

sorting the working and backup nodes of each of the computing resource types necessary for an execution of the service request based on a reliability of each of the working and backup nodes of the computing resource type into lists of nodes of the computing resource types necessary for the execution of the service request,

iterating through each of the lists to find a first working or backup node with a least reliability among the working or backup nodes of each of the computing resource types necessary for the execution of the service request such that a product of the reliabilities of the first nodes of all of the computing resource types necessary for the execution of the service request is equal or higher than a reliability requirement of the service request;

if the first working or backup nodes are found such that the product of the reliabilities of the first working or backup nodes is equal or higher than a reliability requirement of the service request, the first working or backup 10 nodes found execute the service request as its working nodes:

else if the first working or backup nodes are not found such that the product of the reliabilities of the first nodes is equal or higher than a reliability requirement of the service request, then iterating through each of the lists to find a first and second node pair with two least reliabilities among the nodes of each of the computing resource types necessary for the execution of the service request such that a product of the reliabilities of the first and second node pairs each acting as a parallel system of all of the computing resource types necessary for the execution of the service request is equal or higher than a reliability requirement of the service request.

7. The method of claim 6 ,

wherein the reliability model determines the achievable reliability by computing a product of reliabilities of all of the parallel systems;

wherein a reliability is a probability of normal working of a system or a node;

wherein each of the parallel systems comprises a working node r W and a backup node r B of computing resource type r;

wherein the working node r W having a reliability r W ;

wherein the backup node r B having a reliability r B ;

wherein the working node r W and the backup node r B are arranged to form the parallel system of computing resource type r; and

wherein a reliability of the parallel system of computing resource type r is obtained by computing:

1−(1− r W )·(1− r B ).

8. The disaggregated data center (DDC) of claim 5 ,

wherein the reliability model determines the achievable reliability by computing a product of reliabilities of all of the parallel systems;

wherein a reliability is a probability of normal working of a system or a node;

wherein each of the parallel systems comprises a working node r W and a backup node r B of computing resource type r;

wherein the working node r W having a reliability r W ;

wherein the backup node r B having a reliability r B ;

wherein the working node r W and the backup node r B are arranged to form the parallel system of computing resource type r; and

wherein a reliability of the parallel system of computing resource type r is obtained by computing:

1−(1− r W )·(1− r B ).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2022
From: GUO, CHAO; WANG, XINYU; ZUKERMAN, MOSHE
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 058967/0272 →
Continuity (1)
Related Publication 20230244545A1 · Aug 3, 2023
References Cited (55)
US 9479451B1 · Wertheimer · 2016 [cited by examiner]
US 10057339B2 · Yeow · 2018 [cited by examiner]
US 10917321B2 · Schmisseur et al. · 2021 [cited by applicant]
US 20200097358A1 · Mahindru · 2020 [cited by examiner]
De Carlo, Filippo. (2013). Reliability and Maintainability in Operations Management. 10.5772/54161. (Year: 2013). [cited by examiner]
A. B. M. B. Alam, M. Zulkernine and A. Haque, “A Reliability-Based Resource Allocation Approach for Cloud Computing,” 2017 IEEE 7th International Symposium on Cloud and Service Computing (SC2), Kanazawa, Japan, 2017, pp… [cited by examiner]
Q. Zhang, Y. Cai, X. Chen, S. Angel, A. Chen, V. Liu, and B. T. Loo, “Understanding the effect of data center resource disaggregation on production DBMSs,” Proceedings of the VLDB Endowment, vol. 13, No. 9, pp. 1568-158… [cited by applicant]
G. Zervas, H. Yuan, A. Saljoghei, Q. Chen, and V. Mishra, “Optically disaggregated data centers with minimal remote memory latency: Technologies, architectures, and resource allocation,” Journal of Optical Communication… [cited by applicant]
H. M. M. Ali, T. E. El-Gorashi, A. Q. Lawey, and J. M. Elmirghani, “Future energy efficient data centers with disaggregated servers,” Journal of Lightwave Technology, vol. 35, No. 24, pp. 5361-5380, 2017. [cited by applicant]
A. Peters, G. Oikonomou, and G. Zervas, “In compute/memory dynamic packet/circuit switch placement for optically disaggregated data centers,” Journal of Optical Communications and Networking, vol. 10, No. 7, pp. B164-B1… [cited by applicant]
Y. Yan et al., “All-optical programmable disaggregated data centre network realized by FPGA-based switch and interface card,” Journal of Lightwave Technology, vol. 34, No. 8, pp. 1925-1932, 2016. [cited by applicant]
S. Angel, M. Nanavati, and S. Sen, “Disaggregation and the application,” in Proc. 12th {USENIX} Workshop on Hot Topics in Cloud Computing (HotCloud 20), 2020. [cited by applicant]
P. X. Gao et al., “Network requirements for resource disaggregation,” in Proc. 12th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 16), 2016, pp. 249-264. [cited by applicant]
A. D. Papaioannou, R. Nejabati, and D. Simeonidou, “The benefits of a disaggregated data centre: A resource allocation approach,” in Proc. 2016 IEEE Global Communications Conference (GLOBECOM), 2016, pp. 1-7. [cited by applicant]
S. Han, N. Egi, A. Panda, S. Ratnasamy, G. Shi, and S. Shenker, “Network support for resource disaggregation in next-generation datacenters,” in Proc. Proceedings of the Twelfth ACM Workshop on Hot Topics in Networks, 2… [cited by applicant]
C.-S. Li, H. Franke, C. Parris, B. Abali, M. Kesavan, and V. Chang, “Composable architecture for rack scale big data computing,” Future Generation Computer Systems, vol. 67, pp. 180-193, 2017. [cited by applicant]
C. Guo, K. Xu, G. Shen, and M. Zukerman, “Temperature-aware virtual data center embedding to avoid hot spots in data centers,” IEEE Transactions on Green Communications and Networking, 2020, p. 497-511. [cited by applicant]
W. Xia, P. Zhao, Y. Wen, and H. Xie, “A survey on data center networking (DCN): Infrastructure and operations,” IEEE Communications Surveys & Tutorials, vol. 19, No. 1, pp. 640-656, first quarter 2016. [cited by applicant]
K. Lim, J. Chang, T. Mudge, P. Ranganathan, S. K. Reinhardt, and T. F. Wenisch, “Disaggregated memory for expansion and sharing in blade servers,” ACM SIGARCH Computer Architecture News, vol. 37, No. 3, pp. 267-278, Jun… [cited by applicant]
B. Abali, R. J. Eickemeyer, H. Franke, C.-S. Li, and M. A. Taubenblatt, “Disaggregated and optically interconnected memory: When will it be cost effective?,” arXiv preprint arXiv:1503.01416, 2015. [cited by applicant]
X. Guo et al., “RDON: A rack-scale disaggregated data center network based on a distributed fast optical switch,” Journal of Optical Communications and Networking, vol. 12, No. 8, pp. 251-263, 2020. [cited by applicant]
A. Pagès, J. Perelló, F. Agraz, and S. Spadaro, “Optimal VDC service provisioning in optically interconnected disaggregated data centers,” IEEE Communications Letters, vol. 20, No. 7, pp. 1353-1356, 2016. [cited by applicant]
Intel.com, “Intel rack scale design architecture.” intel.com/content/dam/www/public/us/en/documents/white-papers/rack-scale-design-architecture-white-paper.pdf. [cited by applicant]
Z. Wei et al., “High throughput computing data center architecture.” scribd.com/document/532594570/High-Throughput-Computing-Data-Center-Architecture. [cited by applicant]
M. Bielski et al., “dReDBox: Materializing a full-stack rack-scale system prototype of a next-generation disaggregated datacenter,” in Proc. 2018 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2018,… [cited by applicant]
K. Asanovic, “Firebox: A hardware building block for 2020 warehouse-scale computers,” in Proc. {USENIX} Conference on File and Storage Technologies, 2014. [cited by applicant]
K. Lim, Y. Turner, J. R. Santos, A. AuYoung, J. Chang, P. Ranganathan, and T. F. Wenisch, “System-level implications of disaggregated memory,” in Proc. IEEE International Symposium on High-Performance Comp Architecture,… [cited by applicant]
A. Carbonari and I. Beschasnikh, “Tolerating faults in disaggregated datacenters,” in Proc. Proceedings of the 16th ACM Workshop on Hot Topics in Networks, 2017, pp. 164-170. [cited by applicant]
R. Lin, Y. Cheng, M. De Andrade, L. Wosinska, and J. Chen, “Disaggregated data centers: Challenges and trade-offs,” IEEE Communications Magazine, vol. 58, No. 2, pp. 20-26, 2020. [cited by applicant]
O. O. Ajibola, T. El-Gorashi, and J. M. Elmirghani, “Energy efficient placement of workloads in composable data center networks,” Journal of Lightwave Technology, 2021, vol. 39 , No. 10, pp. 3037-3063. [cited by applicant]
A. Pagès, R. Serrano, J. Perelló, and S. Spadaro, “On the benefits of resource disaggregation for virtual data centre provisioning in optical data centres,” Computer Communications, vol. 107, pp. 60-74, Jul. 2017. [cited by applicant]
X. Guo, F. Yan, X. Xue, G. Exarchakos, and N. Calabretta, “Performance assessment of a novel rack-scale disaggregated data center with fast optical switch,” in Proc. Optical Fiber Communications Conference and Exhibitio… [cited by applicant]
X. Guo, F. Yan, G. Exarchakos, X. Xue, B. Pan, and N. Calabretta, “On the workload deployment, resource utilization and operational cost of fast optical switch based rack-scale disaggregated data center network,” in Pro… [cited by applicant]
M. Amaral et al., “DRMaestro: Orchestrating disaggregated resources on virtualized data-centers,” Journal of Cloud Computing, vol. 10, No. 1, pp. 1-20, 2021. [cited by applicant]
A.-D. Lin, C.-S. Li, W. Liao, and H. Franke, “Capacity optimization for resource pooling in virtualized data centers with composable systems,” IEEE Transactions on Parallel and Distributed Systems, vol. 29, No. 2, pp. 3… [cited by applicant]
Y. Shan, Y. Huang, Y. Chen, and Y. Zhang, “LegoOS: A disseminated, distributed {os} for hardware resource disaggregation,” in Proc. 13th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 18), 201… [cited by applicant]
D. Rosendo et al., “Availability analysis of design configurations to compose virtual performance-optimized data center systems in next-generation cloud data centers,” Software: Practice and Experience, vol. 50, No. 6, … [cited by applicant]
Y. Lee, H. A. Maruf, M. Chowdhury, and K. G. Shin, “Mitigating the performance-efficiency tradeoff in resilient memory disaggregation,” arXiv preprint arXiv:1910.09727, 2019. [cited by applicant]
A. Verma et al., “Failure rate prediction of equipment: Can Weibull distribution be applied to automated hematology analyzers?,” Clinical Chemistry and Laboratory Medicine (CCLM), vol. 56, No. 12, pp. 2067-2071, 2018. [cited by applicant]
Intel.com, “MTBF for Intel Xeon CPU E5-2620 v4 @ 2.10ghz.” [Online]. community.intel.com/t5/Processors/MTBF-for-Intel-Xeon-CPU-E5-2620-v4-2-10GHz/td-p/539642. [cited by applicant]
O. Doguc and J. E. Ramirez-Marquez, “A generic method for estimating system reliability using Bayesian networks,” Reliability Engineering & System Safety, vol. 94, No. 2, pp. 542-550, Feb. 2009. [cited by applicant]
A. Bhardwaj et al., “NrOS: Effective replication and sharing in an operating system,” in Proc. 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21). USENIX Association, 2021, pp. 295-312. [cited by applicant]
D. A. Patterson, G. Gibson, and R. H. Katz, “A case for redundant arrays of inexpensive disks (RAID),” in Proc. Proceedings of the 1988 ACM SIGMOD International Conference on Management of Data, 1988, pp. 109-116. [cited by applicant]
J. Ansel, K. Arya, and G. Cooperman, “DMTCP: Transparent checkpointing for cluster computations and the desktop,” In Proc. 2009 IEEE International Symposium on Parallel & Distributed Processing, 2009, pp. 1-12. [cited by applicant]
Gurobi.com, “MIPGap.” gurobi.com/documentation/9.1/refman/mipgap2.html. [cited by applicant]
Y. Cheng, M. De Andrade, L. Wosinska, and J. Chen, “Resource disaggregation versus integrated servers in data centers: Impact of internal transmission capacity limitation,” in Proc. 2018 European Conference on Optical C… [cited by applicant]
A. Pagès, F. Agraz, and S. Spadaro, “On the impact of it resources disaggregation in optically interconnected data centres,” in Proc. 45th European Conference on Optical Communication (ECOC 2019), 2019, pp. 1-4. [cited by applicant]
A. Peters and G. Zervas, “Network synthesis of a topology reconfigurable disaggregated rack scale datacentre for multi-tenancy,” in Proc. Optical Fiber Communications Conference and Exhibition (OFC), 2017, pp. 1-3. [cited by applicant]
A. Pagès, F. Agraz, and S. Spadaro, “Analysis of service blocking reduction strategies in capacity-limited disaggregated datacenters,” in Proc. Optical Fiber Communication Conference, 2020, p. T3K. 2. [cited by applicant]
V. Shrivastav et al., “Shoal: A network architecture for disaggregated racks,” in Proc. 16th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 19), 2019, pp. 255-270. [cited by applicant]
M. Moralis-Pegios, N. Terzenidis, G. Mourgias-Alexandris, K. Vyrsokinos, and N. Pleros, “A low-latency high-port count optical switch with optical delay line buffering for disaggregated data centers,” in Proc. Optical I… [cited by applicant]
N. Terzenidis, M. Moralis-Pegios, G. Mourgias-Alexandris, T. Alexoudi, K. Vyrsokinos, and N. Pleros, “High-port and low-latency optical switches for disaggregated data centers: The hipoλaos switch architecture,” Journal… [cited by applicant]
N. Alachiotis et al., “dReDBox: A disaggregated architectural perspective for data centers,” Hardware Accelerators in Data Centers, pp. 35-56, 2019. [cited by applicant]
H. M. M. Ali, A. Q. Lawey, T. E. El-Gorashi, and J. M. Elmirghani, “Energy efficient disaggregated servers for future data centers,” in Proc. European Conference on Networks and Optical Communications-(NOC), 2015, pp. 1… [cited by applicant]
H. M. M. Ali, A. M. Al-Salim, A. Q. Lawey, T. El-Gorashi, and J. M. Elmirghani, “Energy efficient resource provisioning with VM migration heuristic for disaggregated server design,” in Proc. 2016 18th International Conf… [cited by applicant]