IP Library › Granted Patent US 12,566,629
Granted Patent B2
US 12,566,629 · App. 17/972,570 · Granted Mar 3, 2026

Server and a resource scheduling method for use in a server

Inventors: Chao Guo (Hong Kong, HK); Moshe Zukerman (Hong Kong, HK); Tianjiao Wang (Hong Kong, HK)
Assignee: City University of Hong Kong
G06F9/4881G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,629
App. No.
17/972,570
Granted
Mar 3, 2026
Kind
B2
Abstract

A server and a resource scheduling method for use in a server. The server comprises a plurality of processing modules each having predetermined resources for processing tasks handled by the server, wherein the plurality of processing modules are interconnected by communication links forming a network of processing modules having a Disaggregated Data Centers (DCC) architecture; a DCC hardware monitor arranged to detect hardware information associated with the network of processing modules during an operation of the server, and a task scheduler module arranged to analysis a resource allocation request associated with each respective task and the hardware information, and to facilitate processing of the task by one or more of the processing modules selected based on the analysis.

Claims (41)

1 . A server comprising:

a plurality of processing modules each having predetermined resources for processing tasks handled by the server, wherein the plurality of processing modules are interconnected by communication links forming a network of processing modules having a Disaggregated Data Centers (DDC) architecture;

a scheduler module for scheduling and migrating tasks based on hardware information associated with the processing modules;

a monitor module for detecting topology information and load changes, including failure and repair of processing modules, and report the topology information and load changes to the scheduler module;

a central processing unit (CPU) and a memory storing instructions that when executed by the CPU, cause the CPU to:

by the scheduler module:

detect hardware information associated with the network of processing modules during an operation of the server;

analyze a resource allocation requests associated with a plurality of tasks and the hardware information; and

provide, based on the detection and analysis, a scheduler decision, wherein providing a scheduler decision includes for each task, allocating more than one processing modules in the network to handle the task based on resource availability of the plurality of processing modules and minimizing the blocking probability and number of task failing to complete their service while optimizing physical inter-resource traffic paths based on inter-resource traffic demand among different processing modules involved in handling the respective task of the plurality of tasks, wherein the blocking probability is one minus an acceptance ratio; and

based on changes to the hardware information, modifying the scheduler decision to:

for each task affected by the received change, reallocating more than one processing modules in the network to handle the task based on resource availability of the plurality of processing modules that have not experienced a failure or have been repaired and optimizing physical inter-resource traffic paths based on inter-resource traffic demand among different processing modules involved in handling the respective task of the plurality of tasks;

by the monitor module:

monitor for changes associated with the hardware information, wherein the hardware information includes topology of the DDC architecture, loading of each of the plurality of processing modules, and information related to failure and/or repairing of the network of processing modules; and

based on detected changes associated with the hardware information, notifying the scheduler module of the changes to perform a modification to the scheduler decision.

2 . The server in accordance with claim 1 , wherein the central processing unit (CPU) is further caused to analyze multiple resource allocation requests in a static scenario where the resource allocation requests are serviced in batches.

3 . The server in accordance with claim 2 , wherein the central processing unit (CPU) is further caused to analyze the multiple resource allocation requests based on a mixed-integer linear programming (MILP) method.

4 . The server in accordance with claim 3 , wherein the mixed-integer linear programming (MILP) method includes solving a MILP problem with varied weights in an objective function associated with a single-objective problem with a weighted sum, wherein the single-objective problem is converted from a multi-objective problem associated with multiple resource allocation requests in the static scenario.

5 . The server in accordance with claim 1 , wherein the centrals processing unit (CPU) is further caused to analyze multiple resource allocation requests in a dynamic scenario where the resource allocation requests sequentially and each resource allocation request is served over a predetermined period of time.

6 . The server in accordance with claim 5 , wherein the centrals processing unit (CPU) is further caused to schedule the resource allocation requests arriving in the dynamic scenario, based on the following conditions:

accepting the resource allocation request if sufficient resources are available upon arrival of the request; or

blocking the resource allocation request such that the request leaves the system without re-attempting;

wherein the resources are provided by a single processing module or a group of two or more processing modules involving inter-resource traffic.

7 . The server in accordance with claim 6 , wherein the centrals processing unit (CPU) is further caused to restore an accepted request being interrupted by hardware failure associated with resources allocated for handling the task, by excluding the processing module with hardware failure from the topology of the DDC architecture, and re-allocating resources for handling the accepted request.

8 . The server in accordance with claim 1 , wherein each of the plurality of processing modules includes plurality components of different types of resources, and wherein a single task including a request of more than one resource type is arranged to be processed by components of different types of resources in the plurality of processing modules in a disaggregated manner.

9 . A resource scheduling method for use in a server, wherein the server comprises a plurality of processing modules each having predetermined resources for processing tasks handled by the server, wherein the plurality of processing modules are interconnected by communication links forming a network of processing modules having a Disaggregated Data Centers (DDC) architecture; a scheduler module for scheduling and migrating tasks based on hardware information associated with the processing modules; and a monitor module for detecting topology information and load changes, including failure and repair of processing modules, and report the topology information and load changes to the scheduler module, the method comprising the steps of:

detecting, by the scheduler module, hardware information associated with the network of processing modules during an operation of the server;

analyzing, by the scheduler module, a resource allocation request a plurality of tasks and the hardware information; and

providing, by the scheduler module, based on the detection and analysis, a scheduler decision, wherein providing a scheduler decision include for each task, allocating more than one processing modules in the network to handle the task based on resource availability of the plurality of processing modules and minimizing the blocking probability and number of task failing to complete their service while optimizing physical inter-resource traffic paths based on inter-resource traffic demand among different processing modules involved in handling the respective task of the plurality of tasks, wherein the blocking probability is one minus an acceptance ratio;

subsequent to the providing, monitoring, by the monitoring module, for changes associated with the hardware information, wherein the hardware information includes topology of the DDC architecture, loading of each of the plurality of processing modules, and information related to failure and/or repairing of the network of processing modules;

based on detected changes associated with the hardware information, notifying, by the monitoring module, the scheduler module of the changes to perform a modification to the scheduler decision; and

based on changes to the hardware information, modifying the scheduler decision to: for each task affected by the received change, reallocating more than one processing modules in the network to handle the task based on resource availability of the plurality of processing modules that have not experienced a failure or have been repaired and optimizing physical inter-resource traffic paths based on inter-resource traffic demand among different processing modules involved in handling the respective task of the plurality of tasks.

10 . The resource scheduling method in accordance with claim 9 , wherein the step of analyzing the resource allocation request associated with each respective task and the hardware information comprising the step of analyzing multiple resource allocation requests in a static scenario where the resource allocation requests are serviced in batches.

11 . The resource scheduling method in accordance with claim 10 , wherein the multiple resource allocation requests are analyzed based on a mixed-integer linear programming (MILP) method.

12 . The resource scheduling method in accordance with claim 11 , wherein the mixed-integer linear programming (MILP) method includes solving a MILP problem with varied weights in an objective function associated with a single-objective problem with weighted sum, wherein the single-objective problem is converted from a multi-objective problem associated with multiple resource allocation requests in the static scenario.

13 . The resource scheduling method in accordance with claim 9 , wherein the step of analyzing the resource allocation request associated with each respective task and the hardware information comprising the step of analyzing multiple resource allocation requests in a dynamic scenario where the resource allocation requests are served sequentially and each resource allocation request is served over a predetermined period of time.

14 . The resource scheduling method in accordance with claim 13 , wherein the step of analyzing the resource allocation request associated with each respective task and the hardware information comprising the step of scheduling the resource allocation requests arriving in the dynamic scenario, based on the following conditions:

accepting the resource allocation request if sufficient resources are available upon arrival of the request; or

blocking the resource allocation request such that the request leaves the system without re-attempting;

wherein the resources are provided by a single processing module or a group of two or more processing modules involving inter-resource traffic.

15 . The resource scheduling method in accordance with claim 14 , wherein the step of analyzing the resource allocation request associated with each respective task and the hardware information comprising the step of restoring an accepted request being interrupted by hardware failure associated with resources allocated for handling the task, by excluding the processing module with hardware failure from the topology of the DDC architecture, and re-allocating resources for handling the accepted request.

16 . The resource scheduling method in accordance with claim 11 , wherein each of the plurality of processing modules includes plurality components of different types of resources, and wherein a single task including a request of more than one resource type is arranged to be processed by components of different types of resources in the plurality of processing modules in a disaggregated manner.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2022
From: GUO, CHAO; ZUKERMAN, MOSHE; WANG, TIANJIAO
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 061522/0132 →
Continuity (1)
Related Publication 20240192987A1 · Jun 13, 2024
References Cited (69)
US 9479451B1 · Wertheimer · 2016 [cited by examiner]
US 9760527B2 · Egi et al. · 2017 [cited by applicant]
US 10841367B2 · Bivens · 2020 [cited by examiner]
US 10917321B2 · Schmisseur et al. · 2021 [cited by applicant]
US 10936374B2 · Bivens · 2021 [cited by examiner]
US 10977085B2 · Bivens · 2021 [cited by examiner]
US 11221886B2 · Bivens · 2022 [cited by examiner]
US 20070234363A1 · Ferrandiz · 2007 [cited by examiner]
US 20130297770A1 · Zhang · 2013 [cited by examiner]
US 20160380921A1 · Blagodurov · 2016 [cited by examiner]
US 20170293537A1 · Sun · 2017 [cited by examiner]
US 20180150343A1 · Bernat · 2018 [cited by examiner]
US 20190354412A1 · Bivens · 2019 [cited by examiner]
US 20210117249A1 · Doshi · 2021 [cited by examiner]
US 20210311798A1 · Desai · 2021 [cited by examiner]
Han, Sangjin et al. “Network Support for Resource Disaggregation in Next-Generation Datacenters.” ACM, Nov. 2013, p. 1-7 (Year: 2013). [cited by examiner]
S. Smale, “Global analysis and economics I: Pareto optimum and a generalization of morse theory,” in Dynamical systems: Elsevier, 1973, pp. 531-544. [cited by applicant]
PaloAltoNetworks.com, “What is a data center?” [Online]. Available: https://www.paloaltonetworks.com/cyberpedia/what-is-a-data-center. [cited by applicant]
S. S. Gill et al., “AI for next generation computing: Emerging trends and future directions,” Internet of Things, vol. 19, p. 100514, Aug. 2022. [cited by applicant]
C. Guo, M. Zukerman, and X. Wang, “A Reliability-Aware Resource Allocation Method in Disaggregated Data Centers.” U.S. Appl. No. 17/586,818. [cited by applicant]
Q. Zhang, Y. Cai, X. Chen, S. Angel, A. Chen, V. Liu, and B. T. Loo, “Understanding the effect of data center resource disaggregation on production dbmss,” Proceedings of the VLDB Endowment, vol. 13, No. 9, pp. 1568-158… [cited by applicant]
G. Zervas, H. Yuan, A. Saljoghei, Q. Chen, and V. Mishra, “Optically disaggregated data centers with minimal remote memory latency: Technologies, architectures, and resource allocation,” Journal of Optical Communication… [cited by applicant]
H. M. M. Ali, T. E. El-Gorashi, A. Q. Lawey, and J. M. Elmirghani, “Future energy efficient data centers with disaggregated servers,” Journal of Lightwave Technology, vol. 35, No. 24, pp. 5361-5380, 2017. [cited by applicant]
A. Peters, G. Oikonomou, and G. Zervas, “In compute/memory dynamic packet/circuit switch placement for optically disaggregated data centers,” Journal of Optical Communications and Networking, vol. 10, No. 7, pp. B164-B1… [cited by applicant]
Y. Yan et al., “All-optical programmable disaggregated data centre network realized by fpga-based switch and interface card,” Journal of Lightwave Technology, vol. 34, No. 8, pp. 1925-1932, Apr. 2016. [cited by applicant]
S. Angel, M. Nanavati, and S. Sen, “Disaggregation and the application,” in Proc. 12th {USENIX} Workshop on Hot Topics in Cloud Computing (HotCloud 20), 2020. [cited by applicant]
P. X. Gao, A. Narayan, S. Karandikar, J. Carreira, S. Han, R. Agarwal, S. Ratnasamy, and S. Shenker, “Network requirements for resource disaggregation,” in Proc. 12th {USENIX} Symposium on Operating Systems Design and I… [cited by applicant]
A. D. Papaioannou, R. Nejabati, and D. Simeonidou, “The benefits of a disaggregated data centre: A resource allocation approach,” in Proc. 2016 IEEE Global Communications Conference (GLOBECOM), 2016, pp. 1-7. [cited by applicant]
D. Mourtzis, Design and operation of production networks for mass personalization in the era of cloud technology. Elsevier, 2022. [cited by applicant]
F. Psarommatis, P. A. Dreyfus, and D. Kiritsis, “The role of big data analytics in the context of modeling design and operation of manufacturing systems,” in Design and operation of production networks for mass personal… [cited by applicant]
X. Sun, N. Ansari, and R. Wang, “Optimizing resource utilization of a data center,” IEEE Communications Surveys & Tutorials, vol. 18, No. 4, pp. 2822-2846, Fourthquarter 2016. [cited by applicant]
L. Yin, J. Luo, and H. Luo, “Tasks scheduling and resource allocation in fog computing based on containers for smart manufacturing, ” IEEE Transactions on Industrial Informatics, vol. 14, No. 10, pp. 4712-4721, Oct. 201… [cited by applicant]
D. Fernández-Cerero, J. A. Troyano, A. Jakobik, and A. Fernandez-Montes, “Machine learning regression to boost scheduling performance in hyper-scale cloud-computing data centres,” Journal of King Saud University-Compute… [cited by applicant]
W. Zhang, R. Yadav, Y.-C. Tian, S. K. K. S. Tyagi, I. A. Eelgendy, and O. Kaiwartya, “Two-phase industrial manufacturing service management for energy efficiency of data centers,” IEEE Transactions on Industrial Informa… [cited by applicant]
K. Kaur, S. Garg, G. Kaddoum, E. Bou-Harb, and K.-K. R. Choo, “A big data-enabled consolidated framework for energy efficient software defined data centers in iot setups,” IEEE Transactions on Industrial Informatics, vo… [cited by applicant]
N. Kumar, G. S. Aujla, S. Garg, K. Kaur, R. Ranjan, and S. K. Garg, “Renewable energy-based multi-indexed job classification and container management scheme for sustainability of cloud data centers,” IEEE Transactions o… [cited by applicant]
N. Gholipour, E. Arianyan, and R. Buyya, “A novel energy-aware resource management technique using joint VM and container consolidation approach for green computing in cloud data centers,” Simulation Modelling Practice … [cited by applicant]
Q. Fang, J. Wang, Q. Gong, and M. Song, “Thermal-aware energy management of an HPC data center via two-time-scale control,” IEEE Transactions on Industrial Informatics, vol. 13, No. 5, pp. 2260-2269, Oct. 2017. [cited by applicant]
C. Guo, K. Xu, G. Shen, and M. Zukerman, “Temperature-aware virtual data center embedding to avoid hot spots in data centers,” IEEE Transactions on Green Communications and Networking, vol. 5, No. 1, pp. 497-511, 2020. [cited by applicant]
Y. Shan, Y. Huang, Y. Chen, and Y. Zhang, “LegoOS: A disseminated, distributed {OS} for hardware resource disaggregation,” in Proc. Symposium on Operating Systems Design and Implementation, 2018, pp. 69-87. [cited by applicant]
T. Harvey, “Hype cycle for compute infrastructure, 2021,” Gartner, 2021. [cited by applicant]
A. Moskovsky and P. Lavrenko, “Composable disaggregated environments for HPC workloads,” in Proc. CEUR Workshop, 2020, pp. 43-51. [cited by applicant]
A. S. Al-Harrasi and S. Ali, “Investigating the challenges facing composable/disaggregated infrastructure implementation: A literature review,” in Proc. International Arab Conference on Information Technology, 2021, pp.… [cited by applicant]
ResearchAndMarkets.com, The worldwide composable infrastructure industry is expected to grow at a CAGR of 21% between 2021 to 2027. Available: https://www.globenewswire.com/news-release/2021/10/13/2313556/28124/en/The-W… [cited by applicant]
T. Coughlin, “Digital storage and memory,” Computer, IEEE, vol. 55, No. 1, pp. 20-29, Jan. 2022. [cited by applicant]
X. Guo, X. Xue, F. Yan, B. Pan, G. Exarchakos, and N. Calabretta, “Dacon: A reconfigurable application-centric optical network for disaggregated data center infrastructures,” Journal of Optical Communications and Networ… [cited by applicant]
A. Pagès, F. Agraz, and S. Spadaro, “On the impact of IT resources disaggregation in optically interconnected data centres,” in Proc. ECOC, 2019, pp. 1-4. [cited by applicant]
O. O. Ajibola, T. E. El-Gorashi, and J. M. Elmirghani, “Network topologies for composable data centers,” IEEE Access, vol. 9, pp. 120955-120984, Sep. 2021. [cited by applicant]
Intel.com, “Intel® rack scale design (Intel® RSD) architecture specification (software v2.5).” [Online]. Available: https://www.intel.com/content/www/us/en/architecture-and-technology/rack-scale-design/architecture-spec… [cited by applicant]
A. Pagès, R. Serrano, J. Perello, and S. Spadaro, “On the benefits of resource disaggregation for virtual data centre provisioning in optical data centres,” Computer Communications, vol. 107, pp. 60-74, Jul. 2017. [cited by applicant]
X. Guo, F. Yan, G. Exarchakos, X. Xue, B. Pan, and N. Calabretta, “On the workload deployment, resource utilization and operational cost of fast optical switch based rack-scale disaggregated data center network,” in Pro… [cited by applicant]
C. Zhang, P. Zhang, S. Zheng, Z. Yang, R. Liu, and K. Huang, “An efficient self-healing architecture for improving the RAS characteristics of RISC-V server and its quantitative evaluation method,” IEEE Embedded Systems … [cited by applicant]
C. Guo, X. Wang, G. Shen, S. Bose, J. Xu, and M. Zukerman, “Exploring the benefits of resource disaggregation for service reliability in data centers,” IEEE Transactions on Cloud Computing, (to appear). [cited by applicant]
A. Carbonari and I. Beschasnikh, “Tolerating faults in disaggregated datacenters,” in Proc. HotNets Workshop, 2017, pp. 164-170. [cited by applicant]
O. O. Ajibola, T. E. El-Gorashi, and J. M. Elmirghani, “Energy efficient placement of workloads in composable data center networks,” Journal of Lightwave Technology, vol. 39, No. 10, pp. 3037-3063, May 2021. [cited by applicant]
R. Lin, Y. Cheng, M. De Andrade, L. Wosinska, and J. Chen, “Disaggregated data centers: Challenges and trade-offs,” IEEE Communications Magazine, vol. 58, No. 2, pp. 20-26, 2020. [cited by applicant]
M. Amaral et al., “DRMaestro: Orchestrating disaggregated resources on virtualized data-centers,” Journal of Cloud Computing, vol. 10, No. 1, pp. 1-20, Mar. 2021. [cited by applicant]
L. Ferreira et al., “Optimizing resource availability in composable data center infrastructures,” in Proc. Latin-American Symposium on Dependable Computing, 2019, pp. 1-10. [cited by applicant]
M. Alizadeh and T. Edsall, “On the data path performance of leaf-spine datacenter fabrics,” in Proc. IEEE Annual Symposium on High-Performance Interconnects, 2013, pp. 71-74. [cited by applicant]
B. Cao et al., “Multiobjective 3-D topology optimization of next-generation wireless data center network,” IEEE Transactions on Industrial Informatics, vol. 16, No. 5, pp. 3597-3605, May 2019. [cited by applicant]
C.-S. Li, H. Franke, C. Parris, B. Abali, M. Kesavan, and V. Chang, “Composable architecture for rack scale big data computing,” Future Generation Computer Systems, vol. 67, pp. 180-193, Feb. 2017. [cited by applicant]
V. Shrivastav et al., “Shoal: A network architecture for disaggregated racks,” in Proc. Symposium on Networked Systems Design and Implementation, 2019, pp. 255-270. [cited by applicant]
V. Mishra, J. L. Benjamin, and G. Zervas, “MONet: Heterogeneous memory over optical network for large-scale data center resource disaggregation,” Journal of Optical Communications and Networking, vol. 13, No. 5, pp. 126… [cited by applicant]
Z. Ding, Y.-C. Tian, M. Tang, Y. Li, Y.-G. Wang, and C. Zhou, “Profile-guided three-phase virtual resource management for energy efficiency of data centers,” IEEE Transactions on Industrial Electronics, vol. 67, No. 3, … [cited by applicant]
D. Wang, W. Zhang, H. He, and Y.-C. Tian, “Efficient hybrid central processing unit/input-output resource scheduling for virtual machines,” IEEE Transactions on Industrial Electronics, vol. 68, No. 3, pp. 2714-2724, Mar… [cited by applicant]
L. Liu, Y. Ding, X. Li, H. Wu, and L. Xing, “A container-driven service architecture to minimize the upgrading requirements of user-side smart meters in distribution grids,” IEEE Transactions on Industrial Informatics, … [cited by applicant]
D. S. Johnson, A. Demers, J. D. Ullman, M. R. Garey, and R. L. Graham, “Worst-case performance bounds for simple one-dimensional packing algorithms,” SIAM Journal on Computing, vol. 3, No. 4, pp. 299-325, Dec. 1974. [cited by applicant]
B. LeCun, T. Mautor, F. Quessette, and M.-A. Weisser, “Bin packing with fragmentable items: Presentation and approximations,” Theoretical Computer Science, vol. 602, pp. 50-59, Oct. 2015. [cited by applicant]
X. Jia, J. Zhang, B. Yu, X. Qian, Z. Qi, and H. Guan, “GiantVM: A novel distributed hypervisor for resource aggregation with DSM-aware optimizations,” ACM Transactions on Architecture and Code Optimization (TACO), vol. … [cited by applicant]