IP Library Granted Patent US 12,436,801
Granted Patent B2
US 12,436,801 · App. 17/985,120 · Granted Oct 7, 2025

Deep learning scheduler toolkit

Inventors: Amar Phanishayee (Seattle, WA); Saurabh Agarwal (Madison, WI)
Assignee: Microsoft Technology Licensing, LLC
G06F9/4881G06F9/5077G06F2209/501G06F2209/505
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,801
App. No.
17/985,120
Granted
Oct 7, 2025
Kind
B2
Abstract

The description relates to deep learning cluster scheduler modular toolkits. One example can include generating a deep learning cluster scheduler modular toolkit that includes multiple DL scheduler abstraction modules and interactions between the multiple DL scheduler abstraction modules and allows user composition of the multiple DL scheduler abstraction modules to realize a deep learning scheduler.

Claims (27)

1. A system, comprising:

storage configured to store computer-readable instructions; and

a processor configured to execute the computer-readable instructions to generate a deep learning cluster scheduler modular toolkit that includes multiple deep learning (DL) scheduler abstraction modules and interactions between the multiple DL scheduler abstraction modules and allows user composition of the multiple DL scheduler abstraction modules to realize a DL scheduler and is further configured to allow the user to change composition of an individual DL scheduler abstraction module to change the DL scheduler without having to change other DL scheduler abstraction modules.

2. The system of claim 1 , wherein the deep learning cluster scheduler modular toolkit further includes interaction pathways between pairs of the multiple DL scheduler abstraction modules upon which the interactions occur.

3. The system of claim 1 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning admission policy module.

4. The system of claim 3 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning job wait queue that is configured to receive DL jobs from the deep learning admission policy module.

5. The system of claim 4 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning scheduling policy module configured to receive DL jobs from the deep learning job wait queue.

6. The system of claim 5 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning job placement policy module configured to receive DL job scheduling instructions from the deep learning scheduling policy module.

7. The system of claim 6 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning job preemption launch and restart policy module configured to receive DL job placement instructions from the deep learning job placement policy module.

8. The system of claim 7 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning and machine learning metrics collection module configured to track application and system-level metrics information relating to a cluster that is training the DL jobs.

9. The system of claim 8 , wherein the deep learning cluster scheduler modular toolkit comprises a deep learning cluster management module configured to track resources of the cluster.

10. A device-implemented method, comprising:

providing a deep learning cluster scheduler modular toolkit that includes multiple modular deep learning scheduler abstractions and interaction paths between the multiple modular deep learning scheduler abstractions;

receiving user input for individual modular deep learning scheduler abstractions; and,

composing multiple modular deep learning scheduler abstraction modules to realize a deep learning scheduler from the multiple modular deep learning scheduler abstractions and the user input that follows the interaction paths and is configured to receive user adjustments to an individual modular deep learning scheduler abstraction module without adjusting other individual modular deep learning scheduler abstraction modules.

11. The method of claim 10 , wherein the receiving comprises receiving through an application program interface provided by the deep learning cluster scheduler modular toolkit.

12. The method of claim 11 , wherein the receiving comprises receiving user input that defines values for an individual modular deep learning scheduler abstraction.

13. A system, comprising:

a deep learning central scheduler configured to utilize deep learning scheduler abstractions to make deep learning job scheduling decisions for job submissions;

a deep learning worker manager configured to manage cluster resources for the deep learning job scheduling decisions based upon the deep learning scheduler abstractions;

and,

a deep learning cluster scheduler modular toolkit client library configured to collect application related metrics and provide integration between applications and the deep learning central scheduler and further configured to allow changes to a composition of an individual deep learning scheduler abstraction to change the deep learning central scheduler without having to change other deep learning scheduler abstractions.

14. The system of claim 13 , wherein the deep learning scheduler abstractions of the deep learning central scheduler include a deep learning job admission deep learning scheduler abstraction, a deep learning cluster management deep learning scheduler abstraction, a deep learning and machine learning metrics collection deep learning scheduler abstraction, a deep learning job scheduling deep learning scheduler abstraction, a deep learning job preemption launch and restart deep learning scheduler abstraction, and a deep learning job placement deep learning scheduler abstraction.

15. The system of claim 14 , wherein an instance of the deep learning worker manager launches for each deep learning job running on a cluster.

16. The system of claim 15 , wherein the deep learning cluster scheduler modular toolkit client library comprises a deep learning cluster scheduler modular toolkit (DLCSMT) data loader sub-component.

17. The system of claim 16 , wherein the deep learning cluster scheduler modular toolkit client library comprises a push metrics sub-component.

18. The system of claim 17 , further comprising a DLCSMT throughput predictor model that is configured to utilize hardware features of the cluster to predict total throughput of submitted jobs that are co-located on the cluster.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2023
From: PHANISHAYEE, AMAR; AGARWAL, SAURABH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062427/0492 →
Continuity (1)
Related Publication 20240160471A1 · May 16, 2024
References Cited (71)
US 8650570B2 · Ringseth · 2014 [cited by examiner]
US 10044640B1 · Levine · 2018 [cited by examiner]
US 20050086657A1 · Jason, Jr. · 2005 [cited by examiner]
US 20090300637A1 · Ringseth · 2009 [cited by examiner]
US 20190332422A1 · Liu et al. · 2019 [cited by applicant]
US 20200160171A1 · Rangarajan · 2020 [cited by examiner]
US 20210011762A1 · Lin · 2021 [cited by examiner]
US 20220198296A1 · Liu · 2022 [cited by examiner]
US 20220326915A1 · Lee · 2022 [cited by examiner]
US 20220374775A1 · Liu · 2022 [cited by examiner]
US 20230129998A1 · Jung · 2023 [cited by examiner]
Li et al.; “Aryl: An Elastic Cluster Scheduler for Deep Learning”; ByteDance; arXiv:2202.07896v1 [cs.DC] Feb. 16, 2022; (Li_2022.pdf; pp. 1-20) (Year: 2022). [cited by examiner]
Peng et al.; “Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters”; 2018 Association for Computing Machinery; https://doi.org/10.1145/3190508.3190517; (Peng_2018.pdf) (Year: 2018). [cited by examiner]
Qiao et al.; “Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning”; 15th USENIX Symposium on Operating Systems Design and Implementation; Jul. 14-16, 2021; https://www.usenix.org/conference/osdi21… [cited by examiner]
Zhang et al.; “ReLeS: A Neural Adaptive Multipath Scheduler based on Deep Reinforcement Learning”; 2019 IEEE; (Zhang_2019.pdf; pp. 1648-1656) (Year: 2019). [cited by examiner]
Yao et al.; “Meta-learning with an Adaptive Task Scheduler”; 35th Conference on Neural Information Processing Systems (NeurIPS 2021); (Yao_2021.pdf; pp. 1-13) (Year: 2021). [cited by examiner]
Kohler, et al., “The Click Modular Router”, Laboratory for Computer Science, MIT, ACM SIGOPS Operating Systems Review, vol. 33, Issue 5, Dec. 1999, pp. 217-231. [cited by applicant]
Wu, et al., “Google's neural machine translation system: Bridging the gap between human and machine translation”, In Publication of arXiv preprint arXiv:1609.08144, Sep. 26, 2016, pp. 1-23. [cited by applicant]
Hindman, et al., “Mesos: A Platform for Fine-grained Resource Sharing in the Data Center”, In Proceedings of the 8th USENIX conference on Networked systems design and implementation, Mar. 30, 2011, pp. 1-14. [cited by applicant]
“Artifact for Pollux OSDI 2021”, Retrieved From: https://github.com/petuum/adaptdl/tree/osdi21-artifact, Aug. 23, 2021, 3 Pages. [cited by applicant]
“AWS RDS instance created from snapshot very slow”, Retrieved From: https://stackoverflow.com/questions/47545414/aws-rds-instance-created-from-snapshot-very-slow%20, Dec. 10, 2021, 2 Pages. [cited by applicant]
“Checkpoint-restore in userspace (criu)”, Retrieved From: https://criu.org/Main_Page, May 18, 2022, 3 Pages. [cited by applicant]
“GRPC, A high-performance, open-source universal RPC framework.”, Retrieved from: https://web.archive.org/web/20210416144134/https://grpc.io/, Apr. 16, 2021, 2 Pages. [cited by applicant]
“Kubernetes”, Retrieved From: https://kubernetes.io/, Retrieved on: Oct. 31, 2022, 6 Pages. [cited by applicant]
Acun, et al., “Understanding training efficiency of deep learning recommendation models at scale”, In IEEE International Symposium on High-Performance Computer Architecture (HPCA), Feb. 27, 2021, 13 Pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, In Repository of arXiv:2005.14165, May 28, 2020, 72 Pages. [cited by applicant]
Chaudhary, et al., “Balancing Efficiency and Fairness in Heterogeneous GPU Clusters for Deep Learning”, In Proceedings of the Fifteenth European Conference on Computer Systems, Apr. 15, 2020, 16 Pages. [cited by applicant]
Cranshaw, et al., “Clipper: A Low-Latency Online Prediction Serving System”, In Proceedings of 14th USENIX Symposium on Networked Systems Design and Implementation, Mar. 27, 2017, 17 Pages. [cited by applicant]
Dean, et al., “MapReduce: Simplified Data Processing on Large Clusters”, In Journal of Communications of the ACM Magazine, vol. 51, Issue 1, Jan. 2008, pp. 107-113. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Repository of arXiv:1810.04805v1, Oct. 11, 2018, 14 Pages. [cited by applicant]
Ford, et al., “The Flux OS Toolkit: reusable components for OS implementation”, In Proceedings of the Sixth Workshop on Hot Topics in Operating Systems, May 5, 1997, pp. 1-6. [cited by applicant]
Ford, et al., “The flux OSKit: A substrate for kernel and language research”, In Proceedings of the 16th ACM Symposium on Operating Systems Principles, Oct. 5, 1997, pp. 38-51. [cited by applicant]
Goodfellow, et al., “Generative Adversarial Networks”, In Repository of arXiv: 1406.2661, Jun. 10, 2014, 09 Pages. [cited by applicant]
Gu, et al., “Tiresias: A GPU Cluster Manager for Distributed Deep Learning”, In Proceedings of the 16th USENIX Symposium on Networked Systems Design and Implementation, Feb. 26, 2019, pp. 485-500. [cited by applicant]
Gujarati, et al., “Serving DNNs Like Clockwork: Performance Predictability from the Bottom Up”, In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), Nov. 4, 2020, pp. 443… [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 770-778. [cited by applicant]
Hochreiter, et al., “Long Short-Term Memory”, In Journal of Neural Computation, vol. 9, Issue 8, Nov. 15, 1997., pp. 1735-1780. [cited by applicant]
Hwang, et al., “Elastic resource sharing for distributed deep learning”, In 18th USENIX Symposium on Networked Systems Design and Implementation, May 6, 2021, 19 Pages. [cited by applicant]
Jeon,, et al., “Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads”, In Proceedings of the USENIX Conference on Usenix Annual Technical Conference, Jul. 10, 2019, pp. 947-960. [cited by applicant]
Jouppi, et al., “A domain-specific supercomputer for training deep neural networks”, In Article of Communications of the ACM, vol. 63, Issue 7, Jul. 2020, pp. 67-78. [cited by applicant]
Kingma, et al., “Adam: A method for stochastic optimization”, In Proceedings of the 3rd International Conference on Learning Representations, Dec. 22, 2014, 15 Pages. [cited by applicant]
Kohler, et al., “The click modular router”, In ACM Transactions on Computer Systems, vol. 18, Issue 3, Aug. 1, 2000, pp. 263-297. [cited by applicant]
Krizhevsky, et al., “ImageNet Classification with Deep Convolutional Neural Networks”, In Proceedings of 26th Annual Conference on Neural Information Processing Systems, Dec. 3, 2012, pp. 1-9. [cited by applicant]
Le, et al., “AlloX: compute allocation in hybrid clusters”, In Proceedings of the Fifteenth European Conference on Computer Systems, Apr. 17, 2020, pp. 1-16. [cited by applicant]
Li, et al., “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization”, In Journal of Machine Learning Research, vol. 18, Issue 1, Apr. 1, 2018, pp. 1-52. [cited by applicant]
Mahajan, et al., “THEMIS: Fair and Efficient GPU Cluster Scheduling”, In Proceedings of the 17th USENIX Symposium on Networked Systems Design and Implementation, Feb. 25, 2020, pp. 289-304. [cited by applicant]
Microsoft, “Open platform for ai”, Retrieved From: https://github.com/microsoft/pai, May 18, 2021, 8 Pages. [cited by applicant]
Mnih, et al., “Asynchronous Methods for Deep Reinforcement Learning”, In Proceedings of the 33nd International Conference on Machine Learning, Jun. 16, 2016, 19 Pages. [cited by applicant]
Mohan, et al., “Synergy: Resource sensitive dnn scheduling in multitenant clusters”, In 16th USENIX Symposium on Operating Systems Design and Implementation, Aug. 24, 2022, 21 Pages. [cited by applicant]
Moussawi, “Towards large scale training of autoencoders for collaborative filtering”, In Repository of arXiv:1809.00999, Aug. 30, 2018, 2 Pages. [cited by applicant]
Narayanan, “Heterogeneity-Aware cluster scheduling policies for deep learning workloads”, In 14th USENIX Symposium on Operating Systems Design and Implementation, Nov. 4, 2020, 18 Pages. [cited by applicant]
Peng, et al., “Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters”, In Proceedings of the Thirteenth EuroSys Conference Pattern Recognition, Apr. 23, 2018, 14 Pages. [cited by applicant]
Petuum, “Adaptdl”, Retrieved From: https://github.com/petuum/adaptdl, May 18, 2022, 4 Pages. [cited by applicant]
Qiao, et al., “Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning”, In 15th {USENIX} Symposium on Operating Systems Design and Implementation, May 26, 2021, 18 Pages. [cited by applicant]
Radford, et al., “Language Models are Unsupervised Multitask Learners”, In Journal of OpenAI Blog, vol. 1, Issue 8, Feb. 24, 2019, 24 Pages. [cited by applicant]
Raffel, et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, In Repository of arXiv:1910.10683v2, Oct. 24, 2019, pp. 1-53. [cited by applicant]
Shen, et al., “Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis”, In Proceedings of the 27th ACM Symposium on Operating Systems Principles, Oct. 27, 2019, pp. 332-337. [cited by applicant]
Shoeybi, et al., “Megatron-LM: Training Multi-Billion Parameter Language Models Using GPU Model Parallelism”, In Repository of arXiv:1909.08053v3, Sep. 17, 2019, 15 Pages. [cited by applicant]
Sun, et al., “High-Resolution Representations for Labeling Pixels and Regions”, In Repository of arXiv:1904.04514v1, Apr. 9, 2019, 13 Pages. [cited by applicant]
Vaswani, et al., “Attention is All You Need”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, pp. 1-11. [cited by applicant]
Vavilapalli, et al., “Apache Hadoop YARN: Yet another Resource Negotiator”, In Proceedings of the 4th annual Symposium on Cloud Computing, Oct. 1, 2013, 16 Pages. [cited by applicant]
Verma, et al., “Large-scale cluster management at Google with Borg”, In Proceedings of the Tenth European Conference on Computer Systems, Apr. 21, 2015, 17 Pages. [cited by applicant]
Weng, et al., “MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters”, In 19th USENIX Symposium on Networked Systems Design and Implementation, Apr. 4, 2022, pp. 945-960. [cited by applicant]
Xiao, et al., “AntMan: dynamic scaling on GPU clusters for deep learning”, In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation, Nov. 4, 2020, pp. 533-548. [cited by applicant]
Xiao, et al., “Gandiva: Introspective Cluster Scheduling for Deep Learning”, In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation, Oct. 8, 2018, pp. 595-610. [cited by applicant]
Yoo, et al., “SLURM: Simple Linux Utility for Resource Management”, In Workshop on job scheduling strategies for parallel processing, Jun. 23, 2003, 27 Pages. [cited by applicant]
Zaharia, et al., “Spark: Cluster Computing with Working Sets”, In Proceedings of 2nd USENIX Workshop on Hot Topics in Cloud Computing, Jun. 22, 2010, pp. 1-7. [cited by applicant]
Zhao, et al., “HiveD: sharing a GPU cluster for deep learning with guarantees”, In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation, Nov. 4, 2020, pp. 515-532. [cited by applicant]
Zhu, et al., “Unpaired image-to-image translation using cycle-consistent adversarial networks”, In Proceedings of the IEEE international conference on computer vision, Oct. 22, 2017, pp. 2223-2232. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2023/035359, mailed on Dec. 20, 2023, 11 pages. [cited by applicant]
International preliminary report on patentability Received in European Patent Application No. PCT/US23/035359, mailed on May 22, 2025, 07 pages. [cited by applicant]