IP Library Granted Patent US 12,443,876
Granted Patent B2
US 12,443,876 · App. 17/125,626 · Granted Oct 14, 2025

Context-aware and stateless deep learning autotuning framework

Inventors: Junguk Cho (Milpitas, CA); Diman Zad Tootaghaj (Milpitas, CA); Puneet Sharma (Milpitas, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06N20/00G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,876
App. No.
17/125,626
Granted
Oct 14, 2025
Kind
B2
Abstract

Systems and methods are provided for improving autotuning procedures using stateless processing with a remote key-value store. For example, the system can implement a task launcher, a scheduler, and an agent to launch, schedule, and execute decomposed autotuning stages, respectively. The scheduling policy implemented by the scheduler may perform operations beyond a simple scheduling policy (e.g., a FIFO-based scheduling policy), which produces a high queuing delay. Compared to the traditional systems, by leveraging autotuning specific domain knowledge, queueing delay is reduced and resource utilization is improved.

Claims (42)

1. An autotuning computer system for performing stateless autotuning tasks with respect to a machine learning (ML) model, the autotuning computer system comprising:

a memory storing a scheduler circuit, wherein the scheduler circuit operates in accordance with machine executable instructions; and

one or more processors configured to access the memory and execute the machine executable instructions stored to:

receive, from the scheduler circuit, a scheduling request for an autotuning task, wherein the scheduling request comprises an autotuning stage, an autotuning task ID, and parameters;

receive a set of states of the autotuning task from a key-value store by using the autotuning task ID as a key, wherein the set of states of the autotuning task comprise a current stage and an average runtime per stage;

based at least on the average runtime per stage, determine a hardware resource, the hardware resource configured to generate at least one of the states of the autotuning task through execution of the autotuning task by the hardware resource;

load the states of the autotuning task at the hardware resource;

execute, by the hardware resource, the current stage with the loaded states, which generates new states to be used for executing a next stage; and

store the generated new states as values in the key-value store by using the autotuning task ID as the key, wherein decoupling of the states from computation to the key-value store forms the autotuning computer system as stateless.

2. The autotuning computer system of claim 1 , wherein the autotuning computer system is configured to optimize inference performance of the machine learning (ML) model for a particular hardware configuration.

3. The autotuning computer system of claim 1 , wherein the states of the autotuning task is externalized to the key-value store and sub-procedures load the states from the key-value store when the sub-procedures are scheduled.

4. The autotuning computer system of claim 3 , wherein the sub-procedures are executed on different resources simultaneously.

5. The autotuning computer system of claim 1 , wherein the scheduler circuit utilizes shortest job first (SJF) to prioritize the current stage requiring short completion time over the next stage with longer completion time.

6. The autotuning computer system of claim 1 , wherein a Multi-Process Service (MPS) is implemented to isolate computational resources.

7. The autotuning computer system of claim 1 , wherein the hardware resource comprises a set of loop tiles, and the states of the autotuning task configure the set of loop tiles to execute the autotuning task based on one or more of ordering, caching, and loop unrolling.

8. A computer-implemented method for performing stateless autotuning tasks with respect to a machine learning (ML) model, the method comprising:

receiving, from a scheduler circuit, a scheduling request for an autotuning task, wherein the scheduling request comprises an autotuning stage, an autotuning task ID, and parameters;

receiving a set of states of the autotuning task, wherein the set of states of the autotuning task comprise a current stage and an average runtime per stage;

based at least on the average runtime per stage, determining a hardware resource, the hardware resource configured to generate at least one of the states of the autotuning task through execution of the autotuning task by the hardware resource;

loading the states of the autotuning task at the hardware resource;

executing the current stage with the loaded states, which generates new states to be used for executing a next stage; and

storing the generated new states as values in a key-value store by using the autotuning task ID as a key, wherein decoupling of the states from computation to the key-value store forms a computer system corresponding with the computer-implemented method as stateless.

9. The computer-implemented method of claim 8 , wherein inference performance of the machine learning (ML) model is optimized for a particular hardware configuration.

10. The computer-implemented method of claim 8 , wherein the states of the autotuning task are externalized to the key-value store and sub-procedures load the states from the key-value store when the sub-procedures are scheduled.

11. The computer-implemented method of claim 10 , wherein the sub-procedures are executed on different resources simultaneously.

12. The computer-implemented method of claim 8 , further comprising:

receiving, by the scheduler circuit, a registration request with autotuning options.

13. The computer-implemented method of claim 8 , wherein the scheduler circuit utilizes shortest job first (SJF) to prioritize the current stage requiring short completion time over the next stage with longer completion time.

14. The computer-implemented method of claim 8 , wherein a Multi-Process Service (MPS) is implemented to isolate computational resources.

15. A non-transitory computer-readable storage medium storing a plurality of instructions executable by one or more processors, the plurality of instructions when executed by the one or more processors cause the one or more processors to:

receive, from a scheduler circuit, a scheduling request for an autotuning task, wherein the scheduling request comprises an autotuning stage, an autotuning task ID, and parameters;

receive a set of states of the autotuning task, wherein the set of states of the autotuning task comprise a current stage and an average runtime per stage;

based at least on the average runtime per stage, determine a hardware resource, the hardware resource configured to generate at least one of the states of the autotuning task through execution of the autotuning task by the hardware resource;

load the states of the autotuning task at the hardware resource;

execute the current stage with the loaded states, which generates new states to be used for executing a next stage; and

store the generated new states as values in a key-value store by using the autotuning task ID as a key, wherein decoupling of the states from computation to the key-value store forms an autotuning computer system as stateless.

16. The computer-readable storage medium of claim 15 , wherein inference performance of a machine learning (ML) model implemented by the autotuning computer system is optimized for a particular hardware configuration.

17. The computer-readable storage medium of claim 15 , wherein the states of the autotuning task is externalized to the key-value store and sub-procedures load the states from the key-value store when the sub-procedures are scheduled.

18. The computer-readable storage medium of claim 17 , wherein the sub-procedures are executed on different resources simultaneously.

19. The computer-readable storage medium of claim 15 , the plurality of instructions further cause the one or more processors to:

receive, by the scheduler circuit, a registration request with autotuning options.

20. The computer-readable storage medium of claim 15 , wherein the scheduler circuit utilizes shortest job first (SJF) to prioritize the current stage requiring short completion time over the next stage with longer completion time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2020
From: CHO, JUNGUK; ZAD TOOTAGHAJ, DIMAN; SHARMA, PUNEET
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 054686/0030 →
Continuity (1)
Related Publication 20220198317A1 · Jun 23, 2022
References Cited (72)
US 10228922B2 · Boehm et al. · 2019 [cited by applicant]
US 20100095299A1 · Gupta et al. · 2010 [cited by applicant]
US 20110228922A1 · Dhara et al. · 2011 [cited by applicant]
US 20160078361A1 · Brueckner et al. · 2016 [cited by applicant]
US 20180322462A1 · Jayaraman et al. · 2018 [cited by applicant]
US 20190050756A1 · Dirac et al. · 2019 [cited by applicant]
US 20190205746A1 · Nurvitadhi et al. · 2019 [cited by applicant]
US 20200097847A1 · Convertino et al. · 2020 [cited by applicant]
US 20200134207A1 · Doshi et al. · 2020 [cited by applicant]
US 20200136906A1 · Guim Bernat et al. · 2020 [cited by applicant]
US 20200364088A1 · Ashwathnarayan et al. · 2020 [cited by applicant]
US 20200364508A1 · Gurel et al. · 2020 [cited by applicant]
US 20210144517A1 · Guim et al. · 2021 [cited by applicant]
US 20220101437A1 · Gao et al. · 2022 [cited by applicant]
US 20220188700A1 · Khavronin et al. · 2022 [cited by applicant]
US 20230134690A1 · Chen et al. · 2023 [cited by applicant]
US 20230236947A1 · Frohwitter · 2023 [cited by applicant]
CN 110046046A · 2019 [cited by applicant]
CN 110515735A · 2019 [cited by applicant]
CN 110766090A · 2020 [cited by applicant]
CN 111613234A · 2020 [cited by applicant]
CN 111736463A · 2020 [cited by applicant]
Bae, I., et al., Auto-Tuning CNNs for Coarse-Grained Reconfigurable Array-Based Accelerators, [received on Apr. 1, 2024]. Retrieved from Internet:<https://ieeexplore.ieee.org/abstract/document/8412597> (Year: 2018). [cited by examiner]
Dhakal, A., et al., Spatial Sharing of GPU for Autotuning DNN models, [received on Apr. 1, 2024]. Retrieved from Internet:<https://arxiv.org/abs/2008.03602> (Year: 2020). [cited by examiner]
Karcher, T., et al., Autotuning and Self-Adaptability in Concurrency Libraries, [received on 412024]. Retrieved from Internet:<https://arxiv.org/abs/1405.2918> (Year: 2014). [cited by examiner]
Kwon, S., et al., Serialized Parallel Code Generation Framework for MPSoC, [received on Apr. 1, 2024]. Retrieved from Internet: <https://dl.acm.org/doi/abs/10.1145/1698759.1698761> (Year: 2010). [cited by examiner]
Liu, Y., et al., Optimizing CNN Model Inference on CPUs, [received on Apr. 1, 2024]. Retrieved from Internet:<https://www.usenix.org/conference/atc19/presentation/liu-yizhi> (Year: 2019). [cited by examiner]
Narayanan, D., t al, Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads, [received on Apr. 1, 2024]. Retrieved from Internet:<https://www.usenix.org/conference/osdi20/presentation/narayanan-deep… [cited by examiner]
Xing, Y., et al., An In-depth Comparison of Compilers for Deep Neural Networks on Hardware, [received on Apr. 1, 2024]. Retrieved from Internet:<https://ieeexplore.ieee.org/abstract/document/8782480> (Year: 2019). [cited by examiner]
Gregg, et al., Dynamic Heterogeneous Scheduling Decisions Using Historical Runtime Data, [received Aug. 26, 2024]. Retrieved from Internet:<chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://web.stanford.edu/˜c… [cited by examiner]
Sato, Y., et al., An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation, [received Aug. 27, 2024]. Retrieved from Internet:<https://dl.acm.org/doi/abs/10.1145/3293449> (Year: … [cited by examiner]
Zheng, L. et al, Ansor: Generating High-Performance Tensor Programs for Deep Learning, Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation, Nov. 4-6, 2020, 18 Pgs, USENIX Association. [cited by applicant]
Ahn et al., “Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation,” ICLR 2020, Jan. 23, 2020, pp. 1-17. [cited by applicant]
Ahn et al., “Reinforcement Learning and Adaptive Sampling for Optimized DNN Compilation,” Computer Science, May 30, 2019, 11 pages. [cited by applicant]
Akiba t al., “Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes,” Nov. 12, 2017, pp. 1-4. [cited by applicant]
Apache TVM, “Scale up measurement by using multiple devices,” available online at <https://web.archive.org/web/20200919111851/https://tvm.apache.org/docs/tutorials/autotvm/tune_relay_cuda.html#scale-up-measurement-by-us… [cited by applicant]
Awatramani et al., “Increasing gpu throughput using kernel interleaved thread block scheduling,” In 2013 IEEE 31st International Conference on Computer Design (ICCD), Oct. 6-9, 2013, pp. 503-506. [cited by applicant]
AWS, “Amazon EC2 P2 Instances: Powerful, Scalable GPU instances for high-performance computing,” available online at <https://web.archive.org/web/20201021031252/https://aws.amazon.com/ec2/instance-types/p2/>, Oct. 21, 2… [cited by applicant]
AWS, “Amazon EC2 P3 Instances: Accelerate machine learning and high performance computing applications with powerful GPUs,” available online at <https://web.archive.org/web/20201017163938/https://aws.amazon.com/ec2/inst… [cited by applicant]
AWS, “Amazon SageMaker Neo,” available online at <https://web.archive.org/web/20201018133845/https://aws.amazon.com/sagemaker/neo/>, Oct. 18, 2020, 6 pages. [cited by applicant]
Chen et al., “Learning to Optimize Tensor Programs,” Computer Science, Jan. 8, 2019, pp. 1-16. [cited by applicant]
Chen et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), Oct. 8-10, 2018, 17 pages. [cited by applicant]
Codreanu et al., “Achieving deep learning Training in less than 40 minutes on ImageNet-1K & best accuracy and training time on ImageNet-22K & Places365 with scale-out Intel® Xeon® /Xeon Phi™ architectures,” Surf Communi… [cited by applicant]
Goyal et al., “Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour,” Computer Vision and Pattern Recognition, Apr. 30, 2018, pp. 1-12. [cited by applicant]
Gu, J. et al., Tiresias: A GPU Cluster Manager for Distributed Deep Learning, (Research Paper), 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI '19) Feb. 26-28, 2019 17 Pgs. Boston MA USA. [cited by applicant]
Hazelwood, “Applied machine learning at facebook: A datacenter infrastructure perspective,” In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), Feb. 24-28, 2018, pp. 620-629. [cited by applicant]
He et al., “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Dec. 10, 2015, pp. 1-12. [cited by applicant]
Intel®, “Intel® Math Kernel Library,” available online at <https://software.intel.com/content/www/us/en/develop/tools/math-kernel-library.html>, 2019, 7 pages. [cited by applicant]
Jain et al., “Dynamic Space-Time Scheduling for GPU Inference,” 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Dec. 1, 2018, pp. 1-8. [cited by applicant]
Jeon et al., “Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads,” Proceedings of the 2019 USENIX Annual Technical Conference, Jul. 10-12, 2019, 15 pages. [cited by applicant]
Jeremy Howard, “Training Imagenet in 3 hours for 25; and CIF AR10 for 0.26,” Fast.ai, May 2, 2018, 8 pages. [cited by applicant]
Kirkpatrick et al., “Optimization by Simulated Annealing,” Science, vol. 220, No. 4598, May 13, 1983, pp. 671-680. [cited by applicant]
Liu et al., “Optimizing CNN Model Inference on CPUs,” USENIX ATC'19, Jul. 10-12, 2019, pp. 1025-1039. [cited by applicant]
Ma et al., “Hierarchical task scheduler for interleaving subtasks on heterogeneous multiprocessor platforms,” In Proceedings of the 2005 Asia and South Pacific Design Automation Conference, Feb. 2005, pp. 952-955. [cited by applicant]
NVIDIA Developer, “NVIDIA cuDNN,” available online at <https://web.archive.org/web/20201016224154/https://developer.nvidia.com/cudnn>, Oct. 16, 2020, 5 pages. [cited by applicant]
Nvidia, “Multi-Process Service,” version—vR450, Jun. 2020, 28 pages. [cited by applicant]
Peng, Y. et al., Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters, (Research Paper), EuroSys '18: Thirteenth EuroSys Conference 2018, Apr. 23-26, 2018 14 Pgs. ACM Porto Portugal. [cited by applicant]
Ragan-Kelley et al., “Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines,” PLDI '13, Jun. 16-21, 2013, 12 pages. [cited by applicant]
Rotem et al., “Glow: Graph lowering compiler techniques for neural networks,” May 4, 2018, pp. 1-10. [cited by applicant]
Seagate Technology, “Smart manufacturing moves from autonomous to intelligent: Inside Project Athena: Seagate's internal AI edge platform,” 2019, 4 pages. [cited by applicant]
Sukhoroslov, O. et al., Program Autotuning as a Service: Opportunities and Challenges, (Research Paper), 2016 IEEE/ ACM 9th International Conference on Utility and Cloud Computing, Dec. 6-9, 2016, 8 Pgs., ACM, Shanghai,… [cited by applicant]
Tillenius, M. et al., Resource-Aware Task Scheduling, (Research Paper), ACM Transactions on Embedded Computing Systems, Jan. 2015, 25 Pgs., vol. 14, No. 1, Article 5 ACM. [cited by applicant]
Tomczak et al., “Simulating Execution Time of Tensor Programs using Graph Neural Networks,” Machine Learning, ICLR 2019, Nov. 27, 2019, pp. 1-8. [cited by applicant]
Vasilache et al., “Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions,” Facebook AI Research Technical Report, Feb. 13, 2018, pp. 1-39. [cited by applicant]
Xiao et al., “Gandiva: Introspective cluster scheduling for deep learning,” 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), Oct. 8-10, 2018, pp. 595-610. [cited by applicant]
Yizhi Liu, “Remote Profile and Test Deep Learning Cross Compilation on Mobile Phones with TVM RPC,” Apache TVM, Nov. 8, 2017, 6 pages. [cited by applicant]
You et al., “ImageNet Training in 24 Minutes,” Oct. 16, 2017, 9 pages. [cited by applicant]
Yu et al., “Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications,” Computer Science, Feb. 12, 2019, pp. 1-15. [cited by applicant]
Mullapudi et al., “Automatically Scheduling Halide Image Processing Pipelines”, ACM Trans. Graph., vol. 35, No. 4, Article 83, Publication Date: Jul. 2016, 11 pages. [cited by applicant]
Chen, T., et al., “Xgboost: A scalable tree boosting system.”, Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, Aug. 13-17, 2016, pp. 785-794. [cited by applicant]
Lianmin Zheng “Tuning High Performance Convolution on NVIDIA GPUs.”, Online available at <https://tvm.apache.org/docs/v0.14.0/how_to/tune_with_autotvm/tune_conv2d_cuda.html>, 2023, 10 pages. [cited by applicant]
Mitchell, R., et al., “Accelerating the XGBoost algorithm using GPU computing.”, PeerJ Computer Science, vol. 3, Jul. 24, 2017, pp. 1-37. [cited by applicant]
Cited By (1)
US 12,675,266