IP Library › Granted Patent US 12,579,433
Granted Patent B2
US 12,579,433 · App. 17/783,247 · Granted Mar 17, 2026

Resource usage prediction for deep learning model

Inventors: Yanjie Gao (Redmond, WA); Haoxiang Lin (Redmond, WA); Yu Liu (Redmond, WA); Mao Yang (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,433
App. No.
17/783,247
Granted
Mar 17, 2026
Kind
B2
Abstract

According to implementations of the subject matter described herein, there is provided a solution for predicting the resource usage of the deep learning model. In this solution, information about a deep learning model is obtained, the information comprising first information for describing the deep learning model and second information about an operating environment of a job associated with the deep learning model. The static resource usage of the job is determined based on the first information and a strategy of the job during runtime in the operating environment is determined. Afterwards, resource usage of the job during runtime in the operating environment is predicted based on the strategy and the static resource usage. With this solution, the usage of various resources of the deep learning model, such as computation power consumption, memory consumption, execution time, and the like, under a specific runtime strategy can be accurately predicted.

Claims (12)

1 . A computer-implemented method, comprising: obtaining information about a deep learning model, the information comprising first information for describing the deep learning model and second information about an operating environment of a job associated with the deep learning model; wherein the first information comprises a configuration parameter of the deep learning model and second information comprises at least one of: a framework type of the deep learning model, specifications and number of computing devices for executing the job in the operating environment, and an execution strategy of the job on the computing devices; determining static resource usage of the job based on the first information; wherein determining the static resource usage comprises: generating, based on the first information, a computation graph corresponding to the deep learning model, the computation graph comprising a plurality of nodes corresponding to a plurality of operators in the deep learning model and edges connecting the plurality of nodes indicating dependency among the plurality of operators: predicting, based on the computation graph and respective resource prediction models of the plurality of operators, respective static resource usage of the plurality of operators; and determining the static resource usage of the job based on the respective static resource usage of the plurality of operators: determining, based on the first information and the second information, a strategy of the job during runtime in the operating environment; and predicting, based on the strategy and the static resource usage, resource usage of the job during runtime in the operating environment.

2 . The method of claim 1 , wherein the first information comprises at least one of: a model file of the deep learning model and program codes of the job.

3 . The method of claim 1 , wherein the resource usage comprises at least one of: computation power consumption, memory consumption, I/O resource consumption, execution time, and power consumption.

4 . The method of claim 3 , wherein the resource usage comprises other resource consumption determined based on at least one of the computation power consumption and the memory consumption.

5 . The method of claim 1 , wherein the strategy comprises at least one of: a resource allocation strategy of the deep learning model, and an execution strategy of the job in the operating environment.

6 . The method of claim 5 , wherein predicting the resource usage of the job during runtime in the operating environment comprises: adjusting the static resource usage based on at least one of the resource allocation strategy and the execution strategy, to obtain the resource usage of the job during runtime in the operating environment.

7 . The method of claim 1 , further comprising: generating, using a trained machine learning model, a parameter for optimizing the predicted resource usage; and optimizing, based on the parameter, the predicted resource usage.

8 . An electronic device, comprising: a processing unit; and a memory coupled to the processing unit and comprising instructions stored thereon which, when executed by the processing unit, cause the device to perform acts comprising: obtaining information about a deep learning model, the information comprising first information for describing the deep learning model and second information about an operating environment of a job associated with the deep learning model; wherein the first information comprises a configuration parameter of the deep learning model and second information comprises at least one of: a framework type of the deep learning model, specifications and number of computing devices for executing the job in the operating environment, and an execution strategy of the job on the computing devices; determining static resource usage of the job based on the first information; wherein determining the static resource usage comprises: generating, based on the first information, a computation graph corresponding to the deep learning model, the computation graph comprising a plurality of nodes corresponding to a plurality of operators in the deep learning model and edges connecting the plurality of nodes indicating dependency among the plurality of operators; predicting, based on the computation graph and respective resource prediction models of the plurality of operators, respective static resource usage of the plurality of operators; and determining the static resource usage of the job based on the respective static resource usage of the plurality of operators; determining, based on the first information and the second information, a strategy of the job during runtime in the operating environment; and predicting, based on the strategy and the static resource usage, resource usage of the job during runtime in the operating environment.

9 . The device of claim 8 , wherein the resource usage comprises at least one of: computation power consumption, memory consumption, I/O resource consumption, execution time, and power consumption.

10 . The device of claim 9 , wherein the resource usage comprises other resource consumption determined based on at least one of the computation power consumption and the memory consumption.

11 . The device of claim 8 , wherein determining the static resource usage comprises: generating, based on the first information, a computation graph corresponding to the deep learning model, the computation graph comprising a plurality of nodes corresponding to a plurality of operators in the deep learning model and edges connecting the plurality of nodes indicating dependency among the plurality of operators; predicting, based on the computation graph and respective resource prediction models of the plurality of operators, respective static resource usage of the plurality of operators; and determining the static resource usage of the job based on the respective static resource usage of the plurality of operators.

12 . A computer program product being tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions which, when executed by a device, cause the device to perform acts comprising: obtaining information about a deep learning model, the information comprising first information for describing the deep learning model and second information about an operating environment of a job associated with the deep learning model; wherein the first information comprises a configuration parameter of the deep learning model and second information comprises at least one of: a framework type of the deep learning model, specifications and number of computing devices for executing the job in the operating environment, and an execution strategy of the job on the computing devices: determining static resource usage of the job based on the first information; wherein determining the static resource usage comprises: generating, based on the first information, a computation graph corresponding to the deep learning model, the computation graph comprising a plurality of nodes corresponding to a plurality of operators in the deep learning model and edges connecting the plurality of nodes indicating dependency among the plurality of operators: predicting, based on the computation graph and respective resource prediction models of the plurality of operators, respective static resource usage of the plurality of operators; and determining the static resource usage of the job based on the respective static resource usage of the plurality of operators; determining, based on the first information and the second information, a strategy of the job during runtime in the operating environment; and predicting, based on the strategy and the static resource usage, resource usage of the job during runtime in the operating environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: GAO, YANJIE; LIN, HAOXIANG; LIU, YU; YANG, MAO
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 060154/0130 →
Priority Claims (1)
CN 202010025197.X · Jan 9, 2020 · national
Continuity (1)
Related Publication 20230035451A1 · Feb 2, 2023
References Cited (24)
US 8438120B2 · Raaijmakers · 2013 [cited by applicant]
US 8468109B2 · Moussa et al. · 2013 [cited by applicant]
US 10235625B1 · Walters et al. · 2019 [cited by applicant]
US 10257116B1 · Vadera et al. · 2019 [cited by applicant]
US 12393835B2 · Li · 2025 [cited by examiner]
US 20180365065A1 · Guttmann et al. · 2018 [cited by applicant]
US 20190213099A1 · Schmidt · 2019 [cited by examiner]
US 20190244102A1 · Harvey · 2019 [cited by examiner]
US 20190266015A1 · Chandra et al. · 2019 [cited by applicant]
US 20210286650A1 · Henry · 2021 [cited by examiner]
US 20220147831A1 · Hoang · 2022 [cited by examiner]
US 20250053860A1 · Zhu · 2025 [cited by examiner]
US 20250272138A1 · Shang Guan · 2025 [cited by examiner]
CN 110390387A · 2019 [cited by applicant]
WO 2019168724A1 · 2019 [cited by applicant]
Notice Of Allowance Received for Chinese Application No. 202010025197.X, mailed on Aug. 5, 2025, 04 pages. (English Translation Provided). [cited by applicant]
Second Office Action Received for Chinese Application No. 202010025197.X, mailed on Apr. 4, 2025, 11 pages. (English Translation Provided). [cited by applicant]
First Office Action Received for Chinese Application No. 202010025197.X, mailed on Jul. 19, 2024, 15 pages. (English Translation Provided). [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US23/030881, mailed on Apr. 3, 2025, 06 pages. [cited by applicant]
Justus, et al., “Predicting the Computational Cost of Deep Learning Models”, In Repository of arXiv:1811.11880v1, Nov. 28, 2018, 11 Pages. [cited by applicant]
Li, et al., “Edge AI: On-Demand Accelerating Deep Neural Network Inference Via Edge Computing”, In Journal of IEEE Transactions on Wireless Communications, vol. 19, Issue 1, Jan. 2020, pp. 447-457. [cited by applicant]
Liu, et al., “On-Demand Deep Model Compression for Mobile Devices: A Usage-Driven Model Selection Framework”, In Proceedings of the 16th Annual International Conference on Mobile Systems, Applications, and Services, Jun… [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US20/064139”, Mailed Date: Mar. 16, 2021, 12 Pages. [cited by applicant]
Peng, et al., “Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters”, In Proceedings of the Thirteenth EuroSys Conference, Apr. 23, 2018, 14 Pages. [cited by applicant]