IP Library › Granted Patent US 12,287,721
Granted Patent B2
US 12,287,721 · App. 17/585,695 · Granted Apr 29, 2025

Storage management and usage optimization using workload trends

Inventors: Yan Li (Beijing, CN); Run Qian Bj Chen (Beijing, CN); Chen Guang Zhao (Beijing, CN); Qin Qin Zhou (Beijing, CN); Guang Han Sui (Beijing, CN); Jing Li (Beijing, CN); You Bing Li (Beijing, CN); Yu Xiang Chen (Shanghai, CN)
Assignee: International Business Machines Corporation
G06F11/3442G06F9/505G06F9/5077
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,287,721
App. No.
17/585,695
Granted
Apr 29, 2025
Kind
B2
Abstract

Solutions preparing container images and data for container workloads prior to start times of workloads predicted through workload trend analysis. Local storage space on the node is managed based on workload trends, optimizing local storage of image files without requiring frequent reloading and/or deletion of image files, avoiding network intensive I/O operations when pulling images to local storage by workload scheduling systems. Systems perform collection of historical data including image and workload properties; analyze historical data for workload trends, including predicted start times, image files needed, number of nodes and types of nodes. Based on predicted future workload start times, nodes are selected from an ordered list of node requirements and workload properties. Selected nodes' local storage is managed using predicted future start times of workloads, to avoid removing image files having sooner start times, while removing (as needed) images files predictively utilized for workloads further into the future.

Claims (83)

1. A computer-implemented method for managing and optimizing container image storage, the computer-implemented method comprising:

inputting, by a processor, image and data information of a historical workload record into a machine learning model configured to predict future workload requirements based on workload trends of the historical workload record;

learning, by the processor, image and data requirement trends of the historical workload record using the machine learning model;

outputting, by the processor, predicted image and data requirements for future workloads based on the image and data requirement trends learned from the machine learning model;

engaging, by the processor, a checking cycle wherein a daemon process checks whether an image file for an upcoming future workload as predicted by the predicted image and data requirements needs to be downloaded to one or more nodes prior to running the upcoming future workload; and

triggering, by the processor, a pulling task upon a current local time plus a max recorded download time of the image file plus a pre-defined buffer time is less than or equal to a predicted start time of the upcoming future workload.

2. The computer-implemented method of claim 1 , wherein the pulling task comprises:

selecting, by the processor, a plurality of nodes with properties satisfying image requirements and workload requirements of the upcoming future workload;

ordering, by the processor, the plurality of nodes with the properties satisfying the image requirements and workload requirements by highest availability of one or more computing resources; and

iteratively checking, by the processor, the plurality of nodes sequentially, starting with a first node in the ordering of the plurality of nodes, for whether a selected node has the image file for the upcoming future workload stored by a local storage device of the selected node, wherein upon confirming the image file for the upcoming future workload is locally stored by the local storage device of the selected node, tasking the selected node to run the upcoming future workload on the selected node at the predicted start time of the upcoming future workload.

3. The computer-implemented method of claim 2 , further comprising:

upon confirming the image file for the upcoming future workload is not locally stored by the local storage device of the second node, comparing, by the processor, available space of the local storage device with image size of the image file for deploying the upcoming future workload;

upon comparing the available space of the local storage device with the image size of the image file, finding that the image size exceeds the available space of the local storage device:

triggering, by the processor, a disk utilization process; and

upon comparing the available space of the local storage device with the image size of the image file, finding that the image size is less that the available space on the local storage device:

pulling, by the processor, the image file from a registry; and

storing, by the processor, the image file to the local storage device.

4. The computer-implemented method of claim 3 , wherein finding that the image size is less that the available space on the local storage device further comprises:

determining, by the processor, network input/output (I/O) of the selected node satisfies network I/O requirements of the upcoming future workload.

5. The computer-implemented method of claim 3 , wherein the disk utilization process comprises:

setting, by the processor, a target amount of free space on the local storage device that meets or exceeds the image size;

mapping, by the processor, a relationship between existing image files stored on the local storage device and related workloads predicted to be scheduled to be run at a future point in time;

ordering, by the processor, the existing image files by predicted start times of related workloads, wherein earliest start time is ordered first, and latest start time is last;

while the available space on the local storage device is less than or equal to the target amount of free space, deleting, by the processor, a last image in the ordering of the existing image files until the available space on the local storage device exceeds the target amount of free space on the local storage device; and

pulling, by the processor, the image file from the registry to the local storage device.

6. The computer-implemented method of claim 1 , wherein the image and data requirements are selected from the group consisting of image required, node kind required, number of nodes required, a target deadline, network I/O requirement and a combination thereof.

7. The computer-implemented method of claim 1 , wherein the image and data information of the historical workload record are selected from the group consisting of image name, image size, image version, historical workload start time, download time, workload name, workload type, workload running time, kinds of nodes, number of nodes, and a combination thereof.

8. A computing program product for managing and optimizing container image storage comprising:

one or more computer-readable storage media having computer-readable program instructions stored on the one or more computer-readable storage media, said program instructions executes a computer-implemented method comprising:

inputting, by a processor, image and data information of a historical workload record into a machine learning model configured to predict future workload requirements based on workload trends of the historical workload record;

learning, by the processor, image and data requirement trends of the historical workload record using the machine learning model;

outputting, by the processor, predicted image and data requirements for future workloads based on the image and data requirement trends learned from the machine learning model;

engaging, by the processor, a checking cycle wherein a daemon process checks whether an image file for an upcoming future workload as predicted by the predicted image and data requirements needs to be downloaded to one or more nodes prior to running the upcoming future workload; and

triggering, by the processor, a pulling task upon a current local time plus a max recorded download time of the image file plus a pre-defined buffer time is less than or equal to a predicted start time of the upcoming future workload.

9. The computing program product of claim 8 , wherein the pulling task comprises:

selecting, by a processor, a plurality of nodes with properties satisfying image requirements and workload requirements of the upcoming future workload;

ordering, by the processor, the plurality of nodes with the properties satisfying image requirements and workload requirements by highest availability of one or more computing resources; and

iteratively checking, by the processor, the plurality of nodes sequentially, starting with a first node in the ordering of the plurality of nodes, for whether a selected node has the image file for the upcoming future workload stored by a local storage device of the selected node, wherein upon confirming the image file for the upcoming future workload is locally stored by the local storage device of the selected node, tasking the selected node to run the upcoming future workload on the selected node at the predicted start time of the upcoming future workload.

10. The computing program product of claim 9 , further comprising:

upon confirming the image file for the upcoming future workload is not locally stored by the local storage device of the second node, comparing, by the processor, available space of the local storage device with image size of the image file for the upcoming future workload;

upon comparing the available space of the local storage device with the image size of the image file, finding that the image size exceeds the available space of the local storage device:

triggering, by the processor, a disk utilization process; and

upon comparing the available space of the local storage device with the image size of the image file, finding that the image size is less that the available space on the local storage device:

pulling, by the processor, the image file from a registry; and

storing, by the processor, the image file to the local storage device.

11. The computing program product of claim 10 , wherein finding that the image size is less that the available space on the local storage device further comprises:

determining, by the processor, network input/output (I/O) of the selected node satisfies network I/O requirements of the upcoming future workload.

12. The computing program product of claim 10 , wherein the disk utilization process comprises:

setting, by the processor, a target amount of free space on the local storage device that meets or exceeds the image size;

mapping, by the processor, a relationship between existing image files stored on the local storage device and related workloads predicted to be scheduled to be run at a future point in time;

ordering, by the processor, the existing image files by predicted start times of related workloads, wherein earliest start time is ordered first, and latest start time is last;

while the available space on the local storage device is less than or equal to the target amount of free space, deleting, by the processor, a last image in the ordering of the existing image files until the available space on the local storage device exceeds the target amount of free space on the local storage device; and

pulling, by the processor, the image file from the registry to the local storage device.

13. The computing program product of claim 8 , wherein the image and data requirements are selected from the group consisting of image required, node kind required, number of nodes required, a target deadline, network I/O requirement and a combination thereof.

14. The computing program product of claim 8 , wherein the image and data information of the historical workload record are selected from the group consisting of image name, image size, image version, historical workload start time, download time, workload name, workload type, workload running time, kinds of nodes, number of nodes, and a combination thereof.

15. A computer system for managing and optimizing container image storage comprising:

a processor; and

a computer-readable storage media coupled to the processor, wherein the computer-readable storage media contains program instructions executing, via the processor, a computer-implemented method comprising:

inputting, by the processor, image and data information of a historical workload record into a machine learning model configured to predict future workload requirements based on workload trends of the historical workload record;

learning, by the processor, image and data requirement trends of the historical workload record using the machine learning model;

outputting, by the processor, predicted image and data requirements for future workloads based on the image and data requirement trends learned from the machine learning model;

engaging, by the processor, a checking cycle wherein a daemon process checks whether an image file for an upcoming future workload as predicted by the predicted image and data requirements needs to be downloaded to one or more nodes prior to running the upcoming future workload; and

triggering, by the processor, a pulling task upon a current local time plus a max recorded download time of the image file plus a pre-defined buffer time is less than or equal to a predicted start time of the upcoming future workload.

16. The computer system of claim 15 , wherein the pulling task comprises:

selecting, by a processor, a plurality of nodes with properties satisfying image requirements and workload requirements of the upcoming future workload;

ordering, by the processor, the plurality of nodes with the properties satisfying image requirements and workload requirements by highest availability of one or more computing resources; and

iteratively checking, by the processor, the plurality of nodes sequentially, starting with a first node in the ordering of the plurality of nodes, for whether a selected node has the image file for the upcoming future workload stored by a local storage device of the selected node, wherein upon confirming the image file for the upcoming future workload is locally stored by the local storage device of the selected node, tasking the selected node to run the upcoming future workload on the selected node at the predicted start time of the upcoming future workload.

17. The computer system of claim 16 , further comprising:

upon confirming the image file for the upcoming future workload is not locally stored by the local storage device of the second node, comparing, by the processor, available space of the local storage device with image size of the image file for the upcoming future workload;

upon comparing the available space of the local storage device with the image size of the image file, finding that the image size exceeds the available space of the local storage device:

triggering, by the processor, a disk utilization process; and

upon comparing the available space of the local storage device with the image size of the image file, finding that the image size is less that the available space on the local storage device:

pulling, by the processor, the image file from a registry; and

storing, by the processor, the image file to the local storage device.

18. The computer system of claim 17 , wherein finding that the image size is less that the available space on the local storage device further comprises:

determining, by the processor, network input/output (I/O) of the selected node satisfies network I/O requirements of the upcoming future workload.

19. The computer system of claim 17 , wherein the disk utilization process comprises:

setting, by the processor, a target amount of free space on the local storage device that meets or exceeds the image size;

mapping, by the processor, a relationship between existing image files stored on the local storage device and related workloads predicted to be scheduled to be run at a future point in time;

ordering, by the processor, the existing image files by predicted start times of related workloads, wherein earliest start time is ordered first, and latest start time is last;

while the available space on the local storage device is less than or equal to the target amount of free space, deleting, by the processor, a last image in the ordering of the existing image files until the available space on the local storage device exceeds the target amount of free space on the local storage device; and

pulling, by the processor, the image file from the registry to the local storage device.

20. The computer system of claim 15 , wherein the image and data requirements are selected from the group consisting of image required, node kind required, number of nodes required, a target deadline, network I/O requirement and a combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: LI, YAN; CHEN, RUN QIAN BJ; ZHAO, CHEN GUANG; ZHOU, QIN QIN; SUI, GUANG HAN; LI, JING; LI, YOU BING; CHEN, YU XIANG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058788/0453 →
Continuity (1)
Related Publication 20230236946A1 · Jul 27, 2023
References Cited (29)
US 8352608B1 · Keagy · 2013 [cited by examiner]
US 8732413B2 · Moon · 2014 [cited by applicant]
US 8914515B2 · Alapati · 2014 [cited by applicant]
US 9798827B2 · Liang · 2017 [cited by applicant]
US 10313472B2 · Westberg · 2019 [cited by applicant]
US 20080271038A1 · Rolia · 2008 [cited by applicant]
US 20120324092A1 · Brown · 2012 [cited by applicant]
US 20140215495A1 · Erich · 2014 [cited by examiner]
US 20170118610A1 · Schieman · 2017 [cited by examiner]
US 20170351546A1 · Habak · 2017 [cited by examiner]
US 20180300653A1 · Srinivasan · 2018 [cited by examiner]
US 20190363962A1 · Cimino · 2019 [cited by examiner]
US 20200241917A1 · Chen · 2020 [cited by examiner]
US 20210200593A1 · Allyn · 2021 [cited by examiner]
US 20220206873A1 · He · 2022 [cited by examiner]
US 20220283724A1 · Malamut · 2022 [cited by examiner]
“Configuring kubelet Garbage Collection”, downloaded from the Internet Jun. 10, 2020, 5 pages, <https://kubernetes.io/docs/concepts/cluster-administration/kubelet-garbage-collection/#image-collection>. [cited by applicant]
“Disk utilization in Docker for Mac”, downloaded from the Internet on Aug. 5, 2021, 3 pages, <https://docs.docker.com/docker-for-mac/space/>. [cited by applicant]
“High Performance Innovation | Altair PBS Works”, downloaded from the Internet on Oct. 6, 2021, 9 pages, <https://www.altair.com/pbs-works/>. [cited by applicant]
“Method and Apparatus of Prediction Based Dynamic Workload Mediator”, An IP.com Prior Art Database Technical Disclosure, Authors et al.: Disclosed Anonymously, IP.com No. IPCOM000260647D, IP.com Electronic Publication D… [cited by applicant]
“Monitoring the disk space used by IBM Workload Scheduler”, IBM Workload Automation, Version 9.5, downloaded from the Internet on Jun. 9, 2020, 4 pages, <https://www.ibm.com/support/knowledgecenter/SSGSPN_9.5.0/com.ibm.… [cited by applicant]
“Slurm Workload Manager—Documentation”, SchedMD, Last modified Aug. 20, 2021, 4 pages, <https://slurm.schedmd.com/>. [cited by applicant]
“What Should I Do If Image Re-pull Fails?”, Jun. 1, 2020, Huawei Cloud, 4 pages, <https://support.huaweicloud.com/intl/en-us/cce_faq/cce_faq_00015.html>. [cited by applicant]
Chen et al., “PSO-GA-Based Resource Allocation Strategy for Cloud-Based Software Services With Workload-Time Windows”, IEEE Access, vol. 8, 2020, date of publication Aug. 18, 2020, date of current version Aug. 28, 2020,… [cited by applicant]
Gmach et al., “Workload Analysis and Demand Prediction of Enterprise Data Center Applications”, 2007 IEEE 10th International Symposium on Workload Characterization, Sep. 29, 2007, 10 pages. [cited by applicant]
Kim et al., “Forecasting Cloud Application Workloads with CloudInsight for Predictive Resource Management”, IEEE Transactions on Cloud Computing, 2020, 16 pages. [cited by applicant]
Lu et al., “RVLBPNN: A Workload Forecasting Model for Smart Cloud Computing”, Research Article, Hindawi Publishing Corporation, Scientific Programming, vol. 2016, Article ID 5635673, 9 pages, http://dx.doi.org/10.1155/2… [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Noel et al., “Towards Self-Managing Cloud Storage with Reinforcement Learning”, 2019 IEEE International Conference on Cloud Engineering (IC2E), 10.1109/IC2E.2019.000-9, 11 pages. [cited by applicant]