IP Library Granted Patent US 12,596,570
Granted Patent B2
US 12,596,570 · App. 17/807,145 · Granted Apr 7, 2026

Optimized storage caching for computer clusters using metadata

Inventors: Joseph W. Dain (Tucson, AZ); Simon Lorenz (Biebertal, DE); Piyush Chaudhary (Highland, NY); Gero Friedrich Wolf Schmidt (Mainz, DE); Qais Noorshams (Frankfurt, DE); Gregory T. Kishi (Oro Valley, AZ)
Assignee: International Business Machines Corporation
G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,570
App. No.
17/807,145
Granted
Apr 7, 2026
Kind
B2
Abstract

A method, computer system, and a computer program for managing computer jobs in a queue is provided. This comprises extracting metadata from a new job received for processing and upon determining when a similar enriched metadata exists in a database. A job score and storage footprint may then be determined for the new job from the extracted metadata. It is then determined whether the new job can be grouped for processing with any other jobs already placed on a queue. The new job is then added to the queue based on the new job's score and footprint, and whether it can be grouped with other jobs. The queue is then updated and sent to a scheduler for further processing.

Claims (54)

1 . A method for managing computer jobs in a queue comprising:

extracting current metadata from a new job received for processing;

providing the extracted current metadata from said new job to a self-learning artificial intelligence (AI) system for determining when a similar enriched metadata exists in a database, wherein availability of enriched data is determined by researching said database according to labels used to classify contents of said enriched data;

further comprising generating enriched metadata from said extracted current metadata and updating said database;

determining a job score associated with said received new job by said self-learning AI system based on said extracted current metadata and the generated enriched metadata;

determining a storage footprint for said new job based on said extracted current metadata;

determining whether said new job can be grouped and labelled similarly for processing with a plurality of pending jobs disposed on a queue;

analyzing information about job groupings and an overall storage footprint of the plurality of pending jobs disposed on the queue to provide an optimal solution by said self-learning AI system for processing the plurality of pending jobs,

wherein said provided optimal solution includes both speed and cost considerations of storage and groupings of said plurality of pending jobs;

adding said new job to said queue and sorting said queue based on said provided optimal solution,

wherein the adding of said job to said queue comprises a placement of said new job on said queue is based on a combination of said provided optimal solution, said determined score of the new job, said determined storage footprint, said determined grouping of the new job with the pending jobs disposed on said queue for processing, and said similar enriched metadata;

process said new job and the pending jobs according to the placement of said new job on the queue.

2 . The method of claim 1 , wherein said queue is updated after sorting and said updated queue is sent to a scheduler for further job processing.

3 . The method of claim 2 , wherein said job score is determined from at least one of a job priority, previous job completion record, and/or job complexity.

4 . The method of claim 1 , wherein said similar jobs are also grouped together for processing when said queue is being sorted and updated.

5 . The method of claim 4 , wherein the queue is also sorted based on a plurality of prefetch and data evict needs of each job.

6 . The method of claim 5 , wherein said new job is grouped with other jobs on said queue based on prefetch and data evict needs of said jobs in said queue.

7 . The method of claim 1 , further comprising:

updating and resorting said queue every time one of said jobs on said queue is completed.

8 . The method of claim 7 , wherein after said job completion, evicting data from said storage and determining whether said new job received may be regrouped based on said data eviction in said queue.

9 . The method of claim 8 , further comprising determining an order of said new job received in said queue after any job is completed in said queue by re-determining said score associated with said new job as appropriate and said new job's storage footprint.

10 . The method of claim 8 , wherein said queue is updated and provided to a scheduler for further data and job processing.

11 . A computer system, comprising: one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method for managing computer jobs in a queue comprising:

extracting current metadata from a new job received for processing;

providing the extracted current metadata from said new job to a self-learning artificial intelligence (AI) system for determining when a similar enriched metadata exists in a database, wherein availability of enriched data is determined by researching said database according to labels used to classify contents of said enriched data;

further comprising generating enriched metadata from said extracted current metadata and updating said database;

determining a job score associated with said received new job by said self-learning AI system based on said extracted current metadata and the generating enriched metadata;

determining a storage footprint for said received new job based on said extracted current metadata;

determining whether said new job can be grouped and labelled similarly for processing with a plurality of pending jobs disposed on a queue;

analyzing information about job groupings and an overall storage footprint to provide an

optimal solution by said self-learning AI system for processing of the plurality of pending jobs disposed on the queue, wherein said optimal solution includes both speed and cost considerations and groupings of said plurality of pending jobs;

adding said new job to said queue and sorting said queue based on said provided optimal solution,

wherein the adding of said job to said queue comprises a placement of said new job on said queue is based on a combination of said provided optimal solution and said determined score of the new job, and said determined storage footprint, and said determined grouping of the new job with the pending jobs disposed on said queue for processing and said similar enriched metadata;

process said new job and the pending jobs according to the placement of said new job on the queue.

12 . The computer system of claim 11 , wherein said queue is updated after sorting and said updated queue is sent to a scheduler for further job processing.

13 . The computer system of claim 11 , wherein said job score is determined from at least one of a job priority, previous job completion record, and/or job complexity.

14 . The computer system of claim 11 , wherein said storage footprint is a cache footprint.

15 . The computer system of claim 11 , wherein said similar jobs are also grouped together for processing when said queue is being sorted and updated.

16 . A computer program product, comprising: one or more non-transitory computer-readable storage media and program instructions stored on at least one or more tangible storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

extracting current metadata from a new job received for processing;

providing the extracted current metadata from said new job to a self-learning artificial intelligence (AI) system for determining when a similar enriched metadata exists in a database, wherein availability of enriched data is determined by researching said database according to labels used to classify contents of said enriched data;

further comprising generating enriched metadata from said extracted current metadata and updating said database;

determining a job score associated with said received new job by said self-learning AI system based on said extracted current metadata and the generated enriched metadata, wherein availability of said enriched data is determined by researching said database according to labels used to classify contents of said enriched data;

further comprising generating enriched metadata from said extracted current metadata and updating said database;

determining a storage footprint for said new job based on said extracted current metadata;

determining whether said new job can be grouped and labelled similarly for processing with a plurality of pending jobs disposed on a queue;

analyzing information about job groupings and an overall storage footprint of the plurality of pending jobs disposed on the queue to provide an optimal solution by said self-learning AI system for processing the plurality of pending jobs, wherein provided said optimal solution includes both speed and cost considerations of storage and groupings of said

plurality of pending jobs;

adding said new job to said queue and sorting said queue based on said provided optimal solution,

wherein the adding of said job to said queue comprises a placement of said new job on said queue is based on a combination of said provided optimal solution, said determined score of the new job, said determined storage footprint, and said determined grouping of the new job with the pending jobs disposed on said queue for processing, and said similar enriched metadata;

process said new job and the pending jobs according to the placement of said new job on the queue.

17 . The computer program product of claim 16 , wherein said queue is updated after sorting and said updated queue is sent to a scheduler for further job processing.

18 . The computer program product of claim 16 , wherein said job score is determined from at least one of a job priority, previous job completion record, and/or job complexity.

19 . The computer program product of claim 16 , wherein said storage footprint is a cache footprint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2022
From: DAIN, JOSEPH W.; LORENZ, SIMON; CHAUDHARY, PIYUSH; SCHMIDT, GERO FRIEDRICH WOLF; NOORSHAMS, QAIS; KISHI, GREGORY T.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060218/0043 →
Continuity (1)
Related Publication 20230409384A1 · Dec 21, 2023
References Cited (25)
US 9158540B1 · Tzelnic · 2015 [cited by applicant]
US 9442954B2 · Guha · 2016 [cited by applicant]
US 9456049B2 · Soundararajan · 2016 [cited by applicant]
US 10467569B2 · Voss · 2019 [cited by applicant]
US 10846752B2 · Tsai · 2020 [cited by applicant]
US 11175950B1 · Yang · 2021 [cited by applicant]
US 20160246655A1 · Kimmel · 2016 [cited by examiner]
US 20170262896A1 · Tsai · 2017 [cited by applicant]
US 20180253219A1 · Dotan-Cohen · 2018 [cited by examiner]
US 20190303200A1 · Sitaraman · 2019 [cited by applicant]
US 20190392353A1 · Liu · 2019 [cited by examiner]
US 20210257052A1 · Van Rooyen · 2021 [cited by applicant]
US 20210357155A1 · You · 2021 [cited by examiner]
US 20220083271A1 · Cho · 2022 [cited by examiner]
US 20220179585A1 · Muthiah · 2022 [cited by examiner]
US 20230236759A1 · Lathrop · 2023 [cited by examiner]
Carstens, “High Performance Co-operative Cluster Computing Through Migration of Jobs or Computing Nodes,” IP.com, IP.com No. IPCOM000212445D, IP.com Publication Date: Nov. 14, 2011, 8 pages. [cited by applicant]
Disclosed Anonymously, “Method for Cache Replication During Swap Node Addition with Improved Application Performance in the Storage System,” IP.com, IP.com No. IPCOM000267474D, IP.com Publication Date: Oct. 29, 2021, 11… [cited by applicant]
Disclosed Anonymously, “Method to Improve Application Performance in the Storage Subsystems with Improved Deduplication Metadata Management,” IP.com No. IPCOM000267629D, IP.com Publication Date: Nov. 11, 2021, 10 pages. [cited by applicant]
Helland, “Cosmos—Big Data and Big Challenges,” Microsoft, Jul. 2011, https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/en-us-events-fs2011-helland_cosmos_big_data_and_big_challenges.pdf, 27 pages. [cited by applicant]
Lorenz et al., “Data Accelerator for AI and Analytics,” Redbooks, Jan. 2021, https://www.redbooks.ibm.com/redpapers/pdfs/redp5623.pdf, 88 pages. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Ren et al., “Scaling File System Metadata Performance With Stateless Caching and Bulk Insertion,” Parallel Data aboratory, Carnegie Mellon University, CMU-PDL-14-103, May 2014, 22 pages. [cited by applicant]
Zhu et al, “HBA: Distributed Metadata Management for Large Cluster-Based Storage Systems,” IEEE Transactions on Parallel and Distributed Systems, vol. 19, No. 4, Apr. 2008, 14 pages. [cited by applicant]
IBM, “IBM Spectrum LSF Data Manager,” IBM.com, Last Updated: Feb. 7, 2022, https://www.ibm.com/docs/en/spectrum-lsf/10.1.0?topic=lsf-data-manager, 2 pages. [cited by applicant]