IP Library › Granted Patent US 12,235,795
Granted Patent B2
US 12,235,795 · App. 17/816,276 · Granted Feb 25, 2025

File system provisioning for workload

Inventors: Sagar Venkappa Nyamagouda (Karnataka, IN); Smitha Jayaram (Karnataka, IN); Hiro Rameshlal Lalwani (Karnataka, IN); Rachit Gupta (Karnataka, IN); Sherine Jacob (Karnataka, IN); Anand Andaneppa Ganjihal (Karnataka, IN)
Assignee: Hewlett Packard Enterprise Development LP
G06F16/13G06F11/3414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,235,795
App. No.
17/816,276
Granted
Feb 25, 2025
Kind
B2
Abstract

In some examples, a system receives workload information of a workload collection, and applies a machine learning model on the workload information, the machine learning model trained using training information including features of different types of workloads. The system produces, by the machine learning model, an identification of a first file system from among different types of file systems, the machine learning model producing an output value corresponding to the first file system that is a candidate for use in storing files of the workload collection.

Claims (68)

1. A non-transitory machine-readable storage medium comprising instructions that upon execution cause a system to:

receive workload information of a workload collection;

apply a machine learning model on the workload information, the machine learning model trained using training information including features of different types of workloads;

generate, by the machine learning model based on the workload information, a first file system label selected from among a plurality of file system labels representing different types of file systems;

select, from among the different types of file systems, a first file system represented by the first file system label produced by the machine learning model from the workload information;

provision the first file system in a computing system; and

establish a connection between the provisioned first file system in the computing system and an entity that provided the workload collection to allow reading and writing of files in the provisioned first file system by the entity.

2. The non-transitory machine-readable storage medium of claim 1 , wherein the instructions upon execution cause the system to:

train the machine learning model comprising:

extracting the features from input information, wherein the extracted features comprise features relating to input/output (I/O) requests of the different types of workloads, and

producing feature vectors of the extracted features.

3. The non-transitory machine-readable storage medium of claim 2 , wherein the training information comprises the feature vectors that are labeled with respect to different workload domains.

4. The non-transitory machine-readable storage medium of claim 1 , wherein the instructions upon execution cause the system to:

identify, by the machine learning model based on the workload information, a first type of workload domain from among a plurality of different types of workload domains; and

generate the first file system label based on the identified first type of workload domain using mapping information that maps the different types of workload domains to the different types of file systems.

5. The non-transitory machine-readable storage medium of claim 2 , wherein the features comprise any or some combination of: a size of a read, a size of a write, an indication of whether a workload is read intensive or write intensive, and an I/O pattern.

6. The non-transitory machine-readable storage medium of claim 2 , wherein the training of the machine learning model further comprises:

determining a measure of a performance of the machine learning model, and

responsive to the measure not satisfying a criterion, perform another iteration of the training, wherein the training is re-iterated until a measure of the performance of the machine learning model satisfies the criterion in a current iteration.

7. The non-transitory machine-readable storage medium of claim 6 , wherein each iteration of a plurality of iterations of the training updates parameters of the machine learning model.

8. The non-transitory machine-readable storage medium of claim 6 , wherein the measure of the performance of the machine learning model comprises a loss produced by a loss function.

9. The non-transitory machine-readable storage medium of claim 1 , wherein the instructions upon execution cause the system to:

assign, by the machine learning model, values to respective file system labels of the plurality of file system labels, wherein each value of the values indicates a likelihood of selecting a respective file system of the different types of file systems for the workload collection.

10. The non-transitory machine-readable storage medium of claim 9 , wherein the values comprise probability values, wherein each respective probability value of the probability values represents a likelihood that a corresponding file system label of the of the plurality of file system labels is a correct label.

11. The non-transitory machine-readable storage medium of claim 1 , wherein the instructions upon execution cause the system to:

generate, by the machine learning model based on the workload information, multiple file system labels selected from among the plurality of file system labels representing the different types of file systems, the multiple file system labels comprising the first file system label;

based on the multiple file system labels generated by the machine learning model from the workload information, identify multiple candidate file systems from the different types of file systems;

compare predicted performance measures of the multiple candidate file systems; and

select the first file system from the multiple candidate file systems based on the comparing.

12. The non-transitory machine-readable storage medium of claim 1 , wherein the instructions upon execution cause the system to:

establish the connection by creating a mount point at which the provisioned first file system is accessible.

13. The non-transitory machine-readable storage medium of claim 1 , wherein the instructions upon execution cause the system to:

buffer input/output (I/O) requests of the workload collection in a virtual file system prior to the applying of the machine learning model on the workload information; and

after the provisioning of the first file system, provide the buffered I/O requests to the provisioned first file system.

14. The non-transitory machine-readable storage medium of claim 1 , wherein the provisioning of the first file system is on-demand provisioning of the first file system for the workload collection.

15. A system comprising:

a processor; and

a non-transitory machine-readable storage medium storing instructions executable on the processor to:

receive workload information of a workload collection;

apply a machine learning model on the workload information, the machine learning model trained using training information including features of different types of workloads;

generate, by the machine learning model based on the workload information, a first file system label selected from among a plurality of file system labels representing different types of file systems;

select, from among the different types of file systems, a first file system represented by the first file system label produced by the machine learning model from the workload information;

provision the first file system in a computing system; and

establish a connection between the provisioned first file system in the computing system and an entity that provided the workload collection to allow reading and writing of files in the provisioned first file system by the entity.

16. The system of claim 15 , wherein the instructions are executable on the processor to:

identify, by the machine learning model based on the workload information, a first type of workload domain from among a plurality of different types of workload domains; and

generate the first file system label based on the identified first type of workload domain using mapping information that maps the different types of workload domains to the different types of file systems.

17. The system of claim 15 , wherein the instructions are executable on the processor to:

assign, by the machine learning model, probability values to respective file system labels of the plurality of file system labels, wherein each probability value of the probability values indicates a likelihood of selecting a respective file system of the different types of file systems for the workload collection; and

identify, based on the probability values, that the first file system is optimal for the workload collection.

18. The system of claim 15 , wherein the instructions are executable on the processor to:

generate, by the machine learning model based on the workload information, multiple file system labels selected from among the plurality of file system labels representing the different types of file systems, the multiple file system labels comprising the first file system label;

identify, based on the multiple file system labels generated by the machine learning model, multiple candidate file systems from the different types of file systems;

compare predicted performance measures of the multiple candidate file systems; and

select the first file system from the multiple candidate file systems based on the comparing.

19. A method of a system comprising a hardware processor, comprising:

training a machine learning model using training information including features of different types of workloads;

receiving workload information of a workload collection;

applying the trained machine learning model on the workload information;

generating, by the trained machine learning model based on the workload information, a first file system label selected from among a plurality of file system labels representing different types of file systems;

selecting, from among the different types of file systems, a first file system represented by the first file system label produced by the trained machine learning model from the workload information;

provisioning the first file system in a computing system; and

establishing a connection between the provisioned first file system in the computing system and an entity that provided the workload collection to allow reading and writing of files in the provisioned first file system by the entity.

20. The method of claim 19 , comprising:

generating, by the trained machine learning model based on the workload information, multiple file system labels selected from among the plurality of file system labels representing the different types of file systems, the multiple file system labels comprising the first file system label;

identifying, by the system based on the multiple file system labels generated by the machine learning model, multiple candidate file systems from the different types of file systems;

comparing, by the system, predicted performance measures of the multiple candidate file systems; and

selecting, by the system, the first file system from the multiple candidate file systems based on the comparing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2022
From: VENKAPPA NYAMAGOUDA, SAGAR; JAYARAM, SMITHA; RAMESHLAL LALWANI, HIRO; GUPTA, RACHIT; JACOB, SHERINE; ANDANEPPA GANJIHAL, ANAND
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 060740/0884 →
Continuity (1)
Related Publication 20240037067A1 · Feb 1, 2024
References Cited (32)
US 20050086450A1 · Shiota · 2005 [cited by examiner]
US 20150205637A1 · Lee · 2015 [cited by examiner]
US 20150326500A1 · Jackson · 2015 [cited by examiner]
US 20150379420A1 · Basak · 2015 [cited by examiner]
US 20200104189A1 · Gopalan · 2020 [cited by examiner]
US 20200234144A1 · Such · 2020 [cited by examiner]
US 20200242000A1 · Khosrowpour · 2020 [cited by examiner]
US 20220043752A1 · Hua · 2022 [cited by examiner]
US 20220261295A1 · Nagaraja · 2022 [cited by examiner]
US 20220334944A1 · Martynov · 2022 [cited by examiner]
US 20230022884A1 · Faizian · 2023 [cited by examiner]
US 20230069593A1 · Sherwood, Jr. · 2023 [cited by examiner]
US 20230236946A1 · Li · 2023 [cited by examiner]
US 20230281470A1 · Hwang · 2023 [cited by examiner]
J. Gao, H. Wang and H. Shen, “Machine Learning Based Workload Prediction in Cloud Computing,” 2020 29th International Conference on Computer Communications and Networks (ICCCN), Honolulu, HI, USA, 2020, pp. 1-9, doi: 10… [cited by examiner]
Dryden et al., Clairvoyant Prefetching for Distributed Machine Learning I/O, pp. 1-14, SC'21, Nov. 14-19, 2021, St. Louis, MO. (Year: 2021). [cited by examiner]
Vaghani, S., Virtual machine file system, ACM SIGOPS Operating Systems Reviewvol. 44Issue Dec. 4, 2010pp. 57-70. (Year: 2010). [cited by examiner]
Active State, “What is a Keras model”, Jul. 11, 2022, retrieved from: https://www.activestate.com/resources/quick-reads/what-is-a-keras-model/, pp. 20. [cited by applicant]
Ajagekar, A., “Adam”, Cornell University Computational Optimization Open Textbook last edited on Dec. 16, 2021, retrieved from: https://optimization.cbe.comnell.edu/index.php?title=Adam&oldid=6033, pp. 7. [cited by applicant]
Amazon Web Services; “Choosing an Amazon FSx File System”; retrieved from: https://aws.amazon.com/fsx/when-to-choose-fsx/, retrieved on: Jun. 7, 2022, pp. 8. [cited by applicant]
Amazon Web Services; “Select the Correct Resource Type, Size, and Number”; retrieved from: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/select-the-correct-resource-type-size-and-number.htm… [cited by applicant]
Bannier, Understanding Adam: how loss functions are minimized? Adam : a method for stochastic optimization, Towards Data Science, Aug. 10, 2019 retrieved from: https://towardsdatascience.com/understanding-adam-how-loss-… [cited by applicant]
Google LLC, “Parallel file systems for HPC workloads”, retrieved from: https://cloud.google.com/architecture/parallel-file-systems-for-hpc, retrieved on: Jun. 7, 2022, pp. 7. [cited by applicant]
Jayakumar, N., et al., “Workload Characteristics Impacts on file System Benchmarking”, International Journal of Advanced Research in Computer Science and Software Engineering, vol. 4, Issue 2, Feb. 2014, pp. 39-44. [cited by applicant]
Ozmen, O., et al., “Workload-aware storage layout for database systems”, SIGMOD '10: Proceedings of the 2010 ACM SIGMOD International Conference on Management of data, Jun. 15, 2010, pp. 939-950. [cited by applicant]
Raul Gomez Blog, Understanding Categorical Cross-Entropy Loss, Binary Cross-Entropy Loss, Softmax Loss, Logistic Loss, Focal Loss and all those confusing names, May 23, 2018, pp. 1-12. [cited by applicant]
Seo, B., et al., IO Workload Characterization Revisited: A Data-Mining Approach, IEEE, vol. 63, Issue 12, Dec. 2014, pp. 3026-3038. [cited by applicant]
Wikipedia, “AdvFS”, retrieved from: https://en.wikipedia.org/wiki/AdvFS, retrieved on: Sep. 27, 2022, pp. 3. [cited by applicant]
Wikipedia, “Deep learning”, retrieved from: https://en.wikipedia.org/wiki/Deep_learning, retrieved on; Sep. 27, 2022, pp. 1-40. [cited by applicant]
Wikipedia, “ext4”, retrieved from: https://en.wikipedia.org/wiki/Ext4, retrieved on: Sep. 27, 2022, pp. 9. [cited by applicant]
Yalcin, O.G., “3 Ways to Build Neural Networks in TensorFlow with the Keras API”, Towards Data Science, retrieved from: https://towardsdatascience.com/3-ways-to-build-neural-networks-in-tensorflow-with-the-keras-api-80e… [cited by applicant]
Zhu, Y., “A novel approach to workload prediction using attention-based LSTM encoder-decoder network in cloud environment”, EURASIP Journal on Wireless Communications and Networking, Dec. 17, 2019, pp. 1-18. [cited by applicant]