IP Library › Granted Patent US 12,498,991
Granted Patent B2
US 12,498,991 · App. 17/935,124 · Granted Dec 16, 2025

Methods and systems for identifying multiple workloads in a heterogeneous environment

Inventors: Roshan R. Nair (Karnataka, IN); Arun George (Karnataka, IN); Vishak Guddekoppa (Karnataka, IN)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F9/5083G06F9/5033G06F18/24133
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,498,991
App. No.
17/935,124
Granted
Dec 16, 2025
Kind
B2
Abstract

A method for identifying a plurality of workloads in a heterogenous environment includes collecting a plurality of parameters from at least one layer of a system stack associated with the plurality of workloads, correlating the collected plurality of parameters from different layers of software stack, and creating a feature set based on the correlated plurality of parameters. The method further includes processing the feature set using a successively ordered classifier chain (SOCC) module to identify the presence of the plurality of workloads in the heterogenous environment in a data center.

Claims (63)

1 . A method for identifying a plurality of workloads in a heterogenous environment, the method comprising:

collecting, by a computing device, a plurality of parameters from at least one layer of a system stack associated with the plurality of workloads;

correlating, by the computing device, the collected plurality of parameters;

creating, by the computing device, a feature set based on the correlated plurality of parameters;

processing, by the computing device, the feature set using a successively ordered classifier chain (SOCC) module wherein presence of the plurality of workloads in the heterogenous environment is identified,

training, by the computing device, the SOCC module by forming a plurality of classifier modules in a successively ordered chain,

wherein the training the SOCC module comprises:

creating a dataset with ‘M’ number of labels and ‘N’ number of features, wherein the ‘M’ number of labels corresponds to ‘M’ number of workloads, and wherein the ‘M’ and the ‘N’ are integers;

forming ‘M’ number of classifier modules that correspond to the ‘M’ number of labels with the ‘N’ number of features as input features and enabling the ‘M’ number of classifier modules to identify presence of each of the ‘M’ number of labels;

analyzing an accuracy of each of the ‘M’ number of classifier modules;

selecting a classifier module that has a highest accuracy from the ‘M’ number of classifier modules, wherein a label identified by the selected classifier module is a first label in a chain; and

recursively performing steps of:

selecting a number of labels by discarding the label identified by the selected classifier module in a previous step;

forming a number of classifier modules that correspond to the selected number of labels with the ‘N’ number of features and the label identified by the selected classifier module in the previous step as the input features and enabling the formed number of classifier modules to identify presence of each of the number of labels; and

selecting a classifier module from the formed number of classifier modules that has the highest accuracy, wherein the selected classifier module is a subsequent label in the chain,

until classifier modules are selected for all the ‘M’ number of labels.

2 . The method of claim 1 , wherein the collecting, by the computing device, of the plurality of parameters includes:

selecting the plurality of parameters from a set of parameters associated with the at least one layer of the system stack; and

periodically collecting the plurality of parameters from the at least one layer of the system stack.

3 . The method of claim 1 , wherein the correlating, by the computing device, of the collected plurality of parameters includes:

correlating the plurality of parameters by choosing key transition points and key arguments in each layer of the system stack.

4 . The method of claim 1 , wherein the feature set serves as a signature for each workload.

5 . The method of claim 4 , wherein the feature set includes at least one of a number of overwrites, read-copy updates (RCUs), a warmness of data, continuous points, break points, an average segment length, a standard deviation of a block address, a continued to break point ratio, and an average block size of InputOutputs (IOs), wherein the number of overwrites indicates a number of writes in a given range of a unit of data, and the warmness of data indicates that the data has been recently written or accessed.

6 . The method of claim 4 , wherein the creating of the feature set includes:

obtaining a data point of a trace event from a lower layer that is lower than the at least one layer in the system stack, wherein the data point includes a block address, a number of blocks, and an operation type, if the lower layer is a device driver layer, wherein the data point includes a file name, a file offset and an operation type, if the lower layer is a file system layer, wherein the data point includes a file name, a file offset, an operation type, a block address and a number of blocks, if the lower layer is a file system to a device driver interface layer;

co-relating the obtained data point to data points of trace events collected in at least one of upper layers that are above the lower layer in the system stack; and

recursively comparing spatial and temporal locality of the trace events of the upper layers and adding the trace events to the feature set until the feature set is created by comparing all the upper layers.

7 . A computing device, comprising:

a memory; and

a processor coupled to the memory, wherein the processor is configured to:

collect a plurality of parameters from at least one layer of a system stack associated with a plurality of workloads;

correlate the collected plurality of parameters;

create a feature set based on the correlated plurality of parameters;

process the feature set using a successively ordered classifier chain (SOCC) module wherein presence of the plurality of workloads in a heterogenous environment is identified,

train the SOCC module by forming a plurality of classifier modules in a successively ordered chain,

wherein the processor is configured to:

create a dataset with ‘M’ number of labels and ‘N’ number of features, wherein the ‘M’ number of labels corresponds to ‘M’ number of workloads, and wherein the ‘M’ and the ‘N’ are integers;

form ‘M’ number of classifier modules that correspond to the ‘M’ number of labels with the ‘N’ number of features as input features and enabling the ‘M’ number of classifier modules to identify presence of each of the ‘M’ number of labels;

analyze an accuracy of each of the ‘M’ number of classifier modules;

select a classifier module that has a highest accuracy from the ‘M’ number of classifier modules, wherein a label identified by the selected classifier module is a first label in a chain; and

recursively perform steps of:

selecting a number of labels by discarding the label identified by the selected classifier module in a previous step;

forming a number of classifier modules that correspond to the selected number of labels with the ‘N’ number of features and the label identified by the selected classifier module in the previous step as the input features and enabling the formed number of classifier modules to identify presence of each of the number of labels; and

selecting a classifier module from the formed number of classifier modules that has the highest accuracy, wherein the selected classifier module is a subsequent label in the chain,

until classifier modules are selected for all the ‘M’ number of labels.

8 . The computing device of claim 7 , wherein the processor is configured to:

select the plurality of parameters from a set of parameters associated with the at least one layer of the system stack; and

periodically collect the plurality of parameters from the at least one layer of the system stack.

9 . The computing device of claim 7 , wherein the processor is configured to:

correlate the plurality of parameters by choosing key transition points and key arguments in each layer of the system stack.

10 . The computing device of claim 7 , wherein the feature set serves as a signature for each workload.

11 . The computing device of claim 10 , wherein the feature set includes at least one of a percentage of overwrites, read-copy updates (RCUs), a warmness of data, continuous points, break points, an average segment length, a standard deviation of block address, a continued to break point ratio, and an average block size of InputOutputs (IOs), wherein the number of overwrites indicates a number of writes in a given range of a unit of data, and the warmness of data indicates that the data has been recently written or accessed.

12 . The computing device of claim 10 , wherein the processor is configured to:

obtain a data point of a trace event from a lower layer that is lower than the at least one layer in the system stack, wherein the data point includes a block address, a number of blocks, and an operation type, if the lower layer is a device driver layer, wherein the data point includes a file name, a file offset and an operation type, if the lower layer is a file system layer, wherein the data point includes a file name, a file offset, an operation type, a block address and a number of blocks, if the lower layer is a file system to a device driver interface layer;

co-relate the obtained data point to data points of trace events collected in at least one upper layers that are above the lower layer in the system stack; and

recursively compare spatial and temporal locality of the trace events of the upper layers and co-relate, and add the trace events to the feature set, until the feature set is created by comparing all the upper layers.

13 . A method of training, by a computing device, a successively ordered classifier chain (SOCC) module that identifies presence of a plurality of workloads in a heterogenous environment, the method comprising:

creating a dataset with ‘M’ number of labels and ‘N’ number of features, wherein the ‘M’ number of labels corresponds to ‘M’ number of workloads and wherein the ‘M’ and the ‘N’ are integers;

forming ‘M’ number of classifier modules that correspond to the ‘M’ number of labels with the ‘N’ number of features as input features and enabling the ‘M’ number of classifier modules to identify presence of each of the ‘M’ number of labels;

analyzing an accuracy of each of the ‘M’ number of classifier modules;

selecting a classifier module that has a highest accuracy from the ‘M’ number of classifier modules, wherein a label identified by the selected classifier module is a first label in a chain; and recursively performing steps of: selecting a number of labels by discarding the label identified by the selected classifier module in a previous step;

forming a number of classifier modules that correspond to the selected number of labels with the ‘N’ number of features and the label identified by the selected classifier module in the previous step as the input features and enabling the formed number of classifier modules to identify presence of each of the number of labels; and

selecting a classifier module from the formed number of classifier modules that has the highest accuracy, wherein the selected classifier module is a subsequent label in the chain, until classifier modules are selected for all the ‘M’ number of labels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2022
From: NAIR, ROSHAN R.; GEORGE, ARUN; GUDDEKOPPA, VISHAK
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061206/0466 →
Priority Claims (1)
IN 202241041897 · Jul 21, 2022 · national
Continuity (1)
Related Publication 20240028419A1 · Jan 25, 2024
References Cited (28)
US 10884636B1 · Abrol et al. · 2021 [cited by applicant]
US 20140047095A1 · Breternitz · 2014 [cited by examiner]
US 20150379420A1 · Basak · 2015 [cited by examiner]
US 20160078361A1 · Brueckner · 2016 [cited by examiner]
US 20170126795A1 · Kumar et al. · 2017 [cited by applicant]
US 20180018339A1 · Basak et al. · 2018 [cited by applicant]
US 20190220703A1 · Prakash et al. · 2019 [cited by applicant]
US 20200028935A1 · Sahay et al. · 2020 [cited by applicant]
US 20200099692A1 · Mahindru et al. · 2020 [cited by applicant]
US 20200296182A1 · Khosrowpour · 2020 [cited by examiner]
US 20200410288A1 · Capota · 2020 [cited by examiner]
US 20210037107A1 · Klenk · 2021 [cited by examiner]
US 20210224245A1 · Sharma · 2021 [cited by examiner]
US 20220004320A1 · Matosevich · 2022 [cited by examiner]
US 20230074802A1 · Comer · 2023 [cited by examiner]
US 20230394352A1 · Byrne · 2023 [cited by examiner]
US 20240007414A1 · Jain · 2024 [cited by examiner]
CA 2426439 · 2003 [cited by applicant]
Thang Le Duc, et al., “Machine Learning Methods for Reliable Resource Provisioning in Edge-Cloud Computing: a Survey,” ACM Computing Surveys, vol. 0. No. 0. Artcle 1, Publication Date Jun. 2019. [cited by applicant]
PDAWL: Profile-based Iterative Dynamic Adaptive WorkLoad Balance on Heterogeneous Architectures, JSSPP 2020, Job Scheduling Strategies for Parallel Processing pp. 145-162. [cited by applicant]
Jayanta Basak, Kushal Wadhwani, and Kaladhar Voruganti. 2016. Storage Workload Identification, ACM Transactions on Storage, vol. 12, No. 3, Article 14, Publication date: May 2016, pp. 14.1-14.30. [cited by applicant]
Bumjoon Seo, Sooyong Kang, Jongmoo Choi, Jaehyuk Cha, Youjip Won, and Sungroh Yoon, IO Workload Characterization Revisited: A Data-Mining Approach, IEEE Transactions on Computers, vol. 63, No. 12. Dec. 2014, pp. 3026-30… [cited by applicant]
Keikha, Maryam and Hashemi, Sattar (2016). Ordered Classifier Chains for Multi-label Classification, Journal of Machine Intelligence 1 (1): 7-12, 2016 | https://Isp.institute. [cited by applicant]
NetApp Solutions, Mar. 11, 2022, pp. 1-883. [cited by applicant]
Office Action dated May 6, 2025 issued in corresponding Indian Patent Application No. 202241041897. [cited by applicant]
Bu et al., “Rapid Deployment of Anomaly Detection Models for Large Number of Emerging KPI Streams,” Publication Date: Dec. 31, 2018. [cited by applicant]
Alguliyev et al., “Hybridisation of classifiers for anomaly detection in big data,” Publication Date: Dec. 31, 2019: pp. 11-19. [cited by applicant]
Pham et al., “Recurrent Neural Network for Classifying of HPC Applications,” Publication: Dec. 31, 2019. [cited by applicant]