IP Library › Granted Patent US 12,530,596
Granted Patent B2
US 12,530,596 · App. 16/671,791 · Granted Jan 20, 2026

Adaptive usage of storage resources using data source models and data source representations

Inventors: Sashka T. Davis (Vienna, VA); Naveen Sunkavally (Cary, NC); Zulfikar A. Ramzan (Saratoga, CA)
Assignee: EMC IP Holding Company LLC
G06N5/01G06F16/24553G06F16/285G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,596
App. No.
16/671,791
Filed
Nov 1, 2019
Granted
Jan 20, 2026
Kind
B2
Art Unit
2121
USPC
706/12
Abstract

Techniques are provided for adaptive usage of storage resources using data source models and data source representations generated using the data source models. One method comprises obtaining sampled data generated by sampling source data from a data source; fitting a data model to the sampled data to obtain a representation of the sampled data; obtaining a classification of the sampled data into one of multiple predefined retention models; and adapting a usage of one or more storage resources that store the retained data based on the representation and the classification. The adaptive storage resource usage may comprise, for example: (i) varying a data retention model based on an age of the sampled data; (ii) evicting cache data based on the representation; (iii) moving the retained data to a different storage tier; and (iv) determining an amount of time to store the retained data.

Claims (41)

1 . A method, comprising:

obtaining sampled data that was generated by sampling source data from at least one data source;

obtaining a plurality of predefined retention models, wherein each of the predefined retention models specifies a different amount of retained data that is to be stored in the one or more storage resources, for at least some of the source data from the at least one data source;

fitting at least one data model to at least some of the obtained sampled data to obtain a representation of the obtained sampled data from the at least one data source;

determining, by a processor-based retention model classifier, in response to the obtaining the sampled data, an assignment of at least some of the obtained sampled data from the at least one data source to one of the plurality of predefined retention models, wherein the assignment, by the processor-based retention model classifier, to the one predefined retention model is based at least in part on one or more characteristics of the at least one data source that provides the sampled data;

storing at least a portion, corresponding to the amount of the retained data specified by the one predefined retention model, of the representation of the obtained sampled data from the at least one data source in the one or more storage resources based at least in part on the assignment to the one predefined retention model of the plurality of predefined retention models, wherein the determining is performed prior to the storing of the at least the portion of the representation of the obtained sampled data in the one or more storage resources; and

automatically adapting, by a processor-based storage resource adaptation module, using one or more resource adaptation commands sent to one or more of the storage resources, a usage of the one or more storage resources that store the retained data from the at least one data source based at least in part on the representation and the assignment to the one predefined retention model of the plurality of predefined retention models, wherein the automatically adapting the usage of the one or more storage resources comprises automatically receiving, by the processor-based storage resource adaptation module, subsequent to the storing, from the processor-based retention model classifier, an automatic selection of a different one of the plurality of predefined retention models than the one predefined retention model, wherein the automatic selection is based at least in part on: (i) one or more data life cycle rules and (ii) an age of the obtained sampled data from the at least one data source, wherein the different predefined retention model specifies a different amount of retained data than the one predefined retention model, and wherein the processor-based storage resource adaptation module automatically removes, in response to the automatic selection, at least a portion of the amount of the retained data specified by the one predefined retention model from at least one of the one or more storage resources;

wherein the method is performed by at least one processing device comprising a processor coupled to a memory.

2 . The method of claim 1 , wherein the at least one data model comprises one or more of a parametric model, a non-parametric model, a descriptive statistics model, a time series model, decision trees and an ensemble of decision trees.

3 . The method of claim 2 , wherein the fitting of the parametric model to at least some of the obtained sampled data comprises representing the obtained sampled data using a probability distribution function and determining one or more parameters of the probability distribution function.

4 . The method of claim 2 , wherein the fitting of the non-parametric model to at least some of the obtained sampled data comprises representing the obtained sampled data using one or more of a Gaussian Mixture model, a Gaussian Mixture model that captures temporal variance, and a Kernel Density Estimation model.

5 . The method of claim 2 , wherein the fitting of the descriptive statistics model to at least some of the obtained sampled data comprises recording one or more of predefined summary statistics of the representation of the obtained sampled data and a time stamp of the first and last recorded value in the sampled data.

6 . The method of claim 2 , wherein the fitting of the time series model to at least some of the obtained sampled data comprises recording the source data and corresponding time stamps generated by the at least one data source.

7 . The method of claim 2 , further comprising identifying a distribution drift based at least in part on at least one density function of one or more of the parametric model and the non-parametric model; and performing the fitting when a distribution drift is identified.

8 . The method of claim 1 , further comprising grouping a plurality of data sources into a group and representing the group using one data model.

9 . The method of claim 1 , wherein the automatically adapting the usage of the one or more storage resources further comprises one or more of (i) evicting data from a cache based at least in part on the representation; (ii) moving the retained data from the at least one data source to a different storage tier; and (iii) determining an amount of time to store the retained data from the at least one data source.

10 . The method of claim 1 , wherein the plurality of predefined retention models comprises one or more of a lossy retention model that stores a type of a probability density function, one or more parameters of the probability density function and one or more summary statistics; a subsample retention model that stores a type of a probability density function, one or more parameters of the probability density function, a time interval, one or more summary statistics and a subsample of the source data from the at least one data source; and a complete retention model that stores a type of a probability density function, one or more parameters of the probability density function, the source data from the at least one data source, and a time interval.

11 . The method of claim 1 , wherein the one or more characteristics of the at least one data source comprise one or more of a priority and a criticality of the at least one data source.

12 . An apparatus comprising:

at least one processing device comprising a processor coupled to a memory;

the at least one processing device being configured to implement the following steps:

obtaining sampled data that was generated by sampling source data from at least one data source;

obtaining a plurality of predefined retention models, wherein each of the predefined retention models specifies a different amount of retained data that is to be stored in one or more storage resources, for at least some of the source data from the at least one data source;

fitting at least one data model to at least some of the obtained sampled data to obtain a representation of the obtained sampled data from the at least one data source;

determining, by a processor-based retention model classifier, in response to the obtaining the sampled data, an assignment of at least some of the obtained sampled data from the at least one data source to one of the plurality of predefined retention models, wherein the assignment, by the processor-based retention model classifier, to the one predefined retention model is based at least in part on one or more characteristics of the at least one data source that provides the sampled data;

storing at least a portion, corresponding to the amount of the retained data specified by the one predefined retention model, of the representation of the obtained sampled data from the at least one data source in the one or more storage resources based at least in part on the assignment to the one predefined retention model of the plurality of predefined retention models, wherein the determining is performed prior to the storing of the at least the portion of the representation of the obtained sampled data in the one or more storage resources; and

automatically adapting, by a processor-based storage resource adaptation module, using one or more resource adaptation commands sent to one or more of the storage resources, a usage of the one or more storage resources that store the retained data from the at least one data source based at least in part on the representation and the assignment to the one predefined retention model of the plurality of predefined retention models, wherein the automatically adapting the usage of the one or more storage resources comprises automatically receiving, by the processor-based storage resource adaptation module, subsequent to the storing, from the processor-based retention model classifier, an automatic selection of a different one of the plurality of predefined retention models than the one predefined retention model, wherein the automatic selection is based at least in part on: (i) one or more data life cycle rules and (ii) an age of the obtained sampled data from the at least one data source, wherein the different predefined retention model specifies a different amount of retained data than the one predefined retention model, and wherein the processor-based storage resource adaptation module automatically removes, in response to the automatic selection, at least a portion of the amount of the retained data specified by the one predefined retention model from at least one of the one or more storage resources.

13 . The apparatus of claim 12 , wherein the at least one data model comprises one or more of a parametric model, a non-parametric model, a descriptive statistics model, a time series model, decision trees and an ensemble of decision trees.

14 . The apparatus of claim 12 , further comprising grouping a plurality of data sources into a group and representing the group using one data model.

15 . The apparatus of claim 12 , wherein the automatically adapting the usage of the one or more storage resources further comprises one or more of (i) evicting data from a cache based at least in part on the representation; (ii) moving the retained data from the at least one data source to a different storage tier; and (iii) determining an amount of time to store the retained data from the at least one data source.

16 . The apparatus of claim 12 , wherein the one or more characteristics of the at least one data source comprise one or more of a priority and a criticality of the at least one data source.

17 . A non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes the at least one processing device to perform the following steps:

obtaining sampled data that was generated by sampling source data from at least one data source;

obtaining a plurality of predefined retention models, wherein each of the predefined retention models specifies a different amount of retained data that is to be stored in one or more storage resources, for at least some of the source data from the at least one data source;

fitting at least one data model to at least some of the obtained sampled data to obtain a representation of the obtained sampled data from the at least one data source;

determining, by a processor-based retention model classifier, in response to the obtaining the sampled data, an assignment of at least some of the obtained sampled data from the at least one data source to one of the plurality of predefined retention models, wherein the assignment, by the processor-based retention model classifier, to the one predefined retention model is based at least in part on one or more characteristics of the at least one data source that provides the sampled data;

storing at least a portion, corresponding to the amount of the retained data specified by the one predefined retention model, of the representation of the obtained sampled data from the at least one data source in the one or more storage resources based at least in part on the assignment to the one predefined retention model of the plurality of predefined retention models, wherein the determining is performed prior to the storing of the at least the portion of the representation of the obtained sampled data in the one or more storage resources; and

automatically adapting, by a processor-based storage resource adaptation module, using one or more resource adaptation commands sent to one or more of the storage resources, a usage of the one or more storage resources that store the retained data from the at least one data source based at least in part on the representation and the assignment to the one predefined retention model of the plurality of predefined retention models, wherein the automatically adapting the usage of the one or more storage resources comprises automatically receiving, by the processor-based storage resource adaptation module, subsequent to the storing, from the processor-based retention model classifier, an automatic selection of a different one of the plurality of predefined retention models than the one predefined retention model, wherein the automatic selection is based at least in part on: ( 1 ) one or more data life cycle rules and (ii) an age of the obtained sampled data from the at least one data source, wherein the different predefined retention model specifies a different amount of retained data than the one predefined retention model, and wherein the processor-based storage resource adaptation module automatically removes, in response to the automatic selection, at least a portion of the amount of the retained data specified by the one predefined retention model from at least one of the one or more storage resources.

18 . The non-transitory processor-readable storage medium of claim 17 , wherein the at least one data model comprises one or more of a parametric model, a non-parametric model, a descriptive statistics model, a time series model, decision trees and an ensemble of decision trees.

19 . The non-transitory processor-readable storage medium of claim 17 , wherein the automatically adapting the usage of the one or more storage resources further comprises one or more of (i) evicting data from a cache based at least in part on the representation; (ii) moving the retained data from the at least one data source to a different storage tier; and (iii) determining an amount of time to store the retained data from the at least one data source.

20 . The non-transitory processor-readable storage medium of claim 17 , wherein the one or more characteristics of the at least one data source comprise one or more of a priority and a criticality of the at least one data source.

Assignments (9)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (051302/0528) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.); SECUREWORKS CORP.
Reel/Frame 060438/0593 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053311/0169) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 060438/0742 →
RELEASE OF SECURITY INTEREST AT REEL 051449 FRAME 0728 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.; SECUREWORKS CORP.; EMC CORPORATION
Reel/Frame 058002/0010 →
SECURITY INTEREST Recorded Jun 5, 2020
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 053311/0169 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Dec 31, 2019
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.; SECUREWORKS CORP.; EMC CORPORATION
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 051449/0728 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Dec 16, 2019
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.; SECUREWORKS CORP.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 051302/0528 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2019
From: DAVIS, SASHKA T.; SUNKAVALLY, NAVEEN; RAMZAN, ZULFIKAR A.
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 050892/0079 →
Continuity (1)
Related Publication 20210133211A1 · May 6, 2021
References Cited (15)
US 20130066843A1 · Nair · 2013 [cited by examiner]
US 20130318305A1 · Bivens · 2013 [cited by examiner]
US 20150242133A1 · Smith · 2015 [cited by examiner]
US 20160085469A1 · Cherubini · 2016 [cited by examiner]
US 20160098455A1 · Curtin · 2016 [cited by examiner]
US 20170060769A1 · Wires · 2017 [cited by examiner]
US 20170308482A1 · Alatorre · 2017 [cited by examiner]
US 20200134198A1 · Rudrabhatla · 2020 [cited by examiner]
Tekumalla et al., Copula-HDP-HMM: Non-parametric Modeling of Temporal Multivariate Data for I/O Efficient Bulk Cache Preloading, Proceedings of the 2016 SIAM International Conference on Data Mining (SDM), 2016, pp. 774-… [cited by examiner]
Sithole et al., Cache performance models for quality of service compliance in storage clouds, Journal of Cloud Computing: Advances, Systems and Applications 2013, 2:1, published Jan. 10, 2013, 24 pages (Year: 2013). [cited by examiner]
Wajahat et al., Distribution Fitting and Performance Modeling for Storage Traces, 2019 IEEE 27th International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS), Oct. 21… [cited by examiner]
Bryan Willman, Data Retention Best Practices, Techfino, Sep. 26, 2016. 12 pages. [cited by applicant]
Jun Li et al. Managing Data Retention Policies at Scale, HP Laboratories, 2011, 9 pages. [cited by applicant]
Yinghao Yu, LRC: Dependency-Aware Cache Management for Data Analytics Clusters, arXiv: 1703.08280v1 [cs.DC] Mar. 24, 2017, 9 pages. [cited by applicant]
Ehcache, Cache Eviction Algorithms, https://www.ehcache.org/documentation/2.8/apis/cache-eviction-algorithms.html, last accessed on Oct. 29, 2019. [cited by applicant]