IP Library › Granted Patent US 12,596,481
Granted Patent B2
US 12,596,481 · App. 18/634,288 · Granted Apr 7, 2026

Prefetching data using predictive analysis

Inventor: Sourav Mukherjee (Bangalore, IN)
Assignee: Cohesity, Inc.
G06F3/0611G06F3/0635G06F3/067
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,481
App. No.
18/634,288
Filed
Apr 12, 2024
Granted
Apr 7, 2026
Kind
B2
Art Unit
2132
USPC
711/154
Abstract

Techniques are disclosed for prefetching data using predictive analysis. An example method comprises storing, by a data platform implemented by a computing system, objects of a file system, wherein a first subset of the objects is stored to a first storage tier and a second subset of the objects is stored to a second storage tier, classifying objects into one or more classifications, storing a data access record for the objects, applying a machine learning model to generate a prediction of future data access to one or more objects of the second subset based on the one or more classifications and the data access record, wherein the prediction includes a predicted time for the future data access, and retrieving, based on the prediction, the one or more objects of the second subset from the second storage tier prior to the predicted time.

Claims (58)

1 . A method comprising:

storing, by a data platform implemented by a computing system, a plurality of objects of a file system, wherein a first subset of the plurality of objects is stored to a first storage tier and a second subset of the plurality of objects is stored to a second storage tier;

for each object of the plurality of objects, classifying, by the data platform, the object with one or more classifications of a plurality classifications that each identifies a type of data stored as the object based at least on corresponding metadata identifying a source that created the object;

storing, by the data platform, a data access record indicating previous instances of data access requests for each of the plurality of objects and the one or more classifications of each of the plurality of objects;

applying, by the data platform, a machine learning model to generate a prediction of future data access to one or more objects of the second subset based at least on the one or more classifications of the data access record, wherein the prediction includes a predicted time for the future data access;

retrieving, by the data platform and based on the prediction, the one or more objects of the second subset from the second storage tier prior to the predicted time; and

storing, by the data platform and to the first storage tier, the one or more objects of the second subset retrieved based on the prediction.

2 . The method of claim 1 , further comprising training, by the data platform, the machine learning model using the one or more classifications for each of the plurality of objects and the data access record for each of the plurality of objects.

3 . The method of claim 1 , further comprising receiving, by the data platform, a data access request identifying at least one of the plurality of objects, wherein storing the data access record comprises storing, by the data platform, the data access request in the data access record.

4 . The method of claim 3 ,

wherein classifying the object with the one or more classifications comprises classifying, by the data platform, the at least one of the plurality of objects identified in the data access request with the one or more classifications, and

wherein storing the data access record comprises storing, by the data platform, the data access request in the data access record along with the one or more classifications of the at least one of the plurality of objects identified in the data access request.

5 . The method of claim 1 , further comprising:

determining, by the data platform and based on one or more performance characteristics of the second storage tier, an amount of transfer time for retrieving the one or more objects of the second subset from the second storage tier;

wherein, retrieving, based on the prediction, the one or more objects of the second subset from the second storage tier prior to the predicted time comprises retrieving, by the data platform, the one or more objects of the second subset from the second storage tier at a start time occurring at least the amount of transfer time before the predicted time.

6 . The method of claim 1 ,

wherein the machine learning model is a first machine learning model, and

wherein classifying the object with the one or more classifications comprises applying, by the data platform, a different second machine learning model to classify the object with the one or more classifications.

7 . The method of claim 1 , wherein the first storage tier has a higher data transfer rate than the second storage tier.

8 . The method of claim 1 ,

wherein the plurality of objects each comprise a plurality of chunks, and

wherein retrieving, based on the prediction, the one or more objects of the second subset comprises traversing a tree data structure associated with the one or more objects of the second subset to retrieve one or more chunks of the plurality of chunks corresponding to the one or more objects of the second subset.

9 . The method of claim 1 , further comprising, after storing the one or more objects of the second subset at the first storage tier:

storing, by the data platform, the one or more objects of the second subset back to the second storage tier; and

removing, by the data platform, the one or more objects of the second subset from the first storage tier.

10 . The method of claim 9 , further comprising determining, by the data platform, that the one or more objects of the second subset have been accessed, wherein storing the one or more objects of the second subset back on the second storage tier and removing the one or more objects of the second subset from the first storage tier is responsive to determining that the one or more objects of the second subset have been accessed.

11 . The method of claim 1 , wherein the source that created the object comprises one or more of a user, a department, or a group that created the object.

12 . A computing system comprising:

a memory storing instructions; and

processing circuitry, the instructions configured to cause the processing circuitry to:

store a plurality of objects of a file system, wherein a first subset of the plurality of objects is stored to a first storage tier and a second subset of the plurality of objects is stored to a second storage tier;

for each object of the plurality of objects, classify the object with one or more classifications of a plurality classifications that each identifies a type of data stored as the object based at least on corresponding metadata identifying a source that created the object;

store a data access record indicating previous instances of data access requests for each of the plurality of objects and the one or more classifications of each of the plurality of objects;

apply a machine learning model to generate a prediction of future data access to one or more objects of the second subset based at least on the one or more classifications of the data access record, wherein the prediction includes a predicted time for the future data access;

retrieve, based on the prediction, the one or more objects of the second subset from the second storage tier prior to the predicted time; and

store, to the first storage tier, the one or more objects of the second subset retrieved based on the prediction.

13 . The computing system of claim 12 , wherein the instructions are further configured to cause the processing circuitry to train the machine learning model using the one or more classifications for each of the plurality of objects and the data access record for each of the plurality of objects.

14 . The computing system of claim 12 , wherein the instructions are further configured to cause the processing circuitry to receive a data access request identifying at least one of the plurality of objects, wherein to store the data access record the processing circuitry further executes the instructions to store the data access request in the data access record.

15 . The computing system of claim 14 ,

wherein, to classify the object with the one or more classifications, the instructions are configured to cause the processing circuitry to classify the at least one of the plurality of objects identified in the data access request with the one or more classifications, and

wherein, to store the data access record, the instructions are configured to cause the processing circuitry to store the data access request in the data access record along with the classification one or more classifications of the at least one of the plurality of objects identified in the data access request.

16 . The computing system of claim 12 , wherein the instructions are further configured to cause the processing circuitry to:

determine, based on one or more performance characteristics of the second storage tier, an amount of transfer time for retrieving the one or more objects of the second subset from the second storage tier;

wherein, to retrieve, based on the prediction, the one or more objects of the second subset from the second storage tier prior to the predicted time, the instructions are configured to cause the processing circuitry to retrieve the one or more objects of the second subset from the second storage tier at a start time occurring at least the amount of transfer time before the predicted time.

17 . The computing system of claim 12 ,

wherein the machine learning model is a first machine learning model, and

wherein, to classify the object with the one or more classifications, the instructions are configured to cause the processing circuitry further executes the instructions to apply a different second machine learning model to classify the object with the one or more classifications.

18 . The computing system of claim 12 , wherein the instructions are further configured to cause the processing circuitry to, after storing the one or more objects of the second subset at the first storage tier:

store the one or more objects of the second subset back on the second storage tier; and

remove the one or more objects of the second subset from the first storage tier.

19 . The computing system of claim 18 , wherein the instructions are further configured to cause the processing circuitry to determine that the one or more objects of the second subset have been accessed, wherein storing the one or more objects of the second subset back on the second storage tier and removing the one or more objects of the second subset from the first storage tier is responsive to determining that the one or more objects of the second subset have been accessed.

20 . A computer-readable storage medium comprising instructions that, when executed, cause processing circuitry of a computing system to:

store a plurality of objects of a file system, wherein a first subset of the plurality of objects is stored to a first storage tier and a second subset of the plurality of objects is stored to a second storage tier;

for each object of the plurality of objects, classify the object with one or more classifications of a plurality classifications that each identifies a type of data stored as the object based at least on corresponding metadata identifying a source that created the object;

store a data access record indicating previous instances of data access requests for each of the plurality of objects and the one or more classifications of each of the plurality of objects;

apply a machine learning model to generate a prediction of future data access to one or more objects of the second subset based at least on the one or more classifications of the data access record, wherein the prediction includes a predicted time for the future data access;

retrieve, based on the prediction, the one or more objects of the second subset from the second storage tier prior to the predicted time; and

store, to the first storage tier, the one or more objects of the second subset retrieved based on the prediction.

Assignments (2)
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2024
From: MUKHERJEE, SOURAV
To: COHESITY, INC.
Reel/Frame 067091/0823 →
Continuity (1)
Related Publication 20250321674A1 · Oct 16, 2025
References Cited (8)
US 11061586B1 · Ahuja · 2021 [cited by examiner]
US 20120072672A1 · Anderson · 2012 [cited by examiner]
US 20160381176A1 · Cherubini et al. · 2016 [cited by applicant]
US 20220092022A1 · Agarwal · 2022 [cited by examiner]
US 20240256414A1 · Dar · 2024 [cited by examiner]
U.S. Appl. No. 17/960,515, filed Oct. 5, 2022, naming inventors Lee et al. [cited by applicant]
Extended Search Report from counterpart European Application No. 24187814.9 dated Jan. 8, 2025, 10 pp. [cited by applicant]
Herodotou et al., “Automating Distributed Tiered Storage Management in Cluster Computing”, arXiv, arXiv:1907.02394v1, Jul. 4, 2019, 16 pp. [cited by applicant]