IP Library Granted Patent US 10,606,705
Granted Patent B1
US 10,606,705 · App. 14/953,802 · Granted Mar 31, 2020

Prioritizing backup operations using heuristic techniques

Inventors: Viswesvaran Janakiraman (San Jose, CA); Ashwin Kumar Kayyoor (Sunnyvale, CA)
Assignee: Veritas Technologies LLC
G06F11/1451G06F11/1464G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,606,705
App. No.
14/953,802
Granted
Mar 31, 2020
Kind
B1
Abstract

Various systems, methods, and processes to analyze datasets using heuristic-based data analysis and prioritization techniques to identify, derive, and/or select subsets with important and/or high-priority data for preferential backup are disclosed. A request to perform a backup operation that identifies a dataset to be backed up to a storage device is received. A subset of data is identified and selected from the dataset by analyzing the dataset using one or more prioritization techniques. A backup operation is performed by storing the subset of data in the storage device.

Claims (77)

1. A computer-implemented method comprising:

receiving, at a computing system, a request to perform a backup operation, wherein the request identifies a dataset to be backed up to a storage device;

selecting a first subset of data and a second subset of data from the dataset, wherein the first subset and the second subset each comprise a plurality of data units, and

the selecting comprises

analyzing the dataset by applying, to each of the first subset and the second subset, a social network data analysis technique, a topic modeling technique, and a cluster analysis technique, wherein

the social network data analysis technique provides a first importance metric,

the topic modeling technique provide a second importance metric, wherein

 the topic modeling technique is based, at least in part, on a natural language processing (NLP) methodology, and

the cluster analysis technique provide a third importance metric, and

the social network data analysis technique is based on social network data associated with one or more social network data sources that are distinct from the dataset, and the topic modeling technique and the cluster analysis technique are based on the dataset,

determining an average priority level for each of the first subset and the second subset by averaging the first importance metric, the second importance metric, and the third importance metric, as each of those importance metrics are applied to each data unit in each respective subset,

determining whether each of the first subset of data and the second subset of data meets an importance threshold, and

determining that the first subset of data has a higher priority than the second subset of data based, at least in part, on a comparison of the average priority level for each of the first subset and the second subset; and

performing the backup operation, wherein

the backup operation comprises storing the first subset of data in the storage device prior to storing the second subset of data.

2. The computer-implemented method of claim 1 , comprising:

the social network data analysis technique, the topic modeling technique, and the cluster analysis technique, are applied in order of the social network data analysis technique, the topic modeling technique, and the cluster analysis technique.

3. The computer-implemented method of claim 2 , further comprising:

in response to determining that the first subset of data does not meet the importance threshold, re-analyzing the dataset using any combination of the social network data analysis technique, the cluster analysis technique, and the topic modeling technique.

4. The computer-implemented method of claim 3 , wherein the importance threshold is based on a relative measure or an absolute measure.

5. The computer-implemented method of claim 1 , wherein the selecting the first subset of data from the dataset further comprises:

receiving information indicative of one or more results of analyzing the dataset from another computing device;

determining that the information indicative of the one or more results is associated with at least a portion of the dataset; and

deriving at least the first subset of data based on the information.

6. The computer-implemented method of claim 1 , further comprising:

storing the second subset of data in the storage device for backup after storing the first subset of data, wherein

the second subset of data comprises one or more portions of the dataset that are not responsive to the analyzing.

7. The computer-implemented method of claim 1 , further comprising:

determining whether the request to perform the backup operation specifies a type of backup operation; and

in response to the type of backup operation being specified as an incremental backup operation or a synthetic full backup operation, and based on the analyzing, deriving the first subset of data using only metadata and changed data associated with the first subset of data.

8. A non-transitory computer readable storage medium storing program instructions executable to perform a method comprising:

receiving, at a computing system, a request to perform a backup operation, wherein the request identifies a dataset to be backed up to a storage device;

selecting a first subset of data and a second subset of data from the dataset, wherein the first subset and the second subset each comprise a plurality of data units, and

the selecting comprises

analyzing the dataset by applying, to each of the first subset and the second subset, a social network data analysis technique, a topic modeling technique, and a cluster analysis technique, wherein

the social network data analysis technique provides a first importance metric,

the topic modeling technique provide a second importance metric, wherein

the topic modeling technique is based, at least in part, on a natural language processing (NLP) methodology, and

the cluster analysis technique provide a third importance metric, and

the social network data analysis technique is based on social network data associated with one or more social network data sources that are distinct from the dataset, and the topic modeling technique and the cluster analysis technique are based on the dataset,

determining an average priority level for each of the first subset and the second subset by averaging the first importance metric, the second importance metric, and the third importance metric, as each of those importance metrics are applied to each data unit in each respective subset,

determining whether each of the first subset of data and the second subset of data meets an importance threshold, and

determining that the first subset of data has a higher priority than the second subset of data based, at least in part, on a comparison of the average priority level for each of the first subset and the second subset; and

performing the backup operation, wherein

the backup operation comprises storing the first subset of data in the storage device prior to storing the second subset of data.

9. The non-transitory computer readable storage medium of claim 8 , wherein the method further comprises

the social network data analysis technique, the topic modeling technique, and the cluster analysis technique, are applied in order of the social network data analysis technique, the topic modeling technique, and the cluster analysis technique;

the importance threshold is based on a relative measure or an absolute measure; and

in response to determining that the first subset of data does not meet the importance threshold, re-analyzing the dataset using any combination of the social network data analysis technique, the cluster analysis technique, and the topic modeling technique.

10. The non-transitory computer readable storage medium of claim 8 , wherein

the selecting the first subset of data from the dataset further comprises

receiving information indicative of one or more results of analyzing the dataset from another computing device;

determining that the information indicative of the one or more results is associated with at least a portion of the dataset; and

deriving at least the first subset of data based on the information.

11. A system comprising:

one or more processors; and

a memory coupled to the one or more processors, wherein the memory stores program instructions executable by the one or more processors to perform a method comprising:

receiving, at a computing system, a request to perform a backup operation, wherein the request identifies a dataset to be backed up to a storage device;

selecting a first subset of data and a second subset of data from the dataset, wherein

the first subset and the second subset each comprise a plurality of data units, and

the selecting comprises

analyzing the dataset by applying, to each of the first subset and the second subset, and a social network data analysis technique, a topic modeling technique, and a cluster analysis technique, wherein

the social network data analysis technique provides a first importance metric,

the topic modeling technique provide a second importance metric, wherein

 the topic modeling technique is based, at least in part, on a natural language processing (NLP) methodology, and

the cluster analysis technique provide a third importance metric, and

the social network data analysis technique is based on social network data associated with one or more social network data sources that are distinct from the dataset, and the topic modeling technique and the cluster analysis technique are based on the dataset,

determining an average priority level for each of the first subset and the second subset by averaging the first importance metric, the second importance metric, and the third importance metric, as each of those importance metrics are applied to each data unit in each respective subset,

determining whether each of the first subset of data and the second subset of data meets an importance threshold, and

determining that the first subset of data has a higher priority than the second subset of data based, at least in part, on a comparison of the average priority level for each of the first subset and the second subset; and

perform the backup operation, wherein

the backup operation comprises storing the first subset of data in the storage device prior to storing the second subset of data.

12. The system of claim 11 , wherein

the selecting the first subset of data from the dataset further comprises

receiving information indicative of one or more results of analyzing the dataset from another computing device;

determining that the information indicative of the one or more results is associated with at least a portion of the dataset; and

deriving at least the first subset of data based on the information.

Assignments (13)
AMENDMENT NO. 1 TO PATENT SECURITY AGREEMENT Recorded Apr 8, 2025
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 070779/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2025
From: VERITAS TECHNOLOGIES LLC
To: COHESITY, INC.
Reel/Frame 070335/0013 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2024
From: ACQUIOM AGENCY SERVICES LLC, AS COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC (F/K/A VERITAS US IP HOLDINGS LLC)
Reel/Frame 069712/0090 →
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 069634/0584 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
ASSIGNMENT OF SECURITY INTEREST IN PATENT COLLATERAL Recorded Nov 25, 2024
From: BANK OF AMERICA, N.A., AS ASSIGNOR
To: ACQUIOM AGENCY SERVICES LLC, AS ASSIGNEE
Reel/Frame 069440/0084 →
TERMINATION AND RELEASE OF SECURITY IN PATENTS AT R/F 037891/0726 Recorded Nov 30, 2020
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: VERITAS US IP HOLDINGS, LLC
Reel/Frame 054535/0814 →
SECURITY INTEREST Recorded Aug 20, 2020
From: VERITAS TECHNOLOGIES LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 054370/0134 →
MERGER Recorded Apr 18, 2016
From: VERITAS US IP HOLDINGS LLC
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 038483/0203 →
SECURITY INTEREST Recorded Feb 23, 2016
From: VERITAS US IP HOLDINGS LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 037891/0726 →
SECURITY INTEREST Recorded Feb 23, 2016
From: VERITAS US IP HOLDINGS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 037891/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2016
From: SYMANTEC CORPORATION
To: VERITAS US IP HOLDINGS LLC
Reel/Frame 037693/0158 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2015
From: JANAKIRAMAN, VISWESVARAN; KAYYOOR, ASHWIN KUMAR
To: SYMANTEC CORPORATION
Reel/Frame 037324/0954 →
Cited By (1)
US 12,353,295