IP Library Granted Patent US 9,171,008
Granted Patent B2
US 9,171,008 · App. 13/850,903 · Granted Oct 27, 2015

Performing data storage operations with a cloud environment, including containerized deduplication, data pruning, and data transfer

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,171,008
App. No.
13/850,903
Granted
Oct 27, 2015
Kind
B2
Abstract

Various systems and methods may be used for performing data storage operations, including content-indexing, containerized deduplication, and policy-driven storage, within a cloud environment. The systems support a variety of clients and cloud storage sites that may connect to the system in a cloud environment that requires data transfer over wide area networks, such as the Internet, which may have appreciable latency and/or packet loss, using various network protocols, including HTTP and FTP. Methods for content indexing data stored within a cloud environment may facilitate later searching, including collaborative searching. Methods for performing containerized deduplication may reduce the strain on a system namespace, effectuate cost savings, etc. Methods may identify suitable storage locations, including suitable cloud storage sites, for data files subject to a storage policy. Further, the systems and methods may be used for providing a cloud gateway and a scalable data object store within a cloud environment.

Claims (46)

1. A computer-implemented method for indexing and searching multiple content items, the method comprising:

selecting or accessing, with a secondary copy component of a computing system, at least one secondary copy of the multiple content items,

wherein the secondary copy of the multiple content items is a copy of the multiple content items and is not a primary copy of the multiple content items,

wherein the primary copy is available by the computer system over a local area network, and

wherein the at least one secondary copy is stored at a cloud storage site located geographically remote from the computer system;

for at least some of the multiple content items included in the secondary copy, with a content indexing component of the computing system;

analyzing content of a content item, including analyzing a summary of the content item;

based upon the analysis, generating metadata corresponding to the content item, wherein the metadata includes at least a logical address to the cloud storage site for accessing the content item; and

storing, in a content index, the generated metadata of the content, wherein the content index is not stored at the cloud storage site, but is locally accessible by the computer system; and

identifying, with an index searching component of the computing system, one or more indexed content items based on a search query and the metadata stored within the content index.

2. The method of claim 1 wherein the content index further comprises a file name for the content item, a logical descriptor for a client computer that originated the content item, and a size of the content item.

3. The method of claim 1 wherein the content index further comprises a token to uniquely identify each of the multiple content items for the cloud storage site, and wherein the content index indexes content items accessible via the cloud storage site over an HTTP protocol.

4. The method of claim 1 , further comprising:

identifying and selecting, with the secondary copy component, from an index of secondary copies one secondary copy that is a storage medium of the computing system having higher availability than a secondary copy stored on magnetic tape storage medium.

5. The method of claim 1 , further comprising:

identifying and selecting, with the secondary copy component, from an index of secondary copies an unencrypted secondary copy versus an encrypted secondary copy.

6. A computer system for indexing and searching multiple content items, the computer system comprising:

a processor configured to communicate with components of the computer system;

a memory configured to communicate with the processor;

a secondary copy component configured to select or access at least one secondary copy of the multiple content items,

wherein the secondary copy of the multiple content items is a copy of the multiple content items and is not a primary copy of the multiple content items,

wherein the primary copy is available by the computer system over a local area network, and

wherein the at least one secondary copy is stored at a cloud storage site located geographically remote from the computer system;

a content indexing component configured to, for at least some of the multiple content items included in the secondary copy;

analyze content of a content item, including analyzing a summary of the content item; and

based upon the analysis, generate metadata corresponding to the content item, wherein the metadata includes at least a logical address to the cloud storage site for accessing the content item; and

store in a content index the generated metadata of the content, wherein the content index is not stored at the cloud storage site, but is locally accessible by the computer system; and

an index searching component configured to identify one or more indexed content items based on a search query and the metadata stored within the content index.

7. The computer system of claim 6 wherein the content index further comprises a file name for the content item, a logical descriptor for a client computer that originated the content item, and a size of the content item.

8. The computer system of claim 6 wherein the content index further comprises a token to uniquely identify each of the multiple content items for the cloud storage site, and wherein the content index indexes content items accessible via the cloud storage site over an HTTP protocol.

9. The computer system of claim 6 wherein the secondary copy component is further configured to identify and select from an index of secondary copies one secondary copy that is on storage medium having higher availability than a secondary copy stored on magnetic tape storage medium.

10. The computer system of claim 6 wherein the secondary copy component is further configured to identify and select from an index of secondary copies an unencrypted secondary copy versus an encrypted secondary copy.

11. A computer-readable medium, excluding a transitory propagating signal, that stores instructions that when executed by a computer system cause the computer system to perform operation for indexing and searching multiple content items, the operations comprising:

selecting or accessing at least one secondary copy of the multiple content items,

wherein the secondary copy of the multiple content items is a copy of the multiple content items and is not a primary copy of the multiple content items,

wherein the primary copy is available by the computer system over a local area network, and

wherein the at least one secondary copy is stored at a cloud storage site located geographically remote from the computer system;

for at least some of the multiple content items included in the secondary copy;

analyzing content of a content item, including analyzing a summary of the content item;

based upon the analysis, generate metadata corresponding to the content item, wherein the metadata includes at least a logical address to the cloud storage site for accessing the content item; and

storing, in a content index the generated metadata of the content, wherein the content index is not stored at the cloud storage site, but is locally accessible by the computer system; and

identifying one or more indexed content items based on a search query and the metadata stored within the content index.

12. The computer-readable medium of claim 11 , wherein the content index further comprises a file name for the content item, a logical descriptor for a client computer that originated the content item, and a size of the content item.

13. The computer-readable medium of claim 11 wherein the content index further comprises a token to uniquely identify each of the multiple content items for the cloud storage site, and wherein the content index indexes content items accessible via the cloud storage site over an HTTP protocol.

14. The computer-readable medium of claim 11 , further comprising: identifying and selecting from an index of secondary copies one secondary copy that is a storage medium of the computing system having higher availability than a secondary copy stored on magnetic tape storage medium.

15. The computer-readable medium of claim 11 , further comprising: identifying and selecting from an index of secondary copies an unencrypted secondary copy versus an encrypted secondary copy.

Assignments (5)
SUPPLEMENTAL CONFIRMATORY GRANT OF SECURITY INTEREST IN UNITED STATES PATENTS Recorded Apr 16, 2025
From: COMMVAULT SYSTEMS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 070864/0344 →
SECURITY INTEREST Recorded Dec 13, 2021
From: COMMVAULT SYSTEMS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 058496/0836 →
RELEASE OF SECURITY INTEREST Recorded Jan 6, 2021
From: BANK OF AMERICA, N.A.
To: COMMVAULT SYSTEMS, INC.
Reel/Frame 054913/0905 →
SECURITY INTEREST Recorded Jul 2, 2014
From: COMMVAULT SYSTEMS, INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 033266/0678 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2013
From: PRAHLAD, ANAND; MULLER, MARCUS S.; KOTTOMTHARAYIL, RAJIV; KAVURI, SRINIVAS; GOKHALE, PARAG; VIJAYAN, MANOJ
To: COMMVAULT SYSTEMS, INC.
Reel/Frame 030800/0818 →