IP Library Granted Patent US 11,194,669
Granted Patent B2
US 11,194,669 · App. 16/428,963 · Granted Dec 7, 2021

Adaptable multi-layered storage for generating search indexes

Inventors: Jal Jalali Ekram (Mountain View, CA); David Anthony Terei (San Francisco, CA)
Assignee: RUBRIK, INC.
G06F11/1453G06F11/1464G06F16/128G06F40/10G06F9/45558G06F2009/45583
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,194,669
App. No.
16/428,963
Granted
Dec 7, 2021
Kind
B2
Abstract

Methods and systems for improving data back-up, recovery, and search across different cloud-based applications, services, and platforms are described. A data management and storage system may direct compute and storage resources within a customer's cloud-based data storage account to back-up and restore data while the customer retains full control of their data. The data management and storage system may direct the compute and storage resources within the customer's cloud-based data storage account to generate and store secondary layers that are used for generating search indexes, to generate and store shared space layers and user specific layers to facilitate the deduplication of email attachments and text blocks, to perform a controlled restoration of email snapshots such that sensitive information (e.g., restricted keywords) located within stored snapshots remains protected, and to detect and preserve emails that were received or transmitted and then deleted between two consecutive snapshots.

Claims (55)

1. A method for operating a data management system, the method comprising:

identifying a set of fields associated with a set of electronic messages, the set of electronic messages corresponding to a snapshot that includes a set of data, the set of fields including a subject line field, a sender field, and a message body field associated with a number of lines of text;

detecting that a secondary layer should be generated and stored based on an amount of available disk space within a storage space;

generating the secondary layer based on the set of fields and the set of data included in the snapshot of the set of electronic messages, the generating the secondary layer including extracting a portion of the set of data that is associated with the set of fields and including the portion of the set of data in the secondary layer;

detecting that the set of fields has been modified to comprise a second set of fields different from the set of fields;

regenerating and storing the secondary layer using the second set of fields and the snapshot of the set of electronic messages, a file size of the regenerated secondary layer being less than a file size of the snapshot of the set of electronic messages; and

generating a search index for the snapshot of the set of electronic messages using the regenerated secondary layer.

2. The method of claim 1 , further comprising:

detecting that the amount of available disk space is less than a threshold disk space; and

decreasing the number of lines of text for the message body field in response to detecting that the amount of available disk space is less than the threshold disk space prior to regenerating the secondary layer using the second set of fields.

3. The method of claim 1 , further comprising:

detecting that a keyword search was performed on the snapshot of the set of electronic messages; and

increasing the number of lines of text for the message body field in response to detecting that the keyword search was performed on the snapshot of the set of electronic messages prior to regenerating the secondary layer using the second set of fields.

4. The method of claim 1 , further comprising:

determining a data change rate for the snapshot relative to a prior snapshot of the set of electronic messages; and

detecting that the secondary layer should be generated and stored based on the amount of available disk space and the data change rate.

5. The method of claim 1 , further comprising:

detecting that an amount of available disk space for storing other secondary layers is less than a threshold disk space; and

deleting the secondary layer in response to detecting that the amount of available disk space for storing the other secondary layers is less than the threshold disk space.

6. The method of claim 1 , wherein:

the set of electronic messages comprises a plurality of electronic messages.

7. The method of claim 6 , wherein:

the plurality of electronic messages includes electronic messages from a plurality of email mailboxes.

8. The method of claim 1 , wherein:

the search index includes a word-level inverted index.

9. The method of claim 1 , wherein:

the second set of fields comprises a fewer number of fields than the set of fields.

10. The method of claim 1 , further comprising:

deleting the secondary layer in response to detecting that the search index has not been used for a search of the snapshot for at least a past threshold period of time.

11. A data management system, comprising:

a memory configured to store a snapshot of a set of electronic messages; and

one or more processors in communication with the memory configured to identify a set of fields associated with a set of electronic messages, the set of electronic messages corresponding to a snapshot that includes a set of data, the set of fields including a message body field associated with a number of lines of text, the one or more processors configured to detect that a secondary layer should be generated and stored based on an amount of available disk space within a storage space, the one or more processors configured to generate the secondary layer based on the set of fields and the set of data included in the snapshot of the set of electronic messages via extracting a portion of the set of data that is associated with the set of fields and including the portion of the set of data in the secondary layer, the one or more processors configured to store the secondary layer using the memory, the one or more processors configured to detect that the set of fields has been modified to comprise a second set of fields different from the set of fields, the one or more processors configured to regenerate the secondary layer using the second set of fields and the snapshot of the set of electronic messages, a file size of the regenerated secondary layer being less than a file size of the snapshot of the set of electronic messages, the one or more processors configured to store the regenerated secondary layer using the memory, the one or more processors configured to generate a search index for the snapshot of the set of electronic messages using the regenerated secondary layer and store the search index using the memory.

12. The data management system of claim 11 , wherein:

the one or more processors configured to detect that the amount of available disk space is less than a threshold disk space and decrease the number of lines of text for the message body field in response to detection that the amount of available disk space is less than the threshold disk space prior to regeneration of the secondary layer using the second set of fields.

13. The data management system of claim 11 , wherein:

the one or more processors configured to detect that a keyword search was performed on the snapshot of the set of electronic messages and increase the number of lines of text for the message body field in response to detection that the keyword search was performed on the snapshot of the set of electronic messages prior to regeneration of the secondary layer using the second set of fields.

14. The data management system of claim 11 , wherein:

the one or more processors configured to determine a data change rate for the snapshot relative to a prior snapshot of the set of electronic messages and detect that the secondary layer should be generated and stored based on the amount of available disk space and the data change rate.

15. The data management system of claim 11 , wherein:

the one or more processors configured to detect that an amount of available disk space for storing other secondary layers is less than a threshold disk space and delete the secondary layer in response to detection that the amount of available disk space for storing the other secondary layers is less than the threshold disk space.

16. The data management system of claim 11 , wherein:

the set of electronic messages comprises a plurality of electronic messages.

17. The data management system of claim 16 , wherein:

the plurality of electronic messages includes electronic messages from a plurality of email mailboxes.

18. The data management system of claim 11 , wherein:

the search index includes an inverted index.

19. The data management system of claim 11 , wherein:

the one or more processors configured to delete the secondary layer in response to detection that the search index has not been used for a search of the snapshot for at least a past threshold period of time.

20. One or more storage devices containing processor readable code for programming one or more processors to perform a method for operating a data management system, the processor readable code comprising:

processor readable code configured to identify a set of fields associated with a set of electronic messages, the set of electronic messages corresponding to a snapshot that includes a set of data, the set of fields including a message body field associated with a number of lines of text;

processor readable code configured to detect that a secondary layer should be generated based on an amount of available disk space within a storage space;

processor readable code configured to generate the secondary layer using the set of fields and the snapshot of the set of electronic messages via extraction of portions of the snapshot corresponding with the set of fields;

processor readable code configured to detect that the set of fields has been modified to comprise a second set of fields different from the set of fields, the second set of fields comprising a fewer number of fields than the set of fields;

processor readable code configured to regenerate and store the secondary layer using the second set of fields and the snapshot of the set of electronic messages, a file size of the regenerated secondary layer being less than a file size of the snapshot of the set of electronic messages; and

processor readable code configured to generate a search index for the snapshot of the set of electronic messages using the regenerated secondary layer, the search index comprising an inverted index.

Assignments (3)
RELEASE OF SECURITY INTEREST IN PATENT COLLATERAL AT REEL/FRAME NO. 60333/0323 Recorded Jun 13, 2025
From: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
To: RUBRIK, INC.
Reel/Frame 071565/0602 →
GRANT OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 10, 2022
From: RUBRIK, INC.
To: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
Reel/Frame 060333/0323 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2019
From: EKRAM, JAL JALALI; TEREI, DAVID ANTHONY
To: RUBRIK, INC.
Reel/Frame 049345/0274 →
Cited By (2)
US 12,271,269 US 12,298,941