IP Library Granted Patent US 10,866,874
Granted Patent B1
US 10,866,874 · App. 16/455,026 · Granted Dec 15, 2020

Apparatus and method for sampling large data sets in a distributed data storage system

Inventors: Shaun Ahmadian (San Jose, CA); Sushil Thomas (San Francisco, CA)
Assignee: Cloudera, Inc.
G06F11/3079G06F3/0604G06F3/067G06F3/0646
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,866,874
App. No.
16/455,026
Granted
Dec 15, 2020
Kind
B1
Abstract

A system includes a distributed data storage system disseminated across worker machines connected by a network. A distributed data storage management module has instructions executed by a processor to utilize data block identifiers to track data block accesses to the distributed data storage system. A sampling module with instructions executed by the processor receives a new sample request from a client machine connected to the network. Initial data block samples are gathered from the distributed data storage system during a first time period. A revised sample request is received from the client machine during the first time period. The initial data block samples are gathered. New data block samples are collected from the distributed data storage system. The initial data block samples and the new data block samples are combined to form cumulative data block sample results. The cumulative data block sample results are supplied to the client machine.

Claims (45)

1. A system, comprising;

a distributed database storage (DDS) system disseminated across worker machines connected by a network;

a distributed data storage management module with instructions executed by a processor to assign data block identifiers to different data blocks of the DDS system, and utilize the data block identifiers to track data block accesses to the distributed data storage system; and

a sampling module with instructions executed by the processor to:

receive a new sample request from a client machine connected to the network,

gather initial data block samples from the distributed data storage system during a first time period,

receive a revised sample request from the client machine during the first time period,

collect new data block samples from the distributed data storage system, wherein the data block identifiers are used to select the new data block samples,

combine the initial data block samples and the new data block samples to form cumulative data block sample results, and

supply the cumulative data block sample results to the client machine.

2. The system of claim 1 wherein the new data block samples augment the initial block samples.

3. The system of claim 2 wherein data block identifiers associated with the initial data block samples are used to select new data block identifiers for the new data block samples to prevent sampling again of data samples from the initial data block samples for the new data block samples.

4. The system of claim 3 wherein the data block identifiers comprise configurable hash values.

5. The system of claim 4 wherein a data block size for each of the different data blocks is configurable.

6. The system of claim 1 wherein the revised sample request comprises a request for incrementally more samples than in the new sample request.

7. The system of claim 6 wherein the new sample request comprises a request to sample a first percentage of the data block accesses, and the revised sample request comprises a request to sample a second percentage of the data block accesses that is higher than the first percentage.

8. The system of claim 1 wherein the cumulative data block sample results are used to provide visualization of the data block accesses through a user interface on the client machine, and that can be displayed for different sample requests including the new sample request and the revised sample request.

9. The system of claim 1 wherein the DDS system is deployed in a cloud network.

10. The system of claim 9 wherein the DDS system comprises one of a local file system, a Hadoop Distributed File System, or a large cloud data store.

11. A method comprising:

assigning data block identifiers to different data blocks of a distributed database storage (DDS) system to track data block accesses to the distributed data storage system;

receiving a new sample request from a client machine connected to the network;

gathering initial data block samples from the distributed data storage system during a first time period;

receiving a revised sample request from the client machine during the first time period;

collecting new data block samples from the distributed data storage system, wherein the data block identifiers are used to select the new data block samples;

combining the initial data block samples and the new data block samples to form cumulative data block sample results, and

supplying the cumulative data block sample results to the client machine.

12. The method of claim 11 wherein the new data block samples augment the initial block samples.

13. The method of claim 12 wherein data block identifiers associated with the initial data block samples are used to select new data block identifiers for the new data block samples to prevent sampling again of data samples from the initial data block samples for the new data block samples.

14. The method of claim 13 wherein the data block identifiers comprise configurable hash values, and further wherein a data block size for each of the different data blocks is configurable.

15. The method of claim 11 wherein the revised sample request comprises a request for incrementally more samples than in the new sample request.

16. The method of claim 11 further comprising:

providing, using the cumulative data block sample results, a visualization of the data block accesses through a user interface on the client machine; and

displaying the visualization for different sample requests including the new sample request and the revised sample request.

17. The method of claim 11 wherein the DDS system is disseminated across worker machines connected by a network, and is deployed in a cloud network, and further wherein the DDS system comprises one of a local file system, a Hadoop Distributed File System, or a large cloud data store.

18. A method comprising:

assigning data block identifiers to different data blocks of a distributed database storage (DDS) system to track data block accesses to the distributed data storage system;

receiving a first sample request from a client machine connected to the network;

gathering, in response to the first sample request, initial data block samples from the DDS system during a first time period;

receiving a new sample request from the client machine during the first time period;

collecting, in response to the new sample request, new data block samples from the DDS system, wherein the new data block samples augment the initial data block samples, and further wherein the data block identifiers are used to select the new data block samples to prevent sampling again of any initial data block samples for the new data block samples;

combining the initial data block samples and the new data block samples to form cumulative data block sample results, and

supplying the cumulative data block sample results to the client machine.

19. The method of claim 18 wherein the supplying step supplies the cumulative data block sample results for display of a visualization of the data block accesses for the first sample request and the new sample request.

20. The method of claim 19 wherein the DDS system is disseminated across worker machines connected by a network, and is deployed in a cloud network, and further wherein the DDS system comprises one of a local file system, a Hadoop Distributed File System, or a large cloud data store.

Assignments (6)
RELEASE OF SECURITY INTERESTS IN PATENTS Recorded Oct 14, 2021
From: CITIBANK, N.A.
To: CLOUDERA, INC.; HORTONWORKS, INC.
Reel/Frame 057804/0355 →
FIRST LIEN NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Oct 12, 2021
From: CLOUDERA, INC.; HORTONWORKS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 057776/0185 →
SECOND LIEN NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Oct 12, 2021
From: CLOUDERA, INC.; HORTONWORKS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 057776/0284 →
SECURITY INTEREST Recorded Dec 22, 2020
From: CLOUDERA, INC.; HORTONWORKS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 054832/0559 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2020
From: AHMADIAN, SHAUN; THOMAS, SUSHIL
To: CLOUDERA, INC.
Reel/Frame 053573/0182 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2020
From: ARCADIA DATA, INC.
To: CLOUDERA, INC.
Reel/Frame 052573/0284 →