IP Library Granted Patent US 12,008,404
Granted Patent B2
US 12,008,404 · App. 17/721,175 · Granted Jun 11, 2024

Executing a big data analytics pipeline using shared storage resources

Inventors: Ivan Jibaja (San Jose, CA); Prashant Jaikumar (Sunnyvale, CA); Stefan Dorsett (San Jose, CA); Curtis Pullen (Victoria, CA); Roy Kim (Los Altos, CA)
Assignee: PURE STORAGE, INC.
G06F9/5011G06F9/4856G06F9/505G06F16/2272G06F16/258
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,404
App. No.
17/721,175
Granted
Jun 11, 2024
Kind
B2
Abstract

Executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources, including: receiving, from a data producer, a dataset; storing, within the storage system, the dataset; allocating processing resources to an analytics application; and executing the analytics application on the processing resources, including ingesting the dataset from the storage system.

Claims (69)

1. A method comprising:

receiving a dataset by polling a location to which a data producer writes the dataset;

converting, by a storage system, the dataset for use by an analytics application;

storing, by the storage system, the dataset within shared storage resources of the storage system at a particular location specific to the data producer;

executing the analytics application on first compute resources of the storage system to ingest a first portion of the dataset that is read from the particular location of the shared storage resources;

detecting that the analytics application has ceased executing properly on the first compute resources; and

in response to detecting that the analytics application has ceased executing properly on the first compute resources, executing the analytics application on second compute resources of the storage system to ingest a second portion of the dataset from the particular location of the shared storage resources that is specific to the data producer while bypassing copying the second portion of the dataset to a different location of the shared storage resources.

2. The method of claim 1 , wherein the first compute resources and the second compute resources are located within a physical storage system.

3. The method of claim 1 , further comprising:

allocating additional processing resources to a real-time analytics application; and

executing the real-time analytics application on the additional processing resources, including ingesting at least a portion of the dataset.

4. The method of claim 1 , further comprising:

receiving, by a storage system, an unstructured dataset;

converting, by the storage system, the unstructured dataset into a structured dataset; and

storing the structured dataset within the storage system.

5. The method of claim 1 , further comprising:

detecting that the analytics application needs additional processing resources;

allocating additional processing resources to the analytics application; and

executing the analytics application on the additional processing resources.

6. The method of claim 1 , wherein the dataset includes log files describing one or more execution states of a computing system and wherein executing the analytics application on the compute resources further comprises:

evaluating the log files to identify one or more execution patterns associated with the computing system.

7. The method of claim 6 , wherein evaluating the log files to identify one or more execution patterns associated with the computing system further comprises:

comparing fingerprints associated with known execution patterns to information contained in the log files.

8. An apparatus comprising:

a memory; and

a processing device, operatively coupled to the memory, configured to:

receive a dataset by polling a location to which a data producer writes the dataset;

convert, by a storage system, the dataset for use by an analytics application;

storing, by the storage system, the dataset within shared storage resources of the storage system at a particular location specific to the data producer

execute the analytics application on first compute resources of the storage system to ingest a first portion of the dataset that is read from the particular location of the shared storage resources;

detect that the analytics application has ceased executing properly on the first compute resources; and

in response to detecting that the analytics application has ceased executing properly, execute the analytics application on second compute resources of the storage system to ingest a second portion of the dataset from the particular location that is specific to the data producer while bypassing copying the second portion of the dataset to a different location of the shared storage resources.

9. The apparatus of claim 8 , wherein the first compute resources and the second compute resources are located within a physical storage system.

10. The apparatus of claim 8 , wherein the processing device is further configured to:

allocate additional processing resources to a real-time analytics application; and

execute the real-time analytics application on the additional processing resources, including ingesting at least a portion of the dataset.

11. The apparatus of claim 8 , wherein the processing device is further configured to:

receive, by a storage system, an unstructured dataset;

convert, by the storage system, the unstructured dataset into a structured dataset; and

store the structured dataset within the storage system.

12. The apparatus of claim 8 , wherein the processing device is further configured to:

detect that the analytics application needs additional processing resources;

allocate additional processing resources to the analytics application; and

execute the analytics application on the additional processing resources.

13. The apparatus of claim 8 , wherein the dataset includes log files describing one or more execution states of a computing system and wherein to execute the analytics application on the compute resources, the processing device is further configured to:

evaluate the log files to identify one or more execution patterns associated with the computing system.

14. The apparatus of claim 13 , wherein to evaluate the log files to identify one or more execution patterns associated with the computing system, the processing device is further configured to:

compare fingerprints associated with known execution patterns to information contained in the log files.

15. A non-transitory computer readable storage medium storing instructions which, when executed, cause a processing device to:

receive a dataset by polling a location to which a data producer writes the dataset;

convert, by a storage system, the dataset for use by an analytics application;

store, by the storage system, the dataset within shared storage resources of the storage system at a particular location specific to the data producer;

execute the analytics application on first compute resources of the storage system to ingest a first portion of the dataset that is read from the particular location of the shared storage resources;

detect that the analytics application has ceased executing properly on the first compute resources; and

in response to detecting that the analytics application has ceased executing properly on the first compute resources, execute the analytics application on second compute resources of the storage system to ingest a second portion of the dataset from the particular location of the shared storage resources that is specific to the data producer while bypassing copying the second portion of the dataset to a different location of the shared storage resources.

16. The non-transitory computer readable storage medium of claim 15 , wherein the first compute resources and the second compute resources are located within a physical storage system.

17. The non-transitory computer readable storage medium of claim 15 , wherein the processing device is further configured to:

allocate additional processing resources to a real-time analytics application; and

execute the real-time analytics application on the additional processing resources, including ingesting at least a portion of the dataset.

18. The non-transitory computer readable storage medium of claim 15 , wherein the processing device is further configured to:

receive, by a storage system, an unstructured dataset;

convert, by the storage system, the unstructured dataset into a structured dataset; and

store the structured dataset within the storage system.

19. The non-transitory computer readable storage medium of claim 15 , wherein the processing device is further configured to:

detect that the analytics application needs additional processing resources;

allocate additional processing resources to the analytics application; and

execute the analytics application on the additional processing resources.

20. The non-transitory computer readable storage medium of claim 15 , wherein the dataset includes log files describing one or more execution states of a computing system and wherein to execute the analytics application on the compute resources the processing device is further configured to:

evaluate the log files to identify one or more execution patterns associated with the computing system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2022
From: JIBAJA, IVAN; JAIKUMAR, PRASHANT; DORSETT, STEFAN; PULLEN, CURTIS; KIM, ROY
To: PURE STORAGE, INC.
Reel/Frame 059604/0974 →
Continuity (5)
Continuation 16659798 · Oct 22, 2019
Continuation 15883333 · Jan 30, 2018
Provisional Application 62620286 · Jan 22, 2018
Provisional Application 62574534 · Oct 19, 2017
Related Publication 20220237037A1 · Jul 28, 2022