IP Library Granted Patent US 11,860,820
Granted Patent B1
US 11,860,820 · App. 16/374,175 · Granted Jan 2, 2024

Processing data through a storage system in a data pipeline

Inventors: Ivan Jibaja (San Jose, CA); Curtis Pullen (San Jose, CA); Stefan Dorsett (San Jose, CA); Srinivas Chellappa (Sunnyvale, CA); Prashant Jaikumar (Sunnyvale, CA)
Assignee: PURE STORAGE, INC.
G06F16/156G06F9/45558G06F16/1734G06F16/1858G06F2009/45595
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,860,820
App. No.
16/374,175
Granted
Jan 2, 2024
Kind
B1
Abstract

Processing data through a storage system in a data pipeline including receiving, by the storage system, a dataset from a collector on a data producer, wherein the dataset is disaggregated from metadata for the dataset by the collector; storing the dataset on the storage system; receiving, by the storage system from a data indexer, a request for data from the dataset, wherein the request for the data comprises the metadata gathered by the collector on the data producer; servicing, by the storage system, the request for the data by locating the data using the metadata gathered by the collector on the data producer and received in the request for the data; and receiving, from the data indexer, indexed data indexed using the metadata gathered by the collector on the data producer.

Claims (59)

1. A method of processing data through a storage system in a data pipeline, the method comprising:

receiving, by the storage system, a dataset from a collector on a data producer, wherein the dataset is disaggregated from metadata for the dataset by the collector;

storing the dataset on the storage system;

receiving, by the storage system from a data indexer, a request for data from the dataset, wherein the request for the data comprises the metadata gathered by the collector on the data producer;

servicing, by the storage system, the request for the data by locating the data using the metadata gathered by the collector on the data producer and received in the request for the data; and

receiving, from the data indexer, indexed data indexed using the metadata gathered by the collector on the data producer.

2. The method of claim 1 , wherein receiving, by the storage system, the dataset from the collector on the data producer comprises receiving, by the storage system, the dataset as a continuation of a previous dataset interrupted by a pause in data communications.

3. The method of claim 1 , wherein receiving, by the storage system, the dataset from the collector on the data producer comprises:

receiving the dataset as a line delineated stream; and

organizing the line delineated stream into one or more data objects.

4. The method of claim 1 , wherein storing the dataset on the storage system comprises:

receiving a log line from the dataset;

identifying a log line type for the log line; and

generating a structure for the log line using the log line type.

5. The method of claim 1 , wherein storing the dataset on the storage system comprises organizing the dataset in tiers within the storage system based on previously received requests for data.

6. The method of claim 1 , further comprising:

receiving a request for data, wherein the request comprises a pattern and an index;

servicing the request; and

prefetching additional data based on the pattern and index.

7. The method of claim 1 , wherein the metadata is received by the storage system separately from the dataset and exposed to the data indexer.

8. An apparatus for processing data through a storage system in a data pipeline, the apparatus comprising a computer processor, a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the steps of:

receiving, by the storage system, a dataset from a collector on a data producer, wherein the dataset is disaggregated from metadata for the dataset by the collector;

storing the dataset on the storage system;

receiving, by the storage system from a data indexer, a request for data from the dataset, wherein the request for the data comprises the metadata gathered by the collector on the data producer;

servicing, by the storage system, the request for the data by locating the data using the metadata gathered by the collector on the data producer and received in the request for the data; and

receiving, from the data indexer, indexed data indexed using the metadata gathered by the collector on the data producer.

9. The apparatus of claim 8 , wherein receiving, by the storage system, the dataset from the collector on the data producer comprises receiving, by the storage system, the dataset as a continuation of a previous dataset interrupted by a pause in data communications.

10. The apparatus of claim 8 , wherein receiving, by the storage system, the dataset from the collector on the data producer comprises:

receiving the dataset as a line delineated stream; and

organizing the line delineated stream into one or more data objects.

11. The apparatus of claim 8 , wherein storing the dataset on the storage system comprises:

receiving a log line from the dataset;

identifying a log line type for the log line; and

generating a structure for the log line using the log line type.

12. The apparatus of claim 8 , wherein storing the dataset on the storage system comprises organizing the dataset in tiers within the storage system based on previously received requests for data.

13. The apparatus of claim 8 , further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the steps of:

receiving a request for data, wherein the request comprises a pattern and an index;

servicing the request; and

prefetching additional data based on the pattern and index.

14. The apparatus of claim 8 , wherein the metadata is received by the storage system separately from the dataset and exposed to the data indexer.

15. A computer program product for processing data through a storage system in a data pipeline, the computer program product disposed upon a non-transitory computer readable medium, the computer program product comprising computer program instructions that, when executed, cause a computer to carry out the steps of:

receiving, by the storage system, a dataset from a collector on a data producer, wherein the dataset is disaggregated from metadata for the dataset by the collector;

storing the dataset on the storage system;

receiving, by the storage system from a data indexer, a request for data from the dataset, wherein the request for the data comprises the metadata gathered by the collector on the data producer;

servicing, by the storage system, the request for the data by locating the data using the metadata gathered by the collector on the data producer and received in the request for the data; and

receiving, from the data indexer, indexed data indexed using the metadata gathered by the collector on the data producer.

16. The computer program product of claim 15 , wherein receiving, by the storage system, the dataset from the collector on the data producer comprises receiving, by the storage system, the dataset as a continuation of a previous dataset interrupted by a pause in data communications.

17. The computer program product of claim 15 , wherein receiving, by the storage system, the dataset from the collector on the data producer comprises:

receiving the dataset as a line delineated stream; and

organizing the line delineated stream into one or more data objects.

18. The computer program product of claim 15 , wherein storing the dataset on the storage system comprises:

receiving a log line from the dataset;

identifying a log line type for the log line; and

generating a structure for the log line using the log line type.

19. The computer program product of claim 15 , wherein storing the dataset on the storage system comprises organizing the dataset in tiers within the storage system based on previously received requests for data.

20. The computer program product of claim 15 , further comprising computer program instructions that, when executed, cause the computer to carry out the steps of:

receiving a request for data, wherein the request comprises a pattern and an index;

servicing the request; and

prefetching additional data based on the pattern and index.

Assignments (3)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 11, 2025
From: BARCLAYS BANK PLC, AS ADMINISTRATIVE AGENT
To: PURE STORAGE, INC.
Reel/Frame 071558/0523 →
SECURITY INTEREST Recorded Aug 26, 2020
From: PURE STORAGE, INC.
To: BARCLAYS BANK PLC AS ADMINISTRATIVE AGENT
Reel/Frame 053867/0581 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2019
From: JIBAJA, IVAN; PULLEN, CURTIS; DORSETT, STEFAN; CHELLAPPA, SRINIVAS; JAIKUMAR, PRASHANT
To: PURE STORAGE, INC.
Reel/Frame 048783/0198 →
Continuity (1)
Provisional Application 62729730 · Sep 11, 2018
Cited By (1)
US 12,306,983