IP Library › Granted Patent US 11,775,371
Granted Patent B2
US 11,775,371 · App. 16/881,799 · Granted Oct 3, 2023

Remote validation and preview

Inventor: Madhukar Devaraju (San Jose, CA)
Assignee: StreamSets, Inc.
G06F11/0772G06F3/0482G06F9/544G06F11/0751G06F21/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,371
App. No.
16/881,799
Granted
Oct 3, 2023
Kind
B2
Abstract

Systems and methods are directed to remote validation and preview. An example system receives an indication of a portion of the data pipeline to be processed, generates a data pipeline configuration file describing operations in the portion of the data pipeline, causes a software framework to perform operations corresponding to the portion of the data pipeline, receives results of the operations corresponding to the portion of the data pipeline, and causes presentation of the results on a graphical user interface of a computing device.

Claims (48)

1. A method for generating a remote preview of a data pipeline comprising:

receiving, via a user interface, an indication the data pipeline to be processed, wherein a first portion of the data pipeline is stored on a first computing device and a second portion of the data pipeline is stored on a second computing device;

generating a data pipeline configuration file describing operations in the first portion and the second portion of the data pipeline;

causing a cluster-computing framework, using an application programming interface (API), to perform operations corresponding to the first portion and the second portion of the data pipeline, the causing being based on the data pipeline configuration file;

receiving, from the cluster-computing framework, results of the operations corresponding to the first portion and the second portion of the data pipeline, the results comprising data transformed by the operations; and

causing presentation of the results on a graphical user interface of a third computing device.

2. The method of claim 1 , wherein the first, second and third computing device form a cluster of computing devices, each computing device in the cluster using the cluster-computing framework.

3. The method of claim 1 , further comprising:

transmitting offset information of the data pipeline to the cluster-computing framework, the offset information providing information on how to restart the data pipeline from a last batch of data that was processed.

4. The method of claim 1 , wherein receiving the indication of the data pipeline further comprises:

presentation of a user configuration window, the user configuration window providing user interface elements for specifying the data pipeline to be processed; and

detecting, via the user configuration window, user input specifying the data pipeline to be processed by the cluster-computing framework.

5. The method of claim 1 , wherein the results comprise data transformed by the operations performed by the cluster-computing framework.

6. The method of claim 5 , wherein the results further comprise error notifications due to a failure in the data pipeline.

7. The method of claim 1 , wherein the data pipeline is displayed on the graphical user interface of the computing device, and the causing presentation of the results further comprising:

modifying the graphical user interface of the computing device to cause presentation of the results.

8. A computing apparatus, the computing apparatus comprising:

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving, via a user interface, an indication a data pipeline to be processed, wherein a first portion of the data pipeline is stored on a first computing device and a second portion of the data pipeline is stored on a second computing device;

generating a data pipeline configuration file describing operations in the first portion and the second portion of the data pipeline;

causing a cluster-computing framework, using an application programming interface (API), to perform operations corresponding to the first portion and the second portion of the data pipeline, the causing being based on the data pipeline configuration file;

receiving, from the cluster-computing framework, results of the operations corresponding to the first portion and the second portion of the data pipeline, the results comprising data transformed by the operations; and

causing presentation of the results on a graphical user interface of a third computing device.

9. The computing apparatus of claim 8 , wherein the data pipeline configuration file further comprises data pipeline credential information.

10. The computing apparatus of claim 8 , the operations further comprising:

transmitting offset information of the data pipeline to the cluster-computing framework, the offset information providing information on how to restart the data pipeline from a last batch of data that was processed.

11. The computing apparatus of claim 8 , wherein receiving the indication of the data pipeline further comprises:

presentation of a user configuration window, the user configuration window providing user interface elements for specifying the data pipeline to be processed; and

detecting, via the user configuration window, user input specifying the data pipeline to be processed by the cluster-computing framework.

12. The computing apparatus of claim 8 , wherein the results comprise data transformed by the operations performed by the cluster-computing framework.

13. The computing apparatus of claim 12 , wherein the results further comprise error notifications due to a failure in the data pipeline.

14. The computing apparatus of claim 8 , wherein the data pipeline is displayed on the graphical user interface of the computing device, and the causing presentation of the results further comprises:

modifying the graphical user interface of the computing device to cause presentation of the results.

15. A machine storage medium storing instructions that when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving, via a user interface, an indication a data pipeline to be processed, wherein a first portion of the data pipeline is stored on a first computing device and a second portion of the data pipeline is stored on a second computing device;

generating a data pipeline configuration file describing operations in the first portion and the second portion of the data pipeline;

causing a cluster-computing framework, using an application programming interface (API), to perform operations corresponding to the first portion and the second portion of the data pipeline, the causing being based on the data pipeline configuration file;

receiving, from the cluster-computing framework, results of the operations corresponding to the first portion and the second portion of the data pipeline, the results comprising data transformed by the operations; and

causing presentation of the results on a graphical user interface of a third computing device.

16. The machine storage medium of claim 15 , wherein the data pipeline configuration file further comprises data pipeline credential information.

17. The machine storage medium of claim 15 , the operations further comprising:

transmitting offset information of the data pipeline to the cluster-computing framework, the offset information providing information on how to restart the data pipeline from a last batch of data that was processed.

18. The machine storage medium of claim 15 , wherein receiving the indication of the data pipeline further comprises:

presentation of a user configuration window, the user configuration window providing user interface elements for specifying the data pipeline to be processed; and

detecting, via the user configuration window, user input specifying the data pipeline to be processed by the cluster-computing framework.

19. The machine storage medium of claim 15 , wherein the results further comprise error notifications due to a failure in the data pipeline.

20. The machine storage medium of claim 19 , wherein the results further comprise error notifications due to a failure in the data pipeline.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2024
From: STREAMSETS, INC.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069219/0116 →
RELEASE OF SECURITY INTEREST Recorded May 9, 2022
From: AB PRIVATE CREDIT INVESTORS LLC, AS ADMINISTRATIVE AGENT
To: STREAMSETS, INC.
Reel/Frame 059870/0831 →
SECURITY INTEREST Recorded Nov 25, 2020
From: STREAMSETS, INC.
To: AB PRIVATE CREDIT INVESTORS LLC, AS ADMINISTRATIVE AGENT
Reel/Frame 054472/0345 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2020
From: DEVARAJU, MADHUKAR
To: STREAMSETS, INC.
Reel/Frame 053619/0341 →
Continuity (1)
Related Publication 20210365312A1 · Nov 25, 2021