IP Library Granted Patent US 10,769,167
Granted Patent B1
US 10,769,167 · App. 16/777,637 · Granted Sep 8, 2020

Federated computational analysis over distributed data

Inventors: Pablo Prieto Barja (London, GB); Maria Chatzou (London, GB); Matija Sosic (London, GB); Martin Sosic (London, GB); Diogo Nuno Proenca Silva (London, GB); Bruno Filipe Ribeiro Goncalves (London, GB); Tiago Filipe Salgueiro De Jesus (London, GB); Olga Kruglova (London, GB); Damyan Dobrev (London, GB)
Assignee: LIFEBIT BIOTECH LIMITED
G06F16/256G06F16/285G06F21/6218G06N20/00H04L9/0891G06F16/148G06F16/164G06F16/2255G06F16/2471
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,769,167
App. No.
16/777,637
Granted
Sep 8, 2020
Kind
B1
Abstract

The present disclosure provides for computational data analysis across multiple data sources. A pipeline (or workflow) is imported and a dataset is selected. The dataset resides on a virtual file system and includes data residing on one or more storage locations associated with the virtual file system. One or more compute resources are selected to perform the pipeline analysis based at least on the imported pipeline and the dataset. The one or more compute resources are selected from a plurality of available compute resources associated with the one or more storage locations associated with the virtual file system. The pipeline analysis is performed using the selected compute resources on the dataset in one or more secure clusters. The resulting data generated from the pipeline analysis is submitted to the virtual file system.

Claims (36)

1. A method of performing computational data analysis comprising:

importing a pipeline;

selecting a dataset, the dataset residing on a virtual file system and including data residing on one or more storage locations associated with the virtual file system;

selecting one or more compute resources to perform a pipeline analysis based at least on the imported pipeline and the dataset, the one or more compute resources being selected from a plurality of available compute resources associated with the one or more storage locations associated with the virtual file system;

configuring one or more secure clusters within the virtual file system, the one or more secure clusters including the selected one or more compute resources;

perform the pipeline analysis by streaming the data to the one or more secure clusters within the virtual file system; and

submitting resulting data generated from the pipeline analysis to the virtual file system.

2. The method of claim 1 , wherein performing the pipeline analysis includes streaming the dataset to the one or more secure clusters from one of the one or more storage locations.

3. The method of claim 1 , wherein the virtual file system includes one or more files, the one or more files pointing to the data residing on the one or more storage locations.

4. The method of claim 1 , further comprising:

receiving a selection of a storage location of the one or more storage locations and a location within the virtual file system, wherein the resulting data generated from the pipeline analysis is submitted to the virtual file system in accordance with the selected storage location and the selected location within the virtual file system.

5. The method of claim 1 , wherein the one or more storage locations are cloud storage locations.

6. The method of claim 1 , wherein the one or more storage locations include cloud storage locations and localized user storage.

7. The method of claim 1 , wherein the pipeline analysis is performed on the one or more selected compute resources wherein at least one of the one or more selected compute resources is part of a first secure cluster of the one or more secure clusters and at least a second of the one or more selected compute resources is part of a second secure cluster of the one or more secure clusters.

8. The method of claim 7 , wherein the first secure cluster is correlated with a first storage location associated with the virtual file system and the second secure cluster is correlated with a second storage location associated with the virtual file system.

9. A system for performing computational data analysis comprising:

one or more storage locations including input data, the input data on each of the one or more storage locations being accessible by a computing device using a virtual the system; and

an analysis operating system executing on one or more processors of the computing device, the analysis operating system being configured to select one or more compute resources to perform analysis on the input data using a pipeline and to create one or more secure clusters within the virtual file system including the one or more compute resources, the one or more compute resources being configured to perform the analysis on the input data using the pipeline by streaming the input data to the one or more secure clusters, wherein the one or more secure clusters are not accessible after creation, wherein resulting data generated from the analysis is submitted to the virtual file system.

10. The system of claim 9 , wherein the input data is located on two or more storage locations.

11. The system of claim 10 , wherein the input data is represented by a dataset located on a virtual file system.

12. The system of claim 9 , wherein at least one of the one or more storage locations is a localized user storage location associated with localized user compute resources and wherein the analysis operating system is further configured to create a secure cluster including one or more of the localized user compute resources.

13. The system of claim 9 , wherein the analysis operating system is further configured to store results from the analysis of the input data using the pipeline to a location on the virtual file system.

14. A method of performing computational data analysis comprising:

importing a pipeline;

presenting one or more combinations of one or more compute resources to perform a pipeline analysis based on the pipeline, wherein the one or more compute resources are associated with one or more of a plurality of storage locations communicatively connected using a virtual file system, and wherein the pipeline analysis uses input data located on the plurality of storage locations;

selecting one or more compute resources to perform the pipeline analysis based at least on a user input;

perform the pipeline analysis using the one or more selected compute resources, the one or more compute resources being located within one or more secure clusters on the one or more storage locations; and

submitting resulting data generated from the pipeline analysis to one or more selected storage locations of the plurality of storage locations communicatively connected by the virtual file system.

15. The method of claim 14 , wherein presenting the one or more combinations of one or more compute resources includes presenting an estimated run time of the pipeline analysis for each of the one or more combinations of one or more compute resources.

16. The method of claim 15 , wherein machine learning is used to improve the estimated run time over time.

17. The method of claim 14 , wherein presenting the one or more combinations of one or more compute resources includes presenting an estimated cost of the pipeline analysis for each of the one or more combinations of one or more compute resources.

18. The method of claim 17 , wherein the estimated cost of the pipeline analysis is based on historical pricing data for the one or more compute resources.

19. The method of claim 14 , further comprising:

creating one of the one or more secure clusters on a cloud compute resource, the secure cluster including a monitor configured to return run time data during the analysis.

20. The method of claim 19 , further comprising:

destroying one or more keys to the secure cluster prior to the analysis.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2020
From: BARJA, PABLO PRIETO; CHATZOU, MARIA; SOSIC, MATIJA; SOSIC, MARTIN; PROENCA RICO SILVA, DIOGO NUNO; RIBEIRO GONCALVES, BRUNO FILIPE; SALGUEIRO DE JESUS, TIAGO FILIPE; KRUGLOVA, OLGA; DOBREV, DAMYAN
To: LIFEBIT BIOTECH LIMITED
Reel/Frame 051730/0908 →
Priority Claims (1)
GR 20190100568 · Dec 20, 2019 · national
Cited By (2)
US 12,519,781 US 12,626,005