IP Library Granted Patent US 10,025,791
Granted Patent B2
US 10,025,791 · App. 15/148,848 · Granted Jul 17, 2018

Metadata-driven workflows and integration with genomic data processing systems and techniques

Inventor: Frank N. Lee (Sunset Hills, MO)
Assignee: International Business Machines Corporation
G06F17/3012G06F9/46G06F17/3028G06F17/30106G06F17/30265G06F17/30722G06Q10/06G06Q10/10G06Q50/01G06F8/34G06F8/70G06F17/30017G06F17/30064G06Q10/0633
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,025,791
App. No.
15/148,848
Filed
May 6, 2016
Granted
Jul 17, 2018
Kind
B2
Art Unit
2192
USPC
718/102
Abstract

Systems, methods and computer program products configured to provide and perform metadata-based workflow management are disclosed. The inventive subject matter includes a computer readable storage medium having computer readable program instructions embodied therewith. The computer readable program instructions are configured to: initiate a workflow configured to process data; associate the data with metadata; and drive at least a portion of the workflow based on at least some of the metadata. The metadata include anchoring metadata; common metadata; and custom metadata. Inventive subject matter also encompasses a method for managing genomic data processing workflows using metadata includes: initiating a workflow; receiving a request to manage the workflow using metadata comprising: anchoring metadata, common metadata, and custom metadata, associating the metadata with the data; and driving at least a portion of the workflow based on the metadata. The workflow involves genomic analyses.

Claims (72)

1. A computer program product for driving genomic data processing workflows using metadata, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

initiate, by the processor, a workflow configured to process data, wherein the workflow comprises one or more genomic analysis operations selected from the group consisting of: base calling, variant calling, phylogenetic analysis, primer design, and amplicon design;

receiving, at the processor, a request to manage the workflow using metadata comprising:

anchoring metadata configured to uniquely identify the workflow by using an alphanumeric string;

common metadata comprising one or more characteristics selected from the group consisting of: sample characteristics, processing site characteristics, laboratory characteristics, instrument characteristics, assay characteristics, temporal characteristics, security characteristics and project characteristics; and

custom metadata comprising workflow characteristics and/or data characteristics; and

associate, by the processor, the metadata with the genomic data;

drive, by the processor, at least a portion of the workflow based on the metadata, wherein driving the workflow based at least in part on the metadata comprises:

determining new data and/or at least one new processing setting to use in connection with repeating at least a portion of the workflow; and

repeating the portion of the workflow using the new data and/or the new processing setting, wherein the determining is based at least in part on the common metadata and/or the custom metadata; and

wherein the new data and/or the new processing setting comprise a modified number of permissible gaps in an alignment based at least in part on an average sequence length of input sequence data.

2. The computer program product as recited in claim 1 , further comprising program instructions executable by the processor to cause the processor to:

dynamically create at least some of the metadata in response to driving the portion of the workflow; and

store the dynamically created metadata.

3. The computer program product as recited in claim 1 , wherein the anchoring metadata consists of a single value; and

wherein the custom metadata and the anchoring metadata are characterized by a many-to-one relationship.

4. The computer program product as recited in claim 1 , wherein the workflow characteristics comprise a step or operation number corresponding to a portion of the workflow with which the custom metadata are associated, the step or operation number being embodied as one or more of:

an alphanumeric string defining a relative position of the portion of the workflow in an overall workflow sequence;

a list of users and/or processors that can access the workflow; and

a list of file system directories and/or locations accessible by the workflow.

5. A computer-implemented method for managing genomic data processing workflows using metadata, the method comprising:

initiating a workflow, wherein the workflow comprises one or more genomic analysis operations selected from the group consisting of: base calling, variant calling, phylogenetic analysis, primer design, and amplicon design;

receiving a request to manage the workflow using metadata comprising:

anchoring metadata, wherein the anchoring metadata uniquely identify the workflow by using an alphanumeric string;

common metadata comprising one or more characteristics selected from the group consisting of: sample characteristics, processing site characteristics, laboratory characteristics, instrument characteristics, assay characteristics, temporal characteristics, security characteristics and project characteristics; and

custom metadata comprising workflow characteristics and/or data characteristics; and the method further comprising:

associating the metadata with the genomic data; and

driving at least a portion of the workflow based on the metadata, wherein driving the workflow based at least in part on the metadata comprises:

determining new data and/or at least one new processing setting to use in connection with repeating at least a portion of the workflow; and

repeating the portion of the workflow using the new data and/or the new processing setting, wherein the determining is based at least in part on the common metadata and/or the custom metadata; and

wherein the new data and/or the new processing setting comprise a modified number of permissible gaps in an alignment based at least in part on an average sequence length of input sequence data.

6. The method as recited in claim 5 , further comprising generating at least some of the metadata while performing the workflow;

wherein the generated metadata are associated with the data while performing the workflow; and

wherein the metadata are generated by a file system as part of one or more input/output (I/O) operations.

7. The method as recited in claim 5 , wherein driving the workflow based at least in part on the metadata comprises: parallelizing at least one of the genomic analysis operations across a plurality of dynamically selected compute resources comprising at least one of:

a plurality of processing nodes of a local processing cluster; and

a plurality of processing nodes of a cloud processing pool, and

wherein the compute resources are dynamically selected based on one or more of the common metadata and the custom metadata.

8. The method as recited in claim 5 , wherein the anchoring metadata are configured to uniquely identify the workflow to at least one of a file system and an operating system hosting one or more of the workflow, the data and the metadata, and

wherein the anchoring metadata are a single value that remains unchanged throughout the course of the workflow.

9. The method as recited in claim 5 , wherein all of the custom metadata associated with the workflow correspond to the anchoring metadata configured to uniquely identify the workflow such that the custom metadata and the anchoring metadata have a many-to-one relationship.

10. The method as recited in claim 5 , wherein the custom metadata change at least in part in accordance with each of the genomic analysis operation(s) performed on the data to which the custom metadata correspond.

11. A computer program product for driving genomic data processing workflows using metadata, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

receive, at the processor:

workflow data;

metadata associated with the workflow data, wherein the metadata comprise a plurality of metadata generations, each metadata generation corresponding to at least one operation of the workflow, each metadata generation including:

anchoring metadata configured to uniquely identify the workflow by using an alphanumeric string;

common metadata comprising one or more characteristics selected from: sample characteristics, processing site characteristics, laboratory characteristics, instrument characteristics, assay characteristics, temporal characteristics, security characteristics and project characteristics; and

custom metadata comprising workflow characteristics and/or data characteristics; and

a request to manage a workflow using the metadata;

distribute, by the processor, the workflow data and the associated metadata across a plurality of distributed resources of a cloud computing environment; and

associate the metadata with the workflow data by indexing, using the processor, the workflow data according to the metadata; and

drive at least a portion of the workflow based on the metadata, wherein driving the workflow based at least in part on the metadata comprises:

determining new data and/or at least one new processing setting to use in connection with repeating at least a portion of the workflow; and

repeating the portion of the workflow using the new data and/or the new processing setting, wherein the determining is based at least in part on the common metadata and/or the custom metadata; and

wherein the new data and/or the new processing setting comprise a modified number of permissible gaps in an alignment based at least in part on an average sequence length of input sequence data.

12. The computer program product as recited in claim 11 , further comprising program instructions executable by the processor to cause the processor to:

receive a query corresponding to at least some of the metadata;

locate data corresponding to the query based on the metadata; and

distribute the located data to one or more remote destinations based on the query,

wherein the remote destination(s) are not included in the cloud computing environment.

13. The computer program product as recited in claim 11 , wherein the cloud computing environment comprises:

a plurality of computer readable storage media configured as a cloud storage environment; and

a plurality of processing nodes arranged in at least one cloud processing cluster.

14. The computer program product as recited in claim 11 , wherein the metadata are associated with the workflow data according to a relational structure;

wherein the relational structure is a relational database;

wherein the relational database comprises a plurality of key/value pairs; and

wherein the value of each key/value pair consists of at least some of the metadata.

15. The computer program product as recited in claim 11 , the program instructions executable by the processor to cause the processor to index the workflow metadata further comprising program instructions executable by the processor to cause the processor to dynamically index the workflow data according to each generation of the metadata.

16. The computer program product as recited in claim 11 ,

wherein the anchoring metadata consists of a single value; and

wherein the custom metadata and the anchoring metadata are characterized by a many-to-one relationship.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2016
From: LEE, FRANK N.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 038638/0557 →
Continuity (2)
Continuation 14243301 · Apr 2, 2014
Related Publication 20160253321A1 · Sep 1, 2016
Cited By (4)
US 12,445,519 US 12,536,174 US 12,575,806 US 12,672,844