Genomic infrastructure for on-site or cloud-based DNA and RNA processing and analysis
A system, method and apparatus for executing a sequence analysis pipeline on genetic sequence data includes a integrated circuit formed of a set of hardwired digital logic circuits that are interconnected by physical electrical interconnects. One of the physical electrical interconnects forms an input to the integrated circuit connected with an electronic data source for receiving reads of genomic data. The hardwired digital logic circuits are arranged as a set of processing engines, each processing engine being formed of a subset of the hardwired digital logic circuits to perform one or more steps in the sequence analysis pipeline on the reads of genomic data. Each subset of the hardwired digital logic circuits is formed in a wired configuration to perform the one or more steps in the sequence analysis pipeline.
1 . A method for dynamic configuration and execution of a genomic data processing pipeline based on one or more user-selectable options presented via a graphical user interface (GUI) of a nucleic acid sequencing device, the method comprising:
obtaining, by one or more processors of the nucleic acid sequencing device executing an application programming interface (API) that provides an interface between the GUI of the nucleic acid sequencing device and a programmable logic device, first data representing a selection of one or more of the user-selectable options submitted via the GUI, wherein one of the user-selectable options identifies a particular reference sequence to be used by a genomic data processing pipeline;
configuring, by one or more processors of the nucleic acid sequencing device, one or more programmable hardware resources of the programmable logic device from a first state that does not include hardware resources configured as a genomic analysis pipeline that uses the particular reference sequence identified by the first data into a second state that includes hardware resources that have been configured as a genomic analysis pipeline that uses the particular reference sequence identified by the first data;
providing, to the programmable logic device and by one or more processors of the nucleic acid sequencing device using the API, second data representing a set of genomic data or a set of data derived from genomic data;
instructing, by one or more processors of the nucleic acid sequencing device, the configured genomic data processing pipeline to execute a genomic processing operation on the obtained second data to generate result data;
obtaining, by one or more processors of the nucleic acid sequencing device, the result data that is generated by execution of the configured genomic data processing pipeline on the obtained second data by the programmable logic device; and
providing, by the one or more processors of the nucleic acid sequencing device, output data that is based on the result data.
2 . The method of claim 1 , wherein providing, by the one or more processors of the nucleic acid sequencing device, the output data that is based on the result data comprises:
providing, by the one or more processors of the nucleic acid sequencing device, output data that is based on the result data for display on a display of the nucleic acid sequencing device.
3 . The method of claim 1 , wherein providing, by the one or more processors of the nucleic acid sequencing device, the output data that is based on the result data comprises:
providing, by the one or more processors of the nucleic acid sequencing device, output data that is based on the result data for output by a device that is different from the nucleic acid sequencing device.
4 . The method of claim 1 , wherein the genomic processing operation includes one or more of a read mapping operation, a read alignment operation, a sorting operation, a variant calling operation, or a tertiary analysis operation.
5 . The method of claim 1 , wherein the method further comprises:
obtaining, by one or more processors of the nucleic acid sequencing device, data representing the particular reference sequence; and
storing, by one or more processors of the nucleic acid sequencing device, the obtained data representing the particular reference sequence in a memory device that is accessible by the programmable logic device.
6 . The method of claim 1 , wherein providing, to the programmable logic device and by one or more processors of the nucleic acid sequencing device using the API, second data representing a set of genomic data or a set of data derived from genomic data comprises:
obtaining, by one or more processors of the nucleic acid sequencing device, at least a portion of a FASTQ file generated by the nucleic acid sequencing device; and
storing, by one or more processors of the nucleic acid sequencing device and via the API, the obtained portion of the FASTQ file in a memory device that is accessible by the programmable logic device.
7 . A system for dynamic configuration and execution of a genomic data processing pipeline based on one or more user-selectable options presented via a graphical user interface (GUI) of a nucleic acid sequencing device, the system comprising:
a nucleic acid sequencing device that includes one or more processors and one more memory devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations; and
a second device that includes one or more programmable logic devices that can be configured to execute one or more operations on input data obtained by the second device;
wherein the system is configured to perform operations comprising:
obtaining, by the one or more processors of the nucleic acid sequencing device executing an application programming interface (API) that provides an interface between the GUI of the nucleic acid sequencing device and a programmable logic device, first data representing a selection of one or more of the user-selectable options submitted via the GUI, wherein one of the user-selectable options identifies a particular reference sequence to be used by a genomic data processing pipeline;
configuring, by the one or more processors of the nucleic acid sequencing device, one or more programmable hardware resources of the programmable logic device from a first state that does not include hardware resources configured as a genomic analysis pipeline that uses the particular reference sequence identified by the first data into a second state that includes hardware resources that have been configured as a genomic analysis pipeline that uses the particular reference sequence identified by the first data;
providing, to the programmable logic device and by one or more processors of the nucleic acid sequencing device using the API, second data representing a set of genomic data or a set of data derived from genomic data;
instructing, by the one or more processors of the nucleic acid sequencing device, the configured genomic data processing pipeline to execute a genomic processing operation on the obtained second data to generate result data;
obtaining, by the one or more processors of the nucleic acid sequencing device, the result data that is generated by execution of the configured genomic data processing pipeline on the obtained second data by the programmable logic device; and
providing, by the one or more processors of the nucleic acid sequencing device executing one or more of the instructions, output data that is based on the result data.
8 . The system of claim 7 , wherein providing, by the one or more processors of the nucleic acid sequencing device, the output data that is based on the result data comprises:
providing, by the one or more processors of the nucleic acid sequencing device, output data that is based on the result data for display on a display of the one or more processors of the nucleic acid sequencing device.
9 . The system of claim 7 , wherein providing, by the one or more processors of the nucleic acid sequencing device, the output data that is based on the result data comprises:
providing, by the one or more processors of the nucleic acid sequencing device, output data that is based on the result data for output by a second user device that is different from the nucleic acid sequencing device.
10 . The system of claim 7 , wherein the genomic processing operation includes one or more of a read mapping operation, a read alignment operation, a sorting operation, a variant calling operation, or a tertiary analysis operation.
11 . The system of claim 7 , wherein the operations further comprise:
obtaining, by the one or more processors of the nucleic acid sequencing device, data representing the particular reference sequence; and
storing, by the one or more processors of the nucleic acid sequencing device, the obtained data representing the particular reference sequence in a memory device that is accessible by the programmable logic device.
12 . The system of claim 7 , wherein providing, to the programmable logic device and by the one or more processors of the nucleic acid sequencing device using the API, second data representing a set of genomic data or a set of data derived from genomic data comprises:
obtaining, by the one or more processors of the nucleic acid sequencing device, at least a portion of a FASTQ file generated by the nucleic acid sequencing device; and
storing, by the one or more processors of the nucleic acid sequencing device and via the API, the obtained portion of the FASTQ file in a memory device that is accessible by the programmable logic device.
13 . One or non-transitory more computer readable storage media storing instructions that, when executed by a nucleic acid sequencing device, cause the nucleic acid sequencing device to perform operations for dynamic configuration and execution of a genomic data processing pipeline based on one or more user-selectable options presented via a graphical user interface (GUI) of the nucleic acid sequencing device, the operations comprising:
obtaining, by one or more processors of the nucleic acid sequencing device executing an application programming interface (API) that provides an interface between the GUI of the nucleic acid sequencing device and a programmable logic device, first data representing a selection of one or more of the user-selectable options submitted via the GUI, wherein one of the user-selectable options identify a particular reference sequence to be used by a genomic data processing pipeline;
configuring, by one or more processors of the nucleic acid sequencing device, one or more programmable hardware resources of the programmable logic device from a first state that does not include hardware resources configured as a genomic analysis pipeline that uses the particular reference sequence identified by the first data into a second state that includes hardware resources that have been configured as a genomic analysis pipeline that uses the particular reference sequence identified by the first data;
providing, by one or more processors of the nucleic acid sequencing device, second data representing a set of genomic data or a set of data derived from genomic data; and
instructing, by one or more processors of the nucleic acid sequencing device, the configured genomic data processing pipeline to execute a genomic processing operation on the obtained second data to generate result data;
obtaining, by one or more processors of the nucleic acid sequencing device, the result data that is generated by execution of the configured genomic data processing pipeline on the obtained second data by the programmable logic device; and
providing, by one or more processors of the nucleic acid sequencing device, output data that is based on the result data.
14 . The one or more non-transitory computer readable storage media of claim 13 , wherein providing, by one or more processors of the nucleic acid sequencing device, the output data that is based on the result data comprises:
providing, by one or more processors of the nucleic acid sequencing device, output data that is based on the result data for display on a display of the nucleic acid sequencing device.
15 . The one or more non-transitory computer readable storage media of claim 13 , wherein providing, by one or more processors of the nucleic acid sequencing device, the output data that is based on the result data comprises:
providing, by one or more processors of the nucleic acid sequencing device, output data that is based on the result data for output by a second user device that is different from the nucleic acid sequencing device.
16 . The one or more non-transitory computer readable storage media of claim 13 , wherein the genomic processing operation includes one or more of a read mapping operation, a read alignment operation, a sorting operation, a variant calling operation, or a tertiary analysis operation.
17 . The one or more non-transitory computer readable storage media of claim 13 , wherein the operations further comprise:
obtaining, by one or more processors of the nucleic acid sequencing device, data representing the particular reference sequence; and
storing, by one or more processors of the nucleic acid sequencing device, the obtained data representing the particular reference sequence in a memory device that is accessible by the programmable logic device.
18 . The one or more non-transitory computer readable storage media of claim 13 , wherein obtaining, by one or more processors of the nucleic acid sequencing device, second data representing a set of genomic data or a set of data derived from genomic data comprises:
obtaining, by one or more processors of the nucleic acid sequencing device, at least a portion of a FASTQ file generated by the nucleic acid sequencing device; and
storing, by one or more processors of the nucleic acid sequencing device, the obtained portion of the FASTQ file in a memory device that is accessible by the programmable logic device.