Bioinformatics systems, apparatuses, and methods for performing secondary and/or tertiary processing
Systems, methods, and computer programs for analyzing genetic sequence data is disclosed. In one aspect, the system can include one or more of a first integrated circuit, with each first integrated circuit forming a central processing unit (CPU) that is responsive to one or more software algorithms that are configured to instruct the CPU to perform a first set of genomic processing steps of a sequence analysis pipeline. Additionally, the system can include one or more second integrated circuits, with each second integrated circuit forming a field programmable gate array (FPGA). The FPGA can be configured by firmware to arrange a set of hardwired digital logic circuits to perform a second set of genomic processing stages of the sequence analysis pipeline, the set of hardwired digital logic circuits of each FPGA being arranged as a set of processing engines to perform the second set of genomic processing stages.
1 . A method for constructing and using a chimeric reference standard specific to an individual subject for improving accuracy of analysis of a plurality of portions of a genome of the subject, the method comprising:
comparing a first portion of the genome of the subject to each reference standard of a plurality of reference standards stored in a database, wherein the genome of the subject is based on a biological sample obtained from the subject;
selecting a portion of a first reference standard that satisfies a predetermined matching threshold to a first portion of the genome of the subject based on the comparison of the first portion to the plurality of reference standards;
comparing a second portion of the genome of the subject to each reference standard of the plurality of reference standards;
selecting a portion of a second reference standard that satisfies a predetermined matching threshold to the second portion of the genome of the subject based on the comparison of the second portion to the plurality of references standards;
creating a chimeric reference standard, whereby the creating comprises inserting the selected portions of the first reference standard and the second reference standard into respective portions of a hash table maintained in one or more memory devices; and
using the hash table and the genome of the subject in order to perform one or more of a mapping, aligning, and/or variant calling procedure.
2 . The method of claim 1 , wherein the first reference standard is a linear reference standard, and wherein the second reference standard is a cultural reference standard.
3 . The method of claim 1 , wherein the first reference standard is a linear reference standard, and wherein the second reference standard is a different linear reference standard.
4 . The method of claim 1 , wherein the first reference standard is cultural reference standard, and wherein the second reference standard is a different cultural reference standard.
5 . The method of claim 1 , wherein at least one of the first reference standard and the second reference standard is a previously generated chimeric reference standard.
6 . A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations for constructing and using a chimeric reference standard specific to an individual subject for improving accuracy of analysis of a plurality of portions of a genome of the subject, the operations comprising:
comparing a first portion of the genome of the subject to each reference standard of a plurality of reference standards stored in a database, wherein said genome is based on a biological sample obtained from said subject;
selecting a portion of a first reference standard that satisfies a predetermined matching threshold to a first portion of the genome of the subject based on the comparison of the first portion to the plurality of reference standards;
comparing a second portion of the genome of the subject to each reference standard of the plurality of reference standards;
selecting a portion of a second reference standard that satisfies a predetermined matching threshold to the second portion of the genome of the subject based on the comparison of the second portion to the plurality of references standards;
creating a chimeric reference standard, whereby the creating comprises inserting the selected portions of the first reference standard and the second reference standard into respective portions of a hash table maintained in one or more memory devices; and
using the hash table and the genome of the subject in order to perform one or more of a mapping, aligning, and/or variant calling procedure.
7 . The computer-readable medium of claim 6 , wherein the first reference standard is a linear reference standard, and wherein the second reference standard is a cultural reference standard.
8 . The computer-readable medium of claim 6 , wherein the first reference standard is a linear reference standard, and wherein the second reference standard is a different linear reference standard.
9 . The computer-readable medium of claim 6 , wherein the first reference standard is cultural reference standard, and wherein the second reference standard is a different cultural reference standard.
10 . The computer-readable medium of claim 6 , wherein at least one of the first reference standard and the second reference standard is a previously generated chimeric reference standard.
11 . A system for constructing and using a chimeric reference standard specific to an individual for improving accuracy of nucleic acid sequence analysis of nucleic acid sequences from the subject, the system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations, the operations comprising:
comparing a first portion of a genome of the subject to each reference standard of a plurality of reference standards stored in a database, wherein the genome of the subject is based on a biological sample obtained from the subject;
selecting a portion of a first reference standard that satisfies a predetermined matching threshold to a first portion of the genome of the subject based on the comparison of the first portion to the plurality of reference standards;
comparing a second portion of the genome of the subject to each reference standard of the plurality of reference standards;
selecting a portion of a second reference standard that satisfies a predetermined matching threshold to the second portion of the genome of the subject based on the comparison of the second portion to the plurality of references standards;
creating a chimeric reference standard, whereby the creating comprises inserting the selected portions of the first reference standard and the second reference standard into respective portions of a hash table maintained in one or more memory devices; and
using the hash table and the genome of the subject in order to perform one or more of a mapping, aligning, and/or variant calling procedure.
12 . The system of claim 11 , wherein the first reference standard is a linear reference standard, and wherein the second reference standard is a cultural reference standard.
13 . The system of claim 11 , wherein the first reference standard is a linear reference standard, and wherein the second reference standard is a different linear reference standard.
14 . The system of claim 11 , wherein the first reference standard is cultural reference standard, and wherein the second reference standard is a different cultural reference standard.
15 . The system of claim 11 , wherein at least one of the first reference standard and the second reference standard is a previously generated chimeric reference standard.