TECHNIQUES FOR BIAS CORRECTION IN SEQUENCE DATA
Described herein are various methods of collecting and processing of tumor and/or healthy tissue samples to extract nucleic acid and perform nucleic acid sequencing. Also described herein are various methods of processing nucleic acid sequencing data to remove bias from the nucleic acid sequencing data. Also described herein are various methods of evaluating the quality of nucleic acid sequence information. The identity and/or integrity of nucleic acid sequence data is evaluated prior to using the sequence information for subsequent analysis (for example for diagnostic, prognostic, or clinical purposes). The methods enable a subject, doctor, or user to characterize or classify various types of cancer precisely, and thereby determine a therapy or combination of therapies that may be effective to treat a cancer in a subject based on the precise characterization.
1 . A method, comprising:
obtaining a first biological sample of a first tumor, the first biological sample previously obtained from a subject having, suspected of having or at risk of having cancer;
extracting RNA from the first biological sample of the first tumor to obtain extracted RNA;
enriching the extracted RNA for coding RNA to obtain enriched RNA;
sequencing, using at least one sequencing platform, the enriched RNA to obtain RNA expression data comprising at least 5 kilobases (kb);
using at least one computer hardware processor to perform:
obtaining the RNA expression data using the at least one sequencing platform;
converting the RNA expression data to gene expression data;
determining bias-corrected gene expression data from the gene expression data at least in part by removing, from the gene expression data, expression data for at least one gene that introduces bias in the gene expression data; and
identifying a cancer treatment for the subject using the bias-corrected gene expression data.
2 . The method of claim 1 , further comprising:
administering the identified cancer treatment to the subject.
3 . The method claim 1 , wherein enriching the RNA for coding RNA comprises performing polyA enrichment.
4 . The method claim 1 , wherein the at least one gene that introduces bias in the gene expression data comprises:
a gene having an average transcript length that is higher or lower than an average length of transcripts in the gene expression data;
a gene having at least a threshold variation in average transcript expression level based on transcript expression levels in reference samples;
and/or a gene that has a polyA tail that is at least a threshold amount smaller in length compared to an average length of polyA tails of genes from: the first biological sample from which the RNA expression data was obtained and/or a reference sample.
5 . The method of claim 1 , wherein the at least one gene that introduces bias in the gene expression data belongs to a family of genes selected from the group consisting of: histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, B-cell receptor-encoding genes, and T cell receptor-encoding genes.
6 . The method of claim 5 , wherein the at least one gene comprises at least one histone-encoding gene selected from the group consisting of: HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AL, HIST1H2AM, HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST1H2BG, HIST1H2BH, HIST1H2BI, HIST1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIST1H3A, HIST1H3B, HIST1H3C, HIST1H3D, HIST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, HIST1H4A, HIST1H4B, HIST1H4C, HIST1H4D, HIST1H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST2H2AA3, HIST2H2AA4, HIST2H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, and HIST4H4.
7 . The method of claim 5 , wherein the at least one gene comprises at least one mitochondrial gene selected from the group consisting of: MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, MT-TM, MT-TN, MT-TP, MT-TQ, MT-TR, MT-TS1, MT-TS2, MT-TT, MT-TV, MT-TW, MT-TY, MTRNR2L1, MTRNR2L10, MTRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, and MTRNR2L8.
8 . The method of claim 1 , wherein determining the bias-corrected gene expression data further comprises:
after removing the expression data for the at least one gene that introduces bias in the gene expression data, renormalizing the gene expression data.
9 . The method of claim 1 , wherein converting the RNA expression data to gene expression data comprises:
removing non-coding transcripts from the RNA expression data to obtain filtered RNA expression data; and
after removing the non-coding transcripts, normalizing the filtered RNA expression data to obtain gene expression data in transcripts per million (TPM).
10 . The method of claim 1 , wherein removing the non-coding transcripts from the RNA expression data comprises removing non-coding transcripts that belong to groups selected from the list consisting of: pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, unprocessed pseudogenes, transcribed unitary pseudogenes, IG C pseudogenes, IG J pseudogenes, IG V pseudogenes, transcribed unprocessed pseudogenes, translated unprocessed pseudogene TR J pseudogenes, TR V pseudogenes, snRNA, snoRNA, miRNA, ribozymes, rRNA, Mt tRNA, Mt rRNA, scaRNA, retained introns, sense intronics, sense overlapping RNA, nonsense mediated decay RNA, non stop decay RNA, antisense RNA, lincRNA, macro lncRNA, processed transcripts, 3prime overlapping ncrna, sRNA, misc RNA, vault RNA, and TEC.
11 . The method of claim 1 , further comprising:
prior to performing the removal of the non-coding transcripts,
aligning the RNA expression data to a reference; and
annotating the RNA expression data.
12 . The method of claim 1 , wherein the RNA expression data comprises at least 25 million paired-end reads.
13 . The method of claim 12 , wherein the RNA expression data comprises at least 50 million paired-end reads, with an average read length of at least 100 bp.
14 . The method of claim 1 , wherein identifying the cancer treatment for the subject using the bias-corrected gene expression data comprises:
determining, using the bias-corrected gene expression data, a plurality of gene group expression levels, the plurality of gene group expression levels comprising a gene group expression level for each gene group in a set of gene groups, wherein the set of gene groups comprises at least one gene group associated with cancer malignancy, and at least one gene group associated with cancer microenvironment; and
identifying the cancer treatment using the determined gene group expression levels.
15 . The method of claim 14 , wherein the cancer treatment is selected from the group consisting of a radiation therapy, a surgical therapy, a chemotherapy, and an immunotherapy.
16 . The method of claim 1 , further comprising obtaining a second biological sample of a second tumor, the second biological sample previously obtained from the subject.
17 . The method of claim 16 , further comprising:
combining the first biological sample and the second biological sample to form a combined tumor sample,
wherein extracting the RNA comprises extracting the RNA from the combined tumor sample.
18 . The method of claim 16 , further comprising:
extracting RNA from the second biological sample; and
combining the RNA extracted from the second biological sample with the RNA extracted from the first biological sample to form combined extracted RNA,
wherein enriching the RNA for coding RNA comprises enriching the combined extracted RNA for coding RNA.
19 . The method of claim 1 , wherein the extracted RNA comprises at least 1 μg of RNA upon RNA extraction.
20 . The method of claim 19 , wherein the extracted RNA is at least 1000-6000 ng in total mass, and has a purity corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 2.0.
21 . The method of claim 1 , further comprising performing quality control assessment on the RNA expression data at least in part by:
obtaining asserted information indicating an asserted source and/or an asserted integrity of the RNA expression data;
processing the RNA expression data to obtain determined information indicating a determined source and/or a determined integrity of the RNA expression data; and
determining whether the determined information matches the asserted information.
22 . The method of claim 1 , wherein processing the RNA expression data comprises processing the RNA expression RNA to determine: a tissue type of the first biological sample; a tumor type of the first biological sample; and/or guanine (G) and/or cytosine (C) percentage (%).
23 . A system for identifying a cancer treatment for a subject having, suspected having, or at risk of having cancer, the system comprising:
at least one sequencing platform configured to generate gene expression data from enriched RNA obtained from a first biological sample previously obtained from the subject, wherein the enriched RNA was obtained by: (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA, wherein the RNA expression data comprises at least 5 kilobases (kb);
at least one computer hardware processor; and
at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining the RNA expression data using the at least one sequencing platform;
converting the RNA expression data to gene expression data;
determining bias-corrected gene expression data from the gene expression data at least in part by removing, from the gene expression data, expression data for at least one gene that introduces bias in the gene expression data; and
identifying a cancer treatment for the subject using the bias-corrected gene expression data.
24 . A system for identifying a cancer treatment for a subject having, suspected having, or at risk of having cancer, the system comprising:
at least one computer hardware processor; and
at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining RNA expression data from at least one sequencing platform, the RNA expression data comprising at least 5 kilobases (5 kb), wherein the RNA expression data was obtained, from a first biological sample of a first tumor previously obtained from the subject, at least in part by: (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA;
converting the RNA expression data to gene expression data;
determining bias-corrected gene expression data from the gene expression data at least in part by removing, from the gene expression data, expression data for at least one gene that introduces bias in the gene expression data; and
identifying a cancer treatment for the subject using the bias-corrected gene expression data.
25 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining RNA expression data from at least one sequencing platform, the RNA expression data comprising at least 5 kilobases (5 kb), wherein the RNA expression data was obtained, from a first biological sample of a first tumor previously obtained from a subject having, suspected of having or at risk of having cancer, at least in part by: (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA;
converting the RNA expression data to gene expression data;
determining bias-corrected gene expression data from the gene expression data at least in part by removing, from the gene expression data, expression data for at least one gene that introduces bias in the gene expression data; and
identifying a cancer treatment for the subject using the bias-corrected gene expression data.