Systems and methods for automating RNA expression calls in a cancer prediction pipeline
Systems and methods are provided for performing quality control analysis. The method obtains, in electronic form, a batch dataset comprising, for each respective sample in a batch of samples, a corresponding plurality of sequence reads derived from the respective sample by targeted or whole transcriptome RNA sequencing and corresponding metadata for the respective sample. The method determines for the batch dataset a cohort-matched reference batch, where the cohort-matched reference batch is balanced for tissue site, tumor purity, cancer type, sequencer identity, or date sequenced. The method performs one or more global batch quality control tests on the batch dataset using at least the cohort-matched reference batch. The method removes respective samples from the batch dataset that fail any one of the one or more global batch quality control tests or flagging for manual inspection respective samples that fail any one of the one or more global batch quality control tests.
1 . A method of validating a batch dataset through identification and removal of outliers from the batch dataset, the method comprising:
at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:
a) obtaining, in electronic form, the batch dataset comprising, for each respective biological test sample in a plurality of biological test samples,
a corresponding expression profile comprising a corresponding gene expression value for each respective gene in a first set of genes, wherein obtaining the batch dataset comprises:
obtaining, in electronic form, a corresponding plurality of at least 10,000 sequence reads derived from the respective biological test sample by RNA sequencing; and
determining, from the corresponding plurality of at least 10,000 sequence reads, the corresponding gene expression value for each respective gene in the first set of genes, and
a corresponding set of characteristic values comprising a respective characteristic value for each respective characteristic in a first set of characteristics about the respective biological test sample;
b) determining, for the batch dataset, a reference cohort by a procedure comprising:
(i) selecting a subset of reference samples from a plurality of reference samples, wherein, each respective reference sample in the reference cohort is associated with a corresponding set of characteristic values comprising a corresponding characteristic value for each respective characteristic in a second set of characteristics about the respective reference sample, wherein the second set of characteristics comprises a third set of one or more characteristics that are also present in the first set of characteristics,
(ii) determining that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for a first characteristic in the third set of one or more characteristics, fails to be within a threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic, wherein two respective biological test samples in the plurality of biological test samples have different respective values for the first characteristic and wherein the threshold percentage is five percent,
(iii) determining a subset of biological test samples that cause the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic to fail to be within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic,
(iv) removing the subset of biological test samples from the plurality of biological test samples, and
(v) determining that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic, is within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic;
c) obtaining, in electronic form, for each respective reference sample in the reference cohort,
a corresponding expression profile comprising a corresponding gene expression value for each respective gene in the first set of genes;
d) determining (i) a first distribution for the batch dataset, with the subset of biological test samples removed from the plurality of biological test samples, using the corresponding expression profile for each respective biological test sample in the plurality of biological test samples and (ii) a second distribution for the reference cohort using the corresponding expression profile for each respective reference sample in the subset of reference samples; and
e) validating the batch dataset on the basis that a comparison of the first distribution and the second distribution satisfies a comparison criterion.
2 . The method of claim 1 , wherein the first set of genes comprises at least 10 genes.
3 . The method of claim 1 , wherein the first set of genes comprises at least 10,000 genes.
4 . The method of claim 1 , wherein the first characteristic comprises tissue site, tumor purity, cancer type, sequencer identity, or sequencing date.
5 . The method of claim 1 , wherein the first characteristics comprises a nucleic acid extraction method, a cDNA library preparation method, an RNA sequencing method, a type of reagent used, or a type of equipment used.
6 . The method of claim 1 , wherein the reference cohort comprises at least 100 reference samples.
7 . The method of claim 1 , wherein the reference cohort comprises at least 1000 reference samples.
8 . The method of claim 1 , wherein the threshold percentage is 2.5%.
9 . The method of claim 1 , wherein a combined dataset is formed by a dimension reduction technique that embeds, for each respective biological test sample in the plurality of biological test samples and each respective reference sample in the reference cohort, the corresponding expression profile into a two-dimensional representation.
10 . The method of claim 1 , the method further comprising obtaining the at least 10,000 sequence reads by whole transcriptome RNA sequencing.
11 . The method of claim 1 , the method further comprising performing the first methodology by targeted panel RNA sequencing using a plurality of probes wherein each probe in the plurality of probes uniquely targets a respective portion of a reference transcriptome, and each sequence read in the corresponding plurality of sequence reads corresponds to at least one probe in the plurality of probes.
12 . The method of claim 10 , wherein the whole transcriptome sequencing comprises next-generation sequencing.
13 . The method of claim 1 , further comprising:
performing, for each respective biological test sample in the plurality of test of samples, one or more single sample quality control tests on the respective biological test sample; and
removing respective biological test samples from the plurality of biological test samples that fail any one of the one or more single sample quality control tests or flagging for manual inspection respective biological test samples that fail any one of the one or more single sample quality control tests.
14 . The method of claim 1 , wherein the first characteristics comprises a cancer type.
15 . The method of claim 1 , wherein the first characteristic comprises a tissue site and a cancer type.
16 . The method of claim 1 , wherein the plurality of at least 10,000 sequence reads is at least 100,000 sequence reads.
17 . The method of claim 1 , wherein the plurality of at least 10,000 sequence reads is at least 1,000,000 sequence reads.
18 . The method of claim 1 , wherein each biological test sample in the plurality of biological test samples is processed through an RNA expression pipeline and the method further comprises:
rerunning the subset of biological test samples in the RNA expression pipeline.
19 . A method of validating a change in an RNA expression pipeline through identification and removal of outliers from a batch dataset produced by the RNA expression pipeline, the method comprising:
at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:
a) obtaining, in electronic form, a batch dataset comprising, for each respective biological test sample in a plurality of biological test samples,
a corresponding expression profile prepared using a first methodology, the corresponding expression profile comprising a corresponding gene expression value for each respective gene in a first set of genes, wherein obtaining the batch dataset comprises:
obtaining, in electronic form, a corresponding plurality of at least 10,000 sequence reads derived from the respective biological test sample by RNA sequencing; and
determining, from the corresponding plurality of at least 10,000 sequence reads, the corresponding gene expression value for each respective gene in the first set of genes, and
a corresponding set of characteristic values comprising a respective characteristic value for each respective characteristic in a first set of characteristics about the respective biological test sample;
b) determining, for the batch dataset, a reference cohort by a procedure comprising:
(i) selecting a subset of reference samples from a plurality of reference samples, wherein, each respective reference sample in the reference cohort is associated with a corresponding set of characteristic values comprising a corresponding characteristic value for each respective characteristic in a second set of characteristics about the respective reference sample, wherein the second set of characteristics comprises a third set of one or more characteristics that are also present in the first set of characteristics,
(ii) determining that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for a first characteristic in the third set of one or more characteristics, fails to be within a threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic, wherein two respective biological test samples in the plurality of biological test samples have different respective values for the first characteristic and wherein the threshold percentage is five percent,
(iii) determining a subset of biological test samples that cause the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic to fail to be within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic,
(iv) removing the subset of biological test samples from the plurality of biological test samples, and
(v) determining that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic, is within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic;
c) obtaining, in electronic form, for each respective reference sample in the reference cohort,
a corresponding expression profile comprising a corresponding gene expression value for each respective gene in the first set of genes prepared using a second methodology that is different than the first methodology;
d) determining (i) a first distribution for the batch dataset, with the subset of biological test samples removed from the plurality of biological test samples, using the corresponding expression profile for each respective biological test sample in the plurality of biological test samples and (ii) a second distribution for the reference cohort using the corresponding expression profile for each respective reference sample in the subset of reference samples; and
e) validating the change in the RNA expression pipeline on the basis that a comparison of the first distribution and the second distribution satisfies a comparison criterion.
20 . The method of claim 19 , wherein the plurality of biological test samples comprises at least 10 biological test samples.
21 . The method of claim 19 , wherein the first set of genes comprises at least 10 genes.
22 . The method of claim 19 , wherein the first set of genes comprises at least 10,000 genes.
23 . The method of claim 19 , wherein the first characteristic comprises tissue site, tumor purity, cancer type, sequencer identity, or sequencing date.
24 . The method of claim 19 , wherein the first characteristic comprises a nucleic acid extraction method, a cDNA library preparation method, an RNA sequencing method, a type of reagent used, or a type of equipment used.
25 . The method of claim 19 , wherein the subset of cohort-matched reference samples comprises at least 100 reference samples.
26 . The method of claim 19 , wherein the subset of cohort-matched reference samples comprises at least 1000 reference samples.
27 . The method of claim 19 , wherein the threshold percentage is 2.5%.
28 . The method of claim 19 , wherein a combined dataset is formed by a dimension reduction technique that embeds, for each respective biological test sample in the plurality of biological test samples and each respective reference sample in the reference cohort, the corresponding expression profile into a two-dimensional representation.
29 . The method of claim 19 , the method further comprising obtaining the at least 10,000 sequence reads by a whole transcriptome RNA sequencing.
30 . The method of claim 19 , the method further comprising obtaining the at least 10,000 sequence reads by targeted panel RNA sequencing using a plurality of probes, wherein each probe in the plurality of probes uniquely targets a respective portion of a reference transcriptome, and each sequence read in the corresponding plurality of sequence reads corresponds to at least one probe in the plurality of probes.
31 . The method of claim 29 , wherein the whole transcriptome sequencing comprises next-generation sequencing.
32 . The method of claim 19 , further comprising:
performing, for each respective biological test sample in the batch of test of samples, one or more single sample quality control tests on the respective biological test sample; and
removing respective biological test samples from the plurality of biological test samples that fail any one of the one or more single sample quality control tests or flagging for manual inspection respective biological test samples that fail any one of the one or more single sample quality control tests.
33 . The method of claim 19 , wherein the first characteristic comprises a cancer type.
34 . The method of claim 19 , wherein the first characteristic comprises a tissue site and a cancer type.
35 . The method of claim 19 , wherein the plurality of at least 10,000 sequence reads is at least 100,000 sequence reads.
36 . The method of claim 19 , wherein the plurality of at least 10,000 sequence reads is at least 1,000,000 sequence reads.
37 . The method of claim 19 , wherein each biological test sample in the plurality of biological test samples is processed through the RNA expression pipeline and the method further comprises:
rerunning the subset of biological test samples in the RNA expression pipeline.
38 . A method of adding RNA expression data to a reference database comprising a plurality of reference database samples, the method comprising:
at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:
a) obtaining, in electronic form, an expression dataset comprising, for each respective biological test sample in a plurality of biological test samples,
a corresponding expression profile comprising a corresponding gene expression value for each respective gene in a first set of genes, wherein generating the profile comprises:
obtaining, in electronic form, a corresponding plurality of at least 10,000 sequence reads derived from the respective biological test sample by RNA sequencing; and
determining, from the corresponding plurality of at least 10,000 sequence reads, the corresponding gene expression value for each respective gene in the first set of genes, and
a corresponding set of characteristic values comprising a respective characteristic value for each respective characteristic in a first set of characteristics about the respective biological test sample;
b) determining for the expression dataset, a reference cohort by a procedure comprising:
(i) selecting a subset of reference samples from a plurality of reference samples, wherein, each respective reference sample in the reference cohort is associated with a corresponding set of characteristic values comprising a corresponding characteristic value for each respective characteristic in a second set of characteristics about the respective reference sample, wherein the second set of characteristics comprises a third set of one or more characteristics that are also present in the first set of characteristics,
(ii) determining that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for a first characteristic in the third set of one or more characteristics, fails to be within a threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic, wherein two respective biological test samples in the plurality of biological test samples have different respective values for the first characteristic and wherein the threshold percentage is five percent,
(iii) determining a subset of biological test samples that cause the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic to fail to be within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic,
(iv) removing the subset of biological test samples from the plurality of biological test samples, and
(v) determining that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic, is within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic;
c) obtaining, in electronic form, for each respective reference sample in the reference cohort,
a corresponding profile comprising a corresponding gene expression value for each respective gene in the first set of genes,
each profile corresponding to a reference sample in the subset of reference samples is from the reference database, and
d) determining (i) a first distribution for the expression dataset, with the subset of biological test samples removed from the plurality of biological test samples, using the corresponding expression profile for each respective biological test sample in the plurality of biological test samples and (ii) a second distribution for the reference cohort using the corresponding expression profile for each respective reference sample in the subset of reference samples; and
e) adding the expression dataset to the reference database when a comparison of the first distribution and the second distribution satisfies a comparison criterion, or
when the comparison does not satisfy the comparison criterion:
determining a set of conversion factors for standardizing the expression profiles in the expression dataset against expression profiles in the reference database,
standardizing the expression profiles in the expression dataset using the set of conversion factors, thereby obtaining a standardized expression dataset, and
adding the standardized expression dataset to the reference database.
39 . The method of claim 38 , wherein the plurality of biological test samples comprises at least 10 biological test samples.
40 . The method of claim 38 , wherein the first set of genes comprises at least 10 genes.
41 . The method of claim 38 , wherein the first set of genes comprises at least 10,000 genes.
42 . The method of claim 38 , wherein the first characteristic is tissue site, tumor purity, cancer type, sequencer identity, or sequencing date.
43 . The method of claim 38 , wherein the first characteristic is a nucleic acid extraction method, a cDNA library preparation method, an RNA sequencing method, a type of reagent used, or a type of equipment used.
44 . The method of claim 38 , wherein the subset of reference samples comprises at least 100 reference samples.
45 . The method of claim 38 , wherein the subset of reference samples comprises at least 1000 reference samples.
46 . The method of claim 38 , wherein the threshold percentage is 2.5%.
47 . The method of claim 38 , wherein a combined dataset is formed by a dimension reduction technique that embeds, for each respective biological test sample in the plurality of biological test samples and each respective reference sample in the subset of reference samples, the corresponding expression profile into a two-dimensional representation.
48 . The method of claim 38 , the method further comprising obtaining the at least 10,000 sequence reads by a whole transcriptome RNA sequencing.
49 . The method of claim 38 , the method further comprising obtaining the at least 10,000 sequence reads by targeted panel RNA sequencing using a plurality of probes, wherein each probe in the plurality of probes uniquely targets a respective portion of a reference transcriptome, and each sequence read in the corresponding plurality of sequence reads corresponds to at least one probe in the plurality of probes.
50 . The method of claim 48 , wherein the whole transcriptome sequencing comprises next-generation sequencing.
51 . The method of claim 38 , further comprising:
performing, for each respective biological test sample in the plurality of biological test samples, one or more single sample quality control tests on the respective biological test sample; and
removing respective biological test samples from the plurality of biological test samples that fail any one of the one or more single sample quality control tests or flagging for manual inspection respective biological test samples that fail any one of the one or more single sample quality control tests.
52 . The method of claim 38 , wherein the first characteristic comprises a cancer type.
53 . The method of claim 38 , wherein the first characteristic comprises a tissue site and a cancer type.
54 . The method of claim 38 , wherein the plurality of at least 10,000 sequence reads is at least 100,000 sequence reads.
55 . The method of claim 38 , wherein the plurality of at least 10,000 sequence reads is at least 1,000,000 sequence reads.
56 . The method of claim 38 , wherein each biological test sample in the plurality of biological test samples is processed through an RNA expression pipeline and the method further comprises:
rerunning the subset of biological test samples in the RNA expression pipeline.
57 . A method of processing a batch dataset, the method comprising:
a) obtaining using a computer system, in electronic form, the batch dataset comprising, for each respective biological test sample in a plurality of biological test samples,
a corresponding expression profile comprising a corresponding gene expression value for each respective gene in a first set of genes, wherein obtaining the batch dataset comprises:
obtaining, in electronic form, a corresponding plurality of at least 10,000 sequence reads derived from the respective biological test sample by RNA sequencing; and
determining, from the corresponding plurality of at least 10,000 sequence reads, the corresponding gene expression value for each respective gene in the first set of genes, and
a corresponding set of characteristic values comprising a respective characteristic value for each respective characteristic in a first set of characteristics about the respective biological test sample;
b) determining, for the batch dataset, a reference cohort by a procedure comprising:
(i) selecting, using a computer system, a subset of reference samples from a plurality of reference samples, wherein, each respective reference sample in the reference cohort is associated with a corresponding set of characteristic values comprising a corresponding characteristic value for each respective characteristic in a second set of characteristics about the respective reference sample, wherein the second set of characteristics comprises a third set of one or more characteristics that are also present in the first set of characteristics,
(ii) determining, using a computer system, that the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for a first characteristic in the third set of one or more characteristics, fails to be within a threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic, wherein two respective biological test samples in the plurality of biological test samples have different respective values for the first characteristic and wherein the threshold percentage is five percent,
(iii) determining, using a computer system, a subset of biological test samples that cause the percentage of respective biological test samples, in the plurality of biological test samples, having a respective value for the respective characteristic to fail to be within the threshold percentage of the percentage of respective reference samples, in the subset of reference samples, having the same respective value for the first characteristic,
c) obtaining using a computer system, in electronic form, for each respective reference sample in the reference cohort,
a corresponding expression profile comprising a corresponding gene expression value for each respective gene in the first set of genes;
d) determining, using a computer system, (i) a first distribution for the batch dataset, with the subset of biological test samples removed from the plurality of biological test samples, using the corresponding expression profile for each respective biological test sample in the plurality of biological test samples and (ii) a second distribution for the reference cohort using the corresponding expression profile for each respective reference sample in the subset of reference samples; and
e) validating, using a computer system, the batch dataset on the basis that a comparison of the first distribution and the second distribution satisfies a comparison criterion,
wherein each biological test sample in the plurality of biological test samples is processed through an RNA expression pipeline and the method further comprises:
rerunning the subset of biological test samples in the RNA expression pipeline.