MOLECULAR EVIDENCE PLATFORM FOR AUDITABLE, CONTINUOUS OPTIMIZATION OF VARIANT INTERPRETATION IN GENETIC AND GENOMIC TESTING AND ANALYSIS
Disclosed herein are system, method, and computer program product embodiments for optimizing the determination of a phenotypic impact of a molecular variant identified in molecular tests, samples, or reports of subjects by way of regularly incorporating, updating, monitoring, validating, selecting, and auditing the best-performing evidence models for the interpretation of molecular variants across a plurality of evidence classes.
1 . A computer implemented method, the method comprising:
recording an evidence model comprising evidence data, wherein the evidence data describes a predicted phenotypic impact of a molecular variant for a target entity;
evaluating validation performance data for the evidence model based on production data;
generating a hash value of supporting data for the evidence model, wherein the supporting data comprises the evidence data, and the generation of the hash value enables prospective evaluation of the evidence data in response to receiving test data for the evidence model;
in response to receiving the test data for the evidence model, evaluating test performance data for the evidence model based on the evidence data and the test data;
ranking the evidence model in a set of evidence models for the target entity based on the validation performance data or the test performance data; and
in response to a query for the predicted phenotypic impact of the molecular variant for the target entity from a variant interpretation terminal, providing the predicted phenotypic impact using a best-performing evidence model for the target entity based on the ranking.
2 . The method of claim 1 , wherein the target entity comprises a functional element, molecule, or molecular variant, and a phenotype of interest.
3 . The method of claim 1 or 2 , the recording further comprising:
generating the evidence model based on the production data using a machine learning technique.
4 . The method of any one of claims 1 to 3 , the recording further comprising:
importing the evidence model or the evidence data.
5 . The method of any one claims 1 to 4 , further comprising:
generating the supporting data from at least one of the evidence data, the production data, the test data, the validation performance data, or the test performance data.
6 . The method of any one of claims 1 to 5 , wherein the generation of the hash value enables evaluation of content of the supporting data and a time of creation of the supporting data.
7 . The method of any one of claims 1 to 6 , further comprising:
receiving the production data from a clinical knowledgebase.
8 . The method of any one of claims 1 to 7 , the evaluating the validation performance data further comprising:
calculating, using the evidence model and a model validation technique, a phenotype impact score for the molecular variant of the target entity in the production data; and
generating the validation performance data based on the phenotype impact score using a performance metric of interest.
9 . The method of any one of claims 1 to 8 , the evaluating the test performance data further comprising:
calculating, using the evidence model and a model validation technique, a phenotype impact score for the molecular variant of the target entity in the test data; and
generating the test performance data based on the phenotype impact score using a performance metric of interest.
10 . The method of any one of claims 1 to 9 , further comprising:
storing the hash value of the supporting data in a database, wherein the database associates the hash value with the supporting data.
11 . The method of any one of claims 1 to 10 , further comprising:
inserting the hash value into a distributed data structure.
12 . The method of claim 11 , further comprising:
providing an audit record to a variant interpretation terminal, wherein the audit record references an entry for the supporting data in the distributed data structure, and the audit record enables the variant interpretation terminal to audit content of the supporting data and a time of creation of the supporting data.
13 . The method of claim 11 or claim 12 , wherein the distributed data structure is a blockchain data structure.
14 . The method of any one of claims 11 to 13 , wherein the distributed data structure is a distributed feed.
15 . A variant interpretation terminal system, comprising:
a memory; and
at least one processor coupled to the memory and configured to:
send a support query to a variant interpretation system for supporting data for an evidence model meeting a set of performance metrics for a target entity;
receive the supporting data and an associated auditing record for the supporting data from the variant interpretation system;
send an audit query to a distributed data structure, wherein the audit query comprises the auditing record for the supporting data;
receive a certificate of validation for the auditing record from the distributed database in response to the sending of the audit query; and
determining a data state of the supporting data at a point in time based on the auditing record.
16 . The system of claim 15 , wherein the at least one processor is configured to:
compute a hash value of the supporting data for the evidence model; and
determine the hash value matches a hash value in the auditing record for the supporting data for the evidence model.
17 . The system of claim 15 or claim 16 , wherein the target entity comprises a functional element, molecule, or molecular variant, and a phenotype of interest.