IP Library Patent Application 19185003
Patent Application
App. No. 19/185,003

METHODS AND SYSTEMS FOR MODELING INFERENCE OF MUTATION IMPACT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/185,003
Abstract

Provided herein are systems, methods, computer-readable media, and techniques for analyzing variants, including: (A) obtaining assay data corresponding to an assay for a subject; (B) identifying a plurality of variants in said assay data, wherein said plurality of variants comprise at least about one thousand variants; and (C) selecting a subset of variants from said plurality of variants, wherein said subset of variants comprises a variant most relevant to diagnosing said subject.

Claims (36)

1 .- 22 . (canceled)

23 . A method for analyzing variants, comprising:

(a) performing whole genome sequencing or whole exome sequencing of a biological sample obtained or derived from a subject, thereby generating sequencing data;

(b) processing the sequencing data to determine a plurality of variants in said sequencing data, wherein the computer processing comprises formatting the sequencing data as a table-separated values (TSV) computer data file;

(c) transforming, by a computer, the TSV computer data file to a genomic ordered relational (GOR) computer data format, wherein the transforming comprises relationally ordering one or more columns of the TSV computer data file based on one or more parameters relating to at least a portion of the plurality of variants comprising: information relating to chromosomes, information relating to variant position, information relating to reference alleles, information relating to called alleles, or any combination thereof, thereby generating a GOR case variant computer data file;

(d) storing the GOR case variant computer data file in a computer memory;

(e) automatically annotating, by a computer at least a portion of the plurality of variants of the GOR case variant computer data file with one or more features comprising (i) variant-level features associated with the variant independently of characteristics of the subject and (ii) case-;

(f) transforming, by a computer, the stored GOR case variant computer data file to a list of case records formatted according to an Avro multi-level computer schema;

(g) transforming the list of case records to a serialized Avro file;

(h) prioritizing, utilizing a random forest or a neural network machine learning (ML) model, at least a subset of the annotated variants of the serialized Avro file, wherein the ML model is trained to perform operations comprising:

(1) receiving as input the serialized Avro file,

(2) determining a relevance with respect to a disease;

based at least in part on one or more features of the annotated variants, and

(3) prioritizing the at least portion of the plurality of variants based on the determined relevance with respect to the disease, thereby generating prioritized variants having a rank score in the serialized Avro file,

(i) transforming the prioritized variants from the Avro multi-level computer schema of the Avro file to the GOR computer data format, thereby ordering in the GOR format the prioritized variants based on the rank score; and

(j) selecting, using the one or more computer processors, a subset of the prioritized variants ordered in the GOR computer data format, wherein the selected subset of prioritized variants comprises a diagnostic variant most relevant to indicating a disease of the subject, wherein the selecting of the diagnostic variant has an accuracy rate of at least 90%.

24 . The method of claim 23 , wherein the variant-level features comprise variant population frequency or variant effects.

25 . The method of claim 23 , wherein the case-level features comprise variant segregation data or inheritance data and an association between the variant and one or more phenotypes associated with the subject.

26 . The method of claim 25 , wherein the phenotypes comprise disease phenotypes.

27 . The method of claim 23 , wherein the ML model comprises the random forest ML model.

28 . The method of claim 23 , wherein the ML model comprises the neural network ML model.

29 .- 32 . (canceled)

33 . The method of claim 23 , wherein the plurality of variants comprises at least three thousand variants.

34 . The method of claim 23 , wherein the plurality of variants comprises at least ten thousand variants.

35 . The method of claim 23 , wherein the disease is a Mendelian disease.

36 . The method of claim 23 , wherein the disease is a suspected disease.

37 . The method of claim 23 , further comprising generating a training dataset for training the ML model.

38 . The method of claim 37 , further comprising querying, by the computer, at least a portion of the annotated variants of the GOR case variant computer data file.

39 . The method of claim 38 , further comprising receiving, by the computer in response to the querying, at least a subset of the portion of the annotated variants of the GOR case variant computer data file, thereby generating the training dataset.

40 . The method of claim 39 , wherein the at least subset of the portion of the annotated variants of the GOR case variant computer data file comprise annotations relating to clinical confirmation of the relevance of the at least subset of the portion of the annotated variants with respect to the disease.

41 . The method of claim 40 , further comprising generating a Disease Phenotype Resource (DiPR) case table comprising the clinical confirmation of the relevance of the at least subset of the portion of the annotated variants with respect to the disease.

42 . The method of claim 41 , further comprising querying the DiPR case table to generate the training dataset, wherein the training dataset comprises at least a subset of annotated variants of the DiPR case table.

43 . The method of claim 41 , further comprising updating the DiPR case table.

44 . The method of claim 43 , further comprising updating the DiPR case table with modified annotations of at least a subset of the annotated variants.

45 . The method of claim 43 , further comprising updating the DiPR case table by adding one or more annotated variants comprising annotations relating to the clinical confirmation of the relevance of the one or more annotated variants with respect to the disease.

46 . The method of claim 40 , further comprising utilizing the training dataset to retrain the ML model.