IP Library Patent Application 14470628
Patent Application
App. No. 14/470,628

Identifying Possible Disease-Causing Genetic Variants by Machine Learning Classification

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/470,628
Abstract

The techniques described herein relate identification of disease-causing genetic variant by machine learning classification. The techniques may include receiving a training dataset of predetermined variants associated with disease. A hyperplane is identified having a maximum margin between points of the dataset. Patient input data is received including an observed variant of a gene. Features of the observed variant are selected, and a score is determined The score is determined using Support Vector Machine algorithms based on an observation of a novel non-linear relationship with the selected features of the observed variant. The observed variant may be classified based on the score indicating a distance of the observed variant from the identified hyperplane.

Claims (80)

1 . A method for identifying a possible disease-causing genetic variant by machine learning classification, comprising:

receiving a training dataset of predetermined variants associated with disease;

identifying a hyperplane having a maximum margin between points of the training dataset;

receiving patient input data comprising an observed variant of a gene;

selecting features of the observed variant;

determining a hyperplane score using Support Vector Machine algorithms based on an observation of a novel non-linear relationship with the selected features of the observed variant; and

classifying the observed variant as deleterious or tolerable based on the score indicating a distance of the observed variant from the hyperplane.

2 . The method of claim 1 , wherein the features comprise one or more of:

a value indicating the likelihood that the gene of the observed variant causes disease;

a value or values indicating specific sequence features;

a distance value indicating the distance of the observed variant to a transcription start site;

a likelihood that an amino acid substitution is associated with a disruption of the protein of the observed variant;

a predictive deleteriousness value of an algorithm;

a presence or absence of the observed variant in clinical databases;

a frequency of the observed variant in population databases;

a value indicating whether the variant disrupts intronic sequences controlling the proper splicing of the gene.

3 . The method of claim 1 , wherein the observation of a novel non-linear relationship with the selected features of the observed variant comprises a linear separability derived from an expanded input feature space of one or more kernel functions.

4 . The method of claim 1 , further comprising determining a phenotype adjusted gene score, wherein determining a phenotype score comprises:

identifying the gene containing the observed variant;

identifying occurrences of phenotypes associated with the gene within one or more databases; and

assigning a weight according to the relevance of the association.

5 . The method of claim 1 , further comprising determining a phenotype adjusted score, wherein determining a phenotype adjusted score comprises the square root of the multiplication of the hyperplane score by the phenotype adjusted gene score.

6 . The method of claim 1 , further comprising determining a family adjusted score, wherein determining a family adjusted score comprises:

determining a frequency of the observed variant within a family;

determining a family adjusted score of the observed variant based on a relationship between determined hyperplane score and the determined frequency within the family.

7 . The method of claim 6 , further comprising determining a family adjusted gene score, wherein determining a family adjusted gene score comprises aggregation of the family adjusted score of all variants which locate in the gene.

8 . The method of claim 7 , further comprising determining a gene phenotype combined score, wherein determining the gene phenotype combined score comprises the square root of the multiplication of the family adjusted gene score by the phenotype adjusted gene score.

9 . A system for identifying a possible disease-causing genetic variant by machine learning classification, comprising:

a processing device;

a storage device having instructions thereon that, when executed by the processing device, cause the system to:

receive a training dataset of predetermined variants associated with a disease;

identify a hyperplane having a maximum margin between points of the training dataset;

receive patient input data comprising an observed variant;

select features of the observed variant;

determine a score using Support Vector Machine algorithms based on an observation of a novel non-linear relationship with the selected features of the observed variant; and

classify the observed variant as deleterious or tolerable based on the score indicating a distance of the observed variant from the hyperplane.

10 . The system of claim 1 , wherein the features comprise one or more of:

a value indicating the likelihood that the gene of the observed variant causes disease;

a value or values indicating specific sequence features;

a distance value indicating the distance of the observed variant to a transcription start site;

a likelihood that an amino acid substitution is associated with a disruption of the protein of the observed variant;

a deleteriousness value of an algorithm;

a presence or absence of the observed variant in clinical databases;

a frequency of the observed variant in population databases;

a value indicating whether the variant disrupts intronic sequences controlling the proper splicing of the gene.

11 . The system of claim 10 , wherein the data of the features are based on data of third party databases.

12 . The system of claim 9 , wherein the observation of a novel non-linear relationship with the selected features of the observed variant comprises a linear separability derived from an expanded input feature space of one or more kernel functions.

13 . The system of claim 9 , the storage device further comprising instructions to cause the processing device to determine a phenotype adjusted gene score, wherein determining a phenotype score comprises:

identifying the gene containing the observed variant;

identifying occurrences of phenotypes associated with the gene within one or more databases; and

assigning a weight according to the relevance of the association.

14 . The system of claim 9 , the storage device further comprising instructions to cause the processing device to determine a phenotype adjusted score, wherein determining a phenotype adjusted score comprises the square root of multiplying the hyperplane score by the phenotype adjusted gene score.

15 . The system of claim 9 , the storage device further comprising instructions to cause the processing device to determine a family adjusted score, wherein determining a family adjusted score comprises:

determining a frequency of the observed variant within a family;

determining a family adjusted score of the observed variant based on a relationship between determined hyperplane score and the determined frequency within the family.

16 . The system of claim 15 , the storage device further comprising instructions to cause the processing device to determine a family adjusted gene score, wherein determining a family adjusted gene score comprises aggregation of the family adjusted score of all variants which locate in the gene.

17 . The system of claim 16 , the storage device further comprising instructions to cause the processing device to determine a gene phenotype combined score, wherein determining the gene phenotype combined score comprises the square root of multiplying the family adjusted gene score by the phenotype adjusted gene score.

18 . A non-transitory computer-readable medium for identifying a possible disease-causing genetic variant by machine learning classification, the computer-readable medium comprising processor-executable code to:

receive a training dataset of predetermined variants associated with a disease;

identify a hyperplane having a maximum margin between points of the training dataset;

receive patient input data comprising an observed variant;

select features of the observed variant;

determine a score using Support Vector Machine algorithms based on an observation of a novel non-linear relationship with the selected features of the observed variant; and

classify the observed variant as deleterious or tolerable based on the score indicating a distance of the observed variant from the hyperplane.

19 . The computer-readable medium of claim 18 , wherein the features comprise one or more of:

a value indicating the likelihood that the gene of the observed variant causes disease;

a value or values indicating specific sequence features;

a distance value indicating the distance of the observed variant to a transcription start site;

a likelihood that an amino acid substitution is associated with a disruption of the protein of the observed variant;

a deleteriousness value of an algorithm;

a presence or absence of the observed variant in clinical databases;

a frequency of the observed variant in population databases;

a value indicating whether the variant disrupts intronic sequences controlling the proper splicing of the gene.

20 . The computer-readable medium of claim 18 , wherein the data of the features are based on data of third party databases, wherein the observation of a novel non-linear relationship with the selected features of the observed variant comprises a linear separability derived from an expanded input feature space of one or more kernel functions.

21 . The computer-readable medium of claim 18 , the computer-readable medium further comprising processor-executable code to determine one or more of:

a phenotype adjusted gene score;

a phenotype adjusted score;

a family adjusted score, wherein determining a family adjusted score;

a family adjusted gene score; and

a gene phenotype combined score.

Assignments (7)
SECURITY INTEREST Recorded Aug 4, 2022
From: PIERIANDX, INC.; SEVEN BRIDGES GENOMICS INC.
To: ORBIMED ROYALTY & CREDIT OPPORTUNITIES III, LP
Reel/Frame 061084/0786 →
RELEASE OF SECURITY INTEREST Recorded Aug 3, 2022
From: ORBIMED ROYALTY & CREDIT OPPORTUNITIES III, LP
To: PIERIANDX, INC.
Reel/Frame 060711/0059 →
SECURITY INTEREST Recorded Oct 29, 2021
From: PIERIANDX, INC.
To: ORBIMED ROYALTY & CREDIT OPPORTUNITIES III, LP
Reel/Frame 057964/0895 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2017
From: ROBISON, REID; WANG, KAI
To: TUTE GENOMICS, INC.
Reel/Frame 042000/0124 →
CHANGE OF NAME Recorded Nov 1, 2016
From: TUTE GENOMICS, INC.
To: TPIER, INC.
Reel/Frame 040541/0870 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2016
From: TPIER, INC.
To: PIERIANDX, INC.
Reel/Frame 040188/0204 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2014
From: ROBISON, REID; WANG, KAI
To: TUTE GENOMICS
Reel/Frame 033623/0780 →