IP Library Granted Patent US 10,628,748
Granted Patent B2
US 10,628,748 · App. 15/029,765 · Granted Apr 21, 2020

System and methods for predicting probable relationships between items

Inventor: Hong Yu (Shrewsbury, MA)
Assignee: University of Massachusetts
G06N5/041G06K9/00G06K9/626G06K9/6256G06N7/005G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,628,748
App. No.
15/029,765
Granted
Apr 21, 2020
Kind
B2
Abstract

The present invention relates generally to identifying relationships between items. Certain embodiments of the present invention are configurable to identify the probability that a certain event will occur by identifying relationships between items. Certain embodiments of the present invention provide an improved supervised machine learning system.

Claims (43)

1. A system for automatically predicting a protein-drug interaction comprising:

a processor:

a main memory in communication with said processor via a communication infrastructure and storing instructions that, when executed by said processor, cause the processor to:

access a collection including at least two or more information items;

develop a training set from the at least two or more information items in the collection, the at least two or more information items comprising a protein information item and a drug information item;

analyze the training set using one or more machine learning models selected from the group consisting of a Naive Bayes, a Naive Bayes Multinomial, and a support vector machine, wherein the analyze step further includes a step of identifying one or more interactions between pairs of protein-drug information items and extracting two or more features of the one or more interactions, the two or more extracted features selected from the group consisting of Adamic, numeommonNeighbor, Jaccard, sum Neighbor, sum Pub, simMeSH, sum Glusteringoef, and jaccardArticleeoOccur;

produce a classifier element based upon the analysis of the training set, wherein the classifier element is selected from the group consisting of a Naive Bayes classifier that assumes features are generated independently from each other, a Naive Bayes Multinomial classifier that assumes a conditional probability of features follows multinomial distribution, and a support vector machine classifier defining decision boundaries of features;

establish an evaluation set comprising new pairs of protein-drug information items, each new pair consisting of a new protein information item and new drug information item;

apply the classifier element to the new pairs of protein-drug information items in the evaluation set producing a probability value for each of the new pairs of protein-drug information items;

rank each classified pair of protein-drug information items in a data set from a high probability value to a low probability value; and

predict a likelihood of interaction between the new protein information item and the new drug information item of each ranked pair, wherein a higher probability value of a ranked pair indicates a higher likelihood of interaction between the new protein information item and the new drug information item in the ranked pair; and

deliver an output comprising the likelihood of interaction between the new protein information item and the new drug information item of each new pair in the evaluation set,

wherein said main memory in communication with the processor via the communication infrastructure stores instructions that, when executed by said processor, cause said processor to display a graphical representation configured to convey at least known relationships between information items;

wherein the graphical representation includes a node symbol configured to represent the information items and a link symbol configured to represent any relationships between the information items;

wherein a degree is calculated for the node symbol.

2. The system of claim 1 , wherein the protein information item is a type of protein.

3. The system of claim 2 , wherein the drug information item is a publication.

4. The system of claim 1 , wherein the drug information item is a type of pharmaceutical drug.

5. The system of claim 1 , wherein the two or more information items further include a first publication date of a publication.

6. The system of claim 1 , wherein the two or more information items further include an authorship of a publication.

7. The system of claim 1 , wherein the two or more information items further include a co-authorship of two or more publications.

8. The system of claim 1 , wherein the two or more information items further include an identification tag of a publication.

9. A system of claim 1 , wherein said main memory in communication with said processor via the communication infrastructure stores instructions that, when executed by said processor, cause said processor to display a graphical representation configured to convey predicted relationships between the information items.

10. A system for automatically predicting a protein-drug interaction comprising:

a processor:

a main memory in communication with said processor via a communication infrastructure and storing instructions that, when executed by said processor, cause the processor to:

form a collection including at least two or more information items;

develop a training set from the at least two or more information items in the collection, the two or more information items comprising a protein information item and a drug information item;

analyze the training set using one or more machine learning models selected from the group consisting of a Naive Bayes, a Naive Bayes Multinomial, and a support vector machine, wherein the analyze step further includes a step of identifying one or more interactions between pairs of protein-drug information items and extracting two or more features of the one or more interactions, the two or more extracted features selected from the group consisting of Adamic, numeommonNeighbor, Jaccard, sum Neighbor, sum Pub, simMeSH, sum Glusteringoef, and jaccardArticleeoOccur;

produce a classifier element based upon the analysis of the training set, wherein the classifier element is selected from the group consisting of a Naive Bayes classifier that assumes features are generated independently from each other, a Naive Bayes Multinomial classifier that assumes a conditional probability of features follows multinomial distribution, and a support vector machine classifier defining decision boundaries of features;

establish an evaluation set comprising new pairs of protein-drug information items, each new pair consisting of a new protein information item and a new drug information item;

apply the classifier element to the new pairs of protein-drug information items in the evaluation set producing a probability value for each of the new pairs of protein-drug information items;

rank each classified pair of protein-drug information items in a data set from a high probability value to a low probability value;

predict a likelihood of interaction between the new protein information item and the new drug information item of each ranked pair, wherein a higher probability value of a ranked pair indicates a higher likelihood of interaction between the new protein information item and the new drug information item of that ranked pair; and

deliver an output comprising the likelihood of interaction between the new protein information item and the new drug information item of each of the new pairs in the evaluation set,

wherein said main memory in communication with the processor via the communication infrastructure stores instructions that, when executed by said processor, cause said processor to display a graphical representation configured to convey at least known relationships between information items;

wherein the graphical representation includes a node symbol configured to represent the information items and a link symbol configured to represent any relationships between the information items;

wherein a degree is calculated for the node symbol.

11. The system of claim 10 , wherein the training set is indexed to facilitate analysis of the training set.

12. The system of claim 10 , wherein the evaluation set is indexed to facilitate analysis of the evaluation set.

13. The system of claim 10 , wherein the protein information item is a type of protein.

14. The system of claim 10 , wherein the drug information item is a publication.

15. The system of claim 10 , wherein the drug information item is a type of pharmaceutical drug.

Assignments (2)
CHANGE OF ADDRESS Recorded Dec 21, 2017
From: UNIVERSITY OF MASSACHUSETTS
To: UNIVERSITY OF MASSACHUSETTS
Reel/Frame 044947/0590 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2017
From: YU, HONG
To: UNIVERSITY OF MASSACHUSETTS
Reel/Frame 042626/0187 →
Continuity (2)
Provisional Application 61911066 · Dec 3, 2013
Related Publication 20160239746A1 · Aug 18, 2016