IP Library Granted Patent US 11,164,657
Granted Patent B2
US 11,164,657 · App. 15/826,792 · Granted Nov 2, 2021

Accelerated pharmaceutical repurposing by finding anticorrelations and by text mining

Inventors: Meenakshi Nagarajan (San Jose, CA); Alix Lacoste (Brooklyn, NY)
Assignee: International Business Machines Corporation
G16B40/00C12Q1/025G01N33/5023G06F16/334G16B5/00G16B20/00G16B25/10G16B35/20G16H10/20G16H20/10G16H70/40G01N21/6486G06F2216/03G16B50/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,164,657
App. No.
15/826,792
Granted
Nov 2, 2021
Kind
B2
Abstract

Utilizing a computing device to assist in repurposing of a pharmaceutical. An identification of a pharmaceutical for repurposing study is received by a computing device. A pharmaceutical expression signature is retrieved based upon the identification of the pharmaceutical, the pharmaceutical expression signature indicating differential expressions of a plurality of biomolecules regulated by the pharmaceutical. A plurality of disease expression signatures are retrieved from a disease omics database, each disease expression signature indicating differential expressions of a plurality of biomolecules affected by a disease. A pharmaceutical vector is generated based upon the pharmaceutical expression signature for the pharmaceutical. A plurality of disease vectors are generated based upon the plurality of disease expression signatures for each disease. N hypotheses correlating the pharmaceutical vector and one or more of the plurality of disease vectors are generated, each hypothesis indicating a potential repurposing for the pharmaceutical to treat the disease.

Claims (71)

1. A method comprising:

receiving by a computing device an identification of a pharmaceutical;

retrieving, by the computing device, a pharmaceutical expression signature for the pharmaceutical, the pharmaceutical expression signature indicating differential expressions of biomolecules regulated by the pharmaceutical versus a first control;

retrieving, by the computing device, disease expression signatures from a disease omics database, each disease expression signature indicating a differential expression of at least one biomolecule affected by a respective disease versus a second control;

generating by the computing device a pharmaceutical vector based upon the pharmaceutical expression signature;

generating, by the computing device, disease vectors based upon the disease expression signatures, respectively;

identifying, by the computing device:

a first anticorrelation between the pharmaceutical vector and one of the disease vectors, the first anticorrelation indicating a first potential repurposing for the pharmaceutical to treat a first disease, the first disease corresponding to the one of the disease vectors; and

a second anticorrelation between the pharmaceutical vector and another one of the disease vectors, the second anticorrelation indicating a second potential repurposing for the pharmaceutical to treat a second disease, the second disease corresponding to the other one of the disease vectors;

text mining, by the computing device, a literature database to rank the first and second anticorrelations and the first and second potential repurposings, wherein the text mining comprises identifying references from the literature database, and wherein the text mining comprises:

extracting from the references a first indication of a direct pharmaceutical effect on a first biomolecule via the pharmaceutical and on the first biomolecule via at least one of the first and second diseases, and extracting from the references:

a second indication of an indirect of pharmaceutical effect on a second biomolecule via at least one of the pharmaceutical and the first and second diseases and

a third indication of a pharmaceutical effect on the second biomolecule via another of the pharmaceutical and the first and second diseases;

generating, by the computing device, a first confirmation score for the first potential repurposing and a second confirmation score for the second potential repurposing, wherein the first confirmation score and the second confirmation score are based on the text mining;

presenting, via the computing device, a ranking of the first confirmation score and the second confirmation score; and

testing, via at least one test selected from the group consisting of a chemical assay and a fluorescence assay, the first potential repurposing or the second potential repurposing according to which of the first and the second potential repurposing corresponds to a higher ranking of the first confirmation score and the second confirmation score.

2. The method of claim 1 , wherein the text mining of the literature database uses machine learning and natural language processing.

3. The method of claim 1 , wherein each biomolecule is a gene, a protein, or a metabolite.

4. The method of claim 1 , wherein the differential expressions of the biomolecules regulated by the pharmaceutical and the differential expressions of the biomolecules affected by the diseases indicate one or more biomolecules up-regulated and one or more biomolecules down-regulated with respect to the first control or the second control.

5. The method of claim 1 , wherein the text mining further comprises inferring, from the references, a semantic similarity between the references.

6. The method of claim 5 , wherein the semantic similarity is between a name of a third biomolecule affected by the pharmaceutical and a name of a fourth biomolecule affected by at least one of the first and second diseases, and wherein the names are from the references.

7. The method of claim 5 , wherein the semantic similarity is between a first word associated with a third biomolecule affected by the pharmaceutical and a second word associated with a fourth biomolecule affected by at least one of the first and second diseases, and wherein the first word and the second word are from the references.

8. A computer program product comprising:

one or more non-transitory computer-readable storage media and program instructions stored on the one or more non-transitory computer-readable storage media, wherein the program instructions are configured to cause, when executed by a computing device, the computing device to perform a method comprising:

receiving an identification of a pharmaceutical;

retrieving a pharmaceutical expression signature for the pharmaceutical, the pharmaceutical expression signature indicating differential expressions of biomolecules regulated by the pharmaceutical versus a first control;

retrieving disease expression signatures from a disease omics database, each disease expression signature indicating a differential expression of at least one biomolecule affected by a respective disease versus a second control;

generating a pharmaceutical vector based upon the pharmaceutical expression signature;

generating disease vectors based upon the disease expression signatures, respectively;

identifying:

a first anticorrelation between the pharmaceutical vector and one of the disease vectors, the first anticorrelation indicating a first potential repurposing for the pharmaceutical to treat a first disease, the first disease corresponding to the one of the disease vectors; and

a second anticorrelation between the pharmaceutical vector and another one of the disease vectors, the second anticorrelation indication a second potential repurposing for the pharmaceutical to treat a second disease, the second disease corresponding to the other one of the disease vectors;

text mining a literature database to rank the first and second anticorrelations and the first and second potential repurposings, wherein the text mining comprises identifying references from the literature database, and wherein the text mining comprises:

extracting from the references a first indication of a direct pharmaceutical effect on a first biomolecule via the pharmaceutical and on the first biomolecule via the first and second diseases, and

extracting from the references:

a second indication of an indirect of pharmaceutical effect on a second biomolecule via at least one of the pharmaceutical and the first and second diseases, and

a third indication of a pharmaceutical effect on the second biomolecule via another of the pharmaceutical and the first and second diseases

generating a first confirmation score for the first potential repurposing and a second confirmation score for the second potential repurposing, wherein the first confirmation score and the second confirmation score are based on the text mining;

presenting a ranking of the first confirmation score and the second confirmation score; and

testing, via at least one test selected from the group consisting of a chemical assay and a fluorescence assay, the first potential repurposing or the second potential repurposing according to which of the first and the second potential repurposing corresponds to a higher ranking of the first confirmation score and the second confirmation score.

9. The computer program product of claim 8 , wherein the text mining of the literature database uses machine learning and natural language processing.

10. The computer program product of claim 8 , wherein each biomolecule is a gene, a protein, or a metabolite.

11. The computer program product of claim 8 , wherein the differential expressions of the biomolecules regulated by the pharmaceutical and the differential expressions of the biomolecules affected by the diseases indicate one or more biomolecules up-regulated and one or more biomolecules down-regulated with respect to the first control or the second control.

12. The computer program product of claim 8 , wherein the text mining further comprises inferring, from the references, a semantic similarity between the references.

13. The computer program product of claim 12 , wherein the semantic similarity is between a name of a third biomolecule affected by the pharmaceutical and a name of a fourth biomolecule affected by at least one of the first and second diseases, and wherein the names are from the references.

14. The computer program product of claim 12 , wherein the semantic similarity is between a first word associated with a third biomolecule affected by the pharmaceutical and a second word associated with a fourth biomolecule affected by at least one of the first and second diseases, and wherein the first word and the second word are from the references.

15. A computer system comprising:

one or more computer processors;

one or more computer-readable storage media; and

program instructions stored on the computer-readable storage media and configured to cause the computer system to perform a method comprising:

receiving an identification of a pharmaceutical;

retrieving a pharmaceutical expression signature for the pharmaceutical, the pharmaceutical expression signature indicating differential expressions of biomolecules regulated by the pharmaceutical versus a first control;

retrieving disease expression signatures from a disease omics database, each disease expression signature indicating a differential expression of at least one biomolecule affected by a respective disease versus a second control;

generating a pharmaceutical vector based upon the pharmaceutical expression signature;

generating disease vectors based upon the disease expression signatures, respectively;

identifying:

a first anticorrelation between the pharmaceutical and one of the disease vectors, the first anticorrelation indicating a first potential repurposing for the pharmaceutical to treat a first disease, the first disease corresponding to the one of the disease vectors; and

a second anticorrelation between another one of the pharmaceutical vectors and another one of the disease vectors, the second anticorrelation indicating a second potential repurposing for the pharmaceutical to treat a second disease, the second disease corresponding to the other one of the disease vectors;

text mining a literature database to rank the first and second anticorrelations and the first and second potential repurposings, wherein the text mining comprises identifying references from the literature database, and wherein the text mining comprises:

extracting from the references a first indication of a direct pharmaceutical effect on a first biomolecule via the pharmaceutical and on the first biomolecule via at least one of the first and second diseases, and

extracting from the references:

a second indication of an indirect pharmaceutical effect on a second biomolecule via at least one of the pharmaceutical and the first and second diseases, and

a third indication of a pharmaceutical effect on the second biomolecule via another of the pharmaceutical and the first and second diseases

generating a first confirmation score for the first potential repurposing and a second confirmation score for the second potential repurposing, wherein the first confirmation score and the second confirmation score are based on the text mining;

presenting a ranking of the first confirmation score and the second confirmation score; and

testing, via at least one test selected from the group consisting of a chemical assay and a fluorescence assay, the first potential repurposing or the second potential repurposing according to which of the first and the second potential repurposing corresponds to a higher ranking of the first confirmation score and the second confirmation score.

16. The computer system of claim 15 , wherein the text mining of the literature database uses machine learning and natural language processing.

17. The computer system of claim 15 , wherein each biomolecule is a gene, a protein, or a metabolite.

18. The computer system of claim 15 , wherein the differential expressions of the biomolecules regulated by the pharmaceutical and the differential expressions of the biomolecules affected by the diseases indicate one or more biomolecules up-regulated and one or more biomolecules down-regulated.

19. The computer system of claim 15 , wherein the text mining further comprises inferring, from the references, a semantic similarity between the references.

20. The computer system of claim 19 , wherein the semantic similarity is between a name of a third biomolecule affected by the pharmaceutical and a name of a fourth biomolecule affected by at least one of the first and second diseases, and wherein the names are from the references.

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2017
From: NAGARAJAN, MEENAKSHI; LACOSTE, ALIX
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044256/0308 →