IP Library Granted Patent US 11,080,607
Granted Patent B1
US 11,080,607 · App. 17/166,435 · Granted Aug 3, 2021

Data platform for automated pharmaceutical research using knowledge graph

Inventors: Mikhail Demtchenko (London, GB); Sam Christian Macer (Leicestershire, GB); Artem Krasnoslobodtsev (Frisco, TX); {hacek over (Z)}ygimantas Jo{hacek over (c)}ys (Hove, GB); Roy Tal (Dallas, TX); Charles Dazler Knuff (Dallas, TX)
Assignee: Ro5 Inc.
G06N5/022G06F16/951G06K9/6215G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,080,607
App. No.
17/166,435
Granted
Aug 3, 2021
Kind
B1
Abstract

A system and method for an automated pharmaceutical research data platform comprising a data curation platform which searches for and ingests a plurality of unstructured, heterogenous medical data sources, extracts relevant information from the ingested data sources, and creates a massive, custom-built and intricately related knowledge graph using the extracted data, and a data analysis engine which receives data queries from a user interface, conducts analyses in response to queries, and returns results based on the analyses. The system hosts a suite of modules and tools, integrated with the custom knowledge graph and accessible via the user interface, which may provide a plurality of functions such as statistical and graphical analysis, similarity based searching, and edge prediction among others.

Claims (21)

1. A system for automated pharmaceutical research, comprising:

a computer system comprising a memory and a processor;

a data curation platform, comprising a first plurality of programming instructions stored in the memory of, and operating on the processor, wherein the first plurality of programming instructions, when operating on the processor, cause the computer system to:

scrape data from a plurality of data sources accessible via the internet;

parse the scraped data into a format that may be stored in a relational database;

train an encoder on a subset of the data stored in the relational database to identify the key features that define the data stored in the relational database; and

use the trained encoder to create an abstract latent space represented as a three-dimensional convolutional neural network, wherein the abstract latent space contains vectorized representations of all data points in the relational database; and

a data analysis engine, comprising a second plurality of programming instructions stored in the memory of, and operating on the processor, wherein the second plurality of programming instructions, when operating on the processor, cause the computer system to:

receive a similarity search query, the similarity search query comprising a search term;

input the search term into the trained encoder to create an abstract vector representation of the input search term, existing in the abstract latent space of the three-dimensional convolutional neural network;

compute the distance between two abstract vectors within the latent space;

compare the computed distance to a predetermined threshold of similarity, wherein a computed distance less than the threshold indicates non-similarity between the two abstract vectors and a computed distance greater than the threshold indicates similarity; and

decode the abstract vector that is greater than the threshold and return data related to the decoded vector as a response to the received similarity search query.

2. A method for automated pharmaceutical research, comprising the steps of:

scraping data from a plurality of data sources accessible via the internet;

parsing the scraped data into a format that may be stored in a relational database;

training an encoder on a subset of the data stored in the relational database to identify the key features that define the data stored in the relational database;

using the trained encoder to create an abstract latent space represented as a three-dimensional convolutional neural network, wherein the abstract latent space contains vectorized representations of all data points in the relational database;

computing the distance between two abstract vectors within the latent space;

comparing the computed distance to a predetermined threshold of similarity, wherein a computed distance less than the threshold indicates non-similarity between the two abstract vectors and a computed distance greater that the threshold indicates similarity; and

decoding the abstract vector that is greater than the threshold and return data related to the decoded vector as a response to the received similarity search query.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2024
From: TAL, ROY; KNUFF, CHARLES DAZLER
To: RO5.AI
Reel/Frame 069703/0086 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2021
From: DEMTCHENKO, MIKHAIL; MACER, SAM CHRISTIAN; KRASNOSLOBODTSEV, ARTEM; JOCYS, ZYGIMANTAS
To: RO5.AI
Reel/Frame 055138/0497 →
Continuity (1)
Provisional Application 63126349 · Dec 16, 2020
Cited By (11)
US 12,216,635 US 12,242,796 US 12,288,600 US 12,368,503 US 12,386,621 US 12,406,773 US 12,493,615 US 12,587,274 US 12,603,701 US 12,608,370 US 12,627,372