IP Library Granted Patent US 9,773,091
Granted Patent B2
US 9,773,091 · App. 14/200,520 · Granted Sep 26, 2017

Systems and methods for genomic annotation and distributed variant interpretation

Inventors: Ali Torkamani (San Diego, CA); Nicholas Schork (Solana Beach, CA)
Assignee: The Scripps Research Institute
G06F19/24G06F19/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,773,091
App. No.
14/200,520
Granted
Sep 26, 2017
Kind
B2
Abstract

A computer-based genomic annotation system, including a database configured to store genomic data, non-transitory memory configured to store instructions, and at least one processor coupled with the memory, the processor configured to implement the instructions in order to implement an annotation pipeline and at least one module filtering or analysis of the genomic data.

Claims (28)

1. A computer-based genomic annotation system, comprising:

non-transitory memory configured to store instructions, and

at least one processor coupled with the memory, the processor configured to:

receive a variant file including a plurality of genomic DNA variants relative to a reference genomic DNA sequence;

generate an annotation file containing annotations corresponding to a plurality of predefined annotation categories for each of the plurality of variants in the variant file;

wherein producing the annotation file comprises:

generating data for the annotation file by mapping each of the plurality of genomic DNA variants to at least one known gene, wherein the mapping is created by:

searching the reference genomic DNA sequence for known gene based on a single or multidimensional relative position of the variants with respect to one or more known genes, or on any other spatial or functional relationship relating the variants to known genes, including statistical, regulatory, or related to physical interactions,

identifying at least one known gene nearest to each variant based on the searching, and

generating a gene-based annotation that identifies the at least one known gene mapped to each variant of the plurality of genomic DNA variants;

generating a plurality of annotations in the annotation file for each variant of the plurality of genomic DNA variants, the plurality of annotations corresponding to the plurality of predefined annotation categories, wherein at least one of the plurality of generated annotations includes the gene-based annotation that identifies the at least one known gene, wherein the plurality of generated annotations further comprise additional gene-based annotations associated with the mapped genes, the additional gene based annotations comprising the following levels of annotation information: genomic elements, prediction of impact information, linking element information and prior knowledge; and

populating the annotation file with the plurality of generated annotations, wherein substantially all of the annotated variants in the annotation file contain the gene-based annotation identifying at least one known gene; and

output the annotation file.

2. The system of claim 1 , wherein for each variant of the plurality of variants, the processor is configured to identify a gene that has a start or stop codon closest to the variant, and map the identified known gene to the variant.

3. The system of claim 2 , wherein some of the variants of the plurality of variants are located between the start and stop codons of the variant's associated gene, and some of the variants are located outside the start and stop codons of the variant's associated gene.

4. The system of claim 2 , wherein some of the variants of the plurality of variants are located outside the transcribed portions of the genome.

5. The system of claim 1 , wherein the genomic elements comprise at least one of known genes, protein domains, transcription factor binding sites, conserved elements, microRNA (miRNA), binding sites, splice sites, splicing enhancers, splicing silencers, common single nucleotide polymorphisms (SNPs), Untranslated Regions (UTR) regulatory motifs, post translational modification sites, and custom elements.

6. The system of claim 1 , wherein the prediction of impact information comprises at least one of coding impact, non-synonymous impact prediction, protein domain impact prediction, motif based impact scores, nucleotide conservation, miRNA targets, splicing changes, binding energy, and codon abundance.

7. The system of claim 1 , wherein the linking element information comprises at least one of phase information, molecular information, biological information, protein-protein interactions, co-expression, and genomic context.

8. The system of claim 1 , wherein the prior knowledge comprises at least one of phenotype associations, biological processes, molecular function, drug metabolism, genome-wide association study (GWAS) catalog, allele frequency, expression quantitative trait loci (eQTL) frequency, and text mining information.

9. The system of claim 1 , wherein the gene-based annotations include genomic elements comprising transcription factor binding site motifs, and wherein functional element mapping comprises at least one of the following: scanning the transcription factor binding site motifs against an associated genome, determining the positions of the transcription factor binding site motifs relative to at least one known genomic element, and mapping the variant onto the at least one known genomic element.

10. The system of claim 1 , wherein the processor is further configured to generate synthetic annotations based at least in part on the gene-based annotations.

11. The system of claim 1 , wherein the processor is configured to filter the annotation file to remove variant entries that correspond to variants found in one or more reference genome sequences.

12. The system of claim 11 , wherein the one or more reference genome sequences are derived from subjects having a selected ancestry.

13. The system of claim 11 , wherein the one or more reference genome sequences are derived from subjects of a selected age and health status.

14. The system of claim 1 , wherein the processor is further configured to:

receive or generate a set of seed genes predicted to be associated with a defined phenotype; and

generate a measure of potential functional overlap between each gene in the annotation file to the genes in the set of seed genes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2015
From: TORKAMANI, ALI; SCHORK, NICHOLAS
To: SCRIPPS RESEARCH INSTITUTE, THE
Reel/Frame 034837/0048 →
Continuity (5)
Continuation In Part PCTUS2012062787 · Oct 31, 2012
Provisional Application 61553576 · Oct 31, 2011
Provisional Application 61676885 · Jul 27, 2012
Provisional Application 61852255 · Mar 15, 2013
Related Publication 20140304270A1 · Oct 9, 2014