IP Library Granted Patent US 7,974,788
Granted Patent B2
US 7,974,788 · App. 12/492,992 · Granted Jul 5, 2011

Gene discovery through comparisons of networks of structural and functional relationships among known genes and proteins

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,974,788
App. No.
12/492,992
Granted
Jul 5, 2011
Kind
B2
Abstract

The present invention also relates to natural language processing and extraction of relational information associated with genes and proteins that are found in genomics journal articles. To enable access to information in textual form, the natural language processing system of the present invention provides systems and methods for extracting and structuring information found in the literature in a form appropriate for subsequent applications.

Claims (22)

1. A system for extracting information on biological entities from natural-language text data, comprising:

(i) a computer processing apparatus configured to parse the natural-language text data;

(ii) a computer processing apparatus configured to regularize the parsed text data to form structured word terms; and

(iii) a computer processing apparatus configured to extract interactions between the biological entities from the natural-language text data, wherein the biological entities include genes and/or proteins.

2. The system according to claim 1 , further comprising a computer processing apparatus configured to preprocess the data prior to parsing comprising identifying biological entities.

3. The system according to claim 1 , further comprising a computer processing apparatus configured to referring to an additional parameter which is indicative of the degree to which subphrase parsing is to be carried out.

4. The system according to claim 1 , wherein said parsing further comprises segmenting the text data by sentences.

5. The system according to claim 1 , wherein said parsing further comprises:

segmenting the text data by sentences;

and segmenting each of the sentences at identified words or phrases.

6. The system according to claim 1 , wherein said parsing further comprises:

segmenting the text data by sentences; and

segmenting each of the sentences at a prefix.

7. The system according to claim 1 , wherein said parsing further comprises skipping undefined words.

8. The system according to claim 1 , wherein said parsing further comprises:

identifying one or more binary actions and their relationships; and

identifying one or more arguments associated with the actions.

9. The system according to claim 1 , further comprising a computer processing apparatus configured to perform error recovery when parsing of the text data is unsuccessful.

10. The system according to claim 1 , wherein said error recovery comprises:

segmenting the text data; and

analyzing the segmented text data to achieve at least a partial parsing of the unsuccessfully parsed text data.

11. The system according to claim 1 , wherein said tagging comprises for providing the structured data component in a Standard Generalized Markup Language (SGML) compatible format.

Assignments (1)
CONFIRMATORY LICENSE Recorded Aug 24, 2010
From: COLUMBIA UNIV NEW YORK MORNINGSIDE
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 024878/0006 →
Continuity (5)
Continuation 10921286 · Aug 18, 2004
Division 09549827 · Apr 14, 2000
Continuation In Part 09327983 · Jun 8, 1999
Provisional Application 60129469 · Apr 15, 1999
Related Publication 20100004874A1 · Jan 7, 2010