IP Library Granted Patent US 10,275,424
Granted Patent B2
US 10,275,424 · App. 14/166,160 · Granted Apr 30, 2019

System and method for language extraction and encoding

Inventor: Carol Friedman (New York, NY)
Assignee: The Trustees of Columbia University in the City of New York
G06F17/21G06F17/2765
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,275,424
App. No.
14/166,160
Granted
Apr 30, 2019
Kind
B2
Abstract

Improved systems and methods for extracting information from medical and natural-language text data.

Claims (25)

1. A method for extracting information from medical or natural-language input text, comprising:

receiving, by a computing system, medical or natural-language input text, wherein one or more words or portions of said medical or natural-language input text includes an identification tag;

selecting, by computing system, the input text using the identification tag to determine a relevant text input and an irrelevant text input, wherein the identification tag includes a string value and/or a nested structure value, wherein the identification tag is configured to be customized and recognized by a processor;

utilizing, by the computing system, a lexicon knowledge base to identify and categorize multi-word and single word phrases within sentences of the relevant text input, wherein said lexicon knowledge base is configured to be dynamically customized by a user;

receiving, from a user, filenames having new lexical entries, and modifying the lexicon knowledge based on the filenames;

disambiguating, by the computing system, one or more ambiguous words in the relevant text input using a contextual disambiguation rule, wherein the contextual disambiguation rule is configured to analyze words following or preceding each ambiguous word, words in the same sentence, words in a certain section, and/or words in a certain domain, and wherein the contextual disambiguation rule is configured to be dynamically loaded without compiling the entire computing system;

parsing, by the computing system, said relevant text input to determine a grammatical structure of the relevant text input, said parsing step comprising the step of referring to a domain parameter having a value indicative of a domain from which the text data originated, the domain parameter corresponding to one or more rules of grammar within a knowledge base related to the domain to be applied for parsing the relevant text input;

regularizing, by the computing system, the parsed text data to form a canonical output form;

converting, by the computing system, the canonical output form into controlled vocabulary terms using a table of codes, wherein the table of codes is configured to be dynamically customized without compiling the entire computing system;

tagging, by the computing system, the input text with a structured data component derived from the controlled vocabulary terms; and

outputting the tagged text data to be stored in a database.

2. The method of claim 1 , wherein the identification tag is selected from the group consisting of dates, names, phone numbers, addresses, and locations.

3. The method of claim 1 , wherein the tagged text data is stored in a form compatible with a standard spreadsheet application or relational database.

4. The method of claim 1 , wherein the structure of the identification tag is configured to be updated with an adaptation of the processor.

5. A system for extracting information from medical or natural-language input text, comprising:

a lexicon knowledge base to identify and categorize multi-word and single word phrases within sentences of a relevant input text, wherein said lexicon knowledge base is configured to be dynamically customized by a user by receiving, from the user, filenames having new lexical entries, and the lexicon knowledge is modified based on the filenames;

a processor, coupled to said lexicon knowledge base and receiving said medical or natural-language input text, tagging one or more words or portions of said medical or natural-language input text with an identification tag, and selecting the input test using the identification tag to determine the relevant text input and an irrelevant text input, wherein the identification tag is configured to be customized by a user;

a boundary identifier, coupled to said processor and said lexicon knowledge base and receiving said medical or natural-language input text and dropping the irrelevant text input;

a parser, coupled to said boundary identifier and receiving said relevant input text to determine the grammatical structure of the relevant text input and generating a parsed text wherein one or more ambiguous words in the parsed data are disambiguated using a contextual disambiguation rule, wherein the disambiguation rule is configured to be dynamically loaded without compiling the entire computing system;

a phrase regulator, coupled to said parser and replacing the parsed text with a canonical output form; and

an encoder, coupled to said phrase regulator and receiving the canonical output form, converting the canonical output form into a controlled vocabulary term using a table of code, tagging the input text with a structured data component derived from controlled vocabulary terms, and outputting the tagged text data to be stored in a database, wherein the table of codes is configured to be dynamically customized without compiling the entire computing system.

6. The system of claim 5 , wherein the identification tag is selected from the group consisting of dates, names, phone numbers, addresses, and locations.

7. The system of claim 5 , wherein the tagged text data is stored in a form compatible with a standard spreadsheet application, a xml format, or relational database.

8. The system of claim 5 , wherein the disambiguation rules are configured to be compiled by specifying the filename option, wherein the filename comprises the disambiguation rules.

9. The system of claim 5 , wherein the structure of the identification tag is tag is configured to be updated with an adaptation of the processor.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2016
From: FRIEDMAN, CAROL
To: THE TRUSTEES OF COLUMBIA UNIVERSITY IN THE CITY OF NEW YORK
Reel/Frame 037778/0932 →
CONFIRMATORY LICENSE Recorded Jul 28, 2015
From: COLUMBIA UNIV NEW YORK MORNINGSIDE
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 036202/0050 →
Continuity (3)
Continuation PCTUS2012048251 · Jul 26, 2012
Provisional Application 61513321 · Jul 29, 2011
Related Publication 20140142924A1 · May 22, 2014