IP Library Patent Application 13320308
Patent Application
App. No. 13/320,308

METHODS AND SYSTEMS FOR KNOWLEDGE DISCOVERY

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/320,308
Abstract

In an aspect, provided is a Natural Language Processing (NLP) workflow engine to analyze text. The engine can combine one or more independent NLP components (e.g. Tokenization, Part of Speech Tagging, Named Entity Recognition) into a meaningful processing workflow.

Claims (29)

1 . A method of textual analysis comprising:

analyzing text using a processor comprising a workflow engine, wherein said workflow engine comprises at least a thesaurus component, said thesaurus component comprising a structured datafile of words related to a knowledge field;

creating a knowledge fingerprint of the text using said text analysis.

2 . The method of claim 1 , wherein said workflow engine comprises one or more additional components.

3 . The method of claim 2 , wherein the one or more additional components can include one or more of a tokenization component, a sentence boundary detection component, an abbreviation expansion component, a normalization component, a part-of-speech (POS) tagger component, a noun phrase extraction component, a concept extraction component, a named entity recognition component, a relation extraction component, a quantifier detection component, or an anaphora resolution component.

4 . The method of claim 3 , wherein one or more different knowledge footprints are created by said workflow engine.

5 . The method of claim 3 , wherein a different knowledge footprint is created by each component that comprises said workflow engine.

6 . The method of claim 1 , wherein the thesaurus component comprises a compilation of validated concepts representing a field of knowledge or a piece of knowledge organized into the structured datafile of words related to a knowledge field.

7 . The method of claim 1 , wherein said thesaurus component comprises a structured datafile of normalized words related to a knowledge field.

8 . A system for textual analysis comprised of:

a memory; and

a processor operably connected with said memory, wherein said processor is configured to,

analyze text using a workflow engine, wherein said workflow engine comprises at least a thesaurus component, said thesaurus component comprising a structured datafile of words related to a knowledge field stored in said memory; and

create a knowledge fingerprint of the text using said text analysis.

9 . The system of claim 8 , wherein said workflow engine comprises one or more additional components.

10 . The system of claim 9 , wherein the one or more additional components can include one or more of a tokenization component, a sentence boundary detection component, an abbreviation expansion component, a normalization component, a part-of-speech (POS) tagger component, a noun phrase extraction component, a concept extraction component, a named entity recognition component, a relation extraction component, a quantifier detection component, or an anaphora resolution component.

11 . The system of claim 10 , wherein one or more different knowledge footprints are created by said workflow engine.

12 . The system of claim 10 , wherein a different knowledge footprint is created by each component that comprises said workflow engine.

13 . The system of claim 8 , wherein the thesaurus component comprises a compilation of validated concepts representing a field of knowledge or a piece of knowledge organized into the structured datafile of words related to a knowledge field.

14 . The system of claim 8 , wherein said thesaurus component comprises a structured datafile of normalized words related to a knowledge field.

15 . A computer program product comprising at least one non-transitory computer-readable storage medium having computer-readable program code portions for textual analysis stored therein, said computer-readable program code portions comprising:

a first portion for analyzing text using a processor comprising a workflow engine, wherein said workflow engine comprises at least a thesaurus component, said thesaurus component comprising a structured datafile of words related to a knowledge field; and

a second portion creating a knowledge fingerprint of the text using said text analysis.

16 . The computer program product of claim 15 , wherein said workflow engine comprises one or more additional components.

17 . The computer program product of claim 16 , wherein the one or more additional components can include one or more of a tokenization component, a sentence boundary detection component, an abbreviation expansion component, a normalization component, a part-of-speech (POS) tagger component, a noun phrase extraction component, a concept extraction component, a named entity recognition component, a relation extraction component, a quantifier detection component, or an anaphora resolution component.

18 . The computer program product of claim 17 , wherein one or more different knowledge footprints are created by said workflow engine.

19 . The computer program product of claim 17 , wherein a different knowledge footprint is created by each component that comprises said workflow engine.

20 . The computer program product of claim 15 , wherein the thesaurus component comprises a compilation of validated concepts representing a field of knowledge or a piece of knowledge organized into the structured datafile of words related to a knowledge field.

21 . The computer program product of claim 15 , wherein said thesaurus component comprises a structured datafile of normalized words related to a knowledge field.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2013
From: COLLEXIS B.V.
To: COLLEXIS HOLDINGS, INC.
Reel/Frame 029628/0607 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2013
From: COLLEXIS HOLDINGS, INC.
To: SCIENCE INFORMATION SOLUTIONS LLC
Reel/Frame 029628/0644 →
MERGER Recorded Jan 15, 2013
From: SCIENCE INFORMATION SOLUTIONS LLC
To: ELSEVIER INC.
Reel/Frame 029628/0736 →