IP Library Granted Patent US 9,424,299
Granted Patent B2
US 9,424,299 · App. 14/641,527 · Granted Aug 23, 2016

Method for preserving conceptual distance within unstructured documents

Inventors: John P. Bufe (Washington, DC); Timothy P. Winkler (Clinton, MA)
Assignee: International Business Machines Corporation
G06F17/30336G06F17/278G06F17/2785G06F17/2795G06F17/28G06F17/30011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,424,299
App. No.
14/641,527
Granted
Aug 23, 2016
Kind
B2
Abstract

A method, system and computer-usable medium are disclosed for preserving conceptual distance within unstructured documents by characterizing conceptual relationships. Natural language processing is applied to content in a plurality of documents to identify topics and subjects. Analytic analysis is then applied to the identified topics and subjects to identify concepts. The content in each of the plurality of documents is partitioned into a first structured hierarchy, preserving at least one structure in each document inherent in the each document. Access is then provided to the content through a first index based upon utilizing the first structured hierarchy and through a second index utilizing a second structured hierarchy. The conceptual relationship criteria are based upon a directed graph with weights based upon a similarity and a distance based upon concepts.

Claims (20)

1. A computer-implemented method for characterizing content of documents by conceptual relationships, comprising:

applying natural language processing (NLP) to content in a plurality of documents to identify topics and subjects;

applying analytic analysis to the topics and subjects to identify a conceptual relationships of the content in the plurality of documents;

partitioning the content in each of the plurality of documents into a first structured hierarchy, preserving at least one structure in each document inherent in the each document; and

providing access to content through a first index based upon utilizing the first structured hierarchy and through a second index utilizing a second structured hierarchy; and wherein

the content is characterized by optimizing a vector space model representation of the documents, the optimization performed by a system capable of answering questions, where:

the content from the plurality of documents is ingested by the system;

natural language processing is applied to the content in the plurality of documents to identify terms, topics, subjects and concepts;

the content is partitioned according to a semantic parse distance to identify a context for partitioned content;

the content and context is represented, by the system, utilizing a vector space model;

entries in the vector space model are eliminated based on a difference criteria; and

an iterative genetic algorithm is applied to optimize features of the vector space model.

2. The method of claim 1 , wherein:

the conceptual relationship is based upon a directed graph with weights based upon a similarity and a distance based upon concepts.

3. The method of claim 2 , wherein:

the distance is based upon a topic hierarchy.

4. The method of claim 1 , wherein:

a ground truth is an optimized feature.

5. The method of claim 4 , wherein:

the genetic algorithm determines which features are used during the ingesting and has weighting based on semantic distance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2015
From: BUFE, JOHN P.; WINKLER, TIMOTHY P.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 035111/0709 →
Continuity (2)
Continuation 14508200 · Oct 7, 2014
Related Publication 20160098398A1 · Apr 7, 2016