IP Library Patent Application 12275949
Patent Application
App. No. 12/275,949

METHOD & APPARATUS FOR IDENTIFYING A SECONDARY CONCEPT IN A COLLECTION OF DOCUMENTS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
12/275,949
Abstract

A Methodology for identifying secondary concepts that are included in one or more documents in a collection of documents is disclosed. Training information is manually created from a subset of a collection of documents and used by a primary concept identification function to process textual information contained in the documents included in the collection of documents to identify primary concepts included in the collection of documents. Each of the primary concepts included in the collection of documents is used as input to a secondary concept identification function which results in the identification of secondary concepts included in each of the primary concepts. A query is generated and used as input to both the primary and secondary concept identification functions and the result of both the operation of both of these functions on the query is compared to the identified secondary concepts. The distance between the query and each of the secondary concepts is determined and those secondary concepts that are within a predetermined distance of the query are displayed.

Claims (40)

1 . A method for identifying at least one instance of a secondary concept among a plurality of documents comprising:

creating a primary concept space from primary concept information identified in the plurality of documents;

decomposing the information contained in the primary concept space to create a secondary concept space that includes one or more secondary concepts, each of which secondary concepts is represented in the secondary concept space as a separate vector value;

creating a query and translating the query into the secondary concept space where it is represented as a query vector value;

comparing the query vector value to each of the secondary concept vector values included in the secondary concept space; and

displaying at least one secondary concept that is within a specified distance of the query vector value.

2 . The method of claim 1 wherein the primary concept space is a multidimensional relationship between document terms and primary document topics.

3 . The method of claim 1 wherein the primary concept information is comprised of a plurality of significant terms included in the plurality of documents and one or more primary topics associated with the plurality of documents.

4 . The method of claim 1 wherein decomposing the information contained in the at least one primary concept space is performed by latent semantic analysis.

5 . The method of claim 1 wherein the secondary concept space is comprised of a multidimensional relationship between the one or more secondary concepts and the one or more primary concepts.

6 . The method of claim 1 wherein the query includes one or more selected terms.

7 . The method of claim 1 wherein translating the query into the secondary concept space is comprised of employing a primary concept identification function to generate a relationship between the query terms and one or more of the primary concepts and employing a secondary concept identification function to decompose primary concept-query term relationships.

8 . The method of claim 1 wherein comparing the query vector value to each of the one or more secondary concept vector values is comprised of one or calculating the dot product or the cosine between the query the query vector value and a secondary concept vector value.

9 . A method for identifying at least one instance of a secondary concept in a plurality of documents comprising:

training a primary concept identification function to identify one or more significant terms associated with each of one or more primary concepts in a sub-group of the plurality of documents;

employing the trained primary concept identification function to detect the frequency of substantially all of the significant terms associated with each one of the one or more primary concepts in the plural documents;

defining a relationship between all of the one or more significant terms and at least one of the primary concepts and storing the contents of the defined relationship as a primary concept space;

processing the contents of the stored primary concept space using a secondary concept identification function to identify at least one secondary concept associated with at least one instance of a primary concept and calculating a vector value for it and storing the at least one vector value as a secondary concept vector value in a secondary concept space;

creating a query and translating the query into the secondary concept space and calculating a vector value for it and storing the vector value as a query vector value in the secondary concept space;

comparing the query vector value to each of the at least one secondary concept vector values; and

displaying at least one secondary concept that is within a select distance of the query vector value.

10 . The method of claim 9 wherein training the primary concept identification function includes manually identifying at least one primary concept in a collection of documents and applying one or more natural language processing functions to the at least one manually identified primary concept to identify at least one significant term.

11 . The method of claim 10 wherein the at least one significant term is a word that appears in the text of the primary concept more than a predetermined number of times.

12 . The method of claim 9 wherein the defined relationship is a multidimensional matrix.

13 . The method of claim 9 wherein the primary concept identification function includes at least one natural language processing function.

14 . The method of claim 13 wherein the at least one natural language processing function is one of a stemming function, a part of speech tagging function, a synonym tagging function and a significant word identification function.

15 . The method of claim 9 wherein the secondary concept identification function is a latent semantic indexing process.

16 . The method of claim 9 wherein comparing the query vector value to each of the one or more secondary concept vector values is comprised of one or calculating the dot product or the cosine between the query the query vector value and a secondary concept vector value.

17 . Apparatus for identifying at least one instance of a secondary concept in a plurality of documents comprising:

a processor;

a user interface;

a display device; and

a storage device for storing a secondary concept identification module that operates to create a primary concept space from primary concept information identified in the plurality of documents, decompose the information contained in the primary concept space to create a secondary concept space that includes one or more secondary concepts, each of which secondary concept is represented in the secondary concept space as a separate vector value, create a query and translate the query into the secondary concept space where it is represented as a query vector value, compare the query vector value to each of the secondary concept vector values included in the secondary concept space, and display at least one secondary concept that is within a specified distance of the query vector value.

18 . The apparatus of claim 17 wherein the primary concept space is a multidimensional relationship between document terms and primary document topics.

19 . The apparatus of claim 17 wherein the primary concept information is comprised of a plurality of significant terms included in the plurality of documents and one or more primary topics associated with the plurality of documents.

20 . The apparatus of claim 17 wherein decomposing the information contained in the at least one primary concept space is performed by latent semantic analysis.

21 . The apparatus of claim 17 wherein the secondary concept space is comprised of a multidimensional relationship between the one or more secondary concepts and the one or more primary concepts.

22 . The apparatus of claim 17 wherein the query includes one or more selected terms.

23 . The apparatus of claim 17 wherein translating the query into the secondary concept space is comprised of employing a primary concept identification function to generate a relationship between the query terms and one or more of the primary concepts and employing a secondary concept identification function to decompose primary concept-query term relationships.

24 . The apparatus of claim 17 wherein comparing the query vector value to each of the one or more secondary concept vector values is comprised of one or calculating the dot product or the cosine between the query the query vector value and a secondary concept vector value.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2012
From: EMPTORIS, INC.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 029461/0904 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2008
From: RASKINA, OLGA; JAIMISON, ROBERT; KAMON, AMMIEL
To: EMPTORIS, INC.
Reel/Frame 021975/0136 →