IP Library Granted Patent US 7,219,105
Granted Patent B2
US 7,219,105 · App. 10/664,261 · Granted May 15, 2007

Method, system and computer program product for profiling entities

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,219,105
App. No.
10/664,261
Granted
May 15, 2007
Kind
B2
Abstract

The present invention provides a method, system and computer program product for profiling an entity based on information obtained form at least one information source. Various contexts associated with the entity are identified. This can be achieved by using a clustering algorithm, an ontology, a thesaurus, association rules or manually by an expert. After the classified into various sets and ranked using a ranking algorithm. Thereafter, certain top ranked concepts are presented to a user as the profile of the entity.

Claims (65)

1. A method of profiling an entity, said method comprising:

retrieving information from at least one information source using said entity as a search criteria;

clustering the retrieved information to identify contexts related to said entity, wherein said contexts are represented by a set of documents;

retrieving information corresponding to each identified context from at least one information source;

selecting features from the information retrieved at both of the retrieving steps to identify concepts associated with said entity within each identified context;

structuring the identified concepts within said each identified context by classifying said identified concepts into exactly four classification sets;

ranking said identified concepts within each set; and

profiling said entity by presenting the top ranked concepts within each set to a user,

wherein the contexts are identified by finding a set of the words or phases that occur frequently with said entity and that mutually do not appear together in documents in the information source, and

wherein said contexts are identified by finding prominent nodes, that comprise said entity, in an ontology or a taxonomy.

2. The method of claim 1 , wherein said contexts are identified by using at least one of synonyms, hypernyms, hyponyms, and meronyms of said entity found in a thesaurus.

3. The method of claim 1 , wherein said four classification sets comprise:

a set of concepts that are exclusive to said entity and unrelated to said identified context;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said entity;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said identified context; and

a set of concepts that are exclusive to said identified context and unrelated to said entity.

4. A method of profiling an entity, said method comprising:

identifying contexts associated with said entity, wherein said contexts are represented by a set of documents;

retrieving information corresponding to each identified context from at least one information source;

selecting features from the retrieved information to identify concepts associated with said entity within said each identified context;

structuring the identified concepts within said each identified context by classifying said identified concepts into exactly four classification sets;

ranking said identified concepts within each set; and

profiling said entity by presenting the top ranked concepts within each set to a user,

wherein the contexts are identified by finding a set of the words or phrases that occur frequently with said entity and that mutually do not appear together in documents in the information source, and

wherein said contexts are identified by using at least one of synonyms, hypernyms, hyponyms, and meronyms of said entity found in a thesaurus.

5. The method of claim 4 , said contexts are identified by finding prominent nodes, that comprise said entity, in an ontology or a taxonomy.

6. The method of claim 4 , wherein said four classification sets comprise:

a set of concepts that are exclusive to said entity and unrelated to said identified context;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said entity;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said identified context; and

a set of concepts that are exclusive to said identified context and unrelated to said entity.

7. A program storage device readable by computer, tangibly embodying a program of instructions executable by the computer to perform a method of profiling an entity, said method comprising:

retrieving information from at least one information source using said entity as a search criteria;

clustering the retrieved information to identify contexts related to said entity, wherein said contexts are represented by a set of documents;

retrieving information corresponding to each identified context from at least one information source;

selecting features front the information retrieved at both of the retrieving steps to identify concepts associated with said entity within each identified context;

structuring the identified concepts within said each identified context by classifying said identified concepts into exactly four classification sets;

ranking said identified concepts within each set; and

profiling said entity by presenting the top ranked concepts within each set to a user,

wherein the contexts are identified by finding a set of the words or phrases that occur frequently with said entity and that mutually do not appear together in documents in the information source, and

wherein said contexts are identified by finding prominent nodes, that comprise said entity, in an ontology or a taxonomy.

8. The program storage device of claim 7 , wherein said contexts are identified by using at least one of synonyms, hypernyms, hyponyms, and meronyms of said entity found in a thesaurus.

9. The program storage device of claim 7 , wherein in said method, said four classification sets comprise:

a set of concepts that are exclusive to said entity and unrelated to said identified context;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said entity;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said identified context; and

a set of concepts that are exclusive to said identified context and unrelated to said entity.

10. A system for profiling an entity, said system comprising:

means for retrieving information from at least one information source using said entity as a search criteria;

means for clustering the retrieved information into entity-context pairs in order to identify contexts related to said entity, wherein said contexts are represented by a set of documents;

means for retrieving information corresponding to each identified context from at least one information source;

means for selecting features from the retrieved information in order to identity concepts associated with said entity within said each identified context;

means for structuring the identified concepts within said each identified context by classifying said identified concepts into exactly four classification sets;

means for ranking said identified concepts within each set; and

means for profiling said entity by presenting the top ranked concepts within each set to a user,

wherein the contexts are identified by finding a set of the words or phrases that occur frequently with said entity and that mutually do not appear together in documents in the information source, and

wherein said contexts are identified by finding prominent nodes, that comprise said entity, in an ontology or a taxonomy.

11. The system of claim 10 , wherein said means for structuring the identified concepts within said each identified context comprises:

a classifier operable for classifying said identified concepts into sets with respect to each entity-context pair; and

a component operable for ranking said identified concepts within each set.

12. The system of claim 10 , wherein said four classification sets comprise:

a set of concepts that are exclusive to said entity and unrelated to said identified context;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said entity;

a set of concepts that are common to both said entity and said identified context and more frequently associated with said identified context; and

a set of concepts that are exclusive to said identified context and unrelated to said entity.

Assignments (4)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044127/0735 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2011
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: GOOGLE INC.
Reel/Frame 026664/0866 →
CORRECTIVE ASSIGNMENT REEL 014516/FRAME 0569 CORRECTING INVENTOR KRISHNA KUMMAMURU'S NAME Recorded Oct 24, 2006
From: KUMMAMURU, KRISHNA; KRISHNAPURAM, RAGHURAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 018430/0900 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2003
From: KUMMANURU, KRISHNA; KRISHNAPURAM, RAGHURAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 014516/0569 →
Continuity (1)
Related Publication 20050060170A1 · Mar 17, 2005