IP Library Granted Patent US 9,239,830
Granted Patent B2
US 9,239,830 · App. 13/755,062 · Granted Jan 19, 2016

System and method for building relationship hierarchy

Inventors: Sridhar Gopalakrishnan (Bangalore, IN); Sujatha Raviprasad Upadhyaya (Bangalore, IN)
Assignee: XURMO TECHNOLOGIES PVT. LTD.
G06F17/28G06F17/278G06F17/30539G06F17/30604G06F17/30705G06F17/30734G06N5/02G06N99/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,239,830
App. No.
13/755,062
Granted
Jan 19, 2016
Kind
B2
Abstract

The various embodiments herein provide a method and system for building a relationship hierarchy from big data. The method comprises extracting a plurality of relationships defined between entities from a big data, building relationship recognition models adapted to identify different forms of generic relationships, resolving the relationships by grouping the similar relationships together and separating the relationships which are syntactically and semantically dissimilar and reconciling the resolved relationships to build the relationship hierarchy. The relationship hierarchy comprises groups and subgroups of relationships created based on generic relationship similarity based on a contextual aspect and a specialization aspect using a Language and Domain model. The method of extraction of plurality of relationships from the unstructured data is a self-learning process which uses open information extraction techniques for learning new relationships.

Claims (31)

1. A method for building a relationship hierarchy comprises:

obtaining a plurality of relationships from training samples drawn from big data through open information extraction models;

grouping the obtained relationships under generic relationships based on similarity;

creating sub-groups of the grouped relationships;

tagging the training sample with relationship tags to the occurrences of relationships in training samples;

building relationship extraction models from the tagged training samples using one or more machine learning approaches; and

providing new training samples from the big data;

extracting new relationships using natural language processing based open information extraction models;

reconciling the new relationships under at least one of the existing generic relationships or a new generic relationship;

annotating training samples with the relationships extracted through the open information extraction models;

create one or more new relationship extraction models;

populating the relationship extraction models with new relationship extraction models that can successfully identify new relationships; and

passing the text data through new relationship recognition models to identify new relationships from data using one or more machine Learning models;

extract relationships from the big data using the relationship extraction models;

building relationship recognition models adapted to identify different forms of generic relationships;

resolving the relationships by grouping the similar relationships together and separating the relationships which are syntactically and semantically dissimilar; and reconciling the resolved relationships to build the relationship hierarchy, wherein the relationship hierarchy comprises groups and subgroups of relationships created based on generic relationship similarity based on a contextual aspect and a specialization aspect using a Language and Domain model.

2. The method of claim 1 , wherein the big data comprises structured, unstructured and semi-structured data.

3. The method of claim 2 , wherein extracting the plurality of relationships from the structured data comprises identifying relationships between one or more entities in a structured data, wherein the plurality of relationships defined between the entities comprises:

a ‘hasAttribute’ relationship shared by a table entity with each of non-key attribute entity;

a ‘hasPrimarykey’ relationship shared by a table entity with the primary key attribute;

an ‘IsInstanceOf’ relationship shared by a value entity with a corresponding attribute entity;

a ‘belongsTo’ relationship shared by the table entity with a database entity;

a ‘hasAttributeName’ relationship shared by the entity with respect to a primary key attribute with other attributes;

a ‘hasAttributeName’ relationship shared by each entity with respect to a primary key value and with respect to other values in a topple, where attributeName being the name of the respective attribute; and

a collection of relationships ‘hasAttributeName’ with respect to the primary key value entity which forms an entity by itself and shares an ‘IsInstanceOf’ relationship with the table entity.

4. The method of claim 2 , wherein extracting the plurality of relationships from the unstructured data is a self-learning process.

5. The method of claim 1 , wherein resolving the plurality of relationships comprises grouping the similar relationships, where one or more verb based relationships are brought together on the basis of semantic distance of the words.

6. The method of claim 1 , wherein resolving the plurality of relationships comprises separating the relationships which are syntactically and semantically dissimilar, where the relationships are resolved by capturing specialization aspect of the relationships.

7. The method of claim 1 , wherein the resolving the plurality of relationships comprises at least one of a: word sense disambiguation technique, and contextual resolution technique.

8. The method of claim 1 , wherein the relationship hierarchy is built using the language and domain models which comprise at least one of: language repositories, domain ontology, and knowledge repositories in combination with natural language processing techniques.

9. The method of claim 1 , wherein extracting the plurality of relationships from semi-structured data is a combination of extracting relationships form the structured data and unstructured data.

Continuity (1)
Related Publication 20140046877A1 · Feb 13, 2014