IP Library Granted Patent US 10,643,135
Granted Patent B2
US 10,643,135 · App. 15/242,821 · Granted May 5, 2020

Linkage prediction through similarity analysis

Inventors: Robert G. Farrell (Cornwall, NY); Achille Fokoue-Nkoutche (White Plains, NY); Oktie Hassanzadeh (Port Chester, NY); Mohammad Sadoghi Hamedani (Chappaqua, NY); Meinolf Sellmann (Cortlandt Manor, NY); Ping Zhang (White Plains, NY)
Assignee: International Business Machines Corporation
G06N5/022G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,643,135
App. No.
15/242,821
Granted
May 5, 2020
Kind
B2
Abstract

Methods, systems, and computer program products for linkage prediction through similarity analysis are provided herein. A computer-implemented method includes extracting multiple features from (i) one or more attributes of a set of source nodes within a knowledge graph and (ii) one or more attributes of a set of target nodes within the knowledge graph, wherein at least one extracted feature satisfies a designated complexity level; performing a similarity analysis across the at least one extracted feature by applying one or more similarity measures to the at least one extracted feature; predicting one or more sets of links between the source nodes and the target nodes based on the similarity analysis, wherein one or more sets of predicted links satisfy a pre-determined accuracy threshold; and outputting the one or more sets of predicted links to a user.

Claims (54)

1. A computer-implemented method, comprising:

extracting multiple features from (i) one or more attributes of a set of source nodes within a knowledge graph and (ii) one or more attributes of a set of target nodes within the knowledge graph, wherein at least one extracted feature satisfies a designated complexity level;

performing a similarity analysis across the at least one extracted feature by applying one or more similarity measures to the at least one extracted feature;

predicting one or more sets of links between the source nodes and the target nodes based on the similarity analysis, wherein one or more sets of predicted links satisfy a pre-determined accuracy threshold;

generating a similarity matrix comprising: (i) one row per entity in the set of source nodes, (ii) one column per entity in the set of target nodes, and (iii) a value in row i column j of the similarity matrix indicating a similarity score, based on based on said similarity analysis, between the entity represented in row i and the entity represented in row j;

building a probabilistic model based at least in part on the similarity matrix; and

calculating, using the probabilistic model, a confidence score for each of the one or more sets of predicted links; and

outputting (i) the one or more sets of predicted links and (ii) the calculated confidence scores to a user.

2. The computer-implemented method of claim 1 , comprising:

transforming a collection of raw data; and

integrating the transformed data into the knowledge graph.

3. The computer-implemented method of claim 1 , comprising:

identifying a source entity node type, a target entity node type, and a relationship type.

4. The computer-implemented method of claim 1 , comprising:

setting a complexity level to a value upon a determination that the one or more predicted links do not satisfy the pre-determined accuracy threshold.

5. The computer-implemented method of claim 1 , comprising:

generating a rationale for the one or more sets of predicted links, wherein the rationale comprises identification of one or more similar node pairs.

6. The computer-implemented method of claim 5 , comprising:

outputting a generated rationale to the user.

7. The computer-implemented method of claim 1 , comprising:

generating a ranking of the extracted features.

8. The computer-implemented method of claim 7 , comprising:

outputting the ranking of the extracted features to the user.

9. The computer-implemented method of claim 1 , wherein said extracting further comprises extracting the multiple features from one or more attributes of one or more nodes neighboring the set of source nodes within the knowledge graph.

10. The computer-implemented method of claim 1 , wherein said extracting further comprises extracting the multiple features from one or more attributes of one or more nodes neighboring the set of target nodes within the knowledge graph.

11. The computer-implemented method of claim 1 , wherein said extracting further comprises extracting the multiple features from one or more relationships between nodes within the knowledge graph.

12. The computer-implemented method of claim 1 , wherein the one or more similarity measures comprise one or more data-type based measures.

13. The computer-implemented method of claim 1 , wherein the one or more similarity measures comprise one or more graph-based relational similarity measures.

14. The computer-implemented method of claim 1 , wherein software is provided as a service in a cloud environment.

15. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a device to cause the device to:

extract multiple features from (i) one or more attributes of a set of source nodes within a knowledge graph and (ii) one or more attributes of a set of target nodes within the knowledge graph, wherein at least one extracted feature satisfies a designated complexity level;

perform a similarity analysis across the at least one extracted feature by applying one or more similarity measures to the at least one extracted feature;

predict one or more sets of links between the source nodes and the target nodes based on the similarity analysis, wherein one or more sets of predicted links satisfy a pre-determined accuracy threshold;

generate a similarity matrix comprising: (i) one row per entity in the set of source nodes, (ii) one column per entity in the set of target nodes, and (iii) a value in row i column j of the similarity matrix indicating a similarity score, based on based on said similarity analysis, between the entity represented in row i and the entity represented in row j;

build a probabilistic model based at least in part on the similarity matrix; and

calculate, using the probabilistic model, a confidence score for each of the one or more sets of predicted links; and

output (i) the one or more sets of predicted links and (ii) the calculated confidence scores to a user.

16. The computer program product of claim 15 , wherein the program instructions executable by a computing device further cause the computing device to:

transform a collection of raw data; and

integrate the transformed data into the knowledge graph.

17. A system comprising:

a memory; and

at least one processor coupled to the memory and configured for:

extracting multiple features from (i) one or more attributes of a set of source nodes within a knowledge graph and (ii) one or more attributes of a set of target nodes within the knowledge graph, wherein at least one extracted feature satisfies a designated complexity level;

performing a similarity analysis across the at least one extracted feature by applying one or more similarity measures to the at least one extracted feature;

predicting one or more sets of links between the source nodes and the target nodes based on the similarity analysis, wherein one or more sets of predicted links satisfy a pre-determined accuracy threshold;

generating a similarity matrix comprising: (i) one row per entity in the set of source nodes, (ii) one column per entity in the set of target nodes, and (iii) a value in row i column j of the similarity matrix indicating a similarity score, based on based on said similarity analysis, between the entity represented in row i and the entity represented in row j;

building a probabilistic model based at least in part on the similarity matrix; and

calculating, using the probabilistic model, a confidence score for each of the one or more sets of predicted links; and

outputting (i) the one or more sets of predicted links and (ii) the calculated confidence scores to a user.

18. The system of claim 17 , wherein the at least one processor is further configured for:

transforming a collection of raw data; and

integrating the transformed data into the knowledge graph.

19. The system of claim 17 , wherein software is provided as a service in a cloud environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2016
From: FARRELL, ROBERT G.; FOKOUE-NKOUTCHE, ACHILLE; HASSANZADEH, OKTIE; HAMEDANI, MOHAMMAD SADOGHI; SELLMANN, MEINOLF; ZHANG, PING
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 039497/0103 →
Continuity (1)
Related Publication 20180053096A1 · Feb 22, 2018