IP Library › Granted Patent US 10,331,659
Granted Patent B2
US 10,331,659 · App. 15/257,266 · Granted Jun 25, 2019

Automatic detection and cleansing of erroneous concepts in an aggregated knowledge base

Inventors: Shilpi Ahuja (San Jose, CA); Sheng Hua Bao (San Jose, CA); Rashmi Gangadharaiah (San Jose, CA)
Assignee: International Business Machines Corporation
G06F16/2365G06F16/215G06F16/285G06F16/9024G06F17/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,331,659
App. No.
15/257,266
Filed
Sep 6, 2016
Granted
Jun 25, 2019
Kind
B2
Art Unit
2156
USPC
707/690
Abstract

A mechanism is provided for automatically detecting and cleansing erroneous concepts in an aggregated knowledge base. A graph data structure representing the concept present in a portion of the natural language content is generated. The graph data structure is analyzed to determine whether or not the graph data structure comprises one or more concept conflicts in association with a set of nodes in the graph data structure, the one or more concept conflicts are associated with the set of nodes if two or more nodes represent separate and distinct concepts. Responsive to determining that there are one or more concept conflicts due to there being two or more nodes representing separate and distinct concepts, the two or more nodes are split into separate distinct concepts within the knowledge base.

Claims (61)

1. A method, in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions executed by the at least one processor to cause the at least one processor to implement natural language processing (NLP) system, wherein the method comprises:

receiving, by the NLP system, a portion of natural language content related to a selected concept from a knowledge base;

generating, by the NLP system, a graph data structure representing the concept present in the portion of the natural language content, wherein nodes of the graph data structure comprise a first node representing a name of the concept and a one or more other nodes representing synonyms associated with the first node and wherein the graph data structure further indicates relationships between the first node and the one or more other nodes based on a similarity measure of the other nodes to the first node;

analyzing, by the NLP system, the graph data structure to determine whether or not the graph data structure comprises one or more concept conflicts in association with a set of nodes in the graph data structure, wherein the one or more concept conflicts are associated with the set of nodes responsive to two or more nodes representing separate and distinct concepts; and

responsive to determining that there are one or more concept conflicts due to there being two or more nodes representing separate and distinct concepts, splitting, by the NLP system, the two or more nodes into separate distinct concepts within the knowledge base.

2. The method of claim 1 , wherein analyzing the graph data structure comprises:

automatically generating, by the NLP system, a first cluster of nodes including the first node and one or more other nodes associated with the first node; and

storing, by the NLP system, in the knowledge base, the first cluster of nodes in association with a first separate distinct concept.

3. The method of claim 2 , wherein automatically generating the cluster comprises:

assigning, by the NLP system, other nodes, in the one or more other nodes, to the first cluster of nodes associated with the first node based on the similarity measure of the other nodes to the first node.

4. The method of claim 3 , wherein the similarity measure is a cosine similarity or a Euclidean distance.

5. The method of claim 1 , wherein analyzing the graph data structure comprises:

identifying, by the NLP system, a second node in the graph data structure, wherein the second node is a node that has one or more other nodes connected to the second node by an edge in the graph and has a relatively high visiting probability;

automatically generating, by the NLP system, a second cluster of nodes including the second node and one or more other nodes associated with the second node; and

storing, by the NLP system, in the knowledge base, the second cluster of nodes in association with a second separate distinct concept.

6. The method of claim 5 , wherein identifying the second canonical concept node in the graph comprises:

calculating, by the NLP system, for each node in the graph, a visiting probability indicating a probability that the concept associated with the node will be used by a cognitive system to perform a cognitive operation; and

selecting, by the NLP system, a node having a largest visiting probability as the first node.

7. The method of claim 5 , wherein automatically generating the cluster comprises:

assigning, by the NLP system, other nodes, in the one or more other nodes, to the second cluster of nodes associated with the second node based on a distance of the other nodes to the second node.

8. The method of claim 7 , wherein the similarity measure is a cosine similarity or a Euclidean distance.

9. A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a computing device, causes the computing device to implement a natural language processing (NLP) system which operates to:

receive a portion of natural language content related to a selected concept from a knowledge base;

generate a graph data structure representing the concept present in the portion of the natural language content, wherein nodes of the graph data structure comprise a first node representing a name of the concept and a one or more other nodes representing synonyms associated with the first node and wherein the graph data structure further indicates relationships between the first node and the one or more other nodes based on a similarity measure of the other nodes to the first node;

analyze the graph data structure to determine whether or not the graph data structure comprises one or more concept conflicts in association with a set of nodes in the graph data structure, wherein the one or more concept conflicts are associated with the set of nodes responsive to two or more nodes representing separate and distinct concepts; and

responsive to determining that there are one or more concept conflicts due to there being two or more nodes representing separate and distinct concepts, split the two or more nodes into separate distinct concepts within the knowledge base.

10. The computer program product of claim 9 , wherein the computer readable program to analyze the graph data structure further causes the computing device to implement the NLP system which operates to:

automatically generate a first cluster of nodes including the first node and one or more other nodes associated with the first node; and

store, in the knowledge base, the first cluster of nodes in association with a first separate distinct concept.

11. The computer program product of claim 10 , wherein the computer readable program to automatically generate the cluster further causes the computing device to implement the NLP system which operates to:

assign other nodes, in the one or more other nodes, to the first cluster of nodes associated with the first node based on the similarity measure of the other nodes to the first node.

12. The computer program product of claim 9 , wherein the computer readable program to analyze the graph data structure further causes the computing device to implement the NLP system which operates to:

identify a second node in the graph data structure, wherein the second node is a node that has one or more other nodes connected to the second node by an edge in the graph and has a relatively high visiting probability;

automatically generate a second cluster of nodes including the second node and one or more other nodes associated with the second node; and

store in the knowledge base, the second cluster of nodes in association with a second separate distinct concept.

13. The computer program product of claim 12 , wherein the computer readable program to identify the second canonical concept node in the graph further causes the computing device to implement the NLP system which operates to:

calculate for each node in the graph, a visiting probability indicating a probability that the concept associated with the node will be used by a cognitive system to perform a cognitive operation; and

select a node having a largest visiting probability as the first node.

14. The computer program product of claim 12 , wherein the computer readable program to automatically generate the cluster further causes the computing device to implement the NLP system which operates to:

assign other nodes, in the one or more other nodes, to the second cluster of nodes associated with the second node based on a distance of the other nodes to the second node.

15. An apparatus comprising:

a processor; and

a memory coupled to the processor, wherein the memory comprises instructions which, when executed by the processor, cause the processor to implement a natural language processing (NLP) system which operates to:

receive a portion of natural language content related to a selected concept from a knowledge base;

generate a graph data structure representing the concept present in the portion of the natural language content, wherein nodes of the graph data structure comprise a first node representing a name of the concept and a one or more other nodes representing synonyms associated with the first node and wherein the graph data structure further indicates relationships between the first node and the one or more other nodes based on a similarity measure of the other nodes to the first node;

analyze the graph data structure to determine whether or not the graph data structure comprises one or more concept conflicts in association with a set of nodes in the graph data structure, wherein the one or more concept conflicts are associated with the set of nodes responsive to two or more nodes representing separate and distinct concepts; and

responsive to determining that there are one or more concept conflicts due to there being two or more nodes representing separate and distinct concepts, split the two or more nodes into separate distinct concepts within the knowledge base.

16. The apparatus of claim 15 , wherein the instructions to analyze the graph data structure further cause the processor to implement the NLP system which operates to:

automatically generate a first cluster of nodes including the first node and one or more other nodes associated with the first node; and

store, in the knowledge base, the first cluster of nodes in association with a first separate distinct concept.

17. The apparatus of claim 16 , wherein the instructions to automatically generate the cluster further cause the processor to implement the NLP system which operates to:

assign other nodes, in the one or more other nodes, to the first cluster of nodes associated with the first node based on the similarity measure of the other nodes to the first node.

18. The apparatus of claim 15 , wherein the instructions to analyze the graph data structure further cause the processor to implement the NLP system which operates to:

identify a second node in the graph data structure, wherein the second node is a node that has one or more other nodes connected to the second node by an edge in the graph and has a relatively high visiting probability;

automatically generate a second cluster of nodes including the second node and one or more other nodes associated with the second node; and

store in the knowledge base, the second cluster of nodes in association with a second separate distinct concept.

19. The apparatus of claim 18 , wherein the instructions to identify the second canonical concept node in the graph further cause the processor to implement the NLP system which operates to:

calculate for each node in the graph, a visiting probability indicating a probability that the concept associated with the node will be used by a cognitive system to perform a cognitive operation; and

select a node having a largest visiting probability as the first node.

20. The apparatus of claim 18 , wherein the instructions to automatically generate the cluster further cause the processor to implement the NLP system which operates to:

assign other nodes, in the one or more other nodes, to the second cluster of nodes associated with the second node based on a distance of the other nodes to the second node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2016
From: AHUJA, SHILPI; BAO, SHENG HUA; GANGADHARAIAH, RASHMI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 039639/0282 →
Continuity (1)
Related Publication 20180067981A1 · Mar 8, 2018
Cited By (95)
US 12,206,696 US 12,244,621 US 12,267,345 US 12,309,185 US 12,323,449 US 12,335,286 US 12,335,348 US 12,341,797 US 12,348,545 US 12,355,626 US 12,355,787 US 12,355,793 US 12,363,148 US 12,368,745 US 12,368,746 US 12,368,747 US 12,375,573 US 12,395,573 US 12,401,669 US 12,405,849 US 12,407,701 US 12,407,702 US 12,418,552 US 12,418,555 US 12,425,428 US 12,425,430 US 12,445,474 US 12,452,279 US 12,457,231 US 12,463,995 US 12,463,996 US 12,463,997 US 12,464,003 US 12,470,577 US 12,470,578 US 12,483,576 US 12,489,770 US 12,495,052 US 12,500,910 US 12,500,911 US 12,500,912 US 12,505,126 US 12,506,762 US 12,513,221 US 12,537,836 US 12,537,837 US 12,537,839 US 12,537,840 US 12,537,884 US 12,549,575 US 12,549,577 US 12,556,548 US 12,556,559 US 12,563,060 US 12,563,064 US 12,563,071 US 12,563,072 US 12,580,934 US 12,580,935 US 12,580,936 US 12,580,937 US 12,587,553 US 12,592,950 US 12,598,205 US 12,613,930 US 12,615,271 US 12,621,324 US 12,621,329 US 12,627,686 US 12,627,687 US 12,627,690 US 12,634,312 US 12,634,376 US 12,652,302 US 12,659,325 US 12,659,326 US 12,659,327 US 12,659,333 US 12,676,874 US 12,689,638 US 12,689,640 US 12,695,768 US 12,706,932 US 12,706,933 US 12,712,897 US 12,719,896 US 12,726,495 US 12,730,899 US 12,739,266 US 12,739,267 US 12,743,259 US 12,744,800 US 12,744,802 US 12,750,382 US 12,750,383