IP Library › Granted Patent US 10,496,683
Granted Patent B2
US 10,496,683 · App. 16/183,877 · Granted Dec 3, 2019

Automatically linking text to concepts in a knowledge base

Inventors: Michele M. Franceschini (White Plains, NY); Luis A. Lastras-Montano (Cortlandt Manor, NY); Livio B. Soares (New York, NY); Mark N. Wegman (Ossining, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/313G06F16/3334G06F16/3346G06F17/2235G06N5/003G06N5/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,496,683
App. No.
16/183,877
Granted
Dec 3, 2019
Kind
B2
Abstract

According to an aspect, automatically linking text to concepts in a knowledge base using differential analysis includes receiving a text string and selecting, based on contents of the text string, a plurality of data sources that correspond to concepts in the knowledge base. In a further aspect, automatically linking the text to the concepts includes calculating, for each of the selected data sources, a probability that the text string is output by a language model built using the selected data source, calculating a probability that the text string is output by a generic language model, calculating link confidence scores for each concept based on a differential analysis of the probabilities, and creating a link from the text string to one of the concepts in the knowledge base. The creating is based on a link confidence score of the concept being more than a threshold value away from a prescribed threshold.

Claims (45)

1. A computer program product for automatically linking text to concepts in a knowledge base, the computer program product comprising:

a non-transitory storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method comprising:

receiving, at a computer system, a plurality of text strings;

building a conceptual index that links the text strings to the knowledge base, the building comprising for each of the text strings:

creating an entry in the conceptual index that includes a link between the text string and one of the concepts in the knowledge base, the creating based at least in part on a link confidence score of the concept being more than a first threshold value away from a prescribed threshold;

generating a conceptual inverted index based on entries in the conceptual index, each entry of the conceptual inverted index corresponding to a different one of the concepts in the knowledge base and comprising pointers to at least a subset of text strings of the plurality of text strings linked to the concept in the conceptual index;

receiving a query from an agent external to the computer system, the query specifying a concept in the knowledge base; and

processing the query by the computer system, the processing comprising searching the conceptual inverted index for the concept specified in the query and returning a pointer to a text string in an entry of the conceptual inverted index corresponding to the concept.

2. The computer program product of claim 1 , wherein the method further comprises for each of the text strings:

selecting a plurality of data sources that correspond to at least a subset of the concepts in the knowledge base, the selecting based on contents of the text string; and

calculating link confidence scores for each of the concepts based on a differential analysis of a probability that the text string is output by a language model built using a data source of the plurality of data sources and a probability that the text string is output by a generic language model that is not related to any particular concept in the knowledge base.

3. The computer program product of claim 2 , wherein the method further comprises for each of text strings:

calculating, for each of the selected data sources, the probability that the text string is output by a language model built using the selected data source.

4. The computer program product of claim 2 , wherein the method further comprises for each of the text strings:

calculating the probability that the text string is output by a generic language model that is not related to any particular concept in the knowledge base.

5. The computer program product of claim 2 , wherein the differential analysis compares at least one of:

the probability that the text string is output by a language model built using a data source of the plurality of data sources to the probability that the text string is output by the generic language model; and

the probability that the text string is output by a language model built using a data source of the plurality of data sources to a probability that the text string is output by a language model built using a competing data source.

6. The computer program product of claim 2 , wherein the generic language model is derived from a generic data source not specific to any of the concepts in the knowledge base.

7. The computer program product of claim 1 , wherein the text string is linked to a second one of the concepts in the knowledge base.

8. The computer program product of claim 1 , wherein the link applies to a subset of the text string and the subset is indicated in the link, and words in the subset are not consecutive in the text string.

9. The computer program product of claim 1 , wherein each of the plurality of text strings corresponds to a person and includes a description of their skills, and each of the concepts in the knowledge base is an area of expertise, wherein the query result provides the external agent with a list of possible people having a specified area of expertise.

10. The computer program product of claim 1 , wherein each of the text strings have a version number and the method further comprises periodically, by a garbage collection mechanism, deleting links in the conceptual index to text strings having invalid version numbers.

11. A system for automatically linking text to concepts in a knowledge base, the system comprising:

a memory having computer readable computer instructions; and

at least a processor for executing the computer readable instructions, the computer readable instructions including:

receiving a plurality of text strings;

building a conceptual index that links the text strings to the knowledge base, the building comprising for each of the text strings:

creating an entry in the conceptual index that includes a link between the text string and one of the concepts in the knowledge base, the creating based at least in part on a link confidence score of the concept being more than a first threshold value away from a prescribed threshold;

generating a conceptual inverted index based on entries in the conceptual index, each entry of the conceptual inverted index corresponding to a different one of the concepts in the knowledge base and comprising pointers to at least a subset of text strings of the plurality of text strings linked to the concept in the conceptual index;

receiving a query from an agent external to the computer system, the query specifying a concept in the knowledge base; and

processing the query by the computer system, the processing comprising searching the conceptual inverted index for the concept specified in the query and returning a pointer to a text string in an entry of the conceptual inverted index corresponding to the concept.

12. The system of claim 11 , wherein the computer readable instructions further comprise for each of the text strings:

selecting a plurality of data sources that correspond to at least a subset of the concepts in the knowledge base, the selecting based on contents of the text string; and

calculating link confidence scores for each of the concepts based on a differential analysis of a probability that the text string is output by a language model built using a data source of the plurality of data sources and a probability that the text string is output by a generic language model that is not related to any particular concept in the knowledge base.

13. The system of claim 12 , wherein the computer readable instructions further comprise for each of the text strings:

calculating, for each of the selected data sources, the probability that the text string is output by a language model built using the selected data source.

14. The system of claim 12 , wherein the computer readable instructions further comprise for each of the text strings;

calculating the probability that the text string is output by a generic language model that is not related to any particular concept in the knowledge base.

15. The system of claim 12 , wherein the differential analysis compares at least one of:

the probability that the text string is output by a language model built using a data source of the plurality of data sources to the probability that the text string is output by the generic language model; and

the probability that the text string is output by a language model built using a data source of the plurality of data sources to a probability that the text string is output by a language model built using a competing data source.

16. The system of claim 12 , wherein the generic language model is derived from a generic data source not specific to any of the concepts in the knowledge base.

17. The system of claim 11 , wherein each of the plurality of text strings corresponds to a person and includes a description of their skills, and each of the concepts in the knowledge base is an area of expertise, wherein the query result provides the external agent with a list of possible people having a specified area of expertise.

18. The system of claim 11 , wherein each of the text strings have a version number and the method further comprises periodically, by a garbage collection mechanism, deleting links in the conceptual index to text strings having invalid version numbers.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ZIP CODE FOR RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 047451 FRAME: 0858. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 4, 2018
From: FRANCESCHINI, MICHELE M.; LASTRAS-MONTANO, LUIS A.; SOARES, LIVIO B.; WEGMAN, MARK N.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047714/0095 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2018
From: FRANCESCHINI, MICHELE M.; LASTRAS-MONTANO, LUIS A.; SOARES, LIVIO B.; WEGMAN, MARK N.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047451/0858 →
Continuity (2)
Continuation 14330381 · Jul 14, 2014
Related Publication 20190073414A1 · Mar 7, 2019
Cited By (4)
US 12,353,493 US 12,381,923 US 12,566,804 US 12,737,789