IP Library Granted Patent US 10,162,883
Granted Patent B2
US 10,162,883 · App. 14/657,343 · Granted Dec 25, 2018

Automatically linking text to concepts in a knowledge base

Inventors: Michele M. Franceschini (White Plains, NY); Luis A. Lastras-Montano (Cortlandt Manor, NY); Livio B. Soares (New York, NY); Mark N. Wegman (Ossining, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/30616G06F17/2235G06F17/30663G06F17/30687G06N5/003G06N5/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,162,883
App. No.
14/657,343
Granted
Dec 25, 2018
Kind
B2
Abstract

According to an aspect, automatically linking text to concepts in a knowledge base using differential analysis includes receiving a text string and selecting, based on contents of the text string, a plurality of data sources that correspond to concepts in the knowledge base. In a further aspect, automatically linking the text to the concepts includes calculating, for each of the selected data sources, a probability that the text string is output by a language model built using the selected data source, calculating a probability that the text string is output by a generic language model, calculating link confidence scores for each concept based on a differential analysis of the probabilities, and creating a link from the text string to one of the concepts in the knowledge base. The creating is based on a link confidence score of the concept being more than a threshold value away from a prescribed threshold.

Claims (23)

1. A method for automatically linking text to concepts in a knowledge base using a differential analysis, the method comprising:

receiving, at a computer system, a plurality of text strings;

building a conceptual index that links the text strings to the knowledge base, the building comprising for each of the text strings:

selecting a plurality of data sources that correspond to at least a subset of the concepts in the knowledge base, the selecting based on contents of the text string;

calculating, for each of the selected data sources, a probability that the text string is output by a language model built using the selected data source;

calculating a probability that the text string is output by a generic language model that is not related to any particular concept in the knowledge base;

calculating link confidence scores for each of the at least a subset of the concepts based on a differential analysis of the probabilities; and

creating an entry in the conceptual index that includes a link between the text string and one of the concepts in the knowledge base, the creating based at least in part on a link confidence score of the concept being more than a first threshold value away from a prescribed threshold;

generating a conceptual inverted index based on entries in the conceptual index, each entry of the conceptual inverted index corresponding to a different one of the concepts in the knowledge base and comprising pointers to at least a subset of text strings of the plurality of text strings linked to the concept in the conceptual index;

receiving a query from an agent external to the computer system, the query specifying a concept in the knowledge base;

processing the query by the computer system, the processing comprising searching the conceptual inverted index for the concept specified in the query and returning a pointer to a text string in an entry of the conceptual inverted index corresponding to the concept; and

returning a set of documents to the external agent through the use of the conceptual inverted index, based on the received query.

2. The method of claim 1 , wherein the differential analysis compares the probability that the text string is output by a language model built using a data source to the probability that the text string is output by the generic language model.

3. The method of claim 1 , wherein the differential analysis compares the probability that the text string is output by a language model built using a data source to a probability that the text string is output by a language model built using a competing data source.

4. The method of claim 1 , wherein the generic language model is derived from a generic data source not specific to any of the concepts in the knowledge base.

5. The method of claim 1 , wherein the calculating link confidence scores includes comparing the probabilities to a probability that the text string is contained in a generic data source that is not associated with any of the concepts in the knowledge base.

6. The method of claim 1 , wherein the text string is linked to a second one of the concepts in the knowledge base.

7. The method of claim 1 , wherein the link applies to a subset of the text string and the subset is indicated in the link.

8. The method of claim 7 , wherein words in the subset are not consecutive in the text string.

9. The method of claim 1 , wherein the text string is one of a collection of words, a sentence, a paragraph, and a whole document.

10. The method of claim 1 , wherein each of the selected data sources includes one or more collection of names for the corresponding concept, a description for the corresponding concept, sentences referring to the corresponding concept, and paragraphs referring to the corresponding concept.

11. The method of claim 1 , wherein each of the plurality of text strings corresponds to a person and includes a description of their skills, and each of the concepts in the knowledge base is an area of expertise, wherein the query result provides the external agent with a list of possible people having a specified area of expertise.

12. The method of claim 1 , wherein each of the text strings have a version number and the method further comprises periodically, by a garbage collection mechanism, deleting links in the conceptual index to text strings having invalid version numbers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2015
From: FRANCESCHINI, MICHELE M.; LASTRAS-MONTANO, LUIS A.; SOARES, LIVIO B.; WEGMAN, MARK N.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 035163/0589 →
Continuity (2)
Continuation 14330381 · Jul 14, 2014
Related Publication 20160012336A1 · Jan 14, 2016
Cited By (2)
US 12,406,591 US 12,566,804