IP Library Granted Patent US 10,621,219
Granted Patent B2
US 10,621,219 · App. 15/430,185 · Granted Apr 14, 2020

Techniques for determining a semantic distance between subjects

Inventors: Jennifer Ann English (Redwood City, CA); Malous Melissa Kossarian (San Mateo, CA); Charles E. McManis, Jr. (Santa Clara, CA); Douglas A. Smith (Santa Clara, CA)
Assignee: International Business Machines Corporation
G06F16/3344G06F16/313G06F16/3347G06F17/2785G06F17/2818G06N7/005G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,621,219
App. No.
15/430,185
Granted
Apr 14, 2020
Kind
B2
Abstract

A technique for calculating a semantic distance between subjects includes performing a mathematical operation between each of one or more first topic vectors and each of one or more second topic vectors to generate respective strength values. The first topic vectors are associated with respective first topics of a first subject, the second topic vectors are associated with respective second topics of a second subject, and the respective strength values are indicative of a relative closeness between associated ones of the first and second topics. Relevant ones of the respective strength values are summed to provide an overall strength value between the first subject and the second subject. A semantic distance between the first subject and the second subject is determined based on the overall strength value.

Claims (26)

1. A computer program product configured to determine a semantic distance between subjects, the computer program product comprising:

a computer-readable storage device; and

computer-readable program code embodied on the computer-readable storage device, wherein the computer-readable program code, when executed by a data processing system, causes the data processing system to:

perform a mathematical operation between each of one or more first topic vectors and each of one or more second topic vectors to generate respective strength values, wherein the first topic vectors are associated with respective first topics of a first subject, the second topic vectors are associated with respective second topics of a second subject, and the respective strength values are indicative of a relative closeness between associated ones of the first and second topics;

sum relevant ones of the respective strength values to provide an overall strength value between the first subject and the second subject;

determine a semantic distance between the first subject and the second subject based on the overall strength value; and

utilize both the first and second subjects in a search for information related to the first subject in response to the semantic distance being within a threshold distance value to improve the operation of the data processing system in answering a question about the first subject.

2. The computer program product of claim 1 , wherein the computer-readable program code, when executed by the data processing system, further causes the data processing system to:

generate the first topic vectors and the second topic vectors based on a statistical model analysis.

3. The computer program product of claim 2 , wherein the mathematical operation is a dot product operation.

4. The computer program product of claim 2 , wherein the statistical model analysis is a latent Dirichlet allocation (LDA) analysis.

5. The computer program product of claim 1 , wherein a number of the first topics for the first subject is determined by taking the square root of a number of documents associated with the first subject divided by two.

6. The computer program product of claim 5 , wherein the documents are located using a Web search.

7. The computer program product of claim 1 , wherein the semantic distance between the first subject and the second subject is determined by taking an inverse of the overall strength value.

8. The computer program product of claim 1 , wherein each of the first and second topic vectors have an associated word and an associated strength that is normalized to one.

9. A data processing system, comprising:

a cache memory; and

a processor coupled to the cache memory, wherein the processor is configured to:

perform a mathematical operation between each of one or more first topic vectors and each of one or more second topic vectors to generate respective strength values, wherein the first topic vectors are associated with respective first topics of a first subject, the second topic vectors are associated with respective second topics of a second subject, and the respective strength values are indicative of a relative closeness between associated ones of the first and second topics;

sum relevant ones of the respective strength values to provide an overall strength value between the first subject and the second subject;

determine a semantic distance between the first subject and the second subject based on the overall strength value; and

utilize both the first and second subjects in a search for information related to the first subject in response to the semantic distance being within a threshold distance value to improve the operation of the data processing system in answering a question about the first subject.

10. The data processing system of claim 9 , wherein the processor is further configured to:

generate the first topic vectors and the second topic vectors based on a statistical model analysis.

11. The data processing system of claim 10 , wherein the mathematical operation is a dot product operation, the statistical model analysis is a latent Dirichlet allocation (LDA) analysis, and a number of the first topics for the first subject is determined by taking the square root of a number of documents associated with the first subject divided by two.

12. The data processing system of claim 9 , wherein the semantic distance between the first subject and the second subject is determined by taking an inverse of the overall strength value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2017
From: ENGLISH, JENNIFER ANN; KOSSARIAN, MALOUS MELISSA; MCMANIS, CHARLES E., JR.; SMITH, DOUGLAS A.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 041229/0741 →
Continuity (1)
Related Publication 20180232437A1 · Aug 16, 2018