IP Library › Granted Patent US 10,733,224
Granted Patent B2
US 10,733,224 · App. 15/426,244 · Granted Aug 4, 2020

Automatic corpus selection and halting condition detection for semantic asset expansion

Inventors: Alfredo Alba (Morgan Hill, CA); Clemens Drews (San Jose, CA); Daniel F. Gruhl (San Jose, CA); Linda H. Kato (San Jose, CA); Neal R. Lewis (San Jose, CA); Pablo N. Mendes (San Jose, CA); Meenakshi Nagarajan (San Jose, CA)
Assignee: International Business Machines Corporation
G06F16/36G06F7/08G06F16/353
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,733,224
App. No.
15/426,244
Granted
Aug 4, 2020
Kind
B2
Abstract

A mechanism is provided in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions executed by the at least one processor to cause the at least one processor to implement an automated lexicon expansion for an identified corpus. For a selected corpus in a set of corpora, the mechanism determines an estimated number of new terms in the selected corpus that are not in the lexicon based on a frequency count known terms in the selected corpus. Responsive to the estimated number of new terms in the selected corpus being greater than a threshold, the mechanism performs lexicon expansion using the selected corpus to form an expanded lexicon. Responsive to the estimated number of new terms in the selected corpus not being greater than the threshold, the mechanism halts lexicon expansion.

Claims (50)

1. A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a computing device, causes the computing device to implement an automated lexicon expansion for an identified corpus, wherein the computer readable program causes the computing device to:

for each given corpus in the set of corpora, determine an estimated number of new terms in the given corpus that are not in the lexicon based on a frequency count known terms in the given corpus; and

sort the set of corpora by estimated number of new terms to form a sorted set of corpora;

select a corpus having a highest estimated number of new terms as a selected corpus for lexicon expansion;

determine an estimated number of new terms in the selected corpus that are not in the lexicon based on a frequency count of known terms in the selected corpus, wherein determining the estimated number of new terms in the selected corpus comprises:

identifying a set of known terms in the lexicon;

for each known term in the lexicon, identifying an associated frequency of occurrence of the known term in the selected corpus;

sorting the set of known terms based on the associated frequency of occurrence thereby forming a sorted set of known terms;

fitting a line to a portion of the sorted set of known terms;

determining an X-axis intercept of the line;

determining the estimated number of new terms in the selected corpus based on the X-axis intercept of the line; and

subtracting a number of terms in the sorted set of known terms from the X-axis intercept;

responsive to the estimated number of new terms in the selected corpus being greater than a threshold, perform lexicon expansion using the selected corpus to form an expanded lexicon; and

responsive to the estimated number of new terms in the selected corpus not being greater than the threshold, halt lexicon expansion.

2. The computer program product of claim 1 , wherein the computer readable program further causes the computing device to:

responsive to performing lexicon expansion using the selected corpus, repeat determining the estimated number of new terms in each corpus that are not in the lexicon and sorting the set of corpora by estimated number of new terms to form a new sorted set of corpora.

3. The computer program product of claim 2 , wherein the computer readable program further causes the computing device to select a corpus having a highest estimated number of new terms in the new sorted set of corpora as the selected corpus for a next lexicon expansion.

4. The computer program product of claim 1 , wherein determining the estimated number of new terms in the given corpus comprises:

identifying a set of known terms in the lexicon;

for each known term in the lexicon, identifying an associated frequency of occurrence of the known term in the given corpus;

sorting the set of known terms based on the associated frequency of occurrence thereby forming a sorted set of known terms;

fitting a line to a portion of the sorted set of known terms;

determining an X-axis intercept of the line; and

determining the estimated number of new terms in the given corpus based on the X-axis intercept of the line.

5. An apparatus comprising:

a processor; and

a memory coupled to the processor, wherein the memory comprises instructions which, when executed by the processor, cause the processor to implement an automated lexicon expansion for an identified corpus, wherein the instructions cause the processor to:

for each given corpus in the set of corpora, determine an estimated number of new terms in the given corpus that are not in the lexicon based on a frequency count known terms in the given corpus; and

sort the set of corpora by estimated number of new terms to form a sorted set of corpora;

select a corpus having a highest estimated number of new terms as a selected corpus for lexicon expansion;

determine an estimated number of new terms in the selected corpus that are not in the lexicon based on a frequency count of known terms in the selected corpus, wherein determining the estimated number of new terms in the selected corpus comprises:

identifying a set of known terms in the lexicon;

for each known term in the lexicon, identifying an associated frequency of occurrence of the known term in the selected corpus;

sorting the set of known terms based on the associated frequency of occurrence thereby forming a sorted set of known terms;

fitting a line to a portion of the sorted set of known terms;

determining an X-axis intercept of the line;

determining the estimated number of new terms in the selected corpus based on the X-axis intercept of the line; and

subtracting a number of terms in the sorted set of known terms from the X-axis intercept;

responsive to the estimated number of new terms in the selected corpus being greater than a threshold, perform lexicon expansion using the selected corpus to form an expanded lexicon; and

responsive to the estimated number of new terms in the selected corpus not being greater than the threshold, halt lexicon expansion.

6. The apparatus of claim 5 , wherein the instructions further cause the processor to:

responsive to performing lexicon expansion using the selected corpus, repeat determining the estimated number of new terms in each corpus that are not in the lexicon and sorting the set of corpora by estimated number of new terms to form a new sorted set of corpora.

7. The apparatus of claim 6 , wherein the instructions further cause the processor to select a corpus having a highest estimated number of new terms in the new sorted set of corpora as the selected corpus for a next lexicon expansion.

8. The apparatus of claim 5 , wherein determining the estimated number of new terms in the given corpus comprises:

identifying a set of known terms in the lexicon;

for each known term in the lexicon, identifying an associated frequency of occurrence of the known term in the given corpus;

sorting the set of known terms based on the associated frequency of occurrence thereby forming a sorted set of known terms;

fitting a line to a portion of the sorted set of known terms;

determining an X-axis intercept of the line; and

determining the estimated number of new terms in the given corpus based on the X-axis intercept of the line.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2017
From: ALBA, ALFREDO; DREWS, CLEMENS; GRUHL, DANIEL F.; KATO, LINDA H.; LEWIS, NEAL R.; MENDES, PABLO N.; NAGARAJAN, MEENAKSHI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 041647/0334 →
Continuity (1)
Related Publication 20180225373A1 · Aug 9, 2018