IP Library Granted Patent US 8,527,513
Granted Patent B2
US 8,527,513 · App. 12/869,400 · Granted Sep 3, 2013

Systems and methods for lexicon generation

Inventors: Paul Zhang (Centerville, OH); Harry Silver (Shaker Heights, OH)
Assignee: LexisNexis, a division of Reed Elsevier Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,527,513
App. No.
12/869,400
Granted
Sep 3, 2013
Kind
B2
Abstract

Disclosed herein are embodiments for lexicon generation. More specifically, at least one embodiment of a method includes determining a corpus term from a plurality of documents, generating a candidate term from the corpus term, and selecting a normalized term from the candidate term and the corpus term. Some embodiments include linking the normalized term with the candidate term and providing an electronic search capability for locating a first document, where the electronic search capability receives the candidate term as a search term and utilizes the normalized term to locate the first document.

Claims (48)

1. A method for lexicon generation, comprising the steps of:

determining a corpus term from a plurality of documents;

generating a candidate term from the corpus term, wherein generating the candidate term comprises generating a linguistic variant of the corpus term;

generating a plurality of equivalent terms from the candidate term;

validating the plurality of equivalent terms by comparing the plurality of equivalent terms to frequency of occurrence of the candidate term;

linking each of the plurality of equivalent terms to the candidate term to create respective equivalent term pairs;

determining whether any of the equivalent term pairs are equivalent and, in response to determining that at least two of equivalent term pairs are equivalent, merging the equivalent term pairs to create a group of equivalent terms;

selecting a normalized term from the group of equivalent terms; and

storing the group of equivalent terms.

2. The method of claim 1 , further comprising the steps of:

determining the frequency of occurrence for the candidate term in the first document,

determining whether the frequency of occurrence for the candidate term does not meet a predetermined threshold; and

in response to determining that the frequency of occurrence for the candidate term does not meet the predetermined threshold, removing the candidate term.

3. The method of claim 1 , wherein the normalized term is selected based on a frequency of the candidate term and the corpus term in the plurality of documents.

4. The method of claim 1 , wherein the electronic search capability includes an electronic searching system for locating legal documents.

5. A system for lexicon generation, comprising:

a processor; and

a memory component that stores lexicon generation logic that when executed by the processor, causes a computer to perform at least the following:

determine a corpus term from a plurality of documents;

generate a candidate term from the corpus term, wherein generating the candidate term comprises generating a linguistic variant of the corpus term;

generate a plurality of equivalent terms from the candidate term;

validate the plurality of equivalent terms by comparing the plurality of equivalent terms to a frequency of occurrence of the candidate term;

link each of the plurality of equivalent terms to the candidate term to create respective equivalent term pairs;

determine whether any of the equivalent term pairs are equivalent and, in response to determining that at least two of equivalent term pairs are equivalent, merging the equivalent term pairs to create a group of equivalent terms;

select a normalized term from the group of equivalent terms; and

store the group of equivalent terms.

6. The system of claim 5 , the lexicon generation logic further configured to cause the computer to determine the frequency of occurrence for the candidate term in at least a portion of the plurality of documents.

7. The system of claim 6 , the lexicon generation logic further configured to cause the computer to perform the following:

determine whether the frequency of occurrence for the candidate term meets a predetermined threshold; and

in response to determining that the frequency of occurrence for the candidate term meets the predetermined threshold, remove the candidate term.

8. The system of claim 5 , wherein the normalized term is selected based on a frequency of the candidate term and the corpus term in the plurality of documents.

9. The system of claim 5 , wherein the memory component further stores search logic that is configured to provide an electronic searching system for locating legal documents.

10. A non-transitory computer-readable medium for lexicon generation that stores a program that, when executed by a computer, causes the computer to perform at least the following:

determine a corpus term from a plurality of documents;

generate a candidate term from the corpus term; term, wherein generating the candidate term comprises generating a linguistic variant of the corpus term;

generate a plurality of equivalent terms from the candidate term;

validate the plurality of equivalent terms by comparing the plurality of equivalent terms to a frequency of occurrence of the candidate term;

link each of the plurality of equivalent terms to the candidate term to create respective equivalent term pairs;

determine whether any of the equivalent term pairs are equivalent and, in response to determining that at least two of equivalent term pairs are equivalent, merging the equivalent term pairs to create a group of equivalent terms;

select a normalized term from the group of equivalent terms; and

store the group of equivalent terms.

11. The non-transitory computer-readable medium of claim 10 , the program further causing the computer to determine the frequency of occurrence for the candidate term in at least a portion of the plurality of documents.

12. The non-transitory computer-readable medium of claim 11 , the program further causing the computer to perform at least the following:

determine whether the frequency of occurrence for the candidate term meets a predetermined threshold; and

in response to determining that the frequency of occurrence for the candidate term meets the predetermined threshold, remove the candidate term.

13. The non-transitory computer-readable medium of claim 10 , wherein the normalized term is selected based on a frequency of the candidate term and the corpus term in the plurality of documents.

14. The non-transitory computer-readable medium of claim 10 , the program being further configured to provide an electronic searching system for locating legal documents.

15. The non-transitory computer-readable medium of claim 10 , the program further causing the computer to generate an equivalent term to the candidate term.

Assignments (2)
CHANGE OF NAME Recorded Dec 3, 2019
From: LEXISNEXIS; REED ELSEVIER INC.
To: RELX INC.
Reel/Frame 051198/0325 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2010
From: ZHANG, PAUL; SILVER, HARRY
To: LEXISNEXIS, A DIVISION OF REED ELSEVIER INC.
Reel/Frame 024894/0613 →
Continuity (1)
Related Publication 20120054220A1 · Mar 1, 2012