IP Library Granted Patent US 9,087,122
Granted Patent B2
US 9,087,122 · App. 13/716,446 · Granted Jul 21, 2015

Corpus search improvements using term normalization

Inventors: Joel C. Dubbels (Eyota, MN); Mark G. Megerian (Rochester, MN); Michael W. Schroeder (Rochester, MN); Frances E. Stewart (Palisade, MN)
Assignee: International Business Machines Corporation
G06F17/30672G06F17/30646G06F17/30648
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,087,122
App. No.
13/716,446
Granted
Jul 21, 2015
Kind
B2
Abstract

System and computer program product to perform an operation for query processing based on normalized search terms. The operation begins by, responsive to receiving a query, generating a normalized search term for a concept in the query based on a first language model, of a plurality of language models each having a predefined association with a respective concept. The operation then modifies the query to include the normalized search term, and executes the modified query against an indexed corpus of evidence including a first item of evidence. The operation then, upon determining that the first item of evidence includes the normalized search term, returns the first item of evidence as responsive to the query.

Claims (34)

1. A system, comprising:

one or more computer processors; and

a memory containing a program, which, when executed by the one or more computer processors, performs an operation for query processing based on normalized search terms, the operation comprising:

responsive to receiving a query, generating a normalized search term for a concept in the query based on a first language model, of a plurality of language models each having a predefined association with a respective concept;

modifying the query to include the normalized search term;

executing the modified query against the indexed corpus of evidence, where the corpus of evidence is indexed based on the plurality of language models to include a set of normalized terms for each respective item of evidence in the corpus, wherein the indexed corpus of evidence includes a first item of evidence used to support a first candidate answer, of a plurality of candidate answers; and

upon determining that the set of normalized terms for the first item of evidence includes the normalized search term, returning the first candidate answer as responsive to the query.

2. The system of claim 1 , the operation further comprising generating the plurality of language models, wherein each language model of the plurality corresponds to a respective concept, wherein generating the respective language models comprises:

identifying at least one key term for the respective concept; and

generating a respective normalized search term representing the respective at least one key term.

3. The system of claim 2 , wherein the normalized search term is based on at least one of: (i) at least one variant of the at least one key term, and (ii) a context of the at least one key term.

4. The system of claim 3 , wherein the set of normalized terms comprises two or more normalized terms, wherein the corpus of evidence is further indexed by:

associating the set of normalized terms with the respective item of evidence; and

storing the association.

5. The system of claim 1 , wherein the concept is separately expressed by each of a plurality of variants.

6. The system of claim 5 , wherein the query is received from a requesting entity, wherein the query is processed without requiring the requesting entity to specify any of the variants other than a first variant included in the query.

7. The system of claim 1 , wherein the corpus of evidence is a closed corpus.

8. A computer program product for query processing based on normalized search terms, the computer program product comprising:

a non-transitory computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code comprising:

computer-readable program code configured to, responsive to receiving a query, generate a normalized search term for a concept in the query based on a first language model, of a plurality of language models each having a predefined association with a respective concept;

computer-readable program code configured to modify the query to include the normalized search term;

computer-readable program code configured to execute the modified query against the indexed corpus of evidence, where the corpus of evidence is indexed based on the plurality of language models to include a set of normalized terms for each respective item of evidence in the corpus, wherein the indexed corpus of evidence includes a first item of evidence used to support a first candidate answer, of a plurality of candidate answers; and

computer-readable program code configured to, upon determining that the set of normalized terms for the first item of evidence includes the normalized search term, returning the first candidate answer as responsive to the query.

9. The computer program product of claim 8 , further comprising:

computer-readable program code configured to generate the plurality of language models, wherein each language model of the plurality corresponds to a respective concept, wherein generating the respective language models comprises:

identifying at least one key term for the respective concept; and

generating a respective normalized search term representing the respective at least one key term.

10. The computer program product of claim 9 , wherein the normalized search term is based on at least one of: (i) at least one variant of the at least one key term, and (ii) a context of the at least one key term.

11. The computer program product of claim 10 , wherein the set of normalized terms comprises two or more normalized terms, wherein the corpus of evidence is further indexed by:

associating the set of normalized terms with the respective item of evidence; and

storing the association.

12. The computer program product of claim 8 , wherein the concept is separately expressed by each of a plurality of variants.

13. The computer program product of claim 12 , wherein the query is received from a requesting entity, wherein the query is processed without requiring the requesting entity to specify any of the variants other than a first variant included in the query.

14. The computer program product of claim 8 , wherein the corpus of evidence is a closed corpus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2012
From: DUBBELS, JOEL C.; MEGERIAN, MARK G.; SCHROEDER, MICHAEL W.; STEWART, FRANCES E.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 029480/0165 →
Continuity (1)
Related Publication 20140172904A1 · Jun 19, 2014