IP Library › Granted Patent US 10,769,376
Granted Patent B2
US 10,769,376 · App. 15/803,158 · Granted Sep 8, 2020

Domain-specific lexical analysis

Inventors: Branimir K. Boguraev (Bedford, NY); Esme Manandise (Tallahassee, FL); Benjamin P. Segal (Hyde Park, NY)
Assignee: International Business Machines Corporation
G06F40/284G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,769,376
App. No.
15/803,158
Granted
Sep 8, 2020
Kind
B2
Abstract

A method includes performing, at a device, an analysis of a domain-specific corpus to identify a base term and a modifier term. The modifier term modifies the base term in at least a portion of the domain-specific corpus. The method also includes accessing, by the device, a first entry in lexicon data. The first entry includes core data corresponding to domain-independent lexical information for the base term. The method further includes adding, based on the analysis, non-core data to the first entry. The non-core data corresponds to domain-specific lexical information for the base term. The non-core data identifies the modifier term as a domain-specific modifier of the base term.

Claims (49)

1. A method comprising:

initiating, at an electronic device, an analysis of a domain-specific corpus associated with a particular domain to:

identify the particular domain; and

identify a base term and a modifier term, the base term and the modifier term identified independently from the particular domain, wherein the modifier term: (i) modifies the base term in at least a portion of the domain-specific corpus, and (ii) is designated as a modifier term based, at least in part, upon a co-occurrence statistics that indicate that an identified term appears to modify the base term a minimum number of times;

accessing, by the electronic device, a first entry in lexicon data, wherein the lexicon data is accessible by the electronic device prior to the initiating and is configured for use at the electronic device in a language processing operation, the first entry including core data corresponding to domain-independent lexical information for the base term;

adding, based on the analysis, non-core data to the first entry, the non-core data corresponding to domain-specific lexical information for the base term, wherein the non-core data identifies the modifier term and the particular domain;

determining that a first portion of the non-core data is in a specific position relative to the identified base term;

responsive to the determination that the first portion of the non-core data is in the specific position relative to the identified base term, generating a first collocation rule; and

responsive to the generation of the first collocation rule, parsing, with the first collocation rule, the first entry in lexicon data to determine a first future placement for the first portion of the non-core data.

2. The method of claim 1 , further comprising determining that the modifier term modifies the base term based on at least one of co-occurrence statistics or user input.

3. The method of claim 1 , wherein:

the modifier term includes at least one of an adjectival modifier term, a preposition modifier term, or a nominal modifier term, and

the language processing operation includes processing, by a parser, a language sample based on the lexicon data.

4. The method of claim 1 , further comprising:

identifying, at the electronic device, the modifier term used with the base term as a preferred domain-specific modifier in the domain-specific corpus; and

updating, at the electronic device, the non-core data to identify the modifier term as a preferred domain-specific modifier of the base term.

5. The method of claim 1 , further comprising generating, based on the domain-specific corpus, domain-specific parsing rules for a domain-specific lexically-driven pre-parser.

6. The method of claim 1 wherein the first collocation rule includes information indicative of a set of future placements for the non-core data relative to the identified base term.

7. The method of claim 5 , further comprising:

performing, at the domain-specific lexically-driven pre-parser, a domain-specific analysis of an input text based on the non-core data; and

providing a partially parsed and bracketed version of the input text from the domain-specific lexically-driven pre-parser to a domain-independent rule-based parser.

8. The method of claim 5 , wherein software corresponding to the domain-specific lexically-driven pre-parser is provided as a service in a cloud computing environment.

9. The method of claim 1 , further comprising, based on the analysis of the domain-specific corpus, updating the non-core data to identify one or more second modifier terms as one or more additional domain-specific modifiers of the base term.

10. The method of claim 1 , wherein:

prior to the initiating, the lexicon data is generated by the electronic device, received by the electronic device from another device, provided by a user to the electronic device, or a combination thereof, and

the non-core data includes one or more additional modifier terms for the base term for a second domain that is distinct from a first domain associated with the domain-specific corpus.

11. A method comprising:

analyzing, at an electronic device that includes a parser, a domain-specific corpus associated with a particular domain to:

identify the particular domain; and

identify a base term and a modifier term, the base term and the modifier term identified independently from the particular domain, wherein the modifier term: (i) modifies the base term in at least a portion of the domain-specific corpus, and (ii) is designated as a modifier term based, at least in part, upon a co-occurrence statistics that indicate that an identified term appears to modify the base term a minimum number of times;

accessing a first entry in lexicon data, the lexicon data accessible by the electronic device prior to the analyzing, and the first entry including core data corresponding to domain-independent lexical information for the base term;

adding, based on the analyzing, non-core data to the first entry, the non-core data corresponding to domain-specific lexical information for the base term, and the non-core data identifying the modifier term and the particular domain;

performing a language processing operation at the parser, based, at least in part, upon the lexicon data; and

responsive to the performance of the language processing operation at the parser, generating a first collocation rule.

12. The method of claim 11 , further comprising determining that the modifier term modifies the base term based on at least one of co-occurrence statistics or user input.

13. The method of claim 11 , wherein the modifier term includes at least one of an adjectival modifier term, a preposition modifier term, or a nominal modifier term, and wherein the language processing operation includes processing a language sample.

14. The method of claim 11 , further comprising:

identifying, at the electronic device, the modifier term used with the base term as a preferred domain-specific modifier in the domain-specific corpus; and

updating, at the electronic device, the non-core data to identify the modifier term as a preferred domain-specific modifier of the base term.

15. The method of claim 11 , further comprising generating, based on the domain-specific corpus, domain-specific parsing rules for a domain-specific lexically-driven pre-parser.

16. The method of claim 11 wherein the first collocation rule includes information indicative of a set of future placements for the non-core data relative to the identified base term.

17. The method of claim 15 , wherein the parser includes a domain-independent rule-based parser, and further comprising:

performing, at the domain-specific lexically-driven pre-parser, a domain-specific analysis of an input text based on the non-core data; and

providing a partially parsed and bracketed version of the input text from the domain-specific lexically-driven pre-parser to the domain-independent rule-based parser.

18. The method of claim 15 , wherein software corresponding to the domain-specific lexically-driven pre-parser is provided as a service in a cloud computing environment.

19. The method of claim 11 , further comprising, based on the analyzing, updating the non-core data to identify one or more second modifier terms as one or more additional domain-specific modifiers of the base term.

20. The method of claim 11 , wherein:

the lexicon data is previously generated by the electronic device, received by the electronic device from another electronic device, provided by a user to the electronic device, or a combination thereof; and

the non-core data includes one or more additional modifier terms for the base term for a second domain that is distinct from a first domain associated with the domain-specific corpus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2017
From: BOGURAEV, BRANIMIR K.; MANANDISE, ESME; SEGAL, BENJAMIN P.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044031/0206 →
Continuity (2)
Continuation 15679719 · Aug 17, 2017
Related Publication 20190057078A1 · Feb 21, 2019