IP Library › Granted Patent US 9,424,253
Granted Patent B2
US 9,424,253 · App. 14/812,198 · Granted Aug 23, 2016

Domain specific natural language normalization

Inventors: Shareef Alshinnawi (Durham, NC); Gary D. Cudak (Creedmoor, NC); Edward S. Suffern (Chapel Hill, NC); John M. Weber (Wake Forest, NC)
Assignee: International Business Machines Corporation
G06F17/28G06F17/21G06F17/2795
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,424,253
App. No.
14/812,198
Granted
Aug 23, 2016
Kind
B2
Abstract

Embodiments of the present invention provide a method, system and computer program product for the domain specific normalization of a corpus of text. In an embodiment of the invention, a method for domain specific normalization of a corpus of text is provided, including an industrial, organization, demographic or geographic domain. The method includes loading a corpus of text in memory of a computer and determining a domain for the corpus of text. The method also includes retrieving a lexicon of replacement words for the determined domain. Finally, the method includes text simplifying the corpus of text using the retrieved lexicon. In one aspect of the embodiment, the domain is determined through inference based upon words already presence in the corpus of text. In another aspect of the embodiment, the domain is determined based upon meta-data provided with the corpus of text.

Claims (32)

1. A method for domain specific normalization of a corpus of text, the method comprising:

loading a corpus of text in memory of a computer;

determining by a processor of the computer a domain for the corpus of text by recognizing a presence of words or phrases in the loaded corpus of text that had been previously correlated to the domain;

retrieving from the memory of the computer by the processor of the computer a lexicon of replacement words for the determined domain, the lexicon comprising a set of source terms, at least one of the source terms being mapped to one of multiple different replacement terms having a complexity value aligned with an average complexity value for the multiple different replacement terms; and,

text simplifying by the processor of the computer the corpus of text using the retrieved lexicon by replacing existing words in the corpus of text with the replacement words.

2. The method of claim 1 , wherein the domain is an industrial domain.

3. The method of claim 1 , wherein the domain is an organizational domain.

4. The method of claim 1 , wherein the domain is a demographic domain.

5. The method of claim 1 , wherein the domain is a geographic domain.

6. The method of claim 1 , wherein the domain is determined through inference based upon words already present in the corpus of text.

7. The method of claim 1 , wherein the domain is determined based upon meta-data provided with the corpus of text.

8. A natural language data processing system configured for domain specific normalization of a corpus of text, the system comprising:

a host computing system comprising at least one computer with memory and at least one processor;

a natural language processor providing logic configured for text simplification executing in the memory of the computer; and,

a domain specific normalization module of the natural language processor comprising program code executing in the host computing system enabled to load a corpus of text, to determine a domain for the corpus of text by recognizing a presence of words or phrases in the loaded corpus of text that had been previously correlated to the domain, to retrieve a lexicon of replacement words for the determined domain, the lexicon comprising a set of source terms, at least one of the source terms being mapped to one of multiple different replacement terms having a complexity value aligned with an average complexity value for the multiple different replacement terms, and to direct the natural language processor to text simplify the corpus of text using the retrieved lexicon by replacing existing words in the corpus of text with the replacement words.

9. The system of claim 8 , wherein the domain is an industrial domain.

10. The system of claim 8 , wherein the domain is an organizational domain.

11. The system of claim 8 , wherein the domain is a demographic domain.

12. The system of claim 8 , wherein the domain is a geographic domain.

13. The system of claim 8 , wherein the program code of the module determines the domain through inference based upon words already present in the corpus of text.

14. A computer program product for domain specific normalization of a corpus of text, the computer program product comprising:

a non-transitory computer readable storage medium comprising a device having computer readable program code embodied therewith, the computer readable program code comprising:

computer readable program code for loading a corpus of text in memory of a computer;

computer readable program code for determining a domain for the corpus of text by recognizing a presence of words or phrases in the loaded corpus of text that had been previously correlated to the domain;

computer readable program code for retrieving a lexicon of replacement words for the determined domain, the lexicon comprising a set of source terms, at least one of the source terms being mapped to one of multiple different replacement terms having a complexity value aligned with an average complexity value for the multiple different replacement terms; and,

computer readable program code for text simplifying the corpus of text using the retrieved lexicon by replacing existing words in the corpus of text with the replacement words.

15. The computer program product of claim 14 , wherein the domain is an industrial domain.

16. The computer program product of claim 14 , wherein the domain is an organizational domain.

17. The computer program product of claim 14 , wherein the domain is a demographic domain.

18. The computer program product of claim 14 , wherein the domain is a geographic domain.

19. The computer program product of claim 14 , wherein the domain is determined through inference based upon words already present in the corpus of text.

20. The computer program product of claim 14 , wherein the domain is determined based upon meta-data provided with the corpus of text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2015
From: ALSHINNAWI, SHAREEF; CUDAK, GARY D.; SUFFERN, EDWARD S.; WEBER, JOHN M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 036351/0449 →
Continuity (2)
Continuation 13414687 · Mar 7, 2012
Related Publication 20150331854A1 · Nov 19, 2015