IP Library Granted Patent US 7,945,438
Granted Patent B2
US 7,945,438 · App. 11/695,225 · Granted May 17, 2011

Automated glossary creation

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,945,438
App. No.
11/695,225
Granted
May 17, 2011
Kind
B2
Abstract

A method and device for creating a glossary includes a processor operable for executing computer instructions for identifying, in at least one information source, at least one glossary item identifying a part or a component, determining at least one glossary item form as a canonical form, defining, by using the canonical form, at least one syntactic structure, that includes one of the at least one identified glossary items, for each of at least one semantic classes, and searching a second information source for the at least one syntactic structure of the semantic class.

Claims (100)

1. A computer-implemented method for creating a glossary, the method comprising:

identifying, in at least one information source, a plurality of glossary items, each of the glossary items identifying a part or a component;

identifying a set of glossary items within the plurality of glossary items that comprises a common concept;

selecting one of the glossary items in the set of glossary items as a canonical form for the set of glossary items, the canonical form being a most common way of representing the common concept;

designating each of the other glossary items in the set of glossary items as a variant glossary item with a variant form that varies from the canonical form;

defining, by using the canonical form, at least one syntactic structure for each of at least one semantic classes, the at least one syntactic structure including one of the plurality of glossary items; and

searching, by a processor in at least one of the at least one information source and a second information source, for the at least one syntactic structure of the semantic class.

2. The computer-implemented method according to claim 1 , further comprising:

defining, by using at least one of the variant forms, at least one syntactic structure, that includes one of the plurality of glossary items, for each of at least one semantic classes; and

searching, in at least one of the at least one information source and the second information source, for the at least one of the variant forms of the syntactic structure of the semantic class.

3. The computer-implemented method according to claim 1 , wherein the at least one semantic class is at least one of:

a symptom;

a cause; and

an action.

4. The computer-implemented method according to claim 1 , further comprising:

mapping each variant glossary item onto the canonical form.

5. The computer-implemented method according to claim 1 , wherein the defining is performed by parsing to identify the at least one syntactic structure typical of a semantic class.

6. The computer-implemented method according to claim 1 , further comprising:

defining, by using the canonical form, at least one syntactic structure, that does not include one of the plurality of glossary items, for each of at least one semantic classes.

7. The computer-implemented method according to claim 1 , further comprising:

assigning the at least one semantic class, including the identified syntactic structure, to at least one of the plurality of glossary items.

8. The computer-implemented method according to claim 7 , further comprising:

using parsing rules to assign at least one additional semantic class to at least one of a phrase and a clause in at least one of:

the at least one information source; and

the second information source.

9. The computer-implemented method according to claim 1 , wherein:

at least one of the at least one information source and the second information source is a failure report.

10. The computer-implemented method according to claim 1 , wherein:

the searching is a lexical lookup for text in the second source.

11. A computer-implemented method for creating a glossary, the method comprising:

extracting, from at least one information source, text related to at least one of a part and a component;

identifying a plurality of terms in the text, each of the terms being related to at least one of a part and a component;

identifying a set of terms within the plurality of terms that comprises a common concept;

determining, from the text by a processor, a canonical form for the set of terms, the canonical form being based on a most common term in the set of terms;

determining at least one variant form of the canonical form based on each of the other terms in the set of terms; and

storing the canonical form and the at least one variant form in a first glossary.

12. The computer-implemented method according to claim 11 , further comprising:

extracting, from the at least one information source, text related to symptoms;

searching the text related to symptoms for at least one of the canonical form and the at least one variant form; and

identifying at least one symptom in the text related to symptoms.

13. The computer-implemented method according to claim 12 , further comprising:

determining a canonical form of the symptom;

determining at least one variant form of the canonical form of the symptom; and

storing the canonical form of the symptom and the at least one variant form of the symptom in a second glossary.

14. The computer-implemented method according to claim 11 , further comprising:

extracting, from the at least one information source, text related to causes;

searching the text related to causes for at least one of the canonical form and the at least one variant form; and

identifying at least one cause in the text related to causes.

15. The computer-implemented method according to claim 14 , further comprising:

determining a canonical form of the cause;

determining at least one variant form of the canonical form of the cause; and

storing the canonical form of the cause and the at least one variant form of the cause in a second glossary.

16. The computer-implemented method according to claim 11 , further comprising:

extracting, from the at least one information source, text related to actions;

searching the text related to actions for at least one of the canonical form and the at least one variant form; and

identifying at least one action in the text related to actions.

17. The computer-implemented method according to claim 16 , further comprising:

determining a canonical form of the action;

determining at least one variant form of the canonical form of the action; and

storing the canonical form of the action and the at least one variant form of the action in a second glossary.

18. A device for creating a glossary, the device comprising:

a memory;

a processor communicatively coupled to the memory; and

a glossary creation system adapted to:

identify a plurality of glossary items in a first information source;

identify a set of glossary items within the plurality of glossary items that comprises a common concept;

select one of the glossary items in the set of glossary items as a canonical form for the set of glossary items, the canonical form being a most common way of representing the common concept;

designate each of the other glossary items in the set of glossary items as a variant glossary item with a variant form that varies from the canonical form;

define, by using the canonical form, at least one syntactic structure for each of at least one semantic classes, the at least one syntactic structure including one of the plurality of glossary items; and

searching, in at least one of the at least one information source and a second information source, for the at least one syntactic structure of the semantic class.

19. The device according to claim 18 , wherein the glossary creation system is further adapted to:

define, by using at least one of the variant forms, at least one syntactic structure, that includes one of the plurality of glossary items, for each of at least one semantic classes.

20. The device according to claim 18 , wherein at least one glossary item in the plurality of glossary items is at least one of a component-noun word and a component-noun phrase.

21. A non-transitory computer readable medium encoded with a program for creating a glossary, the program comprising instructions for:

identifying, in at least one information source, a plurality of glossary items, each of the glossary items identifying a part or a component;

identifying a set of glossary items within the plurality of glossary items that comprises a common concept;

selecting one of the glossary items in the set of glossary items as a canonical form for the set of glossary items, the canonical form being a most common way of representing the common concept;

designating each of the other glossary items in the set of glossary items as a variant glossary item with a variant form that varies from the canonical form;

defining, by using the canonical form, at least one syntactic structure for each of at least one semantic classes, the at least one syntactic structure including one of the plurality of glossary items; and

searching, in at least one of the at least one information source and a second information source, for the at least one syntactic structure of the semantic class.

22. The non-transitory computer readable medium according to claim 21 , wherein the program further comprises instructions for:

defining, by using at least one of the variant forms, at least one syntactic structure, that includes one of the plurality of glossary items, for each of at least one semantic classes; and

searching, in at least one of the at least one information source and the second information source, for the at least one of the variant forms of the syntactic structure of the semantic class.

23. The non-transitory computer readable medium according to claim 21 , wherein the at least one semantic class is at least one of:

a symptom;

a cause; and

an action.

24. The non-transitory computer readable medium according to claim 21 , wherein the program further comprises instructions for:

mapping each variant glossary item onto the canonical form.

25. The non-transitory computer readable medium according to claim 21 , wherein the defining is performed by parsing to identify the at least one syntactic structure typical of a semantic class.

26. The non-transitory computer readable medium according to claim 21 , wherein the program further comprises instructions for:

defining, by using the canonical form, at least one syntactic structure, that does not include one of the plurality of glossary items, for each of at least one semantic classes.

27. The non-transitory computer readable medium according to claim 21 , wherein the program further comprises instructions for:

assigning the at least one semantic class, including the identified syntactic structure, to at least one of the plurality of glossary items.

28. The non-transitory computer readable medium according to claim 27 , wherein the program further comprises instructions for:

using parsing rules to assign at least one additional semantic class to at least one of a phrase and a clause in the second information source.

29. The non-transitory computer readable medium according to claim 21 , wherein:

at least one of the first and second information source is a failure report.

30. The non-transitory computer readable medium according to claim 29 , wherein:

the searching is a lexical lookup for text in the failure report.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2007
From: BALMELLI, LAURENT; BYRD, ROY; COHEN, MITCHELL A.; ZENG, SAI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019100/0947 →
Continuity (1)
Related Publication 20080243488A1 · Oct 2, 2008