IP Library Granted Patent US 7,587,308
Granted Patent B2
US 7,587,308 · App. 11/285,090 · Granted Sep 8, 2009

Word recognition using ontologies

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,587,308
App. No.
11/285,090
Granted
Sep 8, 2009
Kind
B2
Abstract

Systems, and associated apparatus, methods, or computer program products, may use ontologies to provide improved word recognition. The ontologies may be applied in word recognition processes to resolve ambiguities in language elements (e.g., words) where the values of some of the characters in the language elements are uncertain. Implementations of the method may use an ontology to resolve ambiguities in an input string of characters, for example. In some implementations, the input string may be received from a language conversion source such as, for example, an optical character recognition (OCR) device that generates a string of characters in electronic form from visible character images, or a voice recognition (VR) device that generates a string of characters in electronic form from speech input. Some implementations may process the generated character strings by using an ontology in combination with syntactic and/or grammatical analysis engines to further improve word recognition accuracy.

Claims (40)

1. A computer-implemented method executed by a processor that performs operations for reducing ambiguities present in electronically stored words, the operations comprising:

receiving a plurality of characters in electronic form, the received plurality of characters corresponding to a sequence of words and including an ambiguous word that has one or more characters whose value is substantially uncertain;

comparing at least some of the words in the sequence to a first ontology, the first ontology defining a plurality of nodes, each node being associated with a word, and each node being connected to at least one other node by a link, each link being associated with a concept that relates the words associated with the nodes connected by the link in a predetermined context, wherein at least some of the nodes are associated with non-textual image information that identifies an enhancement to character-based text.

2. The method of claim 1 , further comprising performing a syntactic analysis on the sequence of words to identify possible candidate words to replace the ambiguous word.

3. The method of claim 2 , further comprising performing a grammatical analysis to determine which of the identified possible candidate words forms a sequence of words that conforms to a set of predefined grammatical rules when the candidate word is substituted for the ambiguous word.

4. The method of claim 1 , wherein each of the uncertain characters is associated with a confidence level that is lower than a predetermined threshold.

5. The method of claim 1 , further comprising selecting a word associated with one of the identified nodes to replace the ambiguous word.

6. The method of claim 5 , wherein the selecting is further based on second information associated with characters or words in the sequence of words.

7. The method of claim 6 , wherein the second information consists of at least one characteristic of written communication selected from among the group consisting of: font style, character size, bold, italics, or underline.

8. The method of claim 1 , wherein the plurality of characters is received from an optical character recognition (OCR) device configured to convert written character image information to character information in electronic form.

9. The method of claim 1 , wherein the operations further comprise:

identifying nodes in the first ontology that correspond to the ambiguous word based on the comparison to the first ontology;

comparing at least some of the words in the sequence to a second ontology, the second ontology defining a plurality of nodes, each node being associated with a word, and each node being connected to at least one other node by a link, each link being associated with a concept that relates the words associated with the nodes connected by the link in a predetermined context, wherein at least some of the nodes are associated with non-textual image information that identifies an enhancement to character-based text;

identifying nodes in the second ontology that correspond to the ambiguous word based on the comparison to the second ontology;

scoring each of the identified nodes from the first ontology and the second ontology, wherein the scoring is at least partially based upon non-textual image information that identifies an enhancement to character-based text and is associated with at least some of the identified nodes; and

selecting a node having the highest score from among the identified nodes in the first ontology and the second ontology.

10. The method of claim 9 , wherein each identified node is scored as a function of the links that connect the node and any other nodes associated with words that correspond to other words in the sequence of words.

11. The method of claim 9 , further comprising replacing the ambiguous word with the word associated with the selected highest scoring node.

12. The method of claim 1 , wherein the non-textual image information that identifies an enhancement to character-based text comprises font type information.

13. The method of claim 1 , wherein the non-textual image information that identifies an enhancement to character-based text comprises character size information.

14. The method of claim 1 , wherein the non-textual image information that identifies an enhancement to character-based text comprises boldface information.

15. The method of claim 1 , wherein the non-textual image information that identifies an enhancement to character-based text comprises italics information.

16. The method of claim 1 , wherein the non-textual image information that identifies an enhancement to character-based text comprises color information.

17. The method of claim 1 , wherein the non-textual image enhancement to character-based text information comprises underlining information.

18. A computer program product tangibly embodied in a computer-readable data storage medium, the computer program product containing instructions that, when executed, cause a processor to perform operations to recognize words, the operations comprising:

receive a plurality of characters in electronic form, the received plurality of characters corresponding to a sequence of language elements and including an ambiguous language element that has one or more characters whose value is substantially uncertain;

perform an analysis on one or more of the language elements in the sequence of language elements according to a first ontology that defines relationships among a predetermined plurality of language elements that includes at least one of the language elements in the received plurality, at least some of the language elements in the predetermined plurality associated with non-textual image information that identifies an enhancement to character-based text.

19. The computer program product of claim 18 , wherein the identified language elements comprise words.

20. The method of claim 18 , wherein the operations further comprise:

perform an analysis on the one or more of the language elements in the sequence of language elements according to a second ontology that defines relationships among a predetermined plurality of language elements that includes at least one of the language elements in the received plurality, at least some of the language elements in the predetermined plurality associated with non-textual image information that identifies an enhancement to character-based text;

identify language elements in the second ontology that match the ambiguous language element based on the analysis of the second ontology;

score each of the identified language elements from the first ontology and the second ontology, wherein the scoring is at least partially based upon non-textual image information that identifies an enhancement to character-based text and is associated with at least some of the identified language elements in the predetermined plurality; and

select from among the identified language elements in the first ontology and the second ontology a language element having the highest score.

21. A computer-implemented method executed by a processor that performs operations to define an ontology for use in word recognition, the operations comprising:

identifying a plurality of language elements that can be used together when a language is used to express ideas in a particular context;

defining at least one link between each language element and another of the language elements, each link indicative of a likelihood that the linked language elements will be used together to express an idea in the particular context, wherein the at least one link is associated with non-textual image information that identifies an enhancement to character-based text;

identifying an ontology for the language elements based on the non-textual image information that identifies an enhancement to character-based text; and

storing the defined links in electronic form in the identified ontology within an information repository.

22. The method of claim 21 , further comprising updating the ontology in response to user input.

23. The method of claim 22 , wherein updating the ontology comprises adding, deleting, or modifying a link in the ontology in response to user input.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2017
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
To: ENT. SERVICES DEVELOPMENT CORPORATION LP
Reel/Frame 041041/0716 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2009
From: ELECTRONIC DATA SYSTEMS, LLC
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 022449/0267 →
CHANGE OF NAME Recorded Mar 24, 2009
From: ELECTRONIC DATA SYSTEMS CORPORATION
To: ELECTRONIC DATA SYSTEMS, LLC
Reel/Frame 022460/0948 →
CHANGE OF NAME Recorded Sep 27, 2006
From: KASRAVI, KASRA
To: KASRAVI, KAS
Reel/Frame 018312/0949 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2006
From: KASRAVI, KAS; RISOV, MARIA; VARADARAJAN, SUNDAR
To: ELECTRONIC DATA SYSTEMS CORPORATION
Reel/Frame 017219/0485 →