IP Library Granted Patent US 7,512,272
Granted Patent B2
US 7,512,272 · App. 10/959,447 · Granted Mar 31, 2009

Method for optical recognition of a multi-language set of letters with diacritics

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,512,272
App. No.
10/959,447
Granted
Mar 31, 2009
Kind
B2
Abstract

A method and system for recognizing alphabetic characters that contain diacritics is described. An image analysis separates the character into its constituent components. The one or more diacritic components are then distinguished and isolated from the base portion of the character. Optical recognition is performed separately on the base portion. The diacritic is recognized through a special image analysis and pattern recognition algorithms. The image analysis extracts geometric information from the one or more diacritic components. The extracted information is used as input for the pattern recognition algorithms. The output is a code that corresponds to a particular diacritic. The recognized base portion and diacritic are combined and a check is performed for acceptable combinations in a chosen language. By separately recognizing the base portion and diacritic, the character sets used by the recognizer can be narrowed, resulting in greater recognition.

Claims (65)

1. A computerized method of identifying characters in a character recognition system having a processor, the method comprising:

a) analyzing, in conjunction with the processor, a character image for separation and extraction of a base character and one or more diacritics;

b) applying, in conjunction with the processor, optical character recognition or intelligent character recognition algorithms to the base character;

c) processing, in conjunction with the processor, the diacritics with at least one of image analysis and pattern recognition algorithms; and

d) combining, in conjunction with the processor, the results of b) and c) so as to check for acceptable combinations of the base character and diacritics with respect to one or more specific languages.

2. The method. of claim 1 , additionally comprising assigning a computer code to an acceptable combination of the base character and diacritics corresponding to the character image.

3. The method of claim 1 , wherein the character image comprises more than one diacritic component.

4. The method of claim 1 , wherein analyzing the character image comprises digitizing a document having printed characters.

5. The method of claim 4 , wherein the digitizing is performed by one of a scanning device, a facsimile machine and a digital camera.

6. The method of claim 1 , wherein analyzing the character image comprises identifying an alphabetic type character.

7. The method of claim 1 , further comprising performing a quantitative image analysis to determine a black pixel count for each of the extracted components.

8. The method of claim 7 , further comprising assigning the extracted component containing the greatest number of black pixels as the base character.

9. The method of claim 1 , wherein the acceptable combinations that exist between the base character and the diacritics are limited by one of the languages.

10. The method of claim 1 , wherein combining the results comprises determining whether the base character and the diacritics are an acceptable combination.

11. The method of claim 10 , wherein determining whether the base character and the diacritics are an acceptable combination comprises using a sequence of non-English languages.

12. A computerized method of identifying characters having a diacritic component in a character recognition system having a processor, the method comprising:

segmenting, in conjunction with the processor, a base component and a diacritic component from a character image;

recognizing, in conjunction with the processor, the base component;

recognizing, in conjunction with the processor, the diacritic component;

performing, in conjunction with the processor, a match analysis to determine whether the base component and the diacritic component is an acceptable combination for one or more particular languages; and

recognizing, in conjunction with the processor, the combination of the base component and the diacritic component in response to the match analysis.

13. The method of claim 12 , wherein the character image comprises more than one diacritic component.

14. The method of claim 12 , additionally comprising assigning a computer code to an acceptable combination of the base and diacritic corresponding to the character image.

15. The method of claim 12 , wherein the one or more particular languages are non-English languages.

16. The method of claim 12 , additionally comprising extracting data from data fields on a document having printed characters.

17. The method of claim 16 , additionally comprising digitizing the extracted data.

18. The method of claim 17 , additionally comprising:

determining if the digitized data is from a non-constrained text field; and

segmenting the non-constrained textual field data.

19. The method of claim 12 , wherein segmenting the character image comprises identifying an alphabetic type character.

20. The method of claim 12 , further comprising performing a quantitative image analysis to determine a black pixel count for each of the segmented components.

21. The method of claim 20 , further comprising assigning the segmented component containing the greatest number of black pixels as the base component.

22. The method of claim 13 , further comprising:

determining a relative location between two or more diacritic components; and

assigning a label to each of the two or more diacritic components based on their respective relative locations.

23. The method of claim 12 , further comprising recognizing the base component using one of optical character recognition and intelligent character recognition.

24. The method of claim 12 , wherein the acceptable combinations that exist between the base component and the diacritic component are limited by one of the languages.

25. The method of claim 12 , wherein performing the match analysis comprises determining whether the base component and the diacritic component is an acceptable combination using a sequence of non-English languages.

26. The method of claim 12 , additionally comprising:

determining during the match analysis whether the base component is one of a plurality of commonly misrecognized base components;

determining if the commonly misrecognized base component can be matched with the recognized diacritic component;

determining the base component that is commonly misrecognized as the commonly misrecognized base component; and

matching the base component to the diacritic component when the commonly misrecognized base component does not match with the diacritic.

27. A system for recognizing characters, the system comprising:

a processor;

a computer readable medium storing algorithms operable to cause the processor to perform:

a) an analysis process configured to analyze a character image so as to segment a base component and one or more diacritic components;

b) a character recognition algorithm configured to recognize the base component;

c) a diacritic recognition algorithm configured to process the diacritic components of the character image; and

d) a diacritic matching algorithm configured to combine the results of b) and c) to check for acceptable combinations of the base component and diacritic components for specific languages.

28. The system of claim 27 , wherein the diacritic matching algorithm is further configured to recognize and assign a computer code to an acceptable combination of the base component and diacritic components corresponding to the character image.

29. A system of identifying characters having a diacritic component, the system comprising:

a processor:

a computer readable medium storing algorithms operable to cause the processor to perform:

a character parts segmentation process configured to segment a character image so as to extract a base component arid a diacritic component;

a character recognition algorithm configured to recognize the base component;

a diacritic recognition algorithm configured to recognize the diacritic component of the character image; and

a diacritic matching algorithm configured to determine whether the base component and the diacritic component are an acceptable combination for a particular language and to recognize the combination.

30. The system of claim 29 , wherein the diacritic matching algorithm is further configured to assign a computer code to an acceptable combination of the base component and the diacritic component.

31. The system of claim 29 , wherein the particular language is a non-English language.

32. A computer readable medium having computer readable program code embodied therein for identifying characters having a diacritic component, the computer readable code comprising:

a character parts segmentation process configured to segment a character image so as to extract a base component and a diacritic component;

a character recognition algorithm configured to recognize the base component;

a diacritic recognition algorithm configured to recognize the diacritic component of the character image; and

a diacritic matching algorithm configured to determine whether the base component and the diacritic component are an acceptable combination for a particular language and to recognize the combination.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2015
From: HEWLETT-PACKARD COMPANY
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 036737/0587 →
MERGER Recorded Jun 26, 2015
From: VERITY, INC.
To: HEWLETT-PACKARD COMPANY
Reel/Frame 035914/0352 →
MERGER Recorded Jun 24, 2015
From: CARDIFF SOFTWARE, INC.
To: VERITY, INC.
Reel/Frame 035964/0753 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2014
From: MAYZLIN, ISAAC; DEERE, EMILY ANN
To: CARDIFF SOFTWARE, INC.
Reel/Frame 032081/0845 →