IP Library Granted Patent US 8,369,612
Granted Patent B2
US 8,369,612 · App. 13/325,789 · Granted Feb 5, 2013

System and methods for Arabic text recognition based on effective Arabic text feature extraction

Inventors: Hussein K. Al-Omari (Riyadh, SA); Mohammad S. Khorsheed (Riyadh, SA)
Assignee: King Abdulaziz City for Science & Technology
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,369,612
App. No.
13/325,789
Granted
Feb 5, 2013
Kind
B2
Abstract

A method for automatically recognizing Arabic text includes digitizing a line of Arabic characters to form a two-dimensional array of pixels each associated with a pixel value, wherein the pixel value is expressed in a binary number, dividing the line of the Arabic characters into a plurality of line images, defining a plurality of cells in one of the plurality of line images, wherein each of the plurality of cells comprises a group of adjacent pixels, serializing pixel values of pixels in each of the plurality of cells in one of the plurality of line images to form a binary cell number, forming a text feature vector according to binary cell numbers obtained from the plurality of cells in one of the plurality of line images, and feeding the text feature vector into a Hidden Markov Model to recognize the line of Arabic characters.

Claims (22)

1. A computer-implemented method for automatically recognizing Arabic text, comprising:

acquiring a text image containing a line of Arabic characters;

digitizing the line of the Arabic characters to form a two-dimensional array of pixels each associated with a pixel value expressed in a binary number, wherein the two-dimensional array of pixels comprises a plurality of rows in a first direction and a plurality of columns in a second direction, wherein the pixel values in the two-dimensional array are expressed in single-bit binary numbers;

counting frequencies of consecutive pixels of a same pixel value in a column of pixels, wherein the step of counting frequencies comprises:

assigning the first frequency count to be “0” when the pixel value of the first one or more pixels in a column is “0”,

wherein the second frequency count is the number of consecutive pixels having a “0” pixel value at the start of the column;

forming a text feature vector using the frequency counts obtained from the column of pixels; and

feeding the text feature vector into a Hidden Markov Model to recognize the line of Arabic characters.

2. The computer-implemented method of claim 1 , wherein the text feature vector is formed by a series of the frequency counts consecutively obtained from the column of pixels.

3. The computer-implemented method of claim 1 , wherein the frequencies of consecutive pixels of a same pixel value are counted up to a predetermined cut-off transition number.

4. The computer-implemented method of claim 3 , wherein the predetermined cut-off transition number is six.

5. A computer program product comprising a non-transitory computer useable medium having computer readable program code functions embedded in said medium for causing a computer to:

acquire a text image containing a line of Arabic characters;

digitize the line of the Arabic characters to form a two-dimensional array of pixels each associated with a pixel value expressed in a binary number, wherein the two-dimensional array of pixels comprises a plurality of rows in a first direction and a plurality of columns in a second direction, wherein the pixel values in the two-dimensional array are expressed in single-bit binary numbers;

count frequencies of consecutive pixels of a same pixel value in a column of pixels, comprising:

assigning the first frequency count to be “0” when the pixel value of the first one or more pixels in a column is “1”,

wherein the second frequency count is the number of consecutive pixels having a “1” pixel value at the start of the column;

form a text feature vector using the frequency counts obtained from the column of pixels; and

feed the text feature vector into a Hidden Markov Model to recognize the line of Arabic characters.

6. The computer program product of claim 5 , wherein the text feature vector is formed by a series of the frequency counts consecutively obtained from the column of pixels.

7. The computer program product of claim 5 , wherein the frequencies of consecutive pixels of a same pixel value are counted up to a predetermined cut-off transition number.

8. The computer program product of claim 7 , wherein the predetermined cut-off transition number is six.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 26, 2012
From: KHORSHEED, MOHAMMAD S; AL-OMARI, HUSSEIN K
To: KING ABDULAZIZ CITY FOR SCIENCE & TECHNOLOGY
Reel/Frame 029349/0119 →
Continuity (2)
Continuation 12430773 · Apr 27, 2009
Related Publication 20120087584A1 · Apr 12, 2012