IP Library Granted Patent US 12,591,739
Granted Patent B2
US 12,591,739 · App. 17/598,633 · Granted Mar 31, 2026

Method and system for diacritizing Arabic text

Inventors: Hamdy S. Mubarak (Doha, QA); Kareem Mohamed Darwish (Doha, QA); Ahmed Abdelali (Doha, QA); Hassan Sajjad (Doha, QA); Younes Samih (Doha, QA)
Assignee: HAMAD BIN KHALIFA UNIVERSITY
G06F40/279G06F40/166G06F40/274G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,739
App. No.
17/598,633
Granted
Mar 31, 2026
Kind
B2
Abstract

The presently disclosed method and system automatically diacritize written Arabic text for use with applications that require verbalizing Arabic text. A method may comprise converting a written sentence into a word sequence and identifying a target source word. The method then may comprise repeatedly overlaying and translating a context window at a plurality of positions in the word sequence to select a plurality of subsets of the word sequence contained within the context window. A diacritized form of the target source word may be generated in each of the word sequence subsets. A final diacritized form of the target source word may be selected from the plurality of diacritized forms based on a voting scheme. The voting scheme may include selecting the diacritized form that is generated the most or may be based on a system of weighting factors.

Claims (42)

1 . A method comprising:

(A) receiving a source sentence;

(B) converting the source sentence into a word sequence, wherein the word sequence comprises a plurality of words including a target source word;

(C) overlaying a context window at a first position within the word sequence to select a first subset of the word sequence contained within the context window at the first position, wherein the first subset of the word sequence contains the target source word;

(D) generating a first diacritized form of the target source word based on the first subset of the word sequence;

(E) translating the context window to a second position within the word sequence to select a second subset of the word sequence contained within the context window at the second position, wherein the second subset of the word sequence contains the target source word;

(F) generating a second diacritized form of the target source word in the second subset of the word sequence;

(G) repeating steps (E) through (F) to generate a plurality of diacritized forms of the target source word; and

(H) selecting a final diacritized form of the target source word from the plurality of diacritized forms of the target source word according to a voting scheme, wherein the context window is a varying number of words in length and the voting scheme weights each diacritized form of the plurality of diacritized forms based on a given length of the context window used to generate a corresponding diacritized form.

2 . The method of claim 1 , wherein each of the plurality of words of the word sequence comprises a sequence of characters.

3 . The method of claim 2 , wherein the sequence of characters comprises a character boundary in between each character within each of the plurality of words and a word boundary between each of the plurality of words in the word sequence.

4 . The method of claim 1 , wherein the voting scheme further comprises, responsive to identifying more than one diacritized form of the target source word generated the most number of times, selecting the diacritized form of the target source word from the segment of words in which the target source word is the center word of the segment of words.

5 . The method of claim 1 , wherein the context window is a fixed number of words in length.

6 . The method of claim 1 , wherein the context window is a fixed number of characters in length.

7 . The method of claim 1 , wherein generating a diacritized form of a target source word comprises inputting the word sequence into at least one machine learning model.

8 . The method of claim 7 , wherein each of the plurality of diacritized forms is generated by the same machine learning model.

9 . The method of claim 7 , wherein the plurality of diacritized forms are generated by more than one machine learning model.

10 . The method of claim 7 , wherein the at least one machine learning model is a sequence to sequence model.

11 . The method of claim 1 , wherein translating the context window further comprises incrementing the position of the context window by a set number of words.

12 . The method of claim 11 , wherein the set number of words is one word.

13 . A system comprising:

a processor; and

a memory storing instructions which, when executed by the processor, cause the processor to:

(A) receive a source sentence;

(B) convert the source sentence into a word sequence, wherein the word sequence comprises a plurality of words including a target source word;

(C) overlay a context window at a first position within the word sequence to select a first subset of the word sequence contained within the context window at the first position, wherein the first subset of the word sequence contains the target source word;

(D) generate a first diacritized form of the target source word based on the first subset of the word sequence;

(E) translate the context window to a second position within the word sequence to select a second subset of the word sequence contained within the context window at the second position, wherein the second subset of the word sequence contains the target source word;

(F) generate a second diacritized form of the target source word in the second subset of the word sequence;

(G) repeat steps (E) through (F) to generate a plurality of diacritized forms of the target source word; and

(H) select a final diacritized form of the target source word from the plurality of diacritized forms of the target source word according to a voting scheme, wherein the context window is a varying number of words in length and the voting scheme weights each diacritized form of the plurality of diacritized forms based on a given length of the context window used to generate a corresponding diacritized form.

14 . The system of claim 13 , wherein each of the plurality of words of the word sequence comprises a sequence of characters.

15 . The system of claim 13 , wherein the context window is a varying number of words in length and the voting scheme comprises selecting the diacritized form of the target source word based on weighting factors assigned to each respective context window of a given length.

16 . A non-transitory, computer-readable medium storing instructions which, when performed by a processor, cause the processor to:

(A) receive a source sentence;

(B) convert the source sentence into a word sequence, wherein the word sequence comprises a plurality of words including a target source word;

(C) overlay a context window at a first position within the word sequence to select a first subset of the word sequence contained within the context window at the first position, wherein the first subset of the word sequence contains the target source word;

(D) generate a first diacritized form of the target source word based on the first subset of the word sequence;

(E) translate the context window to a second position within the word sequence to select a second subset of the word sequence contained within the context window at the second position, wherein the second subset of the word sequence contains the target source word;

(F) generate a second diacritized form of the target source word in the second subset of the word sequence;

(G) repeat steps (E) through (F) to generate a plurality of diacritized forms of the target source word; and

(H) select a final diacritized form of the target source word from the plurality of diacritized forms of the target source word according to a voting scheme, wherein the context window is a varying number of words in length and the voting scheme weights each diacritized form of the plurality of diacritized forms based on a given length of the context window used to generate a corresponding diacritized form.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: QATAR FOUNDATION FOR EDUCATION, SCIENCE & COMMUNITY DEVELOPMENT
To: HAMAD BIN KHALIFA UNIVERSITY
Reel/Frame 069936/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2022
From: MUBARAK, HAMDY S.; DARWISH, KAREEM MOHAMED; ABDELALI, AHMED; SAJJAD, HASSAN; SAMIH, YOUNES
To: QATAR FOUNDATION FOR EDUCATION, SCIENCE AND COMMUNITY DEVELOPMENT
Reel/Frame 060208/0393 →
Continuity (1)
Related Publication 20220188515A1 · Jun 16, 2022
References Cited (21)
US 7802184B1 · Battilana · 2010 [cited by applicant]
US 9298277B1 · Alsabah · 2016 [cited by examiner]
US 9501708B1 · Ahmad · 2016 [cited by examiner]
US 9792271B2 · Keenan · 2017 [cited by applicant]
US 20050192807A1 · Emam · 2005 [cited by examiner]
US 20070225977A1 · Emam · 2007 [cited by examiner]
US 20100082333A1 · Al-Shammari · 2010 [cited by examiner]
US 20120109633A1 · Khorsheed · 2012 [cited by examiner]
US 20130185054A1 · Xiao · 2013 [cited by examiner]
US 20140258852A1 · Sesum · 2014 [cited by examiner]
US 20140380169A1 · Eldawy · 2014 [cited by examiner]
US 20170017854A1 · You · 2017 [cited by examiner]
WO WO2009042861A1 · 2009 [cited by examiner]
WO WO2014138756A1 · 2014 [cited by examiner]
WO WO2014189400A1 · 2014 [cited by examiner]
Rashwan et al, “Deep learning framework with confused sub-set resolution architecture for automatic Arabic diacritization”, 2015, IEEE/ACM Transactions on Audio, Speech, And Language Processing. Feb. 26, 2015;23(3):505-… [cited by examiner]
Preliminary Report on Patentability for related International Application No. PCT/QA2020/050007; report dated Oct. 7, 2021; (6 pages). [cited by applicant]
International Search Report for related International Application No. PCT/QA2019/050007; report dated Oct. 1, 2020; (2 pages). [cited by applicant]
Written Opinion for related International Application No. PCT/QA2019/050007; report dated Oct. 1, 2020; (4 pages). [cited by applicant]
Abandah, et al. “Automatic diacritization of Arabic text using recurrent neural networks”; International Journal on Document Analysis and Recognition (IJDAR), 2015 [online], [retrieved on Jul. 17, 2020]. Retrieved from … [cited by applicant]
Espana-Bonet, et al.; “Discriminative phrase-based models for Arabic machine translation”; ACM Transactions on Asian Language Information Processing, Dec. 2009 [online], [retrieved on Jul. 17, 2020]. Retrieved from the … [cited by applicant]