IP Library Patent Application 12132072
Patent Application
App. No. 12/132,072

SYSTEM AND METHOD FOR DIACRITIZATION OF TEXT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
12/132,072
Abstract

A system and method for restoration of diacritics includes making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration. A best diacritic representation is determined for graphemes in the utterance based upon a best match with the diacritization model. A diacritically restored representation of the utterance is output.

Claims (28)

1 . A method for restoration of diacritics, comprising:

making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration;

determining a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model; and

outputting a diacritically restored representation of the utterance.

2 . The method as recited in claim 1 , wherein the diacritization model includes information for a Semitic language.

3 . The method as recited in claim 1 , further comprising integrating a plurality of sources of information to train the diacritization model.

4 . The method as recited in claim 1 , wherein the diacritization model is trained using a Maximum Entropy technique.

5 . The method as recited in claim 1 , wherein the diacritization model is trained using one or more of support vector machines, boosting, statistical decision trees, and/or their combinations to restore diacritics or to generate a diacritic lattice.

6 . The method as recited in claim 1 , wherein determining a best diacritic representation includes performing a Viterbi search to find the best diacritic representation assigned to the utterance.

7 . The method as recited in claim 1 , wherein the diacritization model includes one or more of previously predicted diacritics, graphemes, morphological units, words, part-of-speech (POS) tags, syntactic information, semantic analysis, and linguistic knowledge.

8 . The method as recited in claim 1 , wherein the diacritization model restores a complete set of diacritics.

9 . The method as recited in claim 1 , further comprising building a Diacritization Parse Tree (DPT) to extract information from in the form of features to train the diacritization model.

10 . The method as recited in claim 9 , wherein the diacritization model includes a statistical model able to learn the DPT, such that each node in the DPT can be predicted.

11 . The method as recited in claim 9 , wherein the diacritization model is constrained to learn leaves of the DPT that correspond to diacritics only.

12 . The method as recited in claim 1 , wherein making classification decisions further comprises building a Diacritization Parse Tree (DPT) for an input utterance to generate a diacritic lattice for each grapheme in the input utterance to determine a best diacritic for each grapheme.

13 . The method as recited in claim 1 , wherein the plurality of information sources include non-overlapping and overlapping sources of information.

14 . The method as recited in claim 1 , further comprising generating a lattice of diacritics for each utterance, and rescoring the diacritics with alternative models to improve the accuracy.

15 . The method as recited in claim 1 , wherein outputting a diacritically restored representation includes outputting one of text and speech.

16 . The method as recited in claim 1 , wherein the utterance includes one or text and speech.

17 . A computer program product for restoration of diacritics comprising a computer useable medium including a computer readable program, wherein the computer readable program when executed on a computer causes the computer to perform:

making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration;

determining a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model; and

outputting a diacritically restored representation of the utterance.

18 . The program product as recited in claim 17 , wherein the diacritization model restores a complete set of diacritics.

19 . The program product as recited in claim 17 , wherein making classification decisions further comprises building a Diacritization Parse Tree (DPT) for an input utterance to generate a diacritic lattice for each grapheme in the input utterance to determine a best diacritic for each grapheme.

20 . A system for restoration of diacritics, comprising:

a diacritization model configured to make classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources for diacritic restoration;

a processing module configured to determine a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model to output a diacritically restored representation of the utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →