IP Library Granted Patent US 7,469,205
Granted Patent B2
US 7,469,205 · App. 10/879,259 · Granted Dec 23, 2008

Apparatus and methods for pronunciation lexicon compression

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,469,205
App. No.
10/879,259
Granted
Dec 23, 2008
Kind
B2
Abstract

A compressed pronunciation lexicon file is generated from a source pronunciation lexicon using a pronunciation prediction algorithm in a multi-output mode. The pronunciation prediction algorithm may generate a deterministic ordered list of phoneme strings from the textual representation of a particular word. The compressed pronunciation lexicon file may include a sorted list of records of compressed textual representations of words and compressed phonetic representations of the words.

Claims (38)

1. A method comprising:

generating a compressed pronunciation lexicon file from a source pronunciation lexicon using a pronunciation prediction algorithm in a multi-output mode;

wherein the source pronunciation lexicon includes a list of textual representations of words and corresponding phonetic representations of the words; and

wherein generating the compressed lexicon file further includes:

sorting the list in alphabetical order of the textual representations to generate a sorted list;

substituting a textual representation of a particular word with i) the length of a common string of initial letters of the textual representation and the textual representation of a preceding word in the list, and ii) an encoded version of the remaining letters of the textual representation of the particular word.

2. The method of claim 1 , wherein generating the compressed lexicon file further includes:

generating a Huffman-encoded version of the remaining letters of the textual representation of the particular word and embedding a Huffman coding table or tree in the compressed file.

3. The method of claim 1 , further comprising:

embedding an index table in the compressed file, the index table containing pointers to selected records of the sorted list such that the selected records are substantially evenly distributed along the sorted list.

4. The method of claim 1 , wherein using the pronunciation prediction algorithm includes using a pronunciation prediction algorithm that is based on a hidden Markov model.

5. The method of claim 1 , wherein using the pronunciation prediction algorithm in the multi-output mode includes:

generating a deterministic ordered list of phoneme strings for a particular one of the words.

6. The method of claim 5 , wherein a particular one of the phoneme strings is identical to the corresponding phonetic representation of the particular word, and generating the compressed pronunciation lexicon file further includes:

substituting the corresponding phonetic representation of the particular word with an index of the particular one of the phoneme strings.

7. The method of claim 5 , wherein none of the phoneme strings is identical to corresponding phonetic representation of the particular word, and generating the compressed pronunciation lexicon file further includes:

substituting the corresponding phonetic representation of the particular word with an index of a most closely matching one of the phoneme strings and an encoded representation of edit operations to convert the most closely matching phoneme string into the corresponding phonetic representation of the particular word.

8. An article comprising a computer-readable storage medium having stored thereon a compressed pronunciation lexicon file comprising:

a sorted list of records of textual representations of words and compressed phonetic representations of the words, the compressed phonetic representations generated using a pronunciation prediction algorithm in a multi-output mode,

wherein the textual representations of words are compressed textual representations including prefix lengths and encoded suffixes, wherein a prefix length of a particular word is the length of a common string of initial letters of a textual representation of the particular word and the textual representation of a preceding word in the sorted list, and wherein an encoded suffix is an encoded version of the remaining letters of the textual representation of the particular word.

9. The article of claim 8 , wherein the compressed pronunciation lexicon further comprises an index table including pointers to selected records of the sorted list such that the selected records are substantially evenly distributed along the sorted list.

10. An article comprising a computer-readable storage medium having stored theron instructions that, when executed by a processor, result in:

generating a compressed pronunciation lexicon file from a source pronunciation lexicon using a pronunciation prediction algorithm in a multi-output mode;

wherein the source pronunciation lexicon includes a list of textual representations of words and corresponding phonetic representations of the words; and

wherein generating the compressed lexicon file further includes:

sorting the list in alphabetical order of the textual representations to generate a sorted list;

substituting a textual representation of a particular word with i) the length of a common string of initial letters of the textual representation and the textual representation of a preceding word in the list, and ii) an encoded version of the remaining letters of the textual representation of the particular word.

11. The article of claim 10 , wherein generating the compressed lexicon file further includes:

generating a Huffman-encoded version of the remaining letters of the textual representation of the particular word and embedding a Huffman coding table or tree in the compressed file.

12. The article of claim 10 , wherein the instructions further result in:

embedding an index table in the compressed file, the index table containing pointers to selected records of the sorted list such that the selected records are substantially evenly distributed along the sorted list.

13. The article of claim 10 , wherein using the pronunciation prediction algorithm includes using a pronunciation prediction algorithm that is based on a hidden Markov model.

14. The article of claim 10 , wherein using the pronunciation prediction algorithm in the multi-output mode includes:

generating a deterministic ordered list of phoneme strings for a particular one of the words.

15. The article of claim 14 , wherein a particular one of the phoneme strings is identical to the corresponding phonetic representation of the particular word, and generating the compressed pronunciation lexicon file further includes:

substituting the corresponding phonetic representation of the particular word with an index of the particular one of the phoneme strings.

16. The article of claim 14 , wherein none of the phoneme strings is identical to corresponding phonetic representation of the particular word, and generating the compressed pronunciation lexicon file further includes:

substituting the corresponding phonetic representation of the particular word with an index of a most closely matching one of the phoneme strings and an encoded representation of edit operations to convert the most closely matching phoneme string into the corresponding phonetic representation of the particular word.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2020
From: CAVIUM INTERNATIONAL
To: MARVELL ASIA PTE, LTD.
Reel/Frame 053475/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2020
From: MARVELL INTERNATIONAL LTD.
To: CAVIUM INTERNATIONAL
Reel/Frame 052918/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2006
From: INTEL CORPORATION
To: MARVELL INTERNATIONAL LTD.
Reel/Frame 018515/0817 →