IP Library Granted Patent US 7,418,387
Granted Patent B2
US 7,418,387 · App. 10/996,732 · Granted Aug 26, 2008

Generic spelling mnemonics

Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,418,387
App. No.
10/996,732
Granted
Aug 26, 2008
Kind
B2
Abstract

A system and method for creating a mnemonics Language Model for use with a speech recognition software application, wherein the method includes generating an n-gram Language Model containing a predefined large body of characters, wherein the n-gram Language Model includes at least one character from the predefined large body of characters, constructing a new language Model (LM) token for each of the at least one character, extracting pronunciations for each of the at least one character responsive to a predefined pronunciation dictionary to obtain a character pronunciation representation, creating at least one alternative pronunciation for each of the at least one character responsive to the character pronunciation representation to create an alternative pronunciation dictionary and compiling the n-gram Language Model for use with the speech recognition software application, wherein compiling the Language Model is responsive to the new Language Model token and the alternative pronunciation dictionary.

Claims (28)

1. A method for creating a mnemonic language model, the method comprising:

generating an n-gram language model from a character string;

constructing a token representing a character from the n-gram language model, the token including a pronunciation representing the character and a pronunciation representing a term meaning “as in”;

extracting a pronunciation from a dictionary for a word, the word beginning with the character;

creating an alternative pronunciation by pre-pending the token to the pronunciation for the word; and

compiling the n-gram language model and the alternative pronunciation to form the mnemonic language model.

2. The method of claim 1 , wherein the character string includes at least one of letters including lower case letters and upper case letters, and numbers, and symbols.

3. The method of claim 2 , wherein at least one of the character, the word, the dictionary and the alternative pronunciation conforms to the English language.

4. The method of claim 1 , wherein the constructing includes constructing a token for each character of the character string.

5. The method of claim 1 , wherein the constructing a token includes appending a long silence representation to the pronunciation of the word to form the alternative pronunciation.

6. The method of claim 1 , wherein, if the character is an upper case character, the constructing the token further includes prepending a representation of a term meaning “capital” to the token to form the alternative pronunciation.

7. The method of claim 1 , wherein the n-gram language model is generated using an ARPA format.

8. The method of claim 1 , wherein computer-executable instructions for carrying out the method are embodied on computer-readable media.

9. The method of claim 1 , wherein at least one of the character, the word, the dictionary and the alternative pronunciation conforms to a spoken language.

10. A method for creating a mnemonic language model, the method comprising:

generating an n-gram language model from a character string, wherein the n-gram language model includes a character from the character string;

constructing a token representing a mnemonic spelling of the character, the token including a pronunciation representing the character and a pronunciation representing a term meaning “as in”;

extracting a pronunciation for the character from a dictionary;

creating an alternative pronunciation for the character using the pronunciation for the character;

extracting a word pronunciation from the dictionary for a word, the word beginning with the character;

pre-pending the token and appending a long silence representation to the word pronunciation to form the alternative pronunciation; and

compiling the n-gram language model and the alternative pronunciation to form the mnemonic language model.

11. The method of claim 10 , wherein the character string includes at least one of letters including lower case letters and upper case letters, and numbers, and symbols.

12. The method of claim 10 , wherein at least one of the character, the dictionary and the alternative pronunciation conforms to the English language.

13. The method of claim 10 , wherein if the character is an upper case character, the constructing the token further includes pre-pending a representation of a term meaning “capital” to the token to form the alternative pronunciation.

14. The method of claim 10 , wherein the n-gram language model is generated using an ARPA format.

15. The method of claim 10 , wherein computer-executable instructions for carrying out the method are embodied on computer-readable media.

16. The method of claim 10 , wherein at least one of the character, the dictionary and the alternative pronunciation conforms to a spoken language.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034543/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2005
From: MOWATT, DAVID; CHAMBERS, ROBERT L.; CHELBA, CIPRIAN; WU, QIANG
To: MICROSOFT CORPORATION
Reel/Frame 016481/0413 →
Continuity (1)
Related Publication 20060111907A1 · May 25, 2006