IP Library Patent Application 12327594
Patent Application
App. No. 12/327,594

SYSTEM AND METHOD FOR GENERATING A PHRASE PRONUNCIATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
12/327,594
Abstract

A system and method for a speech recognition technology that allows language models to be customized through the addition of special pronunciations for components of phrases, which are added to the factory language models during customization. It allows components of a phrase to have different pronunciations inside customer-added phrases than are specified for those isolated components in the factory language models.

Claims (37)

1 .- 17 . (canceled)

18 . A method in a computer system comprising a language model, a background dictionary, and at least one preexisting pron component list, for adding phrase pronunciations to a language model, said method comprising the steps of:

inputting at least one phrase to be added to the language model;

determining if said at least one phrase is contained in said language model;

if said at least one phrase is not contained in said language model, determining if said at least one phrase is contained in said background dictionary, and, if so, adding background dictionary pronunciation to said language model;

if said at least one phrase is not contained in said language model or said background dictionary, parsing said at least one phrase into an ordered set of tokens in accordance with certain rules sequentially associating prons with each said token of said ordered set of tokens, generating a pron derived phrase pronunciation from said prons and adding said pron derived phrase pronunciation to said language model.

19 . A method, in accordance with claim 18 , wherein the step of generating said pron derived phrase pronunciation, further comprises, for each said token of said ordered set of tokens:

a) sequentially determining if each said associated pron is in said preexisting pron component list, and, if so, obtaining pronunciation from said preexisting pron component list;

b) if said associated pron is not in said preexisting pron component list, determining if said associated pron is in a preexisting language model, and, if so, adding said language model pron to said preexisting pron component list;

c) if said associated pron is not in preexisting pron component list or said preexisting language model, determining if said associated pron is in preexisting background dictionary, and, if so, adding said background dictionary pron to said preexisting pron component list;

d) if said associated pron is not in said preexisting pron component list, said preexisting language model or said preexisting background dictionary, generating a guess pron, and adding said guess pron to said preexisting pron component list;

e) if there is an additional token in said ordered set of tokens, repeating steps a) to d); and,

f) if there are no additional tokens in said ordered set of tokens, generating said pron derived phrase pronunciation by combining said associated pron pronunciations as obtained from said preexisting pron component list, in sequence.

20 . A method for adding phrase pronunciations to a language model, in accordance with claim 18 , wherein said pron component list includes punctuations or formatting that is present in the said at least one phrase but is silent in the pronunciation of said at least one phrase.

21 . A method for adding phrase pronunciations to a language model, in accordance with claim 19 , wherein said pron component list selected from a plurality of lists in accordance with the position of the said token within the said at least one phrase.

22 . A method for adding phrase pronunciations to a language model, in accordance with claim 18 , wherein said certain rules comprise breaking up the said phrase into tokens at certain boundaries.

23 . A method for adding phrase pronunciations to a language model, in accordance with claim 22 , wherein said certain boundaries comprise white spaces and/or punctuation.

24 . A method for adding phrase pronunciations to a language model, in accordance with claim 18 , wherein said certain rules comprise looking for the longest match in said preexisting language model or said preexisting background dictionary.

25 . A method for adding phrase pronunciations to a language model, in accordance with claim 19 , wherein said preexisting pron component lists comprise an initial pron component list and a non-initial pron component list.

26 . A tangible computer usable medium having computer readable instructions stored thereon for execution by a processor and comprising a language model, a background dictionary, and at least one preexisting pron component list to perform a method comprising:

inputting at least one phrase to be added to the language model;

determining if said at least one phrase is contained in said language model;

if said at least one phrase is not contained in said language model, determining if said at least one phrase is contained in said background dictionary, and, if so, adding background dictionary pronunciation to said language model;

if said at least one phrase is not contained in said language model or said background dictionary, parsing said at least one phrase into an ordered set of tokens in accordance with certain rules, sequentially associating prons with each said token of said ordered set of tokens, generating a pron derived phrase pronunciation from said prons, and adding said pron derived phrase pronunciation to said language model.

27 . A tangible computer usable medium, in accordance with claim 26 , to perform a method wherein the step of generating said pron derived phrase pronunciation further comprises, for each said token of said ordered set of tokens:

a) sequentially determining if each said associated pron is in said preexisting pron component list, and, if so, obtaining pronunciation from said preexisting pron component list;

b) if said associated pron is not in said preexisting pron component list, determining if said associated pron is in a preexisting language model, and, if so, adding said language model pron to said preexisting pron component list;

c) if said associated pron is not in preexisting pron component list or said preexisting language model, determining if said associated pron is in preexisting background dictionary, and, if so, adding said background dictionary pron to said preexisting pron component list;

d) if said associated pron is not in said preexisting pron component list, said preexisting language model or said preexisting background dictionary, generating a guess pron, and adding said guess pron to said preexisting pron component list;

e) if there is an additional token in said ordered set of tokens, repeating steps a) to d); and,

f) if there are no additional tokens in said ordered set of tokens, generating said pron derived phrase pronunciation by combining said associated pron pronunciations as obtained from said preexisting pron component list, in sequence.

28 . A computer usable medium, in accordance with claim 27 , wherein said pron component list includes punctuations or formatting that is present in the text but is silent in the pronunciation of said at least one phrase.

29 . A computer usable medium, in accordance with claim 27 , wherein said pron component list selected from a plurality of lists in accordance with the position of the said token within the said at least one phrase.

30 . A computer usable medium, in accordance with claim 26 , wherein said certain rules comprise breaking up the said phrase into tokens at certain boundaries.

31 . A computer usable medium, in accordance with claim 30 , wherein said certain boundaries comprise white spaces and/or punctuation.

32 . A computer usable medium, in accordance with claim 26 , wherein said certain rules comprise looking for the longest match in said preexisting language model or said preexisting background dictionary.

33 . A computer usable medium, in accordance with claim 27 , wherein said preexisting pron component lists comprise an initial pron component list and a non-initial pron component list.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2013
From: DICTAPHONE CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 029596/0836 →
MERGER Recorded Sep 13, 2012
From: DICTAPHONE CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028952/0397 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2008
From: COTE, WILLIAM F.; CARRIER, JILL
To: DICTAPHONE CORPORATION
Reel/Frame 021941/0435 →