IP Library Granted Patent US 8,135,590
Granted Patent B2
US 8,135,590 · App. 11/652,451 · Granted Mar 13, 2012

Position-dependent phonetic models for reliable pronunciation identification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,135,590
App. No.
11/652,451
Granted
Mar 13, 2012
Kind
B2
Abstract

A representation of a speech signal is received and is decoded to identify a sequence of position-dependent phonetic tokens wherein each token comprises a phone and a position indicator that indicates the position of the phone within a syllable.

Claims (12)

1. A method comprising:

receiving a representation of a speech signal;

a processor decoding the representation of the speech signal into a sequence of words using a word language model and an acoustic model;

a processor converting each word in the sequence of words into a sequence of position-dependent phonetic tokens using a phonetic token lexicon, which provides a position-dependent phonetic token description of each word in a lexicon, wherein each position-dependent phonetic token comprises a phone and a position indicator that indicates the position of the phone within a syllable; and

a processor determining probabilities for sub-sequences in the sequences of position-dependent phonetic tokens by applying sub-sequences of position-dependent phonetic tokens converted from the sequence of words to a position-dependent phonetic language model that describes probabilities of sequences of position-dependent phonetic tokens comprising a conditional probability of a position-dependent phonetic token given at least two preceding position-dependent phonetic tokens and probabilities of individual position-dependent phonetic tokens.

2. The method of claim 1 further comprising utilizing the at least one probability as a confidence measure.

3. The method of claim 1 wherein at least one position indicator indicates one of a group of syllable positions consisting of onset consonant, and coda consonant.

4. A hardware computer storage medium encoded with a computer program, causing the computer to execute steps comprising:

decoding a representation of a speech signal using a position-dependent phonetic language model that provides probabilities of sequences of position-dependent phonetic tokens wherein each position-dependent phonetic token comprises a phone and a position identifier that identifies a position within a syllable, wherein decoding produces at least one sequence of position-dependent phonetic tokens;

receiving a symbol represented by a portion of the speech signal, wherein the symbol is not part of the language of a lexicon; and

annotating a lexicon by storing a sequence of position-dependent phonetic tokens as a pronunciation for the symbol that is not part of the language of the lexicon, wherein annotating the lexicon further comprises identifying syllable boundaries from the position-dependent phonetic tokens and placing syllable boundaries in the lexicon.

5. The hardware computer storage medium of claim 4 wherein a position identifier is one of a group of position identifiers consisting of onset consonant and coda consonant.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2007
From: LIU, PENG; SHI, YU; SOONG, FRANK KAO-PING
To: MICROSOFT CORPORATION
Reel/Frame 018913/0158 →