IP Library Granted Patent US 8,762,151
Granted Patent B2
US 8,762,151 · App. 13/161,872 · Granted Jun 24, 2014

Speech recognition for premature enunciation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,762,151
App. No.
13/161,872
Granted
Jun 24, 2014
Kind
B2
Abstract

Methods of automatic speech recognition for premature enunciation. In one method, a) a user is prompted to input speech, then b) a listening period is initiated to monitor audio via a microphone, such that there is no pause between the end of step a) and the beginning of step b), and then the begin-speaking audible indicator is communicated to the user during the listening period. In another method, a) at least one audio file is played including both a prompt for a user to input speech and a begin-speaking audible indicator to the user, b) a microphone is activated to monitor audio, after playing the prompt but before playing the begin-speaking audible indicator in step a), and c) speech is received from the user via the microphone.

Claims (24)

1. A method of automatic speech recognition, comprising the steps of:

a) prompting a user to input speech;

b) after step a), initiating a listening period to monitor audio via a microphone, such that there is no pause between the end of step a) and the beginning of step b); and then

c) communicating a begin-speaking audible indicator to the user during the listening period;

d) receiving, via the microphone, signals corresponding to input speech from the user and signals corresponding to the begin-speaking audible indicator; and

e) determining if the input speech was prematurely enunciated;

f) applying a digital filter to filter out signals corresponding to the begin-speaking audible indicator, if it is determined in step e) that the user's input speech was prematurely enunciated;

g) estimating speech energies inadvertently filtered from the speech; and

h) appending the identified speech energies to the received speech signals.

2. The method of claim 1 , wherein step g) includes applying context-dependent phoneme models to estimate the speech energies.

3. The method of claim 2 , wherein the context-dependent phoneme models include maximum mutual information models.

4. The method of claim 2 , further comprising the step of i) decoding the received input speech signals with the speech energies appended in step h) to produce a plurality of hypotheses for the user's input speech.

5. The method of claim 4 , further comprising the step of j) post-processing the plurality of hypotheses to identify one of the hypotheses as the user's input speech.

6. The method of claim 1 , further comprising the step of decoding the received input speech signals to produce a plurality of hypotheses for the user's input speech, if it is determined in step e) that the user's input speech was not prematurely enunciated.

7. The method of claim 6 , further comprising the step of post-processing the plurality of hypotheses to identify one of the hypotheses as the user's input speech.

8. A method of automatic speech recognition, comprising the steps of:

a) determining whether a prompt for a user to input speech is concatenated from at least two audio files;

b) based on the determination of step (a), selecting at least two audio files to be played, wherein one of the audio files includes both a prompt for a user to input speech and a begin-speaking audible indicator to the user and at least one other audio file includes another prompt for the user to input speech and excludes the begin-speaking audible indicator;

c) activating a microphone to monitor audio, after playing the prompts but before playing the begin-speaking audible indicator in step a); and

d) receiving speech from the user via the microphone.

9. The method of claim 8 , further comprising the step of:

d) decoding the speech received from the user using at least one speech recognition model that includes the begin-speaking audible indicator.

10. The method of claim 8 , wherein step a) includes playing the audio file including the begin-speaking audible indicator after the audio file excluding the begin-speaking audible indicator.

11. The method of claim 8 , wherein step a) includes playing a single audio file including the begin-speaking audible indicator.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 7, 2014
From: WILMINGTON TRUST COMPANY
To: GENERAL MOTORS LLC
Reel/Frame 034183/0436 →
SECURITY AGREEMENT Recorded Jun 22, 2012
From: GENERAL MOTORS LLC
To: WILMINGTON TRUST COMPANY
Reel/Frame 028423/0432 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2011
From: CORREIA, JOHN J.; CHENGALVARAYAN, RATHINAVELU; TALWAR, GAURAV; ZHAO, XUFANG
To: GENERAL MOTORS LLC
Reel/Frame 026696/0701 →