IP Library Granted Patent US 7,970,615
Granted Patent B2
US 7,970,615 · App. 12/862,322 · Granted Jun 28, 2011

Turn-taking confidence

Assignee: Enterprise Integration Group, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,970,615
App. No.
12/862,322
Granted
Jun 28, 2011
Kind
B2
Abstract

A method for managing interactive dialog between a machine and a user. In one embodiment, an interaction between the machine and the user is managed by determining at least one likelihood value which is dependent upon a possible speech onset of the user. In another embodiment, the likelihood value can be dependent on a model of a desire of the user for specific items, a model of an attention of the user to specific items, or a model of turn-taking cues. The values can be used to determine a mode confidence value that is used by the system to determine the nature of prompts provided to the user.

Claims (46)

1. A method for managing interactive voice response dialog between a machine comprising automatic speech recognition and a user, said method comprising the steps of:

setting a mode confidence level parameter value to a first value prior to a first input from said user wherein said first input is of a speech input mode;

selecting one of a plurality of audio prompts comprising speech to annunciate to the user from the machine based on said first value of said mode confidence level parameter, wherein said one of said plurality of audio prompts solicits said first input comprising a first semantic response from said user in said speech input mode;

annunciating the at least one of a plurality of audio prompts to said user;

receiving said first input from said user;

determining a first speech recognition confidence level based on said first input;

setting said mode confidence level parameter to a second value based on said first speech recognition confidence level, said second value indicating a lower level of confidence of recognition relative to said first value of said mode confidence level;

selecting an another one of at least one of a plurality of audio prompts based on said mode confidence level; and

annunciating to the user from the machine said another one of least one of a plurality of audio prompts comprising speech based on said second value of said mode confidence level wherein said another one of said plurality of phrases solicits said first semantic response from said user in a DTMF input mode.

2. The method of claim 1 where said another one of at said plurality of audio prompts directs the user to press a telephone keypad in response to said another one of said plurality of audio prompts.

3. The method of claim 2 wherein the step of selecting the another one of said plurality audio prompts comprises the steps of:

a. determining a value of a speech duration confidence parameter based on the onset and offset of said first input user speech, and

b. using the value of the speech duration confidence parameter and said second value of the mode confidence level to select the another one of said plurality of audio prompts.

4. The method of claim 2 wherein said one of a plurality of audio prompts comprises a plurality of segments, and the step of selecting the another one of said plurality of audio prompts comprises the steps of:

detecting an onset of said first input user speech relative to a segment of said audio prompt;

determining a value of a turn-taking confidence level parameter based on detection of said onset; and

using the value of the turn-taking confidence level parameter and said second value of said mode confidence level parameter to select the another one of said plurality of audio prompts.

5. The method of claim 4 where said turn-taking confidence level is based in part on a time period between the beginning of said onset and an end of said segment.

6. The method of claim 4 wherein said first input is recognized as an out-of-grammar utterance.

7. The method of claim 1 further comprising the steps of:

receiving a second input from said user in response to said another one of said plurality of audio prompts;

setting said mode confidence level parameter to a third level based on said second input; and

annunciating a third audio prompt to the user from the machine based on said third level of said mode confidence level parameter wherein said third audio prompt solicits input from said user in said DTMF input mode.

8. The method of claim 1 wherein a speech recognition capability of said machine is disabled after determining said first speech recognition confidence level based on said first input.

9. A method for managing interactive voice response dialog between a machine comprising automatic speech recognition and a user, said method comprising the steps of:

setting a mode confidence level parameter at a first value prior to a first input from said user wherein said first input is of a DTMF input mode;

selecting one of a plurality of audio prompts comprising speech to annunciate to the user from the machine based on said first value of said mode confidence level parameter, wherein said one of said plurality of audio prompts solicits said first input comprising a first semantic response from said user in said DTMF input mode;

annunciating the at least one of a plurality of audio prompts to said user;

receiving said first input from said user;

determining a first speech recognition confidence level based on said first input;

setting said mode confidence level parameter to a second value based on said first speech recognition confidence level, said second value indicating a higher level of confidence of recognition relative to said mode confidence level;

selecting an another one of at least one of a plurality of audio prompts based on said mode confidence level; and

annunciating to the user from the machine said another one of at least one of a plurality of audio prompts comprising speech based on said second value of said mode confidence level wherein said another one of said plurality of audio prompts solicits a different semantic response from said user in a speech input mode.

10. The method of claim 9 wherein said first input is recognized as an in-grammar.

11. A method for managing interactive voice response dialog between a machine comprising automatic speech recognition and a user, said method comprising the steps of:

setting a mode confidence level parameter at a first value prior to a first input from said user wherein said first input is of a speech input mode;

selecting one of a plurality of audio prompts comprising a plurality of speech segments to annunciate to the user from the machine, wherein said one of said plurality of audio prompts solicits said first input comprising a first semantic response from said user in said speech input mode;

annunciating the at least one of the plurality of speech segments to said user;

receiving said first input from said user, wherein said first input comprises speech input mode;

determining a first speech recognition confidence value based on said first input;

determining an onset time of said speech relative to the at least one of the plurality of speech segment for determining a turn confidence value;

determining a speech duration of said first input, said speech duration used for determining a speech duration confidence value;

using said first speech recognition confidence value, said turn confidence value, said speech duration confidence value for setting said mode confidence level parameter at a second value;

selecting another one of at least another one of a plurality of audio prompts based on said mode confidence level parameter; and

annunciating to the user from the machine another one of said plurality of audio prompts wherein said another one of said plurality of audio prompts solicits said first semantic response from said user in a DTMF input mode.

12. The method of claim 11 wherein said speech duration of said first input is based on the duration between speech onset and speech offset of said first input, and said speech duration is longer than a specified range for an expected response.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2021
From: ELOQUI VOICE SYSTEMS, LLC
To: SHADOW PROMPT TECHNOLOGY AG
Reel/Frame 057055/0857 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2019
From: SHADOW PROMPT TECHNOLOGY AG
To: ELOQUI VOICE SYSTEMS, LLC
Reel/Frame 048170/0853 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2013
From: ENTERPRISE INTEGRATION GROUP E.I.G. AG
To: SHADOW PROMPT TECHNOLOGY AG
Reel/Frame 031212/0219 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2013
From: ENTERPRISE INTEGRATION GROUP, INC.
To: ENTERPRISE INTEGRATION GROUP E.I.G. AG
Reel/Frame 030588/0942 →
Continuity (3)
Continuation 11317391 · Dec 22, 2005
Provisional Application 60638431 · Dec 22, 2004
Related Publication 20100324896A1 · Dec 23, 2010