Turn-taking model
A method is claimed for managing interactive dialog between a machine and a user. In one embodiment, an interaction between the machine and the user is managed in response to a timing position of possible speech onset from the user. In another embodiment, the interaction between the machine and the user is dependent upon the timing of a recognition result, which is relative to a cessation of a verbalization of a desired sequence from the machine. In another embodiment, the interaction between the machine and the user is dependent upon a recognition result and whether the desired sequence was ceased or not ceased.
1 . A method for managing interactive dialog between a machine and a user comprising:
verbalizing at least one desired sequence of one or more spoken phrases;
enabling a user to hear the at least one desired sequence of one or more spoken phrases;
receiving audio input from the user or an environment of the user;
determining a timing position of a possible speech onset from the audio input; and
managing an interaction between the at least one desired sequence of spoken phrases and the audio input; in response to the timing position of the possible speech onset from the audio input.
2 . The method of claim 1 further comprising managing the interaction in response to a timing position of a possible speech onset within a plurality of time zones, wherein the at least one desired sequence of one or more spoken phrases comprises the plurality of time zones.
3 . The method of claim 2 , wherein the plurality of time zones are dependent upon a continuous model of onset likelihood.
4 . The method of claim 1 , further comprising adjusting the at least one desired sequence of one or more spoken phrases in response to the timing position of the possible speech onset from the audio input.
5 . The method of claim 4 , further comprising:
stopping the at least one desired sequence of one or more spoken phrases;
restarting the at least one desired sequence of one or more spoken phrases; or
continuing the at least one desired sequence of one or more spoken phrases.
6 . The method of claim 5 , further comprising:
adjusting the timing corresponding to stopping the at least one desired sequence of one or more spoken phrases;
adjusting the timing corresponding to restarting the at least one desired sequence of one or more spoken phrases; or
adjusting the timing corresponding to continuing the at least one desired sequence of one or more spoken phrases.
7 . The method of claim 5 , further comprising:
continuing the at least one desired sequence of one or more spoken phrases for a period of time in response to an interruption of the audio input; and
receiving audio input during the period of time.
8 . The method of claim 1 , wherein a configuration of a process to produce a recognition result from the audio input is dependent upon the timing position of the possible speech onset.
9 . The method of claim 2 , wherein a possible speech onset by the audio input during a beginning portion of one time zone is considered to be in response to a previous time zone.
10 . The method of claim 1 , wherein audio input further comprises user input that corresponds to dual tone multi frequency (“DTMF”).
11 . A method for interactive machine-to-person dialog comprising:
verbalizing at least one desired sequence of one or more spoken phrases;
enabling a user to hear the at least one desired sequence of one or more spoken phrases;
receiving audio input from the user or an environment of the user;
detecting a possible speech onset from the audio input;
ceasing the at least one desired sequence of one or more spoken phrases in response to a detection of the possible speech onset; and
managing an interaction between the at least one desired sequence of one or more spoken phrases and the audio input, wherein the interaction is dependent upon the timing of at least one recognition result relative to a cessation of the at least one desired sequence.
12 . The method of claim 11 , further comprising restarting or not restarting the at least one desired sequence of one or more spoken phrases in response to the timing position of receipt of the recognition result.
13 . The method of claim 12 , wherein restarting the at least one desired sequence of one or more spoken phrases further comprises altering the wording or intonation of the at least one desired sequence of one or more spoken phrases.
14 . The method of claim 12 , wherein restarting the at least one desired sequence of spoken phrases further comprises restarting the at least one desired sequence of spoken phrases from a point that is not a beginning point of the at least one desired sequence of spoken phrases.
15 . The method of claim 12 , wherein restarting the at least one desired sequence of spoken phrases further comprises restarting the at least one desired sequence of spoken phrases from a point that is substantially near to where the desired sequence of one or more spoken phrases ceased.
16 . The method of claim 11 , further comprising adjusting an amplitude of the at least one desired sequence of one or more spoken phrases in response to a possible speech onset, wherein ceasing the at least one desired sequence of one or more phrases is achieved by a modulation of amplitude over time. (D 3 )
17 . A method for interactive machine-to-person dialog comprising:
verbalizing at least one desired sequence of one or more spoken phrases;
enabling a user to hear the at least one desired sequence of one or more spoken phrases;
receiving audio input from the user or an environment of the user;
detecting a possible speech onset from the audio input;
ceasing the at least one desired sequence of one or more spoken phrases in response to a detection of possible speech onset at a point where onset occurred while the desired sequence was being verbalized; and
managing a continuous interaction between the at least one desired sequence of one or more spoken phrases and the audio input, wherein the interaction is dependent upon at least one recognition result and whether the desired sequence of one or more spoken phrases was ceased or not ceased.
18 . The method of claim 17 , wherein in response to a low confidence recognition result, a subsequent desired sequence of one or more spoken phrases does not cease after a detection of a subsequent possible speech onset.
19 . The method of claim 18 , wherein the subsequent desired sequence of one or more spoken phrases is substantially the same as the desired sequence of one or more spoken phrases.
20 . The method of claim 18 , further comprising, in response to a subsequent low confidence recognition result, receiving audio input while continuing to verbalize the at least one desired sequence of one or more spoken phrases, and in response to a subsequent high confidence recognition result, the subsequent desired sequence of one or more spoken phrases ceases after detection of possible speech onset.