CONTROLLING SPEECH DIALOG USING AN ADDITIONAL SENSOR
Methods and systems are provided for managing speech dialog of a speech system. In one embodiment, a method includes: receiving information determined from a non-speech related sensor; using the information in a turn-taking function to confirm at least one of if and when a user is speaking; and generating a command to at least one of a speech recognition module and a speech generation module based on the confirmation.
1 . A method for managing speech dialog of a speech system, comprising:
receiving information determined from a non-speech related sensor;
using the information in a turn-taking function to confirm at least one of if and when a user is speaking; and
generating a command to at least one of a speech recognition module and a speech generation module based on the confirmation.
2 . The method of claim 1 , further comprising determining at least one of if and when a user is speaking based on data received from the non-speech related sensor, and wherein the information is based on the determining.
3 . The method of claim 1 , wherein the using the information comprises using the information to confirm if a particular user is speaking
4 . The method of claim 1 , wherein the using the information comprises using the information to confirm when a user is speaking
5 . The method of claim 1 , wherein the using the information comprises using the information to confirm if and when a user is speaking
6 . The method of claim 1 , wherein the generating the command comprises generating the command to the speech recognition module to at least one of start and stop speech recognition.
7 . The method of claim 1 , wherein the generating the command comprises generating the command to the speech generation module to at least one of start and stop generation of a spoken command.
8 . The method of claim 1 , wherein the turn-taking function is a system start function.
9 . The method of claim 1 , wherein the turn-taking function is a barge-in function.
10 . The method of claim 1 , wherein the turn-taking function is a speech window determination function.
11 . The method of claim 1 wherein the non-speech related sensor is at least one of an image sensor, an ultrasound sensor, and a radar sensor.
12 . A system for managing speech dialog of a speech system, comprising:
a first module that receives information determined from a non-speech related sensor, and that uses the information in a turn-taking function to confirm at least one of if and when a user is speaking; and
a second module that at least one of starts and stops at least one of speech recognition and speech generation based on the confirmation.
13 . The system of claim 12 , further comprising a third module that determines at least one of if and when a user is speaking based on data received from the non-speech related sensor, and generates the information based on the determination.
14 . The system of claim 12 , wherein the first module uses the information to confirm if a particular user is speaking.
15 . The system of claim 12 , wherein the first module uses the information to confirm when a user is speaking.
16 . The system of claim 12 , wherein the first module uses the information to confirm if and when a user is speaking.
17 . The system of claim 12 , wherein the turn-taking function is a system start function.
18 . The system of claim 12 , wherein the turn-taking function is a barge-in function.
19 . The system of claim 12 , wherein the turn-taking function is a speech window determination function.
20 . The system of claim 12 , wherein the non-speech related sensor is at least one of an image sensor, an ultrasound sensor, and a radar sensor.