System and method for continuous multimodal speech and gesture interaction
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing multimodal input. A system configured to practice the method continuously monitors an audio stream associated with a gesture input stream, and detects a speech event in the audio stream. Then the system identifies a temporal window associated with a time of the speech event, and analyzes data from the gesture input stream within the temporal window to identify a gesture event. The system processes the speech event and the gesture event to produce a multimodal command. The gesture in the gesture input stream can be directed to a display, but is remote from the display. The system can analyze the data from the gesture input stream by calculating an average of gesture coordinates within the temporal window.
1. A method comprising:
monitoring an audio stream associated with a non-tactile gesture input stream;
identifying a speech event in an audio stream;
analyzing, via a processor, data from the non-tactile gesture input stream to identify a non-tactile gesture event;
determining, based on the non-tactile gesture event, a temporal window associated with a time of the speech event; and
processing the speech event and the non-tactile gesture event within the temporal window to produce a multimodal command, wherein the temporal window comprises one of a modified temporal window representing a change from an original temporal window associated with the speech event or the original temporal window associated with the speech event and which is unchanged based on a type of the non-tactile gesture event.
2. The method of claim 1 , wherein the type of the non-tactile gesture event is associated with parameters describing specific three-dimensional characteristics of the non-tactile gesture event.
3. The method of claim 1 , wherein the temporal window comprises a sliding time window with a width in time that depends on the type of the non-tactile gesture event.
4. The method of claim 1 , wherein the type of the non-tactile gesture event is chosen from a set of gestures.
5. The method of claim 1 , wherein processing the speech event and the non-tactile gesture event comprises carrying out the multimodal command.
6. The method of claim 1 , wherein determining the temporal window associated with the time of the speech event and based on the type of the non-tactile gesture event is further based on other data separate from speech event data or non-tactile gesture event data.
7. A system comprising:
a processor; and
a computer-readable storage medium storing instructions which, when executed by the processor, cause the processor to perform operations comprising:
monitoring an audio stream associated with a non-tactile gesture input stream;
identifying a speech event in an audio stream;
analyzing data from the non-tactile gesture input stream to identify a non-tactile gesture event;
determining, based on the non-tactile gesture event, a temporal window associated with a time of the speech event; and
processing the speech event and the non-tactile gesture event within the temporal window to produce a multimodal command, wherein the temporal window comprises one of a modified temporal window representing a change from an original temporal window associated with the speech event or the original temporal window associated with the speech event and which is unchanged based on a type of the non-tactile gesture event.
8. The system of claim 7 , wherein the type of the non-tactile gesture event is associated with parameters describing specific three-dimensional characteristics of the non-tactile gesture event.
9. The system of claim 7 , wherein the temporal window comprises a sliding time window with a width in time that depends on the type of the non-tactile gesture event.
10. The system of claim 7 , wherein the type of the non-tactile gesture event is chosen from a set of gestures.
11. The system of claim 7 , wherein processing the speech event and the non-tactile gesture event comprises carrying out the multimodal command.
12. The system of claim 7 , wherein determining the temporal window associated with the time of the speech event and based on the type of the non-tactile gesture event is further based on other data separate from speech event data or non-tactile gesture event data.
13. A computer-readable storage device storing instructions which, when executed by a processor, cause the processor to perform operations comprising:
monitoring an audio stream associated with a non-tactile gesture input stream;
identifying a speech event in an audio stream;
analyzing data from the non-tactile gesture input stream to identify a non-tactile gesture event;
determining, based on the non-tactile gesture event, a temporal window associated with a time of the speech event; and
processing the speech event and the non-tactile gesture event within the temporal window to produce a multimodal command, wherein the temporal window comprises one of a modified temporal window representing a change from an original temporal window associated with the speech event or the original temporal window associated with the speech event and which is unchanged based on a type of the non-tactile gesture event.
14. The computer-readable storage device of claim 13 , wherein the type of the non-tactile gesture event is associated with parameters describing specific three-dimensional characteristics of the non-tactile gesture event.
15. The computer-readable storage device of claim 13 , wherein the temporal window comprises a sliding time window with a width in time that depends on the type of the non-tactile gesture event.