System and method for continuous multimodal speech and gesture interaction
View Patent ↗Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing multimodal input. A system configured to practice the method continuously monitors an audio stream associated with a gesture input stream, and detects a speech event in the audio stream. Then the system identifies a temporal window associated with a time of the speech event, and analyzes data from the gesture input stream within the temporal window to identify a gesture event. The system processes the speech event and the gesture event to produce a multimodal command. The gesture in the gesture input stream can be directed to a display, but is remote from the display. The system can analyze the data from the gesture input stream by calculating an average of gesture coordinates within the temporal window.
1. A method comprising:
monitoring an audio stream associated with a non-tactile gesture input stream;
identifying a speech event, in the audio stream, from a first user;
determining a temporal window associated with a time of the speech event, wherein the temporal window extends forward and backward from the time of the speech event;
analyzing, via a processor, data from the non-tactile gesture input stream within the temporal window to identify, based on the speech event, a non-tactile gesture event;
identifying clarifying information, in the audio stream, about the speech event from a second user;
applying the clarifying information to the speech event to yield a clarification; and
processing, based on the clarification, the speech event and the non-tactile gesture event to produce a multimodal command.
2. The method of claim 1 , wherein a non-tactile gesture in the non-tactile gesture input stream is directed to a display, but is remote from the display.
3. The method of claim 1 , wherein analyzing of the data from the non-tactile gesture input stream further comprises calculating an average of non-tactile gesture coordinates within the temporal window.
4. The method of claim 1 , wherein the speech event comprises a speech command and wherein processing of the speech event and the non-tactile gesture event further comprises:
identifying parameters from the non-tactile gesture event; and
applying the parameters and the clarifying information to the speech command.
5. The method of claim 4 , wherein a gesture filtering module focuses the temporal window based on a timing of specific words in the speech event and the speech command.
6. The method of claim 1 , further comprising executing the multimodal command.
7. The method of claim 1 , wherein one of a length and a position of the temporal window is based on a type of the speech event.
8. The method of claim 1 , wherein the speech event is detected in the audio stream without an explicit user activation via one of a button press and a touch gesture.
9. The method of claim 1 , wherein the non-tactile gesture input stream comprises input to one of a motion detector, a motion capture system, a camera, and an infrared camera.
10. The method of claim 1 , wherein the audio stream comprises input from one of a microphone and an array of microphones.
11. A system comprising:
a processor; and
a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:
monitoring an audio stream associated with a non-tactile gesture input stream;
identifying a speech event, in the audio stream, from a first user;
determining a temporal window associated with a time of the speech event, wherein the temporal window extends forward and backward from the time of the speech event;
analyzing data from the non-tactile gesture input stream within the temporal window to identify, based on the speech event, a non-tactile gesture event;
identifying clarifying information, in the audio stream, about the speech event from a second user;
applying the clarifying information to the speech event to yield a clarification; and
processing, based on the clarification, the speech event and the non-tactile gesture event to produce a multimodal command.
12. The system of claim 11 , wherein a non-tactile gesture in the non-tactile gesture input stream is directed to a display, but is remote from the display.
13. The system of claim 11 , wherein analyzing of the data from the non-tactile gesture input stream further comprises calculating an average of non-tactile gesture coordinates within the temporal window.
14. The system of claim 11 , wherein the speech event comprises a speech command and wherein processing of the speech event and the non-tactile gesture event further comprises:
identifying parameters from the non-tactile gesture event; and
applying the parameters and the clarifying information to the speech command.
15. The system of claim 14 , wherein a gesture filtering module focuses the temporal window based on a timing of specific words in the speech event and the speech command.
16. The system of claim 11 , further comprising executing the multimodal command.
17. The system of claim 11 , wherein one of a length and a position of the temporal window is based on a type of the speech event.
18. The system of claim 11 , wherein the speech event is detected in the audio stream without an explicit user activation via one of a button press and a touch gesture.
19. The system of claim 11 , wherein the non-tactile gesture input stream comprises input to one of a motion detector, a motion capture system, a camera, and an infrared camera.
20. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:
monitoring an audio stream associated with a non-tactile gesture input stream;
identifying a speech event, in the audio stream, from a first user;
determining a temporal window associated with a time of the speech event, wherein the temporal window extends forward and backward from the time of the speech event;
analyzing data from the non-tactile gesture input stream within the temporal window to identify, based on the speech event, a non-tactile gesture event;
identifying clarifying information, in the audio stream, about the speech event from a second user;
applying the clarifying information to the speech event to yield a clarification; and
processing, based on the clarification, the speech event and the non-tactile gesture event to produce a multimodal command.