IP Library Granted Patent US 11,189,288
Granted Patent B2
US 11,189,288 · App. 16/743,117 · Granted Nov 30, 2021

System and method for continuous multimodal speech and gesture interaction

Inventors: Michael Johnston (New York, NY); Derya Ozkan (Playa Vista, CA)
Assignee: Nuance Communications, Inc.
G10L15/22G06F3/017G06F3/167G06F2203/0381G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,189,288
App. No.
16/743,117
Granted
Nov 30, 2021
Kind
B2
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing multimodal input. A system configured to practice the method continuously monitors an audio stream associated with a gesture input stream, and detects a speech event in the audio stream. Then the system identifies a temporal window associated with a time of the speech event, and analyzes data from the gesture input stream within the temporal window to identify a gesture event. The system processes the speech event and the gesture event to produce a multimodal command. The gesture in the gesture input stream can be directed to a display, but is remote from the display. The system can analyze the data from the gesture input stream by calculating an average of gesture coordinates within the temporal window.

Claims (30)

1. A system comprising:

a processor;

a gesture recognition module;

a speech recognition module; and

a computer-readable storage medium storing instructions which, when executed by the processor, cause the processor to perform operations comprising:

receiving, at the speech recognition module, speech from a person within a predefined space;

identifying, by the gesture recognition module, a gesture from the person; and

generating a command based on the speech and the gesture, wherein a temporal window associated with processing the speech is either modified or unmodified from an original temporal window associated with the speech based on the gesture.

2. The system of claim 1 , wherein the predefined space comprises a medical room.

3. The system of claim 1 , wherein the gesture relates to a medical procedure.

4. The system of claim 1 , wherein the temporal window associated with processing the speech is either modified or unmodified from the original temporal window associated with the speech based a type of the gesture.

5. The system of claim 1 , wherein the gesture is associated with parameters describing specific three-dimensional characteristics of the gesture.

6. The system of claim 1 , wherein the temporal window comprises a sliding time window with a width in time that depends on a type of the gesture.

7. The system of claim 1 , wherein the gesture is chosen from a set of gestures.

8. The system of claim 1 , wherein processing the speech and the gesture comprises carrying out the command.

9. The system of claim 1 , wherein the temporal window associated with the speech and the gesture is further associated with other data separate from speech data or gesture data.

10. The system of claim 1 , wherein the gesture comprises a medical procedure gesture or a non-tactile gesture.

11. A method comprising:

receiving, at a speech recognition module configured on a computing device, speech from a person within a predefined space;

identifying, by a gesture recognition module of the computing device, a gesture from the person; and

generating, via a processor of the computing device, a command based on the speech and the gesture, wherein a temporal window associated with processing the speech is either modified or unmodified from an original temporal window associated with the speech based on the gesture.

12. The method of claim 11 , wherein the redefined space comprises a medical room.

13. The method of claim 11 , wherein the gesture relates to a medical procedure.

14. The method of claim 11 , wherein the temporal window associated with processing the speech is either modified or unmodified from the original temporal window associated with the speech based a type of the gesture.

15. The method of claim 11 , wherein the gesture is associated with parameters describing specific three-dimensional characteristics of the gesture.

16. The method of claim 11 , wherein the temporal window comprises a sliding time window with a width in time that depends on a type of the gesture.

17. The method of claim 11 , wherein the gesture is chosen from a set of gestures.

18. The method of claim 11 , wherein processing the speech and the gesture comprises carrying out the command.

19. The method of claim 11 , wherein the temporal window associated with the speech and the gesture is further associated with other data separate from speech data or gesture data.

20. The method of claim 11 , wherein the gesture comprises a medical procedure gesture or a non-tactile gesture.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2021
From: JOHNSTON, MICHAEL; OZKAN, DERYA
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 057540/0626 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2021
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 057540/0896 →
Continuity (4)
Continuation 15651315 · Jul 17, 2017
Continuation 14875105 · Oct 5, 2015
Continuation 13308846 · Dec 1, 2011
Related Publication 20200150921A1 · May 14, 2020