IP Library Granted Patent US 12,105,876
Granted Patent B2
US 12,105,876 · App. 18/358,268 · Granted Oct 1, 2024

System and method for using gestures and expressions for controlling speech applications

Inventors: Sahaj Garg (Foster City, CA); Tanay Kothari (San Francisco, CA); Anthony Leonardo (Broadlands, VA)
Assignee: Wispr AI, Inc.
G06F3/015G06N20/00G06F2203/011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,105,876
App. No.
18/358,268
Granted
Oct 1, 2024
Kind
B2
Abstract

Methods and systems are provided for detecting and processing gestures, expressions (e.g., facial), tone and/or gestures of the user for the purpose of improving the quality and speed of interactions with computer-based systems. Such information may be detected by one or more sensors such as, for example, electromyography (EMG) sensors used to monitor and record electrical activity produced by muscles that are activated. Other sensor types may be used, such as optical, inertial measurement unit (IMU), or other types of bio-sensors. The system may use one or more sensors to detect speech alone or in combination with gestures, expressions (e.g., facial), tone and/or gestures of the user to provide input or control of the system.

Claims (55)

1. A system comprising:

a speech input device wearable on a user and configured to measure a first EMG signal indicative of a prompt provided by the user, wherein the first EMG signal is measured when the user is speaking the prompt; and

at least one processor configured to:

provide text-based or voice-based prompt information associated with the prompt provided by the user to an interactive system to cause the interactive system to take an action to control an application of the interactive system or generate a response to the prompt based on the provided prompt information, wherein the prompt information is determined at least in part based on the first EMG signal indicative of the prompt provided by the user;

receive a second EMG signal responsive to the user making a facial expression responsive to the action or the response;

use a machine learning model and the second EMG signal as input to the machine learning model to determine a feedback signal, wherein:

the machine learning model is configured to determine an indication of the facial expression the user is making responsive to the interactive system taking the action or generating the response based at least in part on the second EMG signal, and

the feedback signal is determined based at least in part on the indication of the facial expression; and

provide the feedback signal to the interactive system to cause the interactive system to take a new action or generate a new response in response to the feedback signal and the provided prompt information.

2. The system according to claim 1 , wherein the feedback signal indicates a degree of confirmation to the response.

3. The system according to claim 2 , wherein the system further comprises a knowledge system, and wherein the degree of confirmation to the response is used to determine whether the knowledge system takes an action that was indicated by the response.

4. The system according to claim 1 , wherein the feedback signal is used by the interactive system to generate a response to the user that includes a question.

5. The system according to claim 4 , wherein the system receives and processes a second prompt provided by the user response to the question.

6. The system according to claim 5 , wherein the first and second prompts are provided as inputs to a knowledge system.

7. The system according to claim 1 , wherein the interactive system is configured to sample a new input or response based on the feedback signal.

8. The system according to claim 2 , wherein the feedback signal includes an indication of a facial or a head gesture.

9. The system according to claim 8 , wherein the feedback signal includes an indication of a frown, a smile, a head nod, or a head shake.

10. The system according to claim 8 , wherein the feedback signal indicates a frown, and the indication of the frown is used to cancel or clarify the prompt.

11. The system according to claim 8 , wherein the feedback signal indicates a smile, and the indication of the smile is used to confirm the action or the response.

12. The system according to claim 1 , wherein the at least one processor is further configured to generate a second prompt and is configured to:

receive a third EMG signal from the speech input device when the user is speaking; and

use a second machine learning model and the third EMG signal as input to the second machine learning model to generate the second prompt.

13. The system according to claim 1 , wherein the at least one processor is configured to:

provide a text prompt to a knowledge system to cause the knowledge system to perform an operation and/or generate a response;

receive a feedback signal responsive to the performed operation or the response;

cause the knowledge system to, based on the feedback signal, take a new operation different from the performed operation or generate a new response different from the generated response.

14. The system according to claim 13 , wherein the feedback comprises receiving a signal from the user that the knowledge system did not perform the operation or provide the response that the user desired.

15. The system according to claim 13 , further comprising:

a speech input device wearable on a user and configured to receive a third EMG signal when the user is speaking;

receive a fourth EMG signal when the user is making a facial expression responsive to the action or the response;

use a machine learning model and the fourth EMG signal as input to the machine learning model to determine the feedback signal.

16. The system of claim 15 , wherein the feedback signal indicates one of a smile, a frown, or a head gesture.

17. A computer-implemented method used in a distributed computer system, the method comprising acts of:

measuring, by a speech input device wearable on a user, a first EMG signal indicative of a prompt provided by the user when the user is speaking the prompt; and

providing text-based or voice-based prompt information associated with the prompt provided by the user to an interactive system to cause the interactive system to take an action to control an application of the interactive system or generate a response to the prompt based on the provided prompt information, wherein the prompt information is determined at least in part based on the first EMG signal indicative of the prompt provided by the user;

receiving a second EMG signal responsive to the user making a facial expression responsive to the action or the response;

using a machine learning model and the second EMG signal as input to the machine learning model to determine a feedback signal by:

determining, with the machine learning model, an indication of the facial expression the user is making responsive to the interactive system taking the action or generating the response based at least in part on the second EMG signal, and

determining the feedback signal at least in part based on the indication of the facial expression; and

providing the feedback signal to the interactive system to cause the interactive system to take a new action or generate a new response in response to the feedback signal and the provided prompt information.

18. The method according to claim 17 , wherein the feedback signal indicates a degree of confirmation to the response.

19. The method according to claim 18 , wherein the system further comprises a knowledge system, and wherein the degree of confirmation to the response is used to determine whether the knowledge system takes an action that was indicated by the response.

20. The method according to claim 19 , further comprising using the feedback signal by the knowledge system to generate a response to the user that includes a question.

21. The system according to claim 20 , further comprising receiving and processing a second prompt provided by the user response to the question.

22. The system according to claim 21 , wherein the first and second prompts are provided as inputs to the knowledge system.

23. The system according to claim 17 , wherein the interactive system is configured to sample a new input or response based on the feedback signal.

24. A non-transitory computer-readable medium containing instruction that, when executed, cause at least one computer hardware processor to perform a method comprising acts of:

measuring, by a speech input device wearable on a user, a first EMG signal indicative of a prompt provided by the user when the user is speaking the prompt; and

providing text-based or voice-based prompt information associated with the prompt provided by the user to an interactive system to cause the interactive system to take an action to control an application of the interactive system or generate a response to the prompt based on the provided prompt information, wherein the prompt information is determined at least in part based on the first EMG signal indicative of the prompt provided by the user;

receiving a second EMG signal responsive to the user making a facial expression responsive to the action or the response;

using a machine learning model and the second EMG signal as input to the machine learning model to determine a feedback signal by:

determining, with the machine learning model, an indication of the facial expression the user is making responsive to the interactive system taking the action or generating the response based at least in part on the second EMG signal, and

determining the feedback signal at least in part based on the indication of the facial expression; and

providing the feedback signal to the interactive system to cause the interactive system to take a new action or generate a new response in response to the feedback signal and the provided prompt information.

25. The system of claim 1 , wherein the new action taken by the interaction system or the new response generated by the interactive system is based at least in part in response to the prompt and the feedback signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2023
From: GARG, SAHAJ; KOTHARI, TANAY; LEONARDO, ANTHONY
To: WISPR AI, INC.
Reel/Frame 065275/0477 →
Continuity (2)
Provisional Application 63437088 · Jan 4, 2023
Related Publication 20240220016A1 · Jul 4, 2024
Cited By (8)
US 12,346,500 US 12,367,784 US 12,374,317 US 12,444,406 US 12,525,240 US 12,586,568 US 12,632,112 US 12,670,910