IP Library Granted Patent US 12,374,317
Granted Patent B2
US 12,374,317 · App. 18/358,236 · Granted Jul 29, 2025

System and method for using gestures and expressions for controlling speech applications

Inventors: Sahaj Garg (San Francisco, CA); Tanay Kothari (San Francisco, CA); Anthony Leonardo (Broadlands, VA)
Assignee: Wispr AI, Inc.
G10L13/027G06F3/011G06F3/012G06F3/015G06F3/017G06N3/092G06N20/00G10L13/033G10L13/047G10L15/18G10L15/22G10L15/24G10L15/25G10L19/012G10L19/04G10L25/18G10L25/78G06F2203/011G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,317
App. No.
18/358,236
Granted
Jul 29, 2025
Kind
B2
Abstract

Methods and systems are provided for detecting and processing gestures, expressions (e.g., facial), tone and/or gestures of the user for the purpose of improving the quality and speed of interactions with computer-based systems. Such information may be detected by one or more sensors such as, for example, electromyography (EMG) sensors used to monitor and record electrical activity produced by muscles that are activated. Other sensor types may be used, such as optical, inertial measurement unit (IMU), or other types of bio-sensors. The system may use one or more sensors to detect speech alone or in combination with gestures, expressions (e.g., facial), tone and/or gestures of the user to provide input or control of the system.

Claims (46)

1. A system comprising:

a wearable device configured to detect electromyography (EMG) signals;

a speech detection unit configured to detect speech from a user based at least in part on the detected EMG signals;

a trained model configured to determine a facial expression, tone, and/or a gesture of the user based on at least one EMG signal of the detected EMG signals wherein the at least one EMG signal is indicative of the facial expression, tone, and/or gesture produced by the user; and

one or more processors configured to receive as input, the detected speech and the determined facial expression, tone, and/or gesture and determine at least one of a control or an output of the system responsive to the detected input speech and the detected facial expression, tone, and/or gesture of the user, wherein:

the control or output of the system comprises a feedback signal to the system based on at least one second detected speech and detected facial expression of the user responsive to the determined control or output; and

the feedback signal is configured to cause the system to update the control or output when the feedback signal provides a negative indication from the user.

2. The system according to claim 1 , wherein the gesture of the user is a facial or head gesture.

3. The system according to claim 1 , wherein the speech detection unit configured to detect the speech from the user is configured to detect silent speech from the user.

4. The system according to claim 1 , wherein the trained model configured to determine the facial expression, tone, and/or gesture of the user is responsive to the at least one EMG signal of the detected EMG signals measured by a sensor in contact with a portion of a face of the user.

5. The system according to claim 4 , wherein the at least one EMG signal of the detected EMG signals is provided as input to the trained model to determine an output indicating the facial expression, tone, and/or gesture of the user.

6. The system according to claim 1 , wherein the trained model configured to determine the facial expression, tone, and/or gesture of the user is responsive to signals from one or more sensors configured to measure signals indicative of the facial expression, tone, and/or gesture of the user.

7. The system according to claim 6 , wherein the one or more sensors comprise one or more sensor types including an optical sensor, an inertial measurement sensor, a camera, or a biosensor.

8. The system according to claim 6 , wherein the one or more sensors are part of the wearable device positioned on the user.

9. The system according to claim 6 , wherein the trained model is a trained machine learning model, and wherein the at least one EMG signal of the detected EMG signals and the measured signals indicative of the facial expression, tone, and/or gesture are provided as input to the trained machine learning model to determine one or more outputs indicating the facial expression, tone, and/or gesture and input speech from a user.

10. The system according to claim 1 , wherein the system is configured to receive an electronic signal indicative of the speech from the user and facial muscle activation patterns of the user when the user is articulating speech.

11. The system according to claim 1 , wherein the trained model includes one or more of a statistical pattern recognition model, an unsupervised learning model, a semi-supervised learning model, a reinforcement learning model, and a machine learning model.

12. The system according to claim 1 , wherein the one or more processors configured to determine at least one of the control or the output of the system responsive to the detected input speech and the detected facial expression and/or gesture of the user are configured to create one or more output symbols responsive to the detected facial expression, tone, and/or gesture of the user.

13. The system according to claim 12 , wherein the one or more output symbols are positioned within an output text sequence generated responsive the detected input speech.

14. The system according to claim 1 , wherein the one or more processors configured to determine at least one of the control or the output of the system responsive to the detected input speech and the detected facial expression and/or gesture of the user are configured to create one or more output audio signals responsive to the detected facial expression, tone, and/or gesture of the user.

15. The system according to claim 1 , wherein the speech detection unit configured to detect input speech from a user and the trained model configured to detect facial expression, a tone, and/or gesture of the user operate substantially simultaneously.

16. The system according to claim 1 , wherein the one or more processors configured to determine at least one of the control or the output of the system are further configured to provide the control or output to an interactive system that produces an output or response that is provided to the user.

17. The system according to claim 16 , wherein the control of the system includes changing a mode of operation of a knowledge system responsive to the detected facial expression, tone and/or gesture of the user.

18. The system according to claim 1 , wherein the control of the system in response to the detected facial expression, tone and/or gesture includes deactivating or canceling operation of the speech detection unit that detects input speech.

19. The system according to claim 1 , wherein the control of the system in response to the detected gesture, facial expression, tone or gesture includes stopping the system while the system is in a process of providing an output or response to the user.

20. The system according to claim 1 , wherein the control of the system in response to the detected gesture, facial expression, tone or gesture is to stop a knowledge system while the knowledge system is determining a response.

21. The system according to claim 1 , wherein the detected facial expression, tone, and/or gesture of the user is determined periodically or continuously in real time.

22. The system according to claim 1 , further comprising a display configured to present a representation of the user based on the detected facial expression, tone, and/or gesture of the user.

23. The system according to claim 1 , wherein the output further indicates a micro- expression of the user.

24. A method comprising acts of:

detecting electromyography (EMG) signal by a wearable device;

detecting speech from a user by at least one processor based at least in part on the detected EMG signals;

determining a facial expression, tone, and/or a gesture of the user by the at least one processor based on at least one EMG signal of the detected EMG signals wherein the at least one EMG signal is indicative of the facial expression, tone, and/or gesture produced by the user; and

determining, using one or more processors configured to receive as input, the detected speech and the determined facial expression, tone, and/or gesture of the user, at least one of a control or an output of a system responsive to the detected input speech and the detected facial expression, tone, and/or gesture of the user, wherein:

the control or output of the system comprises a feedback signal to the system based on at least one second detected speech and detected facial expression of the user responsive to the determined control or output; and

updating the control or output when the feedback signal provides a negative indication from the user.

25. The method according to claim 24 , wherein the act of determining a facial expression, tone, and/or a gesture of the user by the at least one processor includes an act of measuring the at least one EMG signal by a sensor in contact with the user.

26. The method according to claim 25 , further comprising providing at least one model, and providing the at least one EMG signal as an input to the at least one model to determine an output indicating the facial expression, tone, and/or gesture of the user.

27. The method according to claim 25 , further comprising providing at least one trained machine learning model, and providing the at least one EMG signal as an input to the at least one machine learning model to determine one or more outputs indicating the facial expression, tone, and/or gesture and input speech from a user.

28. A non-transitory computer-readable medium containing instruction that, when executed, cause at least one computer hardware processor to perform a method comprising acts of:

detecting electromyography (EMG) signal by a wearable device;

detecting speech from a user by at least one processor based at least in part on the detected EMG signals;

determining a facial expression, tone, and/or a gesture of the user by the at least one processor based on at least one EMG signal of the detected EMG signals wherein the at least one EMG signal is indicative of the facial expression, tone, and/or gesture produced by the user; and

determining, using one or more processors configured to receive as input, the detected speech and the determined facial expression, tone, and/or gesture of the user, at least one of a control or an output of a system responsive to the detected input speech and the detected facial expression, tone, and/or gesture of the user, wherein:

the control or output of the system comprises a feedback signal to the system based on at least one second detected speech and detected facial expression of the user responsive to the determined control or output; and

updating the control or output when the feedback signal provides a negative indication from the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2023
From: GARG, SAHAJ; KOTHARI, TANAY; LEONARDO, ANTHONY
To: WISPR AI, INC.
Reel/Frame 065275/0477 →
Continuity (2)
Provisional Application 63437088 · Jan 4, 2023
Related Publication 20240221753A1 · Jul 4, 2024
References Cited (79)
US 8082149B2 · Schultz et al. · 2011 [cited by applicant]
US 9013264B2 · Parshionikar et al. · 2015 [cited by applicant]
US 10559145B1 · Almehmadi · 2020 [cited by applicant]
US 10878818B2 · Kapur et al. · 2020 [cited by applicant]
US 11709548B2 · Tadi · 2023 [cited by examiner]
US 12105876B2 · Garg · 2024 [cited by examiner]
US 20120290950A1 · Rapaport et al. · 2012 [cited by applicant]
US 20150086052A1 · Park · 2015 [cited by examiner]
US 20150313496A1 · Connor · 2015 [cited by applicant]
US 20160314781A1 · Schultz et al. · 2016 [cited by applicant]
US 20170061034A1 · Ritchey et al. · 2017 [cited by applicant]
US 20170131768A1 · Budavari et al. · 2017 [cited by applicant]
US 20180046851A1 · Kienzle et al. · 2018 [cited by applicant]
US 20180239956A1 · Tadi et al. · 2018 [cited by applicant]
US 20190074012A1 · Kapur et al. · 2019 [cited by applicant]
US 20200258535A1 · Vatanparvar et al. · 2020 [cited by applicant]
US 20210124422A1 · Forsland · 2021 [cited by applicant]
US 20210183383A1 · Volovich et al. · 2021 [cited by applicant]
US 20220137702A1 · Min · 2022 [cited by examiner]
US 20220160296A1 · Rahmani et al. · 2022 [cited by applicant]
US 20220187912A1 · Alcaide · 2022 [cited by examiner]
US 20220208194A1 · Rameau et al. · 2022 [cited by applicant]
US 20230072423A1 · Osborn et al. · 2023 [cited by applicant]
US 20230077010A1 · Zhang · 2023 [cited by examiner]
US 20230078978A1 · Tadi · 2023 [cited by examiner]
US 20230130770A1 · Miller et al. · 2023 [cited by applicant]
US 20230157757A1 · Braido et al. · 2023 [cited by applicant]
US 20230157762A1 · Braido et al. · 2023 [cited by applicant]
US 20240220016A1 · Garg et al. · 2024 [cited by applicant]
US 20240220811A1 · Garg et al. · 2024 [cited by applicant]
US 20240221718A1 · Kothari et al. · 2024 [cited by applicant]
US 20240221719A1 · Kothari et al. · 2024 [cited by applicant]
US 20240221738A1 · Garg et al. · 2024 [cited by applicant]
US 20240221741A1 · Kothari et al. · 2024 [cited by applicant]
US 20240221751A1 · Garg et al. · 2024 [cited by applicant]
US 20240221762A1 · Garg et al. · 2024 [cited by applicant]
US 20250061885A1 · Garg et al. · 2025 [cited by applicant]
CA 2923979A1 · 2014 [cited by applicant]
CA 2998687A1 · 2018 [cited by applicant]
CA 3164001A1 · 2021 [cited by applicant]
CA 2942852C · 2023 [cited by applicant]
CA 2957766C · 2023 [cited by applicant]
CN 101902960A · 2010 [cited by applicant]
CN 112424859A · 2021 [cited by applicant]
CN 114703091A · 2022 [cited by applicant]
KR 1020150104345A · 2015 [cited by applicant]
KR 1020200127150A · 2020 [cited by applicant]
WO WO2019040669A1 · 2019 [cited by applicant]
WO WO2022020968A1 · 2022 [cited by applicant]
WO WO2022245833A2 · 2022 [cited by applicant]
Gaddy et al., An Improved Model for Voicing Silent Speech. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and 11th International Joint Conference on Natural Language Processing. … [cited by applicant]
Gaddy et al., Digital Voicing of Silent Speech. arXiv:2010.02960v1 [eess.AS] Oct. 6, 2020. 10 pages. [cited by applicant]
He et al., Unvoiced Speech Recognition Algorithm Based on Myoelectric Signal. ICMLC. Feb. 15-17, 2020. 7 Pages. [cited by applicant]
Herff et al., Impact of Different Feedback Mechanisms in EMG-based Speech Recognition. Interspeech. Aug. 28-31, 2011:2213-2216. [cited by applicant]
Ishak, Speaker Identification Based on Vocal Cords Vibrations signal: Effect of the Window. International Journal of Digital Information and Wireless Communications. Jan. 2018. Doi:10.17781/P002406. 6 Pages. [cited by applicant]
Jamil et al., A flexible Speech Recognition System for Cerebral Palsy Disabled. ICIEIS. 2011. pp. 42-55. [cited by applicant]
Jou et al., Articulatory Feature Classification Using Surface Electromyography. International Central for Advanced Technologies. IEEE. 2006. 4 Pages. [cited by applicant]
Jou et al., Towards Continuous Speech Recognition Using Surface Electromyography. Interspeech. Sep. 17-21, 2006. 4 Pages. [cited by applicant]
Kapur et al., Alterego: A Personalized Wearable Silent Speech Interface. Session 1B: Multimodel Interfaces. IUI. Mar. 5, 2018. 11 Pages. [cited by applicant]
Kapur, How AI Could Become an Extension of your mind. TED. YouTube. Jun. 6, 2019. [cited by applicant]
Karnjanedecha et al., Signal Modeling for High-Performance Robust Isolated Word Recognition. IEEE Transactions of Speech and Audio Processing. Sep. 2001;9(6):647-654. [cited by applicant]
Maier-Hein et al., Session Independent Non-Audible Speech Recognition Using Surface Electromyography. IEEE Workshop on Automatic Speech Recognition and Understanding. 2005. 6 Pages. [cited by applicant]
Maier-Hein, Speech Recognition Using Surface Electromyography. Diplomarbeit Thesis. Aug. 2005. 129 Pages. [cited by applicant]
Manabe et al., Unvoiced Speech Recognition using EMG—Mime Speech Recognition-.Short Talk: Brains, Eyes and Ears. CHI. Apr. 5-10, 2003. 2 Pages. [cited by applicant]
Meltzner et al., Development of sEMG sensors and algorithms for silent speech recognition. J Neural Eng. Aug. 2018;15(4):046031. doi: 10.1088/1741-2552/aac965. Epub Jun. 1, 2018. PMID: 29855428; PMCID: PMC6168082. [cited by applicant]
Meltzner et al., Silent Speech Recognition as an Alternative Communication Device for Persons with Laryngectomy. IEEE/ACM Trans Audio Speech Lang Process. Dec. 2017;25(12):2386-2398. doi: 10.1109/TASLP.2017.2740000. Epu… [cited by applicant]
Polur et al., Experiments with Fast Fourier Transform, Linear Predictive and Cepstral Coefficients in Dysarthric Speech Recognition Algorithms Using Hidden Markov Model. IEEE Transactions on Neural Systems and Rehabilit… [cited by applicant]
Prajapati et al., A Survey on Isolated Word and Digit Recognition Using Different Techniques. International Journal of Computer Applications (0975-8887). Mar. 2017;161(3). 10 Pages. [cited by applicant]
Purcher, Apple Invents a next-generation AirPods Sensor System that could measure Biosignals and Electrical Activity of a user's Brain. Jul. 20, 2023. 9 pages. https://www.patentlyapple.com/2023/07/apple-invents-a-next-… [cited by applicant]
Rajasekaran et al., Recognition of Speech Under Stress and in Noise. ICASSP. 1986. 4 Pages. [cited by applicant]
Schultz et al., Modeling coarticulation in EMG-based continuous speech recognition. ScienceDirect. Speech Communications 52 (2010). Dec. 2, 2009. 13 Pages. [cited by applicant]
Toruk et al., Short Utterance Speaker Recognition Using Time-Delay Neural Network. 16th International Multi-Conference on Systems, Signals & Devices (SSD'19). 2019. 4 Pages. [cited by applicant]
Wand et al., The EMG-UKA for Electromyographic Speech Processing. Interspeech. Sep. 14-18, 2019. 5 Pages. [cited by applicant]
Wand et al., Towards Speaker-Adaptive Speech Recognition Based on Surface Electromyography. ICBSSP. 2009. 8 Pages. [cited by applicant]
Wand, Advancing Electromyographic Continuous Speech Recognition. Signal Preprocessing and Modeling. Scientific Publishing. Jan. 14, 2014. 256 Pages. [cited by applicant]
Wand, Speaker-Adaptive Speech Recognition Based on Surface Electromyography. IJCBEST. 2009. 15 Pages. [cited by applicant]
Wand, Towards Real-life Application of EMG-based Speech Recognition by using Unsupervised Adaptation. Interspeech. Sep. 14-18, 2014. 5 Pages. [cited by applicant]
Zhou et al., Improved Phoneme-Based Myoelectric Speech Recognition. IEEE Transactions on Biomedical Engineering. Aug. 2009;56(8). 8 Pages. [cited by applicant]
International Search Report and Written Opinion dated Apr. 26, 2024, in connection with International Application No. PCT/US24/10268. [cited by applicant]