IP Library Granted Patent US 12,482,449
Granted Patent B2
US 12,482,449 · App. 18/338,749 · Granted Nov 25, 2025

Systems and methods for using silent speech in a user interaction system

Inventors: Sahaj Garg (San Francisco, CA); Tanay Kothari (San Francisco, CA); Anthony Leonardo (Broadlands, VA)
Assignee: Wispr AI, Inc.
G10L13/027G06F3/011G06F3/012G06F3/015G06F3/017G06N3/092G06N20/00G10L13/033G10L13/047G10L15/18G10L15/22G10L15/24G10L15/25G10L19/012G10L19/04G10L25/18G10L25/78G06F2203/011G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,449
App. No.
18/338,749
Granted
Nov 25, 2025
Kind
B2
Abstract

The techniques described herein relate to computerized methods and systems for integrating with a knowledge system. In some embodiments, a user interaction system may include a speech input device wearable on a user and configured to receive an electronic signal indicative of a user's speech muscle activation patterns when the user is speaking. In some embodiments, the electronic signal may include EMG data received from an EMG sensor on the speech input device. The system may include at least one processor configured to use a speech model and the electronic signal as input to the speech model to generate a text prompt. The at least one processor may use a knowledge system to take an action or generate a response based on the text prompt. In some embodiments, the system may provide context to the knowledge system.

Claims (68)

1 . A user interaction system comprising:

a wearable speech input device configured to measure a signal indicative of a user's speech muscle activation patterns when the user is speaking; and

at least one processor configured to:

use a speech model and the signal as an input to the speech model to generate an output; and

use a knowledge system to take an action or generate a response based on the output.

2 . The system of claim 1 , wherein:

the speech input device includes an electromyography (EMG) sensor; and

the signal is an EMG signal captured from the EMG sensor when the user is silently speaking or whispering.

3 . The system of claim 2 , wherein the signal is recorded non-invasively from one or more regions of face and/or neck of the user via a conductive electrode coupled to an electronic amplifier system.

4 . The system of claim 2 , wherein silently speaking produces a sound volume less than or equal to 30 dB, wherein the sound volume is measured about 10 cm from a mouth of the user.

5 . The system of claim 1 further comprising one or more additional sensors configured to measure respective signals when the user is speaking.

6 . The system of claim 1 , wherein the output of the speech model comprises a text prompt and/or encoded features.

7 . The system of claim 6 , wherein the text prompt corresponds to one or more words spoken by the user.

8 . The system of claim 6 , wherein the text prompt includes two or more candidate transcripts of utterance of the user comprising one or more spoken words, wherein the two or more candidate transcripts are generated by the speech model.

9 . The system of claim 6 , wherein the encoded features include a probability distribution of different text tokens associated with a decoder of the speech model.

10 . The system of claim 1 further comprising an output device configured to output the response to a user, wherein outputting comprises one or more of:

outputting the response in a text format;

outputting the response in an auditory signal;

outputting the response in an image format;

outputting the response in a video format;

outputting the response in augmented reality; or

outputting the response in haptics.

11 . The system of claim 1 , wherein the at least one processor is further configured to:

determine an application based on the output of the speech model; and

cause the knowledge system to use the application to take the action or generate the response based on the output of the speech model.

12 . The system of claim 11 , further comprising a graphical user interface, wherein using the application comprises opening the application on the graphical user interface to receive user input.

13 . The system of claim 1 , wherein the speech model is trained for predicting speech label in a first domain using at least training data in a second domain different from the first domain.

14 . The system of claim 1 , wherein the speech model is trained for predicting speech label using voiced speech training data and silent speech training data.

15 . The system of claim 14 , wherein the knowledge system includes a machine learning foundation model.

16 . The system of claim 15 , wherein the machine learning foundation model includes a large language model.

17 . The system of claim 15 , wherein:

the output of the speech model comprises a text prompt; and

the machine learning foundation model is configured to use the text prompt as input and generate the response by sampling text from a distribution of possible responses.

18 . The system of claim 1 , wherein the at least one processor is configured to take the action by:

causing the knowledge system to use a machine learning foundation model to generate code based on the output of the speech model; and

executing the code to produce the action.

19 . The system of claim 1 , wherein the at least one processor is further configured to:

provide context to the knowledge system; and

use the knowledge system to take the action or generate the response additionally based on the context.

20 . The system of claim 19 , wherein the context includes personalized characteristic of the user, a location of the user, a contact list of the user, an address of the user, an email history of the user, a message history of the user, or a combination thereof.

21 . The system of claim 1 , wherein the action or the response includes prompting the user with a clarifying question.

22 . The system of claim 21 , wherein the at least one processor is further configured to:

receive a user input, the user input responsive to the clarifying question;

use the speech model and the user input to generate a text prompt; and

use the knowledge system to take additional action or generate additional response based on the text prompt.

23 . The system of claim 1 , wherein the knowledge system is configured to interact with a message system to:

generate a message; and

transmit the message over a communication network and/or output the message on an output device.

24 . The system of claim 23 , wherein the message system comprises a text messaging system, an instant messaging system, an email system, and/or a voicemail system.

25 . The system of claim 24 , wherein the at least one processor is further configured to:

provide context to the knowledge system; and

use the knowledge system to interact with the message system to generate the message additionally in combination with the context.

26 . The system of claim 25 , wherein the context comprises one or more of: a user's prior message history with an intended recipient, or a writing style of a user.

27 . The system of claim 1 , wherein the knowledge system is configured to interact with one or more applications comprising a search system, a note-taking system, a payment system, a ride-share system, a programming IDE, a web browser, or a combination thereof.

28 . A computerized method comprising:

receiving a signal indicative of a user's speech muscle activation patterns when the user is speaking;

using a speech model and the signal as an input to the speech model to generate an output; and

using a knowledge system to take an action or generate a response based on the output.

29 . The method of claim 28 , wherein receiving the signal indicative of the user's speech muscle activation patterns when the user is speaking comprises:

receiving the signal from a speech input device including an electromyography (EMG) sensor;

wherein the signal is an EMG signal captured from the EMG sensor when the user is silently speaking.

30 . A non-transitory computer-readable media comprising instructions that, when executed by one or more processors on a computing device, cause the one or more processors to:

receive a signal indicative of a user's speech muscle activation patterns when the user is speaking;

use a speech model and the signal as an input to the speech model to generate an output; and

use a knowledge system to take an action or generate a response based on the output.

31 . The non-transitory computer-readable media of claim 30 , wherein receiving the signal indicative of the user's speech muscle activation patterns when the user is speaking comprises:

receiving the signal from a speech input device including an electromyography (EMG) sensor;

wherein the signal is an EMG signal captured from the EMG sensor when the user is silently speaking.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2023
From: GARG, SAHAJ; KOTHARI, TANAY; LEONARDO, ANTHONY
To: WISPR AI, INC.
Reel/Frame 064737/0893 →
Continuity (2)
Provisional Application 63437088 · Jan 4, 2023
Related Publication 20240221738A1 · Jul 4, 2024
References Cited (83)
US 8082149B2 · Schultz et al. · 2011 [cited by applicant]
US 9013264B2 · Parshionikar et al. · 2015 [cited by applicant]
US 10684686B1 · Milstein · 2020 [cited by examiner]
US 10878818B2 · Kapur et al. · 2020 [cited by applicant]
US 11709548B2 · Tadi et al. · 2023 [cited by applicant]
US 20020077534A1 · DuRousseau · 2002 [cited by examiner]
US 20120290950A1 · Rapaport et al. · 2012 [cited by applicant]
US 20150086052A1 · Park et al. · 2015 [cited by applicant]
US 20150161998A1 · Park · 2015 [cited by examiner]
US 20150288944A1 · Nistico · 2015 [cited by examiner]
US 20160314781A1 · Schultz et al. · 2016 [cited by applicant]
US 20170131768A1 · Budavari et al. · 2017 [cited by applicant]
US 20180239956A1 · Tadi et al. · 2018 [cited by applicant]
US 20190074012A1 · Kapur · 2019 [cited by examiner]
US 20190118834A1 · Wiebel-Herboth · 2019 [cited by examiner]
US 20200258535A1 · Vatanparav et al. · 2020 [cited by applicant]
US 20200289016A1 · Lukyanenko · 2020 [cited by examiner]
US 20210124422A1 · Forsland · 2021 [cited by applicant]
US 20210183383A1 · Volovich et al. · 2021 [cited by applicant]
US 20220137702A1 · Min et al. · 2022 [cited by applicant]
US 20220160296A1 · Rahmani et al. · 2022 [cited by applicant]
US 20220187912A1 · Alcaide et al. · 2022 [cited by applicant]
US 20220208194A1 · Rameau · 2022 [cited by examiner]
US 20220342477A1 · Ross · 2022 [cited by examiner]
US 20230077010A1 · Zhang et al. · 2023 [cited by applicant]
US 20230078978A1 · Tadi et al. · 2023 [cited by applicant]
US 20230130770A1 · Miller et al. · 2023 [cited by applicant]
US 20230157757A1 · Braido et al. · 2023 [cited by applicant]
US 20230157762A1 · Braido et al. · 2023 [cited by applicant]
US 20240220016A1 · Garg et al. · 2024 [cited by applicant]
US 20240220811A1 · Garg et al. · 2024 [cited by applicant]
US 20240221718A1 · Kothari et al. · 2024 [cited by applicant]
US 20240221719A1 · Kothari et al. · 2024 [cited by applicant]
US 20240221738A1 · Garg et al. · 2024 [cited by applicant]
US 20240221741A1 · Kothari et al. · 2024 [cited by applicant]
US 20240221751A1 · Garg et al. · 2024 [cited by applicant]
US 20240221753A1 · Garg et al. · 2024 [cited by applicant]
US 20240221762A1 · Garg et al. · 2024 [cited by applicant]
US 20250061885A1 · Garg et al. · 2025 [cited by applicant]
CA 2923979A1 · 2014 [cited by applicant]
CA 2942852A1 · 2014 [cited by applicant]
CA 298687A1 · 2018 [cited by applicant]
CA 2998687A1 · 2018 [cited by examiner]
CA 3164001A1 · 2021 [cited by applicant]
CA 2942852C · 2023 [cited by applicant]
CA 2957766C · 2023 [cited by applicant]
CN 101902960A · 2010 [cited by applicant]
CN 112424859A · 2021 [cited by applicant]
CN 114703091A · 2022 [cited by applicant]
KR 1020150104345A · 2015 [cited by applicant]
KR 1020200127150A · 2020 [cited by applicant]
WO WO2019040669A1 · 2019 [cited by applicant]
WO WO2022020968A1 · 2022 [cited by applicant]
WO WO2022245833A2 · 2022 [cited by applicant]
Gaddy et al., An Improved Model for Voicing Silent Speech. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and 11th International Joint Conference on Natural Language Processing. … [cited by applicant]
Gaddy et al., Digital Voicing of Silent Speech. arXiv:2010.02960v1 [eess.AS] Oct. 6, 2020. 10 pages. [cited by applicant]
He et al., Unvoiced Speech Recognition Algorithm Based on Myoeletric Signal. ICMLC. Feb. 15-17, 2020. 7 Pages. [cited by applicant]
Herff et al., Impact of Different Feedback Mechanisms in EMG-based Speech Recognition. Interspeech. Aug. 28-31, 2011:2213-2216. [cited by applicant]
Ishak., Speaker Identification Based on Vocal Cords Vibrations signal: Effect of the Window. International Journal of Digital Information and Wireless Communications. Jan. 2018. Doi:10.17781/P002406. 6 Pages. [cited by applicant]
Jamil et al., A flexible Speech Recognition System for Cerebral Palsy Disabled. ICIEIS. 2011. pp. 42-55. [cited by applicant]
Jou et al., Articulatory Feature Classification Using Surface Electromyography. International Central for Advanced Technologies. IEEE. 2006. 4 Pages. [cited by applicant]
Jou et al., Towards Continuous Speech Recognition Using Surface Electromyography. Interspeech. Sep. 17-21, 2006. 4 Pages. [cited by applicant]
Kapur et al., Alterego: A Personalized Wearable Silent Speech Interface. Session 1B: Multimodel Interfaces. IUI. Mar. 5, 2018. 11 Pages. [cited by applicant]
Kapur, How AI Could Become an Extension of your mind. TED. YouTube. Jun. 6, 2019. [cited by applicant]
Karnjanedecha et al., Signal Modeling for High-Performance Robust Isolated Word Recognition. IEEE Transactions of Speech and Audio Processing. Sep. 2001;9(6):647-654. [cited by applicant]
Maier-Hein et al., Session Independent Non-Audible Speech Recognition Using Surface Electromyography. IEEE Workshop on Automatic Speech Recognition and Understanding. 2005. 6 Pages. [cited by applicant]
Maier-Hein, Speech Recognition Using Surface Electromyography. Diplomarbeit Thesis. Aug. 2005. 129 Pages. [cited by applicant]
Manabe et al., Unvoiced Speech Recognition using EMG—Mime Speech Recognition—. Short Talk: Brains, Eyes and Ears. CHI. Apr. 5-10, 2003. 2 Pages. [cited by applicant]
Meltzner et al., Development of sEMG sensors and algorithms for silent speech recognition. J Neural Eng. Aug. 2018;15(4):046031. doi: 10.1088/1741-2552/aac965. Epub Jun. 1, 2018. PMID: 29855428; PMCID: PMC6168082. [cited by applicant]
Meltzner et al., Silent Speech Recognition as an Alternative Communication Device for Persons with Laryngectomy. IEEE/ACM Trans Audio Speech Lang Process. Dec. 2017;25(12):2386-2398. doi: 10.1109/TASLP.2017.2740000. Epu… [cited by applicant]
Polur et al., experiments with Fast Fourier Transform, Linear Predictive and Cepstral Coefficients in Dysarthric Speech Recognition Algorithms Using Hidden Markov Model. IEEE Transactions on Neural Systems and Rehabilit… [cited by applicant]
Prajapati et al., A Survey on Isolated Word and Digit Recognition Using Different Techniques. International Journal of Computer Applications (0975-8887). Mar. 2017; 161(3). 10 Pages. [cited by applicant]
Purcher, Apple Invents a next-generation AirPods Sensor System that could measure Biosignals and Electrical Activity of a user's Brain. Jul. 20, 2023. 9 pages. https://www.patentlyapple.com/2023/07/apple-invents-a-next-… [cited by applicant]
Rajasekaran et al., Recognition of Speech Under Stress and In Noise. ICASSP. 1986. 4 Pages. [cited by applicant]
Schultz et al., Modeling coarticulation in EMG-based continuous speech recognition. ScienceDirect. Speech Communications 52 (2010). Dec. 2, 2009. 13 Pages. [cited by applicant]
Toruk et al., Short Utterance Speaker Recognition Using Time-Delay Neural Network. 16th International Multi-Conference on Systems, Signals & Devices (SSD'19). 2019. 4 Pages. [cited by applicant]
Wand et al., The EMG-UKA for Electromyographic Speech Processing. Interspeech. Sep. 14-18, 2019. 5 Pages. [cited by applicant]
Wand et al., Towards Speaker-Adaptive Speech Recognition Based on Surface Electromyography. ICBSSP. 2009. 8 Pages. [cited by applicant]
Wand, Advancing Electromyographic Continuous Speech Recognition. Signal Preprocessing and Modeling. Scientific Publishing. Jan. 14, 2014. 256 Pages. [cited by applicant]
Wand, Speaker-Adaptive Speech Recognition Based on Surface Electromyography. IJCBEST. 2009. 15 Pages. [cited by applicant]
Wand, Towards Real-life Application of EMG-based Speech Recognition by using Unsupervised Adaptation. Interspeech. Sep. 14-18, 2014. 5 Pages. [cited by applicant]
Zhou et al., Improved Phoneme-Based Myoelectric Speech Recognition. IEEE Transactions on Biomedical Engineering. Aug. 2009;56(8). 8 Pages. [cited by applicant]
International Search Report and Written Opinion dated Apr. 26, 2024, in connection with International Application No. PCT/US24/10268. [cited by applicant]