SYSTEM AND METHOD FOR USING GESTURES AND EXPRESSIONS FOR CONTROLLING SPEECH APPLICATIONS
Methods and systems are provided for detecting and processing gestures, expressions (e.g., facial), tone and/or gestures of the user for the purpose of improving the quality and speed of interactions with computer-based systems. Such information may be detected by one or more sensors such as, for example, electromyography (EMG) sensors used to monitor and record electrical activity produced by muscles that are activated. Other sensor types may be used, such as optical, inertial measurement unit (IMU), or other types of bio-sensors. The system may use one or more sensors to detect speech alone or in combination with gestures, expressions (e.g., facial), tone and/or gestures of the user to provide input or control of the system.
1 . A method for training a model, the method comprising:
receiving an output of a model;
receiving an input signal from a speech input device wearable on a user, wherein the input signal is captured when the user is making a facial expression or gesture or speaking in response to the output;
determining a feedback signal based on the input signal; and
using the feedback signal at least in part to retrain the model.
2 . The method according to claim 1 , wherein the input signal is at least one of a group comprising, an EMG signal, a microphone input signal, an inertial measurement unit, a camera, and a biosensor.
3 . The method according to claim 1 , wherein the feedback signal indicates a frown and/or a head gesture.
4 . The method according to claim 1 , wherein the model is a speech recognition model.
5 . The method according to claim 1 , wherein the model is associated with a digital assistant.
6 . The method according to claim 1 , further comprising determining a dataset comprising a plurality of feedback signals including the feedback signal and retraining the model based on the dataset.
7 . The method according to claim 1 , further comprising converting the feedback signal to a scalar value.
8 . The method according to claim 7 , further comprising using the scalar value representing the feedback signal to retrain the model.
9 . The method according to claim 8 , further comprising training a reward model to predict the scalar value representing the feedback from the input and output of the knowledge system.
10 . The method according to claim 1 , wherein the method used to at least in part retrain the model is based on reinforcement learning.
11 . The method according to claim 1 , further comprising determining content of words spoken by the user based on the feedback signal.
12 . The method according to claim 1 , wherein the output of the model is provided at least in part by a knowledge system configured to interact with the user.
13 . The method according to claim 10 , further comprising:
receiving an input speech signal;
converting the input speech signal to a text output;
providing the text output to the knowledge system as a prompt;
receiving, from the knowledge system, an output to the user, the output being generated by the knowledge system responsive to the provided prompt; and
collecting a feedback signal from the user responsive to the output generated by the knowledge system.
14 . The method according to claim 13 , wherein the knowledge system comprises a machine learning foundation model.
15 . The method according to claim 14 , wherein the machine learning foundation model is retrained at least in part to be personalized to the user based on the user feedback.
16 . The method according to claim 15 , wherein the machine learning foundation model is updated based on aggregated feedback signals collected across a plurality of users.
17 . A non-transitory computer-readable medium containing instruction that, when executed, cause at least one computer hardware processor to perform a method comprising acts of:
receiving an output of a model;
receiving an input signal from a speech input device wearable on a user, wherein the input signal is captured when the user is making a facial expression or gesture or speaking in response to the output;
determining a feedback signal based on the input signal; and
using the feedback signal at least in part to retrain the model.
18 . The computer-readable medium according to claim 17 , wherein the input signal is at least one of a group comprising, an EMG signal, a microphone input signal, an inertial measurement unit, a camera, and a biosensor.
19 . The computer-readable medium according to claim 17 , wherein the feedback signal indicates a frown and/or a head gesture.
20 . The computer-readable medium according to claim 17 , wherein the model is a speech recognition model.