Method and apparatus for recognizing silent speech
View Patent ↗An electronic apparatus includes: a communication device configured to receive a signal from each of a plurality of acceleration sensors attached to a face of a user; a memory configured to store a classification learning model that classifies words based on a plurality of sensor output values; and a processor configured to determine a word corresponding to a mouth shape of the user by input a value of the received signal to the classification learning model, when the signal is received from each of the plurality of acceleration sensors.
1. An electronic apparatus comprising:
a communication device configured to receive a signal from each of a plurality of acceleration sensors attached to a face of a user, wherein the plurality of acceleration sensors are attached to different portions around a mouth of the user, the different portions being those that move the most at the time of speech utterance of the user;
a memory configured to store a classification learning model that classifies a word based on a plurality of sensor output values; and
a processor configured to determine a word corresponding to a mouth shape of the user by inputting signal values of the received signal into the classification learning model, when the signal values are received from each of the plurality of acceleration sensors.
2. The electronic apparatus as claimed in claim 1 , wherein the classification learning model is a model trained by using the value of the signal received from each of the plurality of acceleration sensors in a process of uttering each of a plurality of predetermined words.
3. The electronic apparatus as claimed in claim 1 , wherein the classification learning model is a convolutional neural network-long short-term memory (1D CNN-LSTM) model.
4. The electronic apparatus as claimed in claim 1 , wherein the plurality of acceleration sensors include three to five acceleration sensors.
5. The electronic apparatus as claimed in claim 1 , wherein each of the plurality of acceleration sensors is a 3-axis accelerometer.
6. The electronic apparatus as claimed in claim 1 , wherein the processor is configured to perform an operation corresponding to the determined word.
7. A method for recognizing silent speech, the method comprising:
receiving a signal from each of a plurality of acceleration sensors attached to a face of a user, wherein the plurality of acceleration sensors are attached to different portions around a mouth of the user, the different portions being those that move the most at the time of speech utterance of the user; and
determining a word corresponding to a mouth shape of the user by inputting signal values of the received signal into a classification learning model that classifies a word based on a plurality of sensor output values.
8. The method as claimed in claim 7 , further comprising training the classification learning model by using the value of the signal received from each of the plurality of acceleration sensors in a process of uttering each of a plurality of predetermined words.
9. The method as claimed in claim 7 , wherein the classification learning model is a convolutional neural network-long short-term memory (1D CNN-LSTM) model.
10. The method as claimed in claim 7 , wherein in the receiving, the signal is received from each of the plurality of acceleration sensors attached to different portions around a mouth that move the most at the time of speech utterance of the user.
11. The method as claimed in claim 10 , wherein the plurality of acceleration sensors include three to five acceleration sensors.
12. The method as claimed in claim 7 , further comprising performing an operation corresponding to the determined word.
13. A non-transitory computer-readable recording medium including a program for performing a method for recognizing silent speech, the method including:
receiving a signal from each of a plurality of acceleration sensors attached to a face of a user, wherein the plurality of acceleration sensors are attached to different portions around a mouth of the user, the different portions being those that move the most at the time of speech utterance of the user; and
determining a word corresponding to a mouth shape of the user by inputting signal values of the received signal into a classification learning model that classifies a word based on a plurality of sensor output values.