SYNTHESIZING SPEECH FROM FACIAL SKIN MOVEMENTS
Systems and methods are disclosed for synthesizing speech from minute facial skin movements. In one implementation, a system may include a processor configured to control at least one coherent light source to illuminate a region of a face. The processor may receive from at least one sensor, reflection signals indicative of coherent light reflected from the face. The reflection signals may be analyzed to determine the minute facial skin movements associated with silent speech. Then, based on the determined minute facial skin movements, the processor may determine a sequence of words associated with the silent speech, and synthesize the sequence of words associated with the silent speech into audio signals.
1 - 37 . (canceled)
38 . A system for synthesizing speech from minute facial skin movements, the system comprising:
at least one processor configured to:
control at least one coherent light source to illuminate a region of a face;
receive from at least one sensor, reflection signals indicative of coherent light reflected from the face;
analyze the reflection signals to determine the minute facial skin movements associated with silent speech;
based on the determined minute facial skin movements, determine a sequence of words associated with the silent speech; and
synthesize the sequence of words associated with the silent speech into audio signals.
39 . The system of claim 38 , wherein the at least one processor is further configured to determine the sequence of words by referencing training-derived data stored in memory.
40 . The system of claim 38 , further comprising an earphone and wherein the at least one processor is further configured output the synthesized audio signals for presentation via the earphone.
41 . The system of claim 38 , further comprising a communication interface configured to transmit data associated with the silent speech to a mobile communications device via a communication link.
42 . The system of claim 41 , wherein the at least one processor is further configured to transmit to the mobile communications device a synthetization of the silent speech during a phone call.
43 . The system of claim 38 , wherein synthesizing the sequence of words includes generating a translated sequence of words in a synthesized voice and in a language other than a language of the silent speech.
44 . The system of claim 43 , wherein the synthesized voice in the language other than a language of the silent speech is outputted to occur at a time when the silent speech is anticipated.
45 . The system of claim 38 , wherein synthesizing the sequence of words associated with the silent speech into audio signals includes cleaning background audio signals from the audio signals associated with the silent speech.
46 . The system of claim 38 , wherein determining the sequence of words associated with the silent speech includes extracting multiple candidate phonemes from the minute facial skin movements based on respective probabilities that the minute facial skin movements are associated with the multiple candidate phonemes.
47 . The system of claim 46 , wherein the at least one processor is configured to mix the multiple candidate phonemes based on their respective probabilities to generate an ambiguous audio output.
48 . The system of claim 47 , wherein the generated ambiguous audio output is associated with a same length of time as a corresponding set of the minute facial skin movements.
49 . The system of claim 38 , wherein the at least one processor is configured to determine the silent speech in an absence of vocalization of the sequence of words.
50 . The system of claim 38 , wherein the at least one coherent light source, the at least one sensor, and the at least one processor are part of a wireless headphone.
51 . The system of claim 38 , wherein the at least one coherent light source is configured to direct a plurality of beams of coherent light toward different locations on the region of a face thus creating an array of spots.
52 . The system of claim 38 , further comprising a housing configured to be worn on a head of a user, wherein the at least one sensor is connected to the housing in a manner such that when the housing is worn, the at least one sensor is held at a distance from a skin surface.
53 . The system of claim 52 , wherein the reflection signals are indicative of intended speech by the user.
54 . The system of claim 38 , wherein determining the sequence of words associated with the silent speech involves determining conversation context.
55 . The system of claim 38 , wherein determining the sequence of words includes using an artificial neural network associated with personal training data received from a system user.
56 . The system of claim 55 , wherein the at least one processor is configured to generate one or more network coefficients for the user based on the personal training data.
57 . A method for synthesizing speech from minute facial skin movements, the method comprising:
at least one processor configured to:
control at least one coherent light source to illuminate a region of a face;
receive from at least one sensor, reflection signals indicative of coherent light reflected from the face;
analyze the reflection signals to determine the minute facial skin movements associated with silent speech;
based on the determined minute facial skin movements, determine a sequence of words associated with the silent speech; and
synthesize the sequence of words associated with the silent speech into audio signals.