Transcription based on speech and visual input
A method can include receiving audio input of speech, receiving visual input while receiving the audio input, generating a semantic description based on the visual input, and presenting a transcription of the speech based on the audio input and the semantic description.
1 . A method comprising:
receiving audio input of speech;
receiving visual input while receiving the audio input;
generating a semantic description of a scene based on the visual input, the semantic description of the scene including a list of objects classified based on the visual input, each object in the list of objects being associated with a confidence level indicating a likelihood that the object corresponds to a category included in the list of objects;
generating a transcription of the speech based on the audio input and the semantic description of the scene; and
presenting the transcription of the speech.
2 . The method of claim 1 , wherein the generating the transcription of the speech includes:
generating the transcription of the speech based on the audio input;
determining an error in the transcription based on a confidence level of a word that was transcribed from the speech not satisfying a confidence threshold; and
correcting the error in the transcription by modifying the word based on the semantic description of the scene,
wherein the presenting the transcription includes presenting the transcription with the modified word.
3 . The method of claim 1 , wherein the generating the transcription of the speech includes:
transcribing the speech based on the audio input; and
modifying at least one of a noun, verb, adjective, or adverb that was transcribed from the speech,
wherein the presenting the transcription includes presenting the transcribed speech with the modified noun, verb, adjective, or adverb.
4 . The method of claim 1 , wherein the generating the transcription of the speech includes:
transcribing the speech based on the audio input;
determining that a confidence level of a word that was transcribed from the speech does not satisfy a confidence threshold, the word including one of a noun, verb, adjective, or adverb; and
modifying the word based on the semantic description of the scene,
wherein the presenting the transcription includes presenting the transcribed speech with the modified word.
5 . The method of claim 1 , wherein the generating the transcription of the speech includes:
transcribing the speech based on the audio input;
determining that a word that was transcribed from the speech has a homonym; and
determining that a correspondence between the homonym and the semantic description of the scene is greater than a correspondence between the word that was transcribed from the speech and the semantic description of the scene,
wherein the presenting the transcription includes presenting the transcribed speech with the homonym.
6 . The method of claim 1 , wherein the method is performed by an electronic device worn on a head of a user.
7 . The method of claim 1 , wherein:
the method further comprises receiving accelerometer data; and
the generating the semantic description of the scene includes generating the semantic description of the scene based on the visual input and the accelerometer data.
8 . The method of claim 1 , wherein:
the method further comprises receiving location data; and
the generating the semantic description of the scene includes generating the semantic description of the scene based on the visual input and the location data.
9 . The method of claim 1 , wherein:
the method further comprises receiving input from an audio source other than speech; and
the generating the semantic description of the scene includes generating the semantic description based on the visual input and the audio source other than speech.
10 . The method of claim 1 , wherein the semantic description includes relationships between objects recognized based on the visual input.
11 . A head-worn wearable device, comprising:
a microphone;
a camera;
a display;
at least one processor; and
a non-transitory computer-readable storage medium comprising instructions stored thereon that, when executed by the at least one processor, are configured to cause the head-worn wearable device to:
receive audio input of speech via the microphone;
receive visual input while receiving the audio input via the camera;
generate a semantic description of a scene based on the visual input, the semantic description of the scene including a list of objects classified based on the visual input, each object in the list of objects being associated with a confidence level indicating a likelihood that the object corresponds to a category included in the list of objects;
generate a transcription of the speech based on the audio input and the semantic description of the scene; and
present, via the display, the transcription of the speech.
12 . The head-worn wearable device of claim 11 , wherein the instructions are further configured to cause the head-worn device to modify a word in the transcription of the speech, the modified word including at least one of a noun, a verb, an adjective, or an adverb.
13 . The head-worn wearable device of claim 11 , wherein:
the instructions are further configured to cause the head-worn device to modify a word in the transcription of the speech;
the modified word has a homonym; and
presenting the transcription of the speech including the modified word includes presenting the transcription of the speech with the homonym as the modified word.
14 . A non-transitory computer-readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause an electronic device to:
receive audio input of speech;
receive visual input while receiving the audio input;
generate a semantic description of a scene based on the visual input, the semantic description of the scene including a list of objects classified based on the visual input, each object being associated with a confidence level indicating a likelihood that the object corresponds to a certain category included in the list of objects;
generate a transcription of the speech based on the audio input and the semantic description of the scene; and
present the transcription of the speech.
15 . The non-transitory computer-readable storage medium of claim 14 , wherein the generating the transcription of the speech includes:
generating the transcription of the speech based on the audio input;
determining an error in the transcription based on a confidence level of a word that was transcribed from the speech not satisfying a confidence threshold; and
correcting the error in the transcription by modifying the word based on the semantic description of the scene,
wherein the presenting the transcription includes presenting the transcription with the modified word.
16 . The non-transitory computer-readable storage medium of claim 14 , wherein the generating the transcription of the speech includes:
transcribing the speech based on the audio input; and
modifying at least one of a noun, verb, adjective, or adverb that was transcribed from the speech,
wherein the presenting the transcription includes presenting the transcribed speech with the modified noun, verb, adjective, or adverb.
17 . The non-transitory computer-readable storage medium of claim 14 , wherein the generating the transcription of the speech includes:
transcribing the speech based on the audio input;
determining that a word that was transcribed from the speech has a homonym; and
determining that a correspondence between the homonym and the semantic description of the scene is greater than a correspondence between the word that was transcribed from the speech and the semantic description of the scene,
wherein the presenting the transcription includes presenting the transcribed speech with the homonym.
18 . The non-transitory computer-readable storage medium of claim 14 , wherein:
the instructions are further configured to cause the electronic device to receive input from an audio source other than speech; and
the generating the semantic description of the scene includes generating the semantic description of the scene based on the visual input and the audio source other than speech.