Augmented reality speech-to-text captioning
An approach for improving communications for hearing challenged individuals. The approach captures audio communication and video communication associated with a speaker presentation. The approach converts the audio communication to text. The approach sends the text to an augmented reality (AR) device, associated with a user having hearing challenges, as captions. The approach stores the audio communication, the video communication, and the captions for replay by the user.
1 . A computer-implemented method for improving audio communications, the computer-implemented method comprising:
capturing audio communication and video communication associated with a speaker presentation to a user;
converting the audio communication to text, the converting performed responsive to a machine learning prediction that comprehension of the audio by the user will fall below a threshold level of comprehension;
sending the text to an augmented reality (AR) device as captions for the audio communication; and
displaying the captions along with the video communication on the AR device.
2 . The computer-implemented method of claim 1 , wherein the sending is based on receiving a request from the user to begin sending audio captions to the AR device, the request facilitating the machine learning prediction.
3 . The computer-implemented method of claim 2 , wherein the request is based on at least one of a predetermined eye movement or a predetermined hand swipe gesture.
4 . The computer-implemented method of claim 1 , wherein the converting comprises translating a language of the text generated from a spoken language associated with the speaker presentation to a proficient language associated with the user.
5 . The computer-implemented method of claim 1 , wherein the sending requires valid security credentials.
6 . The computer-implemented method of claim 1 , wherein the AR device comprises at least one of glasses, a mobile phone, a tablet computer or a smart watch.
7 . The computer-implemented method of claim 1 , wherein the converting is further performed based on factors comprising at least one of the audio communication, speaker sentiment, context of the audio communication, and environment of the audio communication.
8 . The computer-implemented method of claim 1 , further comprising:
storing the audio communication, the text, and the video communication for replay.
9 . The computer-implemented method of claim 8 , further comprising:
responsive to a request to replay the audio and video communications, rendering the stored text as audio captions superimposed on top of the video communication.
10 . A computer system for improving audio communications, the computer system comprising:
one or more computer processors;
one or more non-transitory computer readable storage media; and
program instructions stored on the one or more non-transitory computer readable storage media, the program instructions comprising:
program instructions to capture audio communication and video communication associated with a speaker presentation to a user;
program instructions to convert, responsive to a machine learning prediction that comprehension of the audio by the user will fall below a threshold level of comprehension, the audio communication to text;
program instructions to send the text to an augmented reality (AR) device as captions for the audio communication; and
program instructions to display the captions along with the video communication on the AR device.
11 . The computer system of claim 10 , wherein the sending is based on receiving a request from the user to begin sending audio captions to the AR device, the request facilitating the machine learning prediction.
12 . The computer system of claim 11 , wherein the request is based on at least one of a predetermined eye movement or a predetermined hand swipe gesture.
13 . The computer system of claim 10 , wherein the converting comprises translating a language of the text generated from a spoken language associated with the speaker presentation to a written language associated with the user according to the user comprehension of the spoken language.
14 . The computer system of claim 10 , wherein the sending requires valid security credentials.
15 . The computer system of claim 10 , wherein the AR device comprises at least one of glasses, a mobile phone, a tablet computer or a smart watch.
16 . The computer system of claim 10 , wherein the converting is further performed based on factors comprising at least one of the audio communication, speaker sentiment, context of the audio communication, and environment of the audio communication.
17 . A computer program product for improving audio communications, the computer program product comprising:
one or more computer readable storage media; and
program instructions stored on the one or more computer readable storage media, the program instructions comprising:
program instructions to capture audio communication and video communication associated with a speaker presentation to a user;
program instructions to convert, responsive to a machine learning prediction that comprehension of the audio by the user will fall below a threshold level of comprehension, the audio communication to text;
program instructions to send the text to an augmented reality (AR) device as captions for the audio communication; and
program instructions to display the captions along with the video communication on the AR device.