Augmented reality speech balloon system
Disclosed is an augmented reality system to generate and cause display of an augmented reality interface at a client device. Various embodiments may detect speech, identify a source of the speech, transcribe the speech to a text string, generate a speech bubble based on properties of the speech and that includes a presentation of the text string, and cause display of the speech bubble at a location in the augmented reality interface based on the source of the speech.
1. A system comprising:
a memory; and
at least one hardware processor coupled to the memory and comprising instructions that causes the system to perform operations comprising:
causing display of a presentation of image data at a client device, the presentation of the image data comprising a depiction of a set of facial features;
detecting, by the client device, a speech signal that comprises auditory properties;
transcribing the speech signal to a text string based on the auditory properties;
determining an emotional effect of the speech signal based on the set of facial features;
selecting a graphical element based on the emotional effect; and
causing display of the text string within the graphical element within the presentation of the image data.
2. The system of claim 1 , wherein the auditory properties include a volume of the speech signal, and the determining the emotional effect is based on the set of facial features and the volume of the speech signal.
3. The system of claim 1 , wherein the graphical element includes a speech bubble that comprises a set of graphical properties, the graphical properties based on the auditory properties of the auditory signal.
4. The system of claim 1 , wherein the speech signal includes a non-verbal sound, and the operations further comprise:
comparing the non-verbal sound to an onomatopoeia library in response to the detecting the speech signal;
identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and
selecting the graphical element based on at least the auditory properties of auditory signal and the onomatopoeia identified based on the non-verbal sound.
5. The system of claim 1 , wherein the speech signal corresponds with a source within the presentation of the image data, the source comprises a graphical property, and the operations further comprise:
identifying the source of the speech signal within the presentation of the image data based on the auditory properties.
6. The system of claim 5 , wherein the selecting the graphical element is based on the emotional effect and the graphical property of the source of the speech signal.
7. The system of claim 5 , wherein the identifying the source of the speech signal within the presentation of the image data includes:
detecting movement within the presentation of the image data; and
identifying the source of the speech signal based on the movement.
8. A method comprising:
causing display of a presentation of image data at a client device, the presentation of the image data comprising a depiction of a set of facial features;
detecting, by the client device, a speech signal that comprises auditory properties;
transcribing the speech signal to a text string based on the auditory properties;
determining an emotional effect of the speech signal based on the set of facial features;
selecting a graphical element based on the emotional effect; and
causing display of the text string within the graphical element within the presentation of the image data.
9. The method of claim 8 , wherein the auditory properties include a volume of the speech signal, and the determining the emotional effect is based on the set of facial features and the volume of the speech signal.
10. The method of claim 8 , wherein the graphical element includes a speech bubble that comprises a set of graphical properties, the graphical properties based on the auditory properties of the auditory signal.
11. The method of claim 8 , wherein the speech signal includes a non-verbal sound, and the method further comprises:
comparing the non-verbal sound to an onomatopoeia library in response to the detecting the speech signal;
identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and
selecting the graphical element based on at least the auditory properties of auditory signal and the onomatopoeia identified based on the non-verbal sound.
12. The method of claim 8 , wherein the speech signal corresponds with a source within the presentation of the image data, the source comprises a graphical property, and the method further comprises:
identifying the source of the speech signal within the presentation of the image data based on the auditory properties.
13. The method of claim 12 , wherein the selecting the graphical element is based on the emotional effect and the graphical property of the source of the speech signal.
14. The method of claim 12 , wherein the identifying the source of the speech signal within the presentation of the image data includes:
detecting movement within the presentation of the image data; and
identifying the source of the speech signal based on the movement.
15. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations including:
causing display of a presentation of image data at a client device, the presentation of the image data comprising a depiction of a set of facial features;
detecting, by the client device, a speech signal that comprises auditory properties;
transcribing the speech signal to a text string based on the auditory properties;
determining an emotional effect of the speech signal based on the set of facial features;
selecting a graphical element based on the emotional effect; and
causing display of the text string within the graphical element within the presentation of the image data.
16. The non-transitory machine-readable storage medium of claim 15 , wherein the auditory properties include a volume of the speech signal, and the determining the emotional effect is based on the set of facial features and the volume of the speech signal.
17. The non-transitory machine-readable storage medium of claim 15 , wherein the graphical element includes a speech bubble that comprises a set of graphical properties, the graphical properties based on the auditory properties of the auditory signal.
18. The non-transitory machine-readable storage medium of claim 15 , wherein the speech signal includes a non-verbal sound, and the operations further comprise:
comparing the non-verbal sound to an onomatopoeia library in response to the detecting the speech signal;
identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and
selecting the graphical element based on at least the auditory properties of auditory signal and the onomatopoeia identified based on the non-verbal sound.
19. The non-transitory machine-readable storage medium of claim 15 , wherein the speech signal corresponds with a source within the presentation of the image data, the source comprises a graphical property, and the operations further comprise:
identifying the source of the speech signal within the presentation of the image data based on the auditory properties.
20. The non-transitory machine-readable storage medium of claim 19 , wherein the selecting the graphical element is based on the emotional effect and the graphical property of the source of the speech signal.