IP Library › Granted Patent US 11,748,579
Granted Patent B2
US 11,748,579 · App. 17/474,392 · Granted Sep 5, 2023

Augmented reality speech balloon system

Inventors: Piers Cowburn (London, GB); Qi Pan (London, GB); Eitan Pilipski (Los Angeles, CA)
Assignee: Snap Inc.
G06F40/58G06F40/205G06F40/30G06T11/00G06T11/60G06V20/20G06V40/171G10L15/25G10L15/26G10L21/10G10L25/63G06V40/175
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,748,579
App. No.
17/474,392
Granted
Sep 5, 2023
Kind
B2
Abstract

Disclosed is an augmented reality system to generate and cause display of an augmented reality interface at a client device. Various embodiments may detect speech, identify a source of the speech, transcribe the speech to a text string, generate a speech bubble based on properties of the speech and that includes a presentation of the text string, and cause display of the speech bubble at a location in the augmented reality interface based on the source of the speech.

Claims (56)

1. A system comprising:

a memory; and

at least one hardware processor coupled to the memory and comprising instructions that causes the system to perform operations comprising:

causing display of a presentation of image data at a client device, the presentation of the image data comprising a depiction of a set of facial features;

detecting, by the client device, a speech signal that comprises auditory properties;

transcribing the speech signal to a text string based on the auditory properties;

determining an emotional effect of the speech signal based on the set of facial features;

selecting a graphical element based on the emotional effect; and

causing display of the text string within the graphical element within the presentation of the image data.

2. The system of claim 1 , wherein the auditory properties include a volume of the speech signal, and the determining the emotional effect is based on the set of facial features and the volume of the speech signal.

3. The system of claim 1 , wherein the graphical element includes a speech bubble that comprises a set of graphical properties, the graphical properties based on the auditory properties of the auditory signal.

4. The system of claim 1 , wherein the speech signal includes a non-verbal sound, and the operations further comprise:

comparing the non-verbal sound to an onomatopoeia library in response to the detecting the speech signal;

identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and

selecting the graphical element based on at least the auditory properties of auditory signal and the onomatopoeia identified based on the non-verbal sound.

5. The system of claim 1 , wherein the speech signal corresponds with a source within the presentation of the image data, the source comprises a graphical property, and the operations further comprise:

identifying the source of the speech signal within the presentation of the image data based on the auditory properties.

6. The system of claim 5 , wherein the selecting the graphical element is based on the emotional effect and the graphical property of the source of the speech signal.

7. The system of claim 5 , wherein the identifying the source of the speech signal within the presentation of the image data includes:

detecting movement within the presentation of the image data; and

identifying the source of the speech signal based on the movement.

8. A method comprising:

causing display of a presentation of image data at a client device, the presentation of the image data comprising a depiction of a set of facial features;

detecting, by the client device, a speech signal that comprises auditory properties;

transcribing the speech signal to a text string based on the auditory properties;

determining an emotional effect of the speech signal based on the set of facial features;

selecting a graphical element based on the emotional effect; and

causing display of the text string within the graphical element within the presentation of the image data.

9. The method of claim 8 , wherein the auditory properties include a volume of the speech signal, and the determining the emotional effect is based on the set of facial features and the volume of the speech signal.

10. The method of claim 8 , wherein the graphical element includes a speech bubble that comprises a set of graphical properties, the graphical properties based on the auditory properties of the auditory signal.

11. The method of claim 8 , wherein the speech signal includes a non-verbal sound, and the method further comprises:

comparing the non-verbal sound to an onomatopoeia library in response to the detecting the speech signal;

identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and

selecting the graphical element based on at least the auditory properties of auditory signal and the onomatopoeia identified based on the non-verbal sound.

12. The method of claim 8 , wherein the speech signal corresponds with a source within the presentation of the image data, the source comprises a graphical property, and the method further comprises:

identifying the source of the speech signal within the presentation of the image data based on the auditory properties.

13. The method of claim 12 , wherein the selecting the graphical element is based on the emotional effect and the graphical property of the source of the speech signal.

14. The method of claim 12 , wherein the identifying the source of the speech signal within the presentation of the image data includes:

detecting movement within the presentation of the image data; and

identifying the source of the speech signal based on the movement.

15. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations including:

causing display of a presentation of image data at a client device, the presentation of the image data comprising a depiction of a set of facial features;

detecting, by the client device, a speech signal that comprises auditory properties;

transcribing the speech signal to a text string based on the auditory properties;

determining an emotional effect of the speech signal based on the set of facial features;

selecting a graphical element based on the emotional effect; and

causing display of the text string within the graphical element within the presentation of the image data.

16. The non-transitory machine-readable storage medium of claim 15 , wherein the auditory properties include a volume of the speech signal, and the determining the emotional effect is based on the set of facial features and the volume of the speech signal.

17. The non-transitory machine-readable storage medium of claim 15 , wherein the graphical element includes a speech bubble that comprises a set of graphical properties, the graphical properties based on the auditory properties of the auditory signal.

18. The non-transitory machine-readable storage medium of claim 15 , wherein the speech signal includes a non-verbal sound, and the operations further comprise:

comparing the non-verbal sound to an onomatopoeia library in response to the detecting the speech signal;

identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and

selecting the graphical element based on at least the auditory properties of auditory signal and the onomatopoeia identified based on the non-verbal sound.

19. The non-transitory machine-readable storage medium of claim 15 , wherein the speech signal corresponds with a source within the presentation of the image data, the source comprises a graphical property, and the operations further comprise:

identifying the source of the speech signal within the presentation of the image data based on the auditory properties.

20. The non-transitory machine-readable storage medium of claim 19 , wherein the selecting the graphical element is based on the emotional effect and the graphical property of the source of the speech signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2021
From: COWBURN, PIERS; PAN, QI; PILIPSKI, EITAN
To: SNAP INC.
Reel/Frame 057801/0731 →
Continuity (4)
Continuation 16749678 · Jan 22, 2020
Continuation 16014193 · Jun 21, 2018
Continuation 15437018 · Feb 20, 2017
Related Publication 20210407533A1 · Dec 30, 2021
Cited By (2)
US 12,197,884 US 12,394,127