IP Library Granted Patent US 10,074,381
Granted Patent B1
US 10,074,381 · App. 15/437,018 · Granted Sep 11, 2018

Augmented reality speech balloon system

Inventors: Piers Cowburn (London, GB); Qi Pan (London, GB); Eitan Pilipski (Los Angeles, CA)
Assignee: Snap Inc.
G10L21/10G06F17/2705G06F17/289G06K9/00281G06T11/60G10L15/25G10L15/265G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,074,381
App. No.
15/437,018
Granted
Sep 11, 2018
Kind
B1
Abstract

Disclosed is an augmented reality system to generate and cause display of an augmented reality interface at a client device. Various embodiments may detect speech, identify a source of the speech, transcribe the speech to a text string, generate a speech bubble based on properties of the speech and that includes a presentation of the text string, and cause display of the speech bubble at a location in the augmented reality interface based on the source of the speech.

Claims (62)

1. A system comprising:

a memory; and

at least one hardware processor coupled to the memory and comprising instructions that causes the system to perform operations comprising:

causing display at a client device, of a presentation of a space, the presentation of the space including at least a first person;

detecting, by the client device, speech that includes speech properties;

identifying the first person as a source of the speech;

transcribing the speech to a text string based on the speech properties;

determining an emotional effect of the speech based on the speech properties;

selecting a speech bubble from a speech bubble library based on the emotional effect of the speech; and

causing display of the speech bubble at a position adjacent to the first person within the presentation of the space in response to the identifying the first person as the source of the speech, the speech bubble containing a display of the text string.

2. The system of claim 1 wherein the determining the emotional effect of the speech includes: parsing the text string to a set of words; determining a definition of each word among the set of words; comparing the definition of each word to an emotional effect library; and selecting the emotional effect from the emotional effect library based on the comparison.

3. The system of claim 1 wherein the first person has a set of facial landmarks, and wherein the determining the emotional effect of the speech includes: capturing an image of the first person in response to identifying the first person as the source of the speech; extracting the set of facial landmarks from the image of the first person; and determining the emotional effect based on the set of facial landmarks.

4. The system of claim 1 wherein the speech properties include at least a volume, and wherein the determining the emotional effect of the speech includes: determining the volume of the speech based on the speech properties; and determining the emotional effect based on the volume.

5. The system of claim 1 , wherein the instructions cause the system to perform operations further comprising:

detecting a non-verbal sound;

comparing the non-verbal sound to an onomatopoeia library in response to the detecting the non-verbal sound;

identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and

causing display of a graphical element based on the onomatopoeia within the presentation of the space.

6. The system of claim 1 , wherein the instructions cause the system to perform operations further comprising:

detecting movement of a mouth of the first person; and

identifying the first person as the source of the speech in response to the detecting the movement of the mouth.

7. The system of claim 1 , wherein the instructions cause the system to perform operations further comprising:

identifying a first language of the speech based on the speech properties; and

wherein the transcribing the speech to the text string includes translating the speech from the first language to at least a second language.

8. The system of claim 7 , wherein the instructions cause the system to perform operations further comprising:

accessing a user profile of a first user associated with the client device in response to the detecting the speech, wherein the user profile includes a language preference, the language preference specifying the second language.

9. The system of claim 1 , wherein the text string has a length based on the speech properties, and wherein the generating the speech bubble includes:

determining a size of the speech bubble based on the length of the text string.

10. A method including:

causing display at a client device, of a presentation of a space, the presentation of the space including at least a first person;

detecting, by the client device, speech that includes speech properties;

identifying the first person as a source of the speech;

transcribing the speech to a text string based on the speech properties;

determining an emotional effect of the speech based on the speech properties;

selecting a speech bubble from a speech bubble library based on the emotional effect of the speech; and

causing display of the speech bubble at a position adjacent to the first person within the presentation of the space in response to the identifying the first person as the source of the speech, the speech bubble containing a display of the text string.

11. The method of claim 10 , wherein the determining the emotional effect of the speech includes:

parsing the text string to a set of words;

determining a definition of each word among the set of words;

comparing the definition of each word to an emotional effect library; and

selecting the emotional effect from the emotional effect library based on the comparison.

12. The method of claim 10 wherein the first person has a set of facial landmarks, and wherein the determining the emotional effect of the speech includes: capturing an image of the first person in response to identifying the first person as the source of the speech; extracting the set of facial landmarks from the image of the first person; and determining the emotional effect based on the set of facial landmarks.

13. The method of claim 10 wherein the speech properties include at least a volume, and wherein the determining the emotional effect of the speech includes: determining the volume of the speech based on the speech properties; and determining the emotional effect based on the volume.

14. The method of claim 10 , further comprising:

detecting a non-verbal sound;

comparing the non-verbal sound to an onomatopoeia library in response to the detecting the non-verbal sound;

identifying an onomatopoeia from the onomatopoeia library based on the non-verbal sound; and

causing display of a graphical element based on the onomatopoeia within the presentation of the space.

15. The method of claim 10 , further comprising:

detecting movement of a mouth of the first person; and

identifying the first person as the source of the speech in response to the detecting the movement of the mouth.

16. The method of claim 10 , further comprising:

identifying a first language of the speech based on the speech properties; and

wherein the transcribing the speech to the text string includes translating the speech from the first language to at least a second language.

17. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations including:

causing display at a client device, of a presentation of a space, the presentation of the space including at least a first person;

detecting, by the client device, speech that includes speech properties;

identifying the first person as a source of the speech;

transcribing the speech to a text string based on the speech properties;

determining an emotional effect of the speech based on the speech properties;

selecting a speech bubble from a speech bubble library based on the emotional effect of the speech; and

causing display of the speech bubble at a position adjacent to the first person within the presentation of the space in response to the identifying the first person as the source of the speech, the speech bubble containing a display of the text string.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2018
From: COWBURN, PIERS; PAN, QI; PILIPSKI, EITAN
To: SNAP INC.
Reel/Frame 046629/0866 →
Cited By (10)
US 12,197,884 US 12,293,759 US 12,315,495 US 12,340,475 US 12,361,934 US 12,394,127 US 12,475,893 US 12,563,159 US 12,567,418 US 12,700,192