IP Library › Granted Patent US 11,227,442
Granted Patent B1
US 11,227,442 · App. 16/721,459 · Granted Jan 18, 2022

3D captions with semantic graphical elements

Inventors: Kyle Goodrich (Venice, CA); Samuel Edward Hare (Los Angeles, CA); Maxim Maximov Lazarov (Culver City, CA); Tony Mathew (Los Angeles, CA); Andrew James McPhee (Culver City, CA); Daniel Moreno (Los Angeles, CA); Wentao Shang (Los Angeles, CA)
Assignee: Snap Inc.
G06T19/006G06F3/012G06F3/04883G06T15/80G06T19/20G06T2219/2004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,227,442
App. No.
16/721,459
Granted
Jan 18, 2022
Kind
B1
Abstract

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing at least one program and method for performing operations comprising: receiving, by a messaging application, a video feed from a camera of a user device that depicts a face; receiving a request to add a 3D caption to the video feed; identifying a graphical element that is associated with context of the 3D caption; and displaying the 3D caption and the identified graphical element in the video feed at a position in 3D space of the video feed proximate to the face depicted in the video feed.

Claims (71)

1. A system comprising:

at least one hardware processor;

a memory storing instructions which, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:

receiving, by a messaging application, a video feed from a camera of a user device that depicts a face;

receiving a request to add a 3D caption to the video feed;

identifying a graphical element that is associated with context of the 3D caption;

displaying the 3D caption and the identified graphical element in the video feed at a position in 3D space of the video feed proximate to the face depicted in the video feed;

detecting input indicating that a user tapped on screen the 3D caption that is displayed in the video feed; and

in response to detecting the input indicating that the user tapped the 3D caption, presenting text of the 3D caption in 2D to enable the user to modify text of the 3D caption.

2. The system of claim 1 , wherein the graphical element is an emoji or avatar, wherein the graphical element comprises a first emoji or avatar and a second emoji or avatar, wherein the first and second emojis or avatars are identical, wherein the first identical emoji or avatar is placed on a left side of the 3D caption and the second identical emoji or avatar is placed on a right side of the 3D caption.

3. The system of claim 1 , wherein the operations further comprise:

determining that the camera from which the video feed is received is a front-facing camera of the user device; and

in response to determining that the camera from which the video feed is received is the front-facing camera of the user device, automatically populating the 3D caption with the identified graphical element.

4. The system of claim 1 , wherein the operations further comprise:

determining that the camera from which the video feed is received is a rear-facing camera of the user device; and

in response to determining that the camera from which the video feed is received is the rear-facing camera of the user device, populating the 3D caption with the identified graphical element in response to a user request.

5. The system of claim 1 , wherein the operations further comprise:

after the face is detected in the video feed, determining that the face is no longer detected in the video feed; and

in response to determining that the face is no longer detected in the video feed, positioning the 3D caption on a surface of an object depicted in the video feed without displaying the graphical element.

6. The system of claim 1 , wherein the operations further comprise orienting the 3D caption and the graphical element above the face.

7. The system of claim 1 , wherein the operations further comprise curving the 3D caption and graphical element around a top of the face or around a bottom of the face.

8. The system of claim 1 , wherein the operations further comprise:

receiving input that selects the identified graphical element that is displayed;

in response to receiving the input:

presenting text of the 3D caption in two-dimensions (2D) in place of the 3D caption while maintaining display of the identified graphical element in the video feed at the position in the 3D space; and

presenting a list of alternate graphical elements while maintaining display of the identified graphical element in the video feed at the position in the 3D space;

receiving a selection of a given alternate graphical element from the list of alternate graphical elements; and

replacing the identified graphical element that is displayed in the video feed with the given alternate graphical element, wherein the 3D caption and the given alternate graphical element are displayed at the position in the 3D space in response to the selection of the given alternate graphical element.

9. The system of claim 1 , wherein the operations further comprise searching a database of graphical elements based on a word or phrase from the 3D caption to identify the graphical element that is associated with the word or phrase from the 3D caption.

10. The system of claim 1 , wherein the operations further comprise:

determining, by the messaging application, context associated with the video feed; and

automatically populating text of the 3D caption based on the context.

11. The system of claim 10 , wherein the context is determined based on a current date and time or day of the week.

12. The system of claim 1 , wherein the operations further comprise:

prior to detecting the face in the video feed, displaying the 3D caption on a surface of an object depicted in the video feed; and

in response to detecting the face:

moving the 3D caption from being displayed on the surface of the object to being displayed proximate to the face; and

adding the graphical element to the video feed in close proximity to the 3D caption.

13. The system of claim 1 , wherein the operations further comprise dimming the screen in which the text is presented to focus the user on the text.

14. The system of claim 1 , wherein the operations further comprise:

determining that a user tapped on the screen at a location between two characters of the text; and

positioning a cursor to modify the text starting at the location between the two characters of the text in response to determining that the user tapped on the screen at the location between the two characters of the text.

15. The system of claim 1 , wherein the 3D caption includes words or phrases, and wherein the operations further comprise selecting a default graphical element to be displayed as the identified graphical element when none of the words or phrases in the 3D caption matches words or phrases stored in a database that associates words or phrases with graphical elements.

16. The system of claim 1 , wherein the operations further comprise:

in response to receiving the request to add the 3D caption, presenting a two-dimensional (2D) text entry interface;

receiving a 2D text string from the text entry interface;

identifying the graphical element based on the 2D text string;

converting the 2D text string to the 3D caption; and

displaying the 3D caption and the identified graphical element in response to receiving indication of completion of entry of the text string.

17. A method comprising:

receiving, by one or more processors that implement a messaging application, a video feed from a camera of a user device that depicts a face;

receiving a request to add a 3D caption to the video feed;

identifying a graphical element that is associated with context of the 3D caption;

displaying the 3D caption and the identified graphical element in the video feed at a position in 3D space of the video feed proximate to the face depicted in the video feed;

detecting input indicating that a user tapped on screen the 3D caption that is displayed in the video feed; and

in response to detecting the input indicating that the user tapped the 3D caption, presenting text of the 3D caption in 2D to enable the user to modify text of the 3D caption.

18. The method of claim 17 , further comprising:

prior to detecting the face in the video feed, displaying the 3D caption on a surface of an object depicted in the video feed; and

in response to detecting the face:

moving the 3D caption from being displayed on the surface of the object to being displayed proximate to the face; and

adding the graphical element to the video feed in close proximity to the 3D caption.

19. The method of claim 17 , wherein the graphical element is an emoji or avatar, wherein the graphical element comprises first and second identical emojis or avatars, wherein the first identical emoji or avatar is placed on a left side of the 3D caption and the second identical emoji or avatar is placed on a right side of the 3D caption.

20. A non-transitory machine-readable medium storing instructions which, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving, by one or more processors that implement a messaging application, a video feed

from a camera of a user device that depicts a face;

receiving a request to add a 3D caption to the video feed;

identifying a graphical element that is associated with context of the 3D caption;

displaying the 3D caption and the identified graphical element in the video feed at a position

in 3D space of the video feed proximate to the face depicted in the video feed;

detecting input indicating that a user tapped on screen the 3D caption that is displayed in the video feed; and

in response to detecting the input indicating that the user tapped the 3D caption, presenting text of the 3D caption in 2D to enable the user to modify text of the 3D caption.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2021
From: GOODRICH, KYLE; HARE, SAMUEL EDWARD; LAZAROV, MAXIM MAXIMOV; MATHEW, TONY; MCPHEE, ANDREW JAMES; MORENO, DANIEL; SHANG, WENTAO
To: SNAP INC.
Reel/Frame 058137/0844 →
Cited By (15)
US 1,058,584 US 1,094,423 US 1,109,163 US 1,147,004 US 12,211,159 US 12,217,374 US 12,277,632 US 12,293,433 US 12,347,045 US 12,387,436 US 12,444,138 US 12,482,131 US 12,488,548 US 12,541,929 US 12,580,784