IP Library Granted Patent US 12,430,812
Granted Patent B2
US 12,430,812 · App. 17/950,945 · Granted Sep 30, 2025

Text-guided cameo generation

Inventors: Arnab Ghosh (Oxford, GB); Jian Ren (Marina Del Ray, CA); Pavel Savchenkov (London, GB); Sergey Tulyakov (Marina del Rey, CA)
Assignee: Snap Inc.
G06T11/00G06F40/289G06F40/35G06V40/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,812
App. No.
17/950,945
Granted
Sep 30, 2025
Kind
B2
Abstract

A method of generating an image for use in a conversation taking place in a messaging application is disclosed. Conversation input text is received from a user of a portable device that includes a display. Model input text is generated from the conversation input text, which is processed with a text-to-image model to generate an image based on the model input text. The coordinates of a face in the image are determined, and the face of the user or another person is added to the image at the location. The final image is displayed on the portable device, and user input is received to transmit the image to a remote recipient.

Claims (52)

1. A computer-implemented method of generating an image including an existing representation of a face of a person, for use in a conversation taking place in a messaging application, the method comprising:

receiving conversation input text from a user of a portable device that includes a display;

generating model input text from the conversation input text;

generating an image based on the model input text using a text-to-image model;

determining coordinates of a face in the generated image;

applying the existing representation of the face of the person to the generated image based on the coordinates of the face in the generated image, to generate an updated image including the existing representation of the face of the person;

displaying the updated image on the display of the portable device;

receiving user input to transmit the updated image in a message; and

transmitting, in response to receiving the user input, the updated image to a remote recipient.

2. The method of claim 1 , wherein the generating of the model input text comprises generating additional text using a creative caption function.

3. The method of claim 2 , wherein the generating of the model input text further comprises:

extracting key phrases from the additional text.

4. The method of claim 1 , further comprising:

processing the conversation input text with a safety filter to determine suitability of the conversation input text for image generation.

5. The method of claim 1 , wherein the text-to-image model has been generated from a large scale image dataset and refined by an existing collection of images for use in a conversation taking place in a messaging application.

6. The method of claim 5 , wherein text associated with images in the existing collection of images is expanded using an image-to-text module.

7. The method of claim 5 , further comprising:

animating the face of the person in the updated image.

8. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to perform operations for generating an image including an existing representation of a face of a person, for use in a conversation taking place in a messaging application, the operations comprising:

receiving conversation input text from a user of a portable device that includes a display;

generating model input text from the conversation input text;

generating an image based on the model input text using a text-to-image model;

determining coordinates of a face in the generated image;

applying the existing representation of the face of the person to the generated image based on the coordinates of the face in the generated image, to generate an updated image including the existing representation of the face of the person;

displaying the updated image on the display of the portable device;

receiving user input to transmit the updated image in a message; and

transmitting, in response to receiving the user input, the updated image to a remote recipient.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the generating of the model input text comprises generating additional text using a creative caption function.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the generating of the model input text further comprises:

extracting key phrases from the additional text.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the operations further comprise:

processing the conversation input text with a safety filter to determine suitability of the conversation input text for image generation.

12. The non-transitory computer-readable storage medium of claim 8 , wherein the text-to-image model has been generated from a large scale image dataset and refined by an existing collection of images for use in a conversation taking place in a messaging application.

13. The non-transitory computer-readable storage medium of claim 12 , wherein text associated with images in the existing collection of images is expanded using an image-to-text module.

14. The non-transitory computer-readable storage medium of claim 13 , wherein text associated with the images in the existing collection of images is filtered for relevance of the text to the images in the existing collection of images, by using an image-text relevance model.

15. A computing apparatus comprising:

a processor; and

a memory storing instructions that, when executed by the processor, configure the apparatus to perform operations for generating an image including an existing representation of a face of a person, for use in a conversation taking place in a messaging application, the operations comprising:

receiving conversation input text from a user of a portable device that includes a display;

generating model input text from the conversation input text;

generating an image based on the model input text using a text-to-image model;

determining coordinates of a face in the generated image;

applying the existing representation of the face of the person to the generated image based on the coordinates of the face in the generated image, to generate an updated image including the existing representation of the face of the person;

displaying the updated image on the display of the portable device;

receiving user input to transmit the updated image in a message; and

transmitting, in response to receiving the user input, the updated image to a remote recipient.

16. The computing apparatus of claim 15 , wherein the generating of the model input text comprises generating additional text using a creative caption function.

17. The computing apparatus of claim 16 , wherein the generating of the model input text further comprises:

extracting key phrases from the additional text.

18. The computing apparatus of claim 15 , wherein the text-to-image model has been generated from a large scale image dataset and refined by an existing collection of images for use in a conversation taking place in a messaging application.

19. The computing apparatus of claim 18 , wherein text associated with images in the existing collection of images is expanded using an image-to-text module.

20. The computing apparatus of claim 18 , wherein text associated with the images in the existing collection of images is filtered for relevance of the text to the images in the existing collection of images, by using an image-text relevance model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: GHOSH, ARNAB; REN, JIAN; SAVCHENKOV, PAVEL; TULYAKOV, SERGEY
To: SNAP, INC.
Reel/Frame 061188/0905 →
Continuity (1)
Related Publication 20240104789A1 · Mar 28, 2024
References Cited (68)
US 10242477B1 · Charlton et al. · 2019 [cited by applicant]
US 10432559B2 · Baldwin et al. · 2019 [cited by applicant]
US 10467792B1 · Roche et al. · 2019 [cited by applicant]
US 10788900B1 · Brendel et al. · 2020 [cited by applicant]
US 11468883B2 · Ribas Machado Das Neves et al. · 2022 [cited by applicant]
US 20020194006A1 · Challapali · 2002 [cited by examiner]
US 20030050778A1 · Nguyen et al. · 2003 [cited by applicant]
US 20060143569A1 · Kinsella et al. · 2006 [cited by applicant]
US 20080120258A1 · Shin et al. · 2008 [cited by applicant]
US 20090252435A1 · Wen · 2009 [cited by examiner]
US 20100141662A1 · Storey · 2010 [cited by examiner]
US 20120130717A1 · Xu et al. · 2012 [cited by applicant]
US 20150100537A1 · Grieves et al. · 2015 [cited by applicant]
US 20160092410A1 · Martin · 2016 [cited by applicant]
US 20160292148A1 · Aley et al. · 2016 [cited by applicant]
US 20170154314A1 · Mones et al. · 2017 [cited by applicant]
US 20170300462A1 · Cudworth et al. · 2017 [cited by applicant]
US 20180083898A1 · Pham · 2018 [cited by examiner]
US 20180113587A1 · Allen et al. · 2018 [cited by applicant]
US 20180136794A1 · Cassidy et al. · 2018 [cited by applicant]
US 20180189822A1 · Kulkarni et al. · 2018 [cited by applicant]
US 20180210874A1 · Fuxman et al. · 2018 [cited by applicant]
US 20200106728A1 · Grantham et al. · 2020 [cited by applicant]
US 20200175061A1 · Penta et al. · 2020 [cited by applicant]
US 20200403817A1 · Daredia et al. · 2020 [cited by applicant]
US 20210192800A1 · Dutta et al. · 2021 [cited by applicant]
US 20210335350A1 · Ribas Machado Das Neves et al. · 2021 [cited by applicant]
US 20210385179A1 · Heikkinen et al. · 2021 [cited by applicant]
US 20220012929A1 · Blackstock et al. · 2022 [cited by applicant]
US 20220019747A1 · Guo et al. · 2022 [cited by applicant]
US 20220109646A1 · Lakshmipathy · 2022 [cited by applicant]
US 20220284884A1 · Tongya · 2022 [cited by examiner]
US 20240062008A1 · Ghosh et al. · 2024 [cited by applicant]
CN 107977928 · 2018 [cited by applicant]
CN 110136216 · 2019 [cited by applicant]
CN 110163220 · 2019 [cited by applicant]
CN 110554782 · 2019 [cited by applicant]
CN 114187405 · 2022 [cited by applicant]
KR 20060125333 · 2006 [cited by applicant]
KR 20200095781 · 2020 [cited by applicant]
WO 2021137942 · 2021 [cited by applicant]
WO 2024039957 · 2024 [cited by applicant]
WO 2024064806 · 2024 [cited by applicant]
“U.S. Appl. No. 17/820,437, Non Final Office Action mailed Jan. 20, 2023”, 21 pgs. [cited by applicant]
“U.S. Appl. No. 17/820,437, Response filed Apr. 19, 2023 to Non Final Office Action mailed Jan. 20, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/820,437, Non Final Office Action mailed Sep. 21, 2023”, 20 pgs. [cited by applicant]
“International Application Serial No. PCT US2023 071003, International Search Report mailed Oct. 6, 2023”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT US2023 071003, Written Opinion mailed Oct. 6, 2023”, 4 pgs. [cited by applicant]
U.S. Appl. No. 17/820,437, filed Aug. 17, 2022, Text-Guided Sticker Generation. [cited by applicant]
“International Application Serial No. PCT/US2023/074762, International Search Report mailed Jan. 8, 2024”, 5 pgs. [cited by applicant]
“International Application Serial No. PCT/US2023/074762, Written Opinion mailed Jan. 8, 2024”, 10 pgs. [cited by applicant]
“U.S. Appl. No. 17/820,437, Response filed Jan. 22, 2024 to Non Final Office Action mailed Sep. 21, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/820,437, Final Office Action mailed Feb. 23, 2024”, 14 pgs. [cited by applicant]
Wang, Xingyao, “An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Sep. 2… [cited by applicant]
“U.S. Appl. No. 17/820,437, Response filed Apr. 23, 2024 to Final Office Action mailed Feb. 23, 2024”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/820,437, Advisory Action mailed May 10, 2024”, 5 pgs. [cited by applicant]
“U.S. Appl. No. 17/820,437, Response filed May 23, 2024 to Advisory Action mailed May 10, 2024”, 11 pgs. [cited by applicant]
Brown, Tom B, et al., “Language Models are Few-Shot Learners”, arXiv:2005.14165v4 [cs.CL], (75 pgs), Jul. 22, 2020. [cited by applicant]
Crowson, Katherine, et al., “VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance”, arXiv:2204.08583 [cs.CV], (Apr. 18, 2022), 31 pgs. [cited by applicant]
Esser, Patrick, et al., “Taming Transformers for High-Resolution Image Synthesis”, arXiv:2012.09841v3 [cs.CV], (Jun. 23, 2021), 52 pgs. [cited by applicant]
Li, Junnan, et al., “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation”, arXiv: 2201.12086v2 [cs.CV], (Feb. 15, 2022), 12 pgs. [cited by applicant]
Radford, Alec, et al., “CLIP: Connecting Text and Images”, OpenAI Blog, [Online] Retrieved from the Internet: <URL: https://openai.com/blog/clip/>, [Retrieved on Jul. 22, 2022], (Jan. 15, 2021), 16 pgs. [cited by applicant]
Radford, Alec, et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020v1 [cs.CV], (Feb. 26, 2021), 48 pgs. [cited by applicant]
Rombach, Robin, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv:2112.10752v2 [cs.CV], (Apr. 13, 2022), 45 pgs. [cited by applicant]
Stanton, William, “How to Do Cameos in Snapchat”, https://www.alphr.com/snapchat-cameos/, (Jul. 1, 2021). [cited by applicant]
Vougioukas, Konstantinos, et al., “End-to-End Speech-Driven Facial Animation with Temporal GANs”, arXiv:1805.09313v4 [eess.AS], (Jul. 19, 2018), 14 pgs. [cited by applicant]
Navali, et al., “Sentence Generation Using Selective Text Prediction”, Comp, y Sist. vol. 23 No.3 Ciudad de MA©xico Jul./Sep, [Online] Retrieved from the internet: <https://doi.org/10.13053/cys-23-3-3252>, (2019), 8 pgs. [cited by applicant]
Ying, Hua Tan, et al., “Phrase-based image caption generator with hierarchical LSTM network”, Neurocomputing, vol. 333, ISSN 0925-2312, [Online] Retrieved from the internet: <https://doi.Org/10.1016/j.neucom.2018.12.026… [cited by applicant]