IP Library Granted Patent US 12,737,997
Granted Patent B2
US 12,737,997 · App. 18/609,960 · Granted Sep 15, 2026

Generating lip sync augmented reality effects

Inventors: Leonid Gorkin (Chappaqua, NY); Matthew Mahar (San Francisco, CA); Hanbo Chen (Kirkland, WA); Dmytro Barbaruk (Los Angeles, CA)
Assignee: Snap Inc.
G06T19/006G06T19/20G06T2219/2004G06T2219/2016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,997
App. No.
18/609,960
Granted
Sep 15, 2026
Kind
B2
Abstract

An audio track with vocals is played back using a device with a display screen that displays a video feed from a camera. A location of a mouth depicted in the video feed is detected. A timestamp of playback of the audio track is compared to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback. A viseme is positioned at the detected location of the mouth in the video feed.

Claims (56)

1 . A system comprising:

at least one processor;

at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

playing back an audio track;

detecting a location of a mouth depicted in a video feed captured by a camera of the system;

comparing a timestamp of playback of the audio track to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback;

scaling the viseme based on a distance between two facial features detected in the video feed, the distance between the two facial features being fixed in real life, thereby to decouple a size of the viseme from a size of the mouth depicted in the video feed; and

positioning the viseme at the detected location of the mouth in the video feed.

2 . The system of claim 1 , wherein the operations further comprise:

detecting an updated location of the mouth depicted in the video feed; and

positioning the viseme at the detected updated location of the mouth in the video feed.

3 . The system of claim 1 , wherein the operations further comprise:

comparing an updated timestamp of the playback of the audio track to the viseme-timestamp data for the audio track to identify an updated viseme corresponding to the updated timestamp of the audio playback; and

positioning the updated viseme at the detected location of the mouth in the video feed.

4 . The system of claim 1 , wherein the detection of the location of the mouth in the video feed comprises detecting corners of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises positioning corners of a mouth depicted in the viseme along a line between the corners of the mouth in the video feed.

5 . The system of claim 1 , wherein the detection of the location of the mouth in the video feed comprises detecting an angle of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises rotating the viseme by the detected angle.

6 . The system of claim 1 , wherein the operations further comprise:

receiving user selection of a set of visemes from a plurality of available sets of visemes, each available set of visemes having a different visual style; and

retrieving the viseme from the set of visemes indicated by the user selection.

7 . The system of claim 1 , wherein the viseme-timestamp data comprises first and second sets of viseme-timestamps for two vocal tracks, and wherein the operations further comprise:

detecting first and second mouths in the video feed;

positioning a first viseme at a location of the first mouth in the video feed based on the first set of viseme-timestamps; and

positioning a second viseme at a location of the second mouth in the video feed based on the second set of viseme-timestamps.

8 . The system of claim 7 , wherein the first set of viseme-timestamps corresponds to a lead vocal and the first mouth is larger in the video feed than the second mouth.

9 . The system of claim 7 , wherein assignment of the first set of viseme-timestamps to a mouth is done randomly.

10 . A method, executed by one or more processors, the method comprising:

playing back an audio track;

detecting a location of a mouth depicted in a video feed captured by a camera;

comparing a timestamp of playback of the audio track to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback;

scaling the viseme based on a distance between two facial features detected in the video feed, the distance between the two facial features being fixed in real life, thereby to decouple a size of the viseme from a size of the mouth depicted in the video feed; and

positioning the viseme at the detected location of the mouth in the video feed.

11 . The method of claim 10 , further comprising:

detecting an updated location of the mouth depicted in the video feed; and

positioning the viseme at the detected updated location of the mouth in the video feed.

12 . The method of claim 10 , further comprising:

comparing an updated timestamp of the playback of the audio track to the viseme-timestamp data for the audio track to identify an updated viseme corresponding to the updated timestamp of the audio playback; and

positioning the updated viseme at the detected location of the mouth in the video feed.

13 . The method of claim 10 , wherein the viseme-timestamp data comprises first and second sets of viseme-timestamps for two vocal tracks, the method further comprising:

detecting first and second mouths in the video feed;

positioning a first viseme at a location of the first mouth in the video feed based on the first set of viseme-timestamps; and

positioning a second viseme at a location of the second mouth in the video feed based on the second set of viseme-timestamps.

14 . The method of claim 13 , wherein the first set of viseme-timestamps corresponds to a lead vocal and the first mouth is larger in the video feed than the second mouth.

15 . The method of claim 13 , wherein assignment of the first set of viseme-timestamps to a mouth is done randomly.

16 . A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

playing back an audio track;

detecting a location of a mouth depicted in a video feed captured by a camera;

comparing a timestamp of playback of the audio track to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback;

scaling the viseme based on a distance between two facial features detected in the video feed, the distance between the two facial features being fixed in real life, thereby to decouple a size of the viseme from a size of the mouth depicted in the video feed; and

positioning the viseme at the detected location of the mouth in the video feed.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the detection of the location of the mouth in the video feed comprises detecting an angle of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises rotating the viseme by the detected angle.

18 . The non-transitory computer-readable storage medium of claim 16 , wherein the detection of the location of the mouth in the video feed comprises detecting corners of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises positioning corners of a mouth depicted in the viseme along a line between the corners of the mouth in the video feed.

19 . The non-transitory computer-readable storage medium of claim 16 , wherein the viseme-timestamp data comprises first and second sets of viseme-timestamps for two vocal tracks, and wherein the operations further comprise:

detecting first and second mouths in the video feed;

positioning a first viseme at a location of the first mouth in the video feed based on the first set of viseme-timestamps; and

positioning a second viseme at a location of the second mouth in the video feed based on the second set of viseme-timestamps.

20 . The non-transitory computer-readable storage medium of claim 19 , wherein assignment of the first set of viseme-timestamps to a mouth is done randomly.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2024
From: GORKIN, LEONID; MAHAR, MATTHEW; CHEN, HANBO; BARBARUK, DMYTRO
To: SNAP INC.
Reel/Frame 066835/0569 →
Continuity (1)
Related Publication 20250299449A1 · Sep 25, 2025
References Cited (57)
US 10242477B1 · Charlton et al. · 2019 [cited by applicant]
US 10432559B2 · Baldwin et al. · 2019 [cited by applicant]
US 10467792B1 · Roche et al. · 2019 [cited by applicant]
US 10770092B1 · Adams et al. · 2020 [cited by applicant]
US 10788900B1 · Brendel et al. · 2020 [cited by applicant]
US 11468883B2 · Ribas Machado et al. · 2022 [cited by applicant]
US 20020097380A1 · Moulton · 2002 [cited by examiner]
US 20020194006A1 · Challapali · 2002 [cited by applicant]
US 20030050778A1 · Nguyen et al. · 2003 [cited by applicant]
US 20060143569A1 · Kinsella et al. · 2006 [cited by applicant]
US 20080120258A1 · Shin et al. · 2008 [cited by applicant]
US 20100141662A1 · Storey et al. · 2010 [cited by applicant]
US 20100332229A1 · Aoyama · 2010 [cited by examiner]
US 20120130717A1 · Xu et al. · 2012 [cited by applicant]
US 20130307856A1 · Keane et al. · 2013 [cited by applicant]
US 20150100537A1 · Grieves et al. · 2015 [cited by applicant]
US 20160012853A1 · Cabanilla · 2016 [cited by examiner]
US 20160292148A1 · Aley et al. · 2016 [cited by applicant]
US 20170154314A1 · Mones et al. · 2017 [cited by applicant]
US 20170300462A1 · Cudworth et al. · 2017 [cited by applicant]
US 20180083898A1 · Pham · 2018 [cited by applicant]
US 20180113587A1 · Allen et al. · 2018 [cited by applicant]
US 20180136794A1 · Cassidy et al. · 2018 [cited by applicant]
US 20180210874A1 · Fuxman et al. · 2018 [cited by applicant]
US 20180253881A1 · Edwards et al. · 2018 [cited by applicant]
US 20180253895A1 · Arumugam · 2018 [cited by applicant]
US 20200106728A1 · Grantham et al. · 2020 [cited by applicant]
US 20200125322A1 · Wilde · 2020 [cited by applicant]
US 20200175061A1 · Penta et al. · 2020 [cited by applicant]
US 20210192800A1 · Dutta et al. · 2021 [cited by applicant]
US 20210335350A1 · Ribas Machado Das Neves et al. · 2021 [cited by applicant]
US 20210385179A1 · Heikkinen et al. · 2021 [cited by applicant]
US 20220012929A1 · Blackstock et al. · 2022 [cited by applicant]
US 20220019747A1 · Guo et al. · 2022 [cited by applicant]
US 20220109646A1 · Lakshmipathy · 2022 [cited by applicant]
US 20220284884A1 · Tongya · 2022 [cited by applicant]
US 20240062008A1 · Ghosh et al. · 2024 [cited by applicant]
US 20240104789A1 · Ghosh et al. · 2024 [cited by applicant]
CN 107977928 · 2018 [cited by applicant]
CN 110136216 · 2019 [cited by applicant]
CN 110163220 · 2019 [cited by applicant]
CN 110554782 · 2019 [cited by applicant]
CN 114187405 · 2022 [cited by applicant]
KR 20060125333 · 2006 [cited by applicant]
KR 20200095781 · 2020 [cited by applicant]
WO 2021137942 · 2021 [cited by applicant]
WO 2023137557 · 2023 [cited by applicant]
WO 2024039957 · 2024 [cited by applicant]
WO 2024064806 · 2024 [cited by applicant]
“2D Animated TTS”, [Online]. Retrieved from the Internet: <URL: https://docs.snap.com/lens-studio/references/templates/audio/2d-animated-tts#guide>, (Accessed Mar. 19, 2024), 17 pgs. [cited by applicant]
Huggins-Daines, David, “GitHub—cmusphinx/pocketsphinx: A small speech recognizer”, [Online]. Retrieved from the Internet: <URL: https://github.com/cmusphinx/pocketsphinx/releases>, (Accessed Jan. 10, 2024), 5 pgs. [cited by applicant]
Moussallam, Manuel, “Releasing Spleeter: Deezer Research source separation engine”, [Online]. Retrieved from the Internet: <URL: https://deezer.io/releasing-spleeter-deezer-r-d-source-separation-engine-2b88985e797e>, (N… [cited by applicant]
Zhou, Yang, et al., “VisemeNet: Audio-Driven Animator-Centric Speech Animation”, ACM Transactions on Graphics, vol. 37, No. 4, Article 161, (Aug. 2018), 10 pgs. [cited by applicant]
Vougioukas, Konstantinos, “End-to-End Speech-Driven Facial Animation with Temporal GANs”, arXiv:1805.09313v4 [eess.AS], (Jul. 19, 2018), 14 pgs. [cited by applicant]
Wang, Xingyao, “An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Sep. 2… [cited by applicant]
“International Application Serial No. PCT US2025 015704, International Search Report mailed Jun. 12, 2025”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT US2025 015704, Written Opinion mailed Jun. 12, 2025”, 5 pgs. [cited by applicant]