Generating lip sync augmented reality effects
An audio track with vocals is played back using a device with a display screen that displays a video feed from a camera. A location of a mouth depicted in the video feed is detected. A timestamp of playback of the audio track is compared to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback. A viseme is positioned at the detected location of the mouth in the video feed.
1 . A system comprising:
at least one processor;
at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
playing back an audio track;
detecting a location of a mouth depicted in a video feed captured by a camera of the system;
comparing a timestamp of playback of the audio track to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback;
scaling the viseme based on a distance between two facial features detected in the video feed, the distance between the two facial features being fixed in real life, thereby to decouple a size of the viseme from a size of the mouth depicted in the video feed; and
positioning the viseme at the detected location of the mouth in the video feed.
2 . The system of claim 1 , wherein the operations further comprise:
detecting an updated location of the mouth depicted in the video feed; and
positioning the viseme at the detected updated location of the mouth in the video feed.
3 . The system of claim 1 , wherein the operations further comprise:
comparing an updated timestamp of the playback of the audio track to the viseme-timestamp data for the audio track to identify an updated viseme corresponding to the updated timestamp of the audio playback; and
positioning the updated viseme at the detected location of the mouth in the video feed.
4 . The system of claim 1 , wherein the detection of the location of the mouth in the video feed comprises detecting corners of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises positioning corners of a mouth depicted in the viseme along a line between the corners of the mouth in the video feed.
5 . The system of claim 1 , wherein the detection of the location of the mouth in the video feed comprises detecting an angle of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises rotating the viseme by the detected angle.
6 . The system of claim 1 , wherein the operations further comprise:
receiving user selection of a set of visemes from a plurality of available sets of visemes, each available set of visemes having a different visual style; and
retrieving the viseme from the set of visemes indicated by the user selection.
7 . The system of claim 1 , wherein the viseme-timestamp data comprises first and second sets of viseme-timestamps for two vocal tracks, and wherein the operations further comprise:
detecting first and second mouths in the video feed;
positioning a first viseme at a location of the first mouth in the video feed based on the first set of viseme-timestamps; and
positioning a second viseme at a location of the second mouth in the video feed based on the second set of viseme-timestamps.
8 . The system of claim 7 , wherein the first set of viseme-timestamps corresponds to a lead vocal and the first mouth is larger in the video feed than the second mouth.
9 . The system of claim 7 , wherein assignment of the first set of viseme-timestamps to a mouth is done randomly.
10 . A method, executed by one or more processors, the method comprising:
playing back an audio track;
detecting a location of a mouth depicted in a video feed captured by a camera;
comparing a timestamp of playback of the audio track to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback;
scaling the viseme based on a distance between two facial features detected in the video feed, the distance between the two facial features being fixed in real life, thereby to decouple a size of the viseme from a size of the mouth depicted in the video feed; and
positioning the viseme at the detected location of the mouth in the video feed.
11 . The method of claim 10 , further comprising:
detecting an updated location of the mouth depicted in the video feed; and
positioning the viseme at the detected updated location of the mouth in the video feed.
12 . The method of claim 10 , further comprising:
comparing an updated timestamp of the playback of the audio track to the viseme-timestamp data for the audio track to identify an updated viseme corresponding to the updated timestamp of the audio playback; and
positioning the updated viseme at the detected location of the mouth in the video feed.
13 . The method of claim 10 , wherein the viseme-timestamp data comprises first and second sets of viseme-timestamps for two vocal tracks, the method further comprising:
detecting first and second mouths in the video feed;
positioning a first viseme at a location of the first mouth in the video feed based on the first set of viseme-timestamps; and
positioning a second viseme at a location of the second mouth in the video feed based on the second set of viseme-timestamps.
14 . The method of claim 13 , wherein the first set of viseme-timestamps corresponds to a lead vocal and the first mouth is larger in the video feed than the second mouth.
15 . The method of claim 13 , wherein assignment of the first set of viseme-timestamps to a mouth is done randomly.
16 . A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
playing back an audio track;
detecting a location of a mouth depicted in a video feed captured by a camera;
comparing a timestamp of playback of the audio track to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback;
scaling the viseme based on a distance between two facial features detected in the video feed, the distance between the two facial features being fixed in real life, thereby to decouple a size of the viseme from a size of the mouth depicted in the video feed; and
positioning the viseme at the detected location of the mouth in the video feed.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein the detection of the location of the mouth in the video feed comprises detecting an angle of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises rotating the viseme by the detected angle.
18 . The non-transitory computer-readable storage medium of claim 16 , wherein the detection of the location of the mouth in the video feed comprises detecting corners of the mouth in the video feed, and wherein positioning the viseme at the detected location of the mouth in the video feed comprises positioning corners of a mouth depicted in the viseme along a line between the corners of the mouth in the video feed.
19 . The non-transitory computer-readable storage medium of claim 16 , wherein the viseme-timestamp data comprises first and second sets of viseme-timestamps for two vocal tracks, and wherein the operations further comprise:
detecting first and second mouths in the video feed;
positioning a first viseme at a location of the first mouth in the video feed based on the first set of viseme-timestamps; and
positioning a second viseme at a location of the second mouth in the video feed based on the second set of viseme-timestamps.
20 . The non-transitory computer-readable storage medium of claim 19 , wherein assignment of the first set of viseme-timestamps to a mouth is done randomly.