AUTOMATICALLY SEGMENTING VIDEO FOR REACTIVE PROFILE PORTRAITS
A reactive profile picture brings a profile image to life by displaying short video segments of the target user expressing a relevant emotion in reaction to an action by a viewing user that relates to content associated with the target user in an online system such as a social media web site. The viewing user therefore experiences a real-time reaction in a manner similar to a face-to-face interaction. The reactive profile picture can be automatically generated from either a video input of the target user or from a single input image of the target user.
1 . A method comprising:
receiving, by a server of an online system, an input video depicting a portrait of a target individual;
detecting locations of facial feature points of the target individual in each frame of the input video;
obtaining, from the input video, an idle frame depicting the target individual in a neutral expression;
comparing baseline locations of the facial feature points of the target individual in the idle frame to locations of the facial feature points in each non-idle frame of the input video to generate respective distance metrics between each of the non-idle frames and the idle frame;
identifying a first peak expression frame at which the respective distance metrics reach a first local peak;
identifying a first start frame before the first peak expression frame and a first end frame after the first peak expression;
generating a first emotion segment comprising a first range of frames beginning at the first start frame and ending at the first end frame; and
storing the first emotion segment to a storage medium.
2 . The method of claim 1 , further comprising:
identifying a second peak expression frame at which the respective distance metrics reach a second local peak;
identifying a second start frame before the second peak expression frame and a second end frame after the second peak expression;
generating a second emotion segment comprising a second range of frames beginning at the second start frame and ending at the second end frame; and
storing the second emotion segment to the storage medium.
3 . The method of claim 1 , wherein storing the first emotion segment to the storage medium comprises:
determining a time location associated with the first peak expression frame;
identifying, from a lookup table, an expected emotion associated with the time location;
generating a metadata tag representing the expected emotion associated with the first emotion segment; and
storing the metadata tag in association with the first emotion segment.
4 . The method of claim 1 , wherein storing the first emotion segment to the storage medium comprises:
performing a facial analysis to identify an emotion associated with the first emotion segment;
generating a metadata tag representing the emotion associated with the first emotion segment; and
storing the metadata tag in association with the first emotion segment.
5 . The method of claim 1 , wherein obtaining the idle frame in the video comprises:
identifying an idle segment comprising a range of frames;
detecting a frame within the idle segment meeting having facial feature points in locations meeting a predefined criteria; and
assigning the frame meeting the predefined criteria as the idle frame.
6 . The method of claim 1 , wherein obtaining the idle frame in the video comprises:
identifying an idle segment comprising a range of frames; and
synthesizing the idle frame by averaging the range of frames in the idle segment.
7 . The method of claim 1 , wherein identifying the first start frame and the first end frame comprises:
identifying a starting range of frames within a predefined range prior to the first peak expression frame;
selecting the first start frame having a best match to the idle frame from the starting range of frames;
identifying an end range of frames within a predefined range after the first peak expression frame; and
selecting the first end frame having a best match to the idle frame from the end range of frames.
8 . A non-transitory computer-readable storage medium storing instructions executable by a processor, the instructions when executed causing the processor to perform steps including:
receiving, by a server of an online system, an input video depicting a portrait of a target individual;
detecting locations of facial feature points of the target individual in each frame of the input video;
obtaining, from the input video, an idle frame depicting the target individual in a neutral expression;
comparing baseline locations of the facial feature points of the target individual in the idle frame to locations of the facial feature points in each non-idle frame of the input video to generate respective distance metrics between each of the non-idle frames and the idle frame;
identifying a first peak expression frame at which the respective distance metrics reach a first local peak;
identifying a first start frame before the first peak expression frame and a first end frame after the first peak expression;
generating a first emotion segment comprising a first range of frames beginning at the first start frame and ending at the first end frame; and
storing the first emotion segment to a storage medium.
9 . The non-transitory computer-readable storage medium of claim 8 , the instructions when executed further causing the processor to perform steps including:
identifying a second peak expression frame at which the respective distance metrics reach a second local peak;
identifying a second start frame before the second peak expression frame and a second end frame after the second peak expression;
generating a second emotion segment comprising a second range of frames beginning at the second start frame and ending at the second end frame; and
storing the second emotion segment to the storage medium.
10 . The non-transitory computer-readable storage medium of claim 8 , wherein storing the first emotion segment to the storage medium comprises:
determining a time location associated with the first peak expression frame;
identifying, from a lookup table, an expected emotion associated with the time location;
generating a metadata tag representing the expected emotion associated with the first emotion segment; and
storing the metadata tag in association with the first emotion segment.
11 . The non-transitory computer-readable storage medium of claim 8 , wherein storing the first emotion segment to the storage medium comprises:
performing a facial analysis to identify an emotion associated with the first emotion segment;
generating a metadata tag representing the emotion associated with the first emotion segment; and
storing the metadata tag in association with the first emotion segment.
12 . The non-transitory computer-readable storage medium of claim 8 , wherein obtaining the idle frame in the video comprises:
identifying an idle segment comprising a range of frames;
detecting a frame within the idle segment meeting having facial feature points in locations meeting a predefined criteria; and
assigning the frame meeting the predefined criteria as the idle frame.
13 . The non-transitory computer-readable storage medium of claim 8 , wherein obtaining the idle frame in the video comprises:
identifying an idle segment comprising a range of frames; and
synthesizing the idle frame by averaging the range of frames in the idle segment.
14 . The non-transitory computer-readable storage medium of claim 8 , wherein identifying the first start frame and the first end frame comprises:
identifying a starting range of frames within a predefined range prior to the first peak expression frame;
selecting the first start frame having a best match to the idle frame from the starting range of frames;
identifying an end range of frames within a predefined range after the first peak expression frame; and
selecting the first end frame having a best match to the idle frame from the end range of frames.
15 . A computer system comprising:
a processor; and
a non-transitory computer-readable storage medium storing instructions executable by the processor, the instructions when executed causing the processor to perform steps including:
receiving an input video depicting a portrait of a target individual;
detecting locations of facial feature points of the target individual in each frame of the input video;
obtaining, from the input video, an idle frame depicting the target individual in a neutral expression;
comparing baseline locations of the facial feature points of the target individual in the idle frame to locations of the facial feature points in each non-idle frame of the input video to generate respective distance metrics between each of the non-idle frames and the idle frame;
identifying a first peak expression frame at which the respective distance metrics reach a first local peak;
identifying a first start frame before the first peak expression frame and a first end frame after the first peak expression;
generating a first emotion segment comprising a first range of frames beginning at the first start frame and ending at the first end frame; and
storing the first emotion segment to a storage medium.
16 . The computer system of claim 15 , the instructions when executed further causing the processor to perform steps including:
identifying a second peak expression frame at which the respective distance metrics reach a second local peak;
identifying a second start frame before the second peak expression frame and a second end frame after the second peak expression;
generating a second emotion segment comprising a second range of frames beginning at the second start frame and ending at the second end frame; and
storing the second emotion segment to the storage medium.
17 . The computer system of claim 15 , wherein storing the first emotion segment to the storage medium comprises:
determining a time location associated with the first peak expression frame;
identifying, from a lookup table, an expected emotion associated with the time location;
generating a metadata tag representing the expected emotion associated with the first emotion segment; and
storing the metadata tag in association with the first emotion segment.
18 . The computer system of claim 15 , wherein storing the first emotion segment to the storage medium comprises:
performing a facial analysis to identify an emotion associated with the first emotion segment;
generating a metadata tag representing the emotion associated with the first emotion segment; and
storing the metadata tag in association with the first emotion segment.
19 . The computer system of claim 15 , wherein obtaining the idle frame in the video comprises:
identifying an idle segment comprising a range of frames;
detecting a frame within the idle segment meeting having facial feature points in locations meeting a predefined criteria; and
assigning the frame meeting the predefined criteria as the idle frame.
20 . The computer system of claim 15 , wherein identifying the first start frame and the first end frame comprises:
identifying a starting range of frames within a predefined range prior to the first peak expression frame;
selecting the first start frame having a best match to the idle frame from the starting range of frames;
identifying an end range of frames within a predefined range after the first peak expression frame; and
selecting the first end frame having a best match to the idle frame from the end range of frames.