Generating segmentation masks for objects in digital videos using pose tracking data
The present disclosure relates to systems, methods, and non-transitory computer-readable media that generate joint-based segmentation masks for digital objects portrayed in digital videos. In particular, in one or more embodiments, the disclosed systems utilize a video masking model having a pose tracking neural network and a segmentation neural network to generate the joint-based segmentation masks. To illustrate, in some embodiments, the disclosed systems utilize the pose tracking neural network to identify a set of joints of the digital object across the frames of the digital video. The disclosed systems further utilize the segmentation neural network to generate joint-based segmentation masks for the video frames that portray the object using the identified joints. In some cases, the segmentation neural network includes a multi-layer perceptron mixer layer for mixing visual features propagated via convolutional layers.
1 . A method comprising:
generating a down-sampled frame corresponding to a frame of a digital video;
determining, using a global pose tracking neural network and from the down-sampled frame, a set of joint coordinates corresponding to a digital object portrayed in the frame of the digital video;
generating, using the set of joint coordinates, a joint heat map corresponding to the digital object portrayed in the frame of the digital video, the joint heat map having values that distinguish between joint points corresponding to the set of joint coordinates and areas of the frame of the digital video that surround the joint points;
determining, using one or more convolutional layers of a local segmentation neural network, visual features from the joint heat map and the frame of the digital video;
generating, using at least one multi-perceptron mixer layer of the local segmentation neural network, mixer encodings that mix the visual features across feature channels by generating the mixer encodings from the visual features using layer normalization and an activation function of the at least one multi-perceptron mixer layer; and
generating, using the local segmentation neural network and from the mixer encodings, a joint-based segmentation mask that corresponds to the digital object portrayed in the frame.
2 . The method of claim 1 , wherein determining, using the global pose tracking neural network and the down-sampled frame, the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video comprises determining the set of joint coordinates by using the global pose tracking neural network to globally analyze the down-sampled frame and an additional down-sampled frame corresponding to a preceding frame of the digital video.
3 . The method of claim 1 , wherein generating the joint-based segmentation mask that corresponds to the digital object comprises generating the joint-based segmentation mask from a bounding box associated with the digital object for the frame of the digital video.
4 . The method of claim 1 , wherein generating, using the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask comprises generating, using one or more additional convolutional layers of the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask.
5 . The method of claim 3 , further comprising:
determining the bounding box associated with the digital object for the frame of the digital video based on the set of joint coordinates; and
generating, from the frame of the digital video, a cropped frame that includes the digital object using the bounding box,
wherein generating the joint-based segmentation mask from the bounding box associated with the digital object comprises generating the joint-based segmentation mask from the cropped frame.
6 . The method of claim 1 ,
further comprising extracting, from the digital video, the frame of the digital video and a preceding frame of the digital video,
wherein determining the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video comprises determining the set of joint coordinates using the frame of the digital video and the preceding frame of the digital video.
7 . The method of claim 1 , wherein generating the joint heat map using the set of joint coordinates comprises centering a Gaussian distribution at each joint point associated with the set of joint coordinates.
8 . The method of claim 1 , further comprising modifying the frame of the digital video utilizing the joint-based segmentation mask.
9 . A non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
generating a down-sampled frame corresponding to a frame of a digital video;
determining, using a global pose tracking neural network and from the down-sampled frame, a set of joint coordinates corresponding to a digital object portrayed in the frame of the digital video;
generating, using the set of joint coordinates, a joint heat map corresponding to the digital object portrayed in the frame of the digital video, the joint heat map having values that distinguish between joint points corresponding to the set of joint coordinates and areas of the frame of the digital video that surround the joint points;
determining, using one or more convolutional layers of a local segmentation neural network, visual features from the joint heat map and the frame of the digital video;
generating, using at least one multi-perceptron mixer layer of the local segmentation neural network, mixer encodings that mix the visual features across feature channels by generating the mixer encodings from the visual features using layer normalization and an activation function of the at least one multi-perceptron mixer layer; and
generating, using the local segmentation neural network and from the mixer encodings, a joint-based segmentation mask that corresponds to the digital object portrayed in the frame.
10 . The non-transitory computer-readable medium of claim 9 , wherein:
determining, utilizing the global pose tracking neural network and the down-sampled frame, the set of joint coordinates comprises determining the set of joint coordinates by globally analyzing the down-sampled frame using the global pose tracking neural network; and
determining the visual features from the joint heat map and the frame of the digital video using the one or more convolutional layers of the local segmentation neural network comprises determining the visual features by locally analyzing a portion of the frame of the digital video that contains the digital object based on the joint heat map using the one or more convolutional layers of the local segmentation neural network.
11 . The non-transitory computer-readable medium of claim 10 , wherein:
the operations further comprise generating, from the frame of the digital video, a cropped frame that includes the digital object based on the set of joint coordinates; and
determining the visual features by locally analyzing the portion of the frame of the digital video that contains the digital object based on the joint heat map utilizing the one or more convolutional layers of the local segmentation neural network comprises determining the visual features based on the cropped frame and the joint heat map utilizing the one or more convolutional layers of the local segmentation neural network.
12 . The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise determining, utilizing the global pose tracking neural network, a bounding box associated with the digital object and a tracking identifier that distinguishes the digital object from other digital objects portrayed in the frame of the digital video.
13 . The non-transitory computer-readable medium of claim 9 , wherein determining, utilizing the global pose tracking neural network, the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video comprises determining, utilizing the global pose tracking neural network, the set of joint coordinates associated with a set of pre-determined human joints corresponding to a person portrayed in the frame of the digital video.
14 . The non-transitory computer-readable medium of claim 9 , wherein generating, using the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask comprises generating, using one or more additional convolutional layers of the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask.
15 . The non-transitory computer-readable medium of claim 10 , wherein the operations further comprise:
determining, utilizing the global pose tracking neural network, an additional set of joint coordinates corresponding to an additional digital object portrayed in the frame of the digital video; and
generating, utilizing the local segmentation neural network, an additional joint-based segmentation mask corresponding to the additional digital object based on the additional set of joint coordinates.
16 . The non-transitory computer-readable medium of claim 10 , wherein the operations further comprise:
determining intermediate feature values associated with the digital object portrayed in the frame of the digital video utilizing the global pose tracking neural network; and
providing the intermediate feature values from the global pose tracking neural network to the local segmentation neural network for generating the joint-based segmentation mask via a skip link.
17 . A system comprising:
at least one memory device; and
at least one server device configured to cause the system to:
generate a down-sampled frame corresponding to a frame of a digital video;
determine, utilizing a global pose tracking neural network and from the down-sampled frame, a set of joint coordinates corresponding to a digital object portrayed in the frame of the digital video;
generate, using the set of joint coordinates, a joint heat map corresponding to the digital object portrayed in the frame of the digital video, the joint heat map having values that distinguish between joint points corresponding to the set of joint coordinates and areas of the frame of the digital video that surround the joint points;
determine, utilizing an encoder of a segmentation neural network, mixer encodings for the digital object based on the joint heat map, the encoder comprising:
a plurality of convolutional layers that determine a set of features using the joint heat map; and
a multi-layer perceptron mixer layer that determines the mixer encodings by mixing the set of features using layer normalization and an activation function; and
generate, utilizing a decoder of the segmentation neural network, a joint-based segmentation mask for the digital object portrayed in the frame of the digital video based on the mixer encodings.
18 . The system of claim 17 , wherein the at least one server device is configured to cause the system to determine, utilizing the global pose tracking neural network, the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video by determining, utilizing the global pose tracking neural network, the set of joint coordinates based on the down-sampled frame and an additional down-sampled frame that corresponds to a preceding frame of the digital video that portrays the digital object.
19 . The system of claim 17 , wherein the at least one server device is further configured to cause the system to:
determine intermediate feature values associated with the digital object utilizing the multi-layer perceptron mixer layer of the encoder; and
provide the intermediate feature values from the encoder of the segmentation neural network to the decoder of the segmentation neural network for generating the joint-based segmentation mask via a skip link.
20 . The system of claim 17 , wherein the at least one server device is configured to cause the system to generate, utilizing the decoder of the segmentation neural network, the joint-based segmentation mask based on the mixer encodings by:
generating, using one or more convolutional layers of the decoder and from the mixer encodings, the joint-based segmentation mask.