Transformer-based image segmentation on mobile devices
The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating segmentation masks for a digital visual media item. In particular, in one or more embodiments, the disclosed systems generate, utilizing a neural network encoder, high-level features of a digital visual media item. Further, the disclosed systems generate, utilizing the neural network encoder, low-level features of the digital visual media item. In some implementations, the disclosed systems generate, utilizing a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features. Moreover, the disclosed systems generate, utilizing the neural network decoder, a refined segmentation mask of the digital visual media item from the initial segmentation mask and the high-level features.
1 . A computer-implemented method comprising:
generating, utilizing a neural network encoder, high-level features of a digital visual media item;
generating, utilizing the neural network encoder, low-level features of the digital visual media item;
generating, utilizing a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features; and
generating, utilizing the neural network decoder, a refined segmentation mask of the digital visual media item by passing the initial segmentation mask and the high-level features through the neural network decoder.
2 . The computer-implemented method of claim 1 , wherein the neural network encoder comprises a transformer block; and further comprising:
down-sampling a portion of the low-level features to match a dimension of a lower-level set of the low-level features; and
generating a transformed set of features by passing the down-sampled portion of the low- level features through the transformer block.
3 . The computer-implemented method of claim 1 , wherein generating the initial segmentation mask of the digital visual media item from the low-level features comprises decoding the low-level features, without the high-level features.
4 . The computer-implemented method of claim 3 , further comprising:
up-sampling a portion of the low-level features to match a dimension of a higher-level set of the low-level features; and
combining the up-sampled portion of the low-level features with the higher-level set of the low-level features utilizing a series sum operation.
5 . The computer-implemented method of claim 1 , wherein generating the refined segmentation mask of the digital visual media item from the initial segmentation mask and the high-level features comprises decoding the initial segmentation mask and the high-level features by:
generating, utilizing a multilayer perceptron, refined high-level features from the initial segmentation mask and the high-level features;
up-sampling the refined high-level features; and
combining the up-sampled refined high-level features.
6 . The computer-implemented method of claim 5 , wherein combining the up-sampled refined high-level features comprises utilizing a concatenation operation.
7 . The computer-implemented method of claim 5 , wherein combining the up-sampled refined high-level features comprises utilizing a series sum operation.
8 . The computer-implemented method of claim 1 , further comprising generating, utilizing a feature refinement head, a feature refinement segmentation mask of the digital visual media item.
9 . The computer-implemented method of claim 1 , wherein at least one of generating the initial segmentation mask or generating the refined segmentation mask comprises segmenting a portrait into a plurality of semantic regions.
10 . The computer-implemented method of claim 1 , wherein generating the initial segmentation mask and generating the refined segmentation mask are performed on a mobile device.
11 . A non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
generating, utilizing a neural network encoder, high-level features of a digital visual media item;
generating, utilizing the neural network encoder, low-level features of the digital visual media item;
generating, utilizing a first segmentation head of a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features; and
generating a refined segmentation mask of the digital visual media item by processing the initial segmentation mask and the high-level features utilizing a second segmentation head of the neural network decoder.
12 . The non-transitory computer-readable medium of claim 11 , wherein the neural network encoder comprises a transformer block; and wherein the operations further comprise:
down-sampling a portion of the low-level features to match a dimension of a lower-level set of the low-level features; and
generating a transformed set of features by passing the down-sampled portion of the low-level features through the transformer block.
13 . The non-transitory computer-readable medium of claim 11 , wherein generating the initial segmentation mask of the digital visual media item from the low-level features comprises decoding the low-level features without the high-level features.
14 . The non-transitory computer-readable medium of claim 11 , wherein the instructions cause a processor of a mobile device to perform the operations.
15 . A system comprising:
at least one processor; and
at least one memory device coupled to the at least one processor that causes the system to perform operations comprising:
generating, utilizing a neural network encoder, high-level features of a digital visual media item;
generating, utilizing the neural network encoder, low-level features of the digital visual media item;
generating, utilizing a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features; and
generating, utilizing the neural network decoder, a refined segmentation mask of the digital visual media item by processing the initial segmentation mask and the high-level features through the neural network decoder.
16 . The system of claim 15 , wherein the neural network encoder comprises a transformer block; and further comprising:
down-sampling a portion of the low-level features to match a dimension of a lower-level set of the low-level features; and
generating a transformed set of features by passing the down-sampled portion of the low-level features through the transformer block.
17 . The system of claim 15 , wherein generating the initial segmentation mask of the digital visual media item from the low-level features comprises decoding the low-level features, without the high-level features.
18 . The system of claim 17 , further comprising:
up-sampling a portion of the low-level features to match a dimension of a higher-level set of the low-level features; and
combining the up-sampled portion of the low-level features with the higher-level set of the low-level features utilizing a series sum operation.
19 . The system of claim 15 , wherein generating the refined segmentation mask of the digital visual media item from the initial segmentation mask and the high-level features comprises decoding the initial segmentation mask and the high-level features by:
generating, utilizing a multilayer perceptron, refined high-level features from the initial segmentation mask and the high-level features;
up-sampling the refined high-level features; and
combining the up-sampled refined high-level features.
20 . The system of claim 19 , wherein combining the up-sampled refined high-level features comprises utilizing a concatenation operation.