Text refinement network
Systems and methods for text segmentation are described. Embodiments of the inventive concept are configured to receive an image including a foreground text portion and a background portion, classify each pixel of the image as foreground text or background using a neural network that refines a segmentation prediction using a key vector representing features of the foreground text portion, wherein the key vector is based on the segmentation prediction, and identify the foreground text portion based on the classification.
1. A method for text segmentation, comprising:
receiving an image including a foreground text portion and a background portion;
encoding the image to produce a feature map;
generating a segmentation map of the image based on the feature map, wherein the segmentation map partially identifies the foreground text portion and the background portion;
generating a key vector representing features of the foreground text portion based on the segmentation map and the feature map;
combining the key vector and the feature map to produce an attention map; and
generating, using a neural network, a refined segmentation map by classifying each pixel of the image as foreground text or background based on the attention map and the feature map.
2. The method of claim 1 , further comprising:
decoding the feature map to produce a segmentation prediction; and
identifying the key vector based on the segmentation prediction.
3. The method of claim 2 , wherein decoding the feature map comprises:
applying a convolutional layer to the feature map;
applying a first bias to an output of the convolutional layer; and
applying a first softmax to an output of the first bias.
4. The method of claim 2 , wherein identifying the key vector comprises:
computing a cosine similarity of the segmentation prediction;
applying a second bias based on the cosine similarity;
applying a second softmax to an output of the second bias;
combining the second softmax with the feature map; and
applying a pooling layer to produce the key vector.
5. The method of claim 1 , further comprising:
combining the attention map and the feature map to produce a combined feature map; and
decoding the combined feature map to produce a refined segmentation prediction, wherein the foreground text portion is identified based on the refined segmentation prediction.
6. The method of claim 1 , further comprising:
modifying a texture of the foreground text portion to produce a modified image.
7. A method for text segmentation, comprising:
receiving an image including a foreground text portion and a background portion;
encoding the image to produce a feature map;
decoding the feature map to produce a segmentation prediction;
identifying a key vector based on the segmentation prediction, wherein the key vector represents features of the foreground text portion;
combining the key vector and the feature map to produce an attention map;
combining the attention map and the feature map to produce a combined feature map;
decoding the combined feature map to produce a refined segmentation prediction; and
identifying the foreground text portion based on the refined segmentation prediction.
8. The method of claim 7 , further comprising:
classifying each pixel of the image as foreground text or background using a neural network that refines the segmentation prediction using the key vector.
9. The method of claim 7 , wherein decoding the feature map comprises:
applying a convolutional layer to the feature map;
applying a first bias to an output of the convolutional layer; and
applying a first softmax to an output of the first bias.
10. The method of claim 7 , wherein identifying the key vector comprises:
computing a cosine similarity of the segmentation prediction;
applying a second bias based on the cosine similarity;
applying a second softmax to an output of the second bias;
combining the second softmax with the feature map; and
applying a pooling layer to produce the key vector.
11. The method of claim 7 , further comprising:
re-thresholding the segmentation prediction to obtain a modified segmentation prediction, wherein the key vector is identified based on the modified segmentation prediction.
12. The method of claim 11 , further comprising:
computing a weighted sum based on the feature map and the modified segmentation prediction, wherein the key vector is identified based on the weighted sum.
13. The method of claim 11 , wherein:
the modified segmentation prediction comprises a foreground prediction of the foreground text portion.
14. An apparatus for text segmentation, comprising:
one or more processors; and
one or more memories including instructions executable by the one or more processors to:
receive an image including a foreground text portion and a background portion;
encode the image to produce a feature map;
generate a segmentation map of the image based on the feature map, wherein the segmentation map partially identifies the foreground text portion and the background portion;
generate a key vector representing features of the foreground text portion based on the segmentation map and the feature map;
combine the key vector and the feature map to produce an attention map; and
generate a refined segmentation map by classifying each pixel of the image as foreground text or background based on the attention map and the feature map.
15. The apparatus of claim 14 , the instructions being further executable to:
decode the feature map to produce a segmentation prediction; and
identify the key vector based on the segmentation prediction.
16. The apparatus of claim 15 , wherein decoding the feature map comprises:
applying a convolutional layer to the feature map;
applying a first bias to an output of the convolutional layer; and
applying a first softmax to an output of the first bias.
17. The apparatus of claim 15 , wherein identifying the key vector comprises:
computing a cosine similarity of the segmentation prediction;
applying a second bias based on the cosine similarity;
applying a second softmax to an output of the second bias;
combining the second softmax with the feature map; and
applying a pooling layer to produce the key vector.
18. The apparatus of claim 14 , the instructions being further executable to:
combine the attention map and the feature map to produce a combined feature map; and
decode the combined feature map to produce a refined segmentation prediction, wherein the foreground text portion is identified based on the refined segmentation prediction.
19. The apparatus of claim 14 , the instructions being further executable to:
modify a texture of the foreground text portion to produce a modified image.