Methods and apparatus to perform dense prediction using transformer blocks
Methods, apparatus, systems and articles of manufacture disclosed herein perform dense prediction of an input image using transformers at an encoder stage and at a reassembly stage of an image processing system. A disclosed apparatus includes an encoder with an embedder to convert an input image to a plurality of tokens representing features extracted from the input image. The tokens are embedded with a learnable position embedding. The encoder also includes one or more transformers configured in a sequence of stages to relate the tokens to each other. The apparatus further includes a decoder that includes one or more of reassemblers to assemble the tokens into feature representations, one or more of fusion blocks to combine the feature representations to generate a final feature representation, and an output head to generate a dense prediction based on the final feature representation and based on an output task.
1 . An apparatus for assigning labels to pixels of an image, the apparatus comprising:
an embedder to generate a plurality of embedded tokens from the image;
one or more transformers to process the plurality of embedded tokens through a plurality of transformer stages, wherein an output of a first transformer stage is used as an input of a second transformer stage; and
a fusion block, wherein the fusion block comprises:
a first convolution unit, the first convolution unit to generate a first feature map based on a feature representation corresponding to an output of the first transformer stage,
an adder, the adder to receive the first feature map from the first convolution unit, to receive a second feature map corresponding to an output of the second transformer stage, and to generate an output by adding the first feature map with the second feature map, and
a second convolution unit, the second convolution unit to process the output of the adder.
2 . The apparatus of claim 1 , wherein the embedder comprises a convolutional neural network to apply to the image.
3 . The apparatus of claim 1 , wherein the second feature map is generated by another fusion block from another feature representation corresponding to the output of the second transformer stage.
4 . The apparatus of claim 1 , wherein a transformer comprises a multi-head self-attention block.
5 . The apparatus of claim 4 , wherein the transformer further comprises a plurality of normalizers and adders.
6 . The apparatus of claim 1 , wherein the decoder comprises-a reassembler, is coupled with the fusion block, wherein the feature representation is generated by the reassembler from the output of the first transformer stage.
7 . The apparatus of claim 1 , wherein the labels identify at least one category associated with the pixels in the image.
8 . One or more non-transitory machine-readable media having instructions stored thereon, the instructions executable by a machine to implement:
an embedder to generate a plurality of embedded tokens from the image;
one or more transformers to process the plurality of embedded tokens through a plurality of transformer stages, wherein an output of a first transformer stage is used as an input of a second transformer stage; and
a fusion block, wherein the fusion block comprises:
a first convolution unit, the first convolution unit to generate a first feature map based on a feature representation corresponding to an output of the first transformer stage,
an adder, the adder to receive the first feature map from the first convolution unit, to receive a second feature map corresponding to an output of the second transformer stage, and to generate an output by adding the first feature map with the second feature map, and
a second convolution unit, the second convolution unit to process the output of the adder.
9 . The one or more non-transitory machine-readable media of claim 8 , wherein the embedder comprises a convolutional neural network to apply to the image.
10 . The one or more non-transitory machine-readable media of claim 8 , wherein the second feature map is generated by another fusion block from another feature representation corresponding to the output of the second transformer stage.
11 . The one or more non-transitory machine-readable media of claim 8 , wherein a transformer comprises a multi-head self-attention block.
12 . The one or more non-transitory machine-readable media of claim 11 , wherein the transformer further comprises a plurality of normalizers and adders.
13 . The one or more non-transitory machine-readable media of claim 8 , wherein a reassembler is coupled with the fusion block, wherein the feature representation is generated by the reassembler from the output of the first transformer stage.
14 . The one or more non-transitory machine-readable media of claim 8 , wherein the instructions are executable by the machine to further implement a head, the head to assign labels to pixels of the image based on an output of the decoder, the labels identifying at least one category associated with the pixels in the image.
15 . A computing device comprising:
a camera to capture an image;
a memory to store instructions; and
a processor coupled to the memory to execute the instructions to implement:
an embedder to generate a plurality of embedded tokens from the image,
one or more transformers to process the plurality of embedded tokens through a plurality of transformer stages, wherein an output of a first transformer stage is used as an input of a second transformer stage, and
a fusion block, wherein the fusion block comprises:
a first convolution unit, the first convolution unit to generate a first feature map based on a feature representation corresponding to an output of the first transformer stage,
an adder, the adder to receive the first feature map from the first convolution unit, to receive a second feature map corresponding to an output of the second transformer stage, and to generate an output by adding the first feature map with the second feature map, and
a second convolution unit, the second convolution unit to process the output of the adder.
16 . The computing device of claim 15 , wherein the embedder comprises a convolutional neural network to apply to the image.
17 . The computing device of claim 15 , wherein the second feature map is generated by another fusion block from another feature representation corresponding to the output of the second transformer stage.
18 . The computing device of claim 15 , wherein a transformer comprises a multi-head self-attention block.
19 . The computing device of claim 18 , wherein the transformer further comprises a plurality of normalizers and adders.
20 . The computing device of claim 15 , wherein a reassembler is coupled with the fusion block, wherein the feature representation is generated by the reassembler from the output of the first transformer stage.
21 . The computing device of claim 15 , wherein the processor is to execute the instructions to further implement a head, the head to assign labels to pixels of the image based on an output of the decoder, the labels identifying at least one category associated with the pixels in the image.