IP Library › Granted Patent US 12,646,343
Granted Patent B2
US 12,646,343 · App. 17/855,763 · Granted Jun 2, 2026

Methods and apparatus to perform dense prediction using transformer blocks

Inventors: Rene Ranftl (Munich, DE); Alexey Bochkovskiy (Podolsk, RU); Vladlen Koltun (Santa Clara, CA)
Assignee: Intel Corporation
G06V20/70G06N3/04G06T3/4046G06T5/00G06T2207/20016G06T2207/20021G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,343
App. No.
17/855,763
Granted
Jun 2, 2026
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture disclosed herein perform dense prediction of an input image using transformers at an encoder stage and at a reassembly stage of an image processing system. A disclosed apparatus includes an encoder with an embedder to convert an input image to a plurality of tokens representing features extracted from the input image. The tokens are embedded with a learnable position embedding. The encoder also includes one or more transformers configured in a sequence of stages to relate the tokens to each other. The apparatus further includes a decoder that includes one or more of reassemblers to assemble the tokens into feature representations, one or more of fusion blocks to combine the feature representations to generate a final feature representation, and an output head to generate a dense prediction based on the final feature representation and based on an output task.

Claims (42)

1 . An apparatus for assigning labels to pixels of an image, the apparatus comprising:

an embedder to generate a plurality of embedded tokens from the image;

one or more transformers to process the plurality of embedded tokens through a plurality of transformer stages, wherein an output of a first transformer stage is used as an input of a second transformer stage; and

a fusion block, wherein the fusion block comprises:

a first convolution unit, the first convolution unit to generate a first feature map based on a feature representation corresponding to an output of the first transformer stage,

an adder, the adder to receive the first feature map from the first convolution unit, to receive a second feature map corresponding to an output of the second transformer stage, and to generate an output by adding the first feature map with the second feature map, and

a second convolution unit, the second convolution unit to process the output of the adder.

2 . The apparatus of claim 1 , wherein the embedder comprises a convolutional neural network to apply to the image.

3 . The apparatus of claim 1 , wherein the second feature map is generated by another fusion block from another feature representation corresponding to the output of the second transformer stage.

4 . The apparatus of claim 1 , wherein a transformer comprises a multi-head self-attention block.

5 . The apparatus of claim 4 , wherein the transformer further comprises a plurality of normalizers and adders.

6 . The apparatus of claim 1 , wherein the decoder comprises-a reassembler, is coupled with the fusion block, wherein the feature representation is generated by the reassembler from the output of the first transformer stage.

7 . The apparatus of claim 1 , wherein the labels identify at least one category associated with the pixels in the image.

8 . One or more non-transitory machine-readable media having instructions stored thereon, the instructions executable by a machine to implement:

an embedder to generate a plurality of embedded tokens from the image;

one or more transformers to process the plurality of embedded tokens through a plurality of transformer stages, wherein an output of a first transformer stage is used as an input of a second transformer stage; and

a fusion block, wherein the fusion block comprises:

a first convolution unit, the first convolution unit to generate a first feature map based on a feature representation corresponding to an output of the first transformer stage,

an adder, the adder to receive the first feature map from the first convolution unit, to receive a second feature map corresponding to an output of the second transformer stage, and to generate an output by adding the first feature map with the second feature map, and

a second convolution unit, the second convolution unit to process the output of the adder.

9 . The one or more non-transitory machine-readable media of claim 8 , wherein the embedder comprises a convolutional neural network to apply to the image.

10 . The one or more non-transitory machine-readable media of claim 8 , wherein the second feature map is generated by another fusion block from another feature representation corresponding to the output of the second transformer stage.

11 . The one or more non-transitory machine-readable media of claim 8 , wherein a transformer comprises a multi-head self-attention block.

12 . The one or more non-transitory machine-readable media of claim 11 , wherein the transformer further comprises a plurality of normalizers and adders.

13 . The one or more non-transitory machine-readable media of claim 8 , wherein a reassembler is coupled with the fusion block, wherein the feature representation is generated by the reassembler from the output of the first transformer stage.

14 . The one or more non-transitory machine-readable media of claim 8 , wherein the instructions are executable by the machine to further implement a head, the head to assign labels to pixels of the image based on an output of the decoder, the labels identifying at least one category associated with the pixels in the image.

15 . A computing device comprising:

a camera to capture an image;

a memory to store instructions; and

a processor coupled to the memory to execute the instructions to implement:

an embedder to generate a plurality of embedded tokens from the image,

one or more transformers to process the plurality of embedded tokens through a plurality of transformer stages, wherein an output of a first transformer stage is used as an input of a second transformer stage, and

a fusion block, wherein the fusion block comprises:

a first convolution unit, the first convolution unit to generate a first feature map based on a feature representation corresponding to an output of the first transformer stage,

an adder, the adder to receive the first feature map from the first convolution unit, to receive a second feature map corresponding to an output of the second transformer stage, and to generate an output by adding the first feature map with the second feature map, and

a second convolution unit, the second convolution unit to process the output of the adder.

16 . The computing device of claim 15 , wherein the embedder comprises a convolutional neural network to apply to the image.

17 . The computing device of claim 15 , wherein the second feature map is generated by another fusion block from another feature representation corresponding to the output of the second transformer stage.

18 . The computing device of claim 15 , wherein a transformer comprises a multi-head self-attention block.

19 . The computing device of claim 18 , wherein the transformer further comprises a plurality of normalizers and adders.

20 . The computing device of claim 15 , wherein a reassembler is coupled with the fusion block, wherein the feature representation is generated by the reassembler from the output of the first transformer stage.

21 . The computing device of claim 15 , wherein the processor is to execute the instructions to further implement a head, the head to assign labels to pixels of the image based on an output of the decoder, the labels identifying at least one category associated with the pixels in the image.

Continuity (2)
Continuation 17485349 · Sep 25, 2021
Related Publication 20230113271A1 · Apr 13, 2023
References Cited (19)
US 20220012848A1 · Ranftl et al. · 2022 [cited by applicant]
US 20220391635A1 · Lian · 2022 [cited by examiner]
Carion et al. “End-to end object detection with transformers.” In Proc. Eur. Conf. Comp. Vis., 2020. (Year: 2020). [cited by examiner]
Wang et al., “Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions, ” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 2021, pp. 548-558, doi:… [cited by examiner]
Wu et al., “Fully Transformer Networks for Semantic Image Segmentation,” https://arxiv.org/abs/2106.04108. (Year: 2021). [cited by examiner]
Xu, Weijian et al. “Co-Scale Conv-Attentional Image Transformers.” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Apr. 13, 2021 : 9961-9970. (Year: 2021). [cited by examiner]
Xie, Enze, et al. “SegFormer: Simple and efficient design for semantic segmentation with transformers.” Advances in neural information processing systems 34 (2021): 12077-12090. (Year: 2021). [cited by examiner]
Lin, Tsung-Yi, et al. “Feature pyramid networks for object detection.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. (Year: 2017). [cited by examiner]
Chen, Chun-Fu Richard, Quanfu Fan, and Rameswar Panda. “Crossvit: Cross-attention multi-scale vision transformer for image classification.” Proceedings of the IEEE/CVF international conference on computer vision. 2021. … [cited by examiner]
Strudel, Robin, et al. “Segmenter: Transformer for semantic segmentation.” Proceedings of the IEEE/CVF international conference on computer vision. 2021. (Year: 2021). [cited by examiner]
Rakhimov et al., “Latent Video Transformer,” arXiv:2006.10704v1 [cs.CV], Jun. 18, 2020, retrieved from https://arxiv.org/pdf/2006.10704.pdf, 18 pages. [cited by applicant]
Touvron et al. “Training data-efficient image transformers & distillation through attention,” International Conference on Machine Learning, PMLR, arXiv:2012.12877v2 [cs.CV], Jan. 15, 2021, 22 pages. [cited by applicant]
isl-org/DPT, “Dense Prediction Transformers,” GitHub, Mar. 2021, retrieved from https://github.com/intel-isl/DPT, 4 pages. [cited by applicant]
Ranftl et al., “Vision Transformers for Dense Prediction,” arXiv:2103.13413v1 [cs.CV], Mar. 24, 2021, retrieved from https://arxiv.org/abs/2103.13413, 15 pages. [cited by applicant]
Liu et al., “Swin transformer: Hierarchical vision transformer using shifted windows,” Proceedings of the IEEE/CVF International Conference on Computer Vision, arXiv:2103.14030v2 [cs.CV], Aug. 17, 2021, 14 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/485,349, dated Jan. 12, 2024, 16 pages. [cited by applicant]
Vaswani, A. et al. “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), 11 pages. provided in parent U.S. Appl. No. 17/485,349. [cited by applicant]
Lin, G. et al. “RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation”, CVPR, 2017, pp. 1925-1934 (10 pages). provided in parent U.S. Appl. No. 17/485,349. [cited by applicant]
Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” In International Conference on Learning Representations, 2021. (Year: 2021). [cited by applicant]