IP Library Granted Patent US 12,639,848
Granted Patent B2
US 12,639,848 · App. 18/474,347 · Granted May 26, 2026

Pixelwise positional embeddings for medical images in vision transformers

Inventors: Gengyan Zhao (Plainsboro, NJ); Badhan Kumar Das (Erlangen, DE); Eli Gibson (Plainsboro, NJ); Dorin Comaniciu (Princeton, NJ)
Assignee: Siemens Healthineers AG
G06T7/74G16H30/40G06T2207/20081G06T2207/20084G06T2207/30016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,848
App. No.
18/474,347
Granted
May 26, 2026
Kind
B2
Abstract

Systems and methods for performing a medical imaging analysis task based on pixelwise positionally encoded features are provided. One or more input medical images are received. One or more pixelwise positional embedding images are generated for the one or more input medical images using a spatially varying function. Patches are extracted from the one or more input medical images and the one or more pixelwise positional embedding images. The patches extracted from the one or more input medical images are encoded with corresponding ones of the patches extracted from the one or more pixelwise positional embedding images into pixelwise positionally encoded features. A medical imaging analysis task is performed using a machine learning based network based on the pixelwise positionally encoded features. Results of the medical imaging analysis task are output.

Claims (51)

1 . A computer-implemented method comprising:

receiving one or more input medical images;

generating one or more pixelwise positional embedding images for the one or more input medical images using a spatially varying function;

extracting patches from the one or more input medical images and the one or more pixelwise positional embedding images;

encoding the patches extracted from the one or more input medical images with corresponding ones of the patches extracted from the one or more pixelwise positional embedding images into pixelwise positionally encoded features by:

combining each patch extracted from the one or more input medical images with its corresponding patch extracted from the one or more pixelwise positional embedding images, and

separately encoding the combined patches;

performing a medical imaging analysis task using a machine learning based network based on the pixelwise positionally encoded features; and

outputting results of the medical imaging analysis task.

2 . The computer-implemented method of claim 1 , wherein generating one or more pixelwise positional embedding images for the one or more input medical images using a spatially varying function comprises:

sampling a spatially varying function at a location of each pixel of the one or more input medical images.

3 . The computer-implemented method of claim 1 , wherein the spatially varying function is a sinusoidal function.

4 . The computer-implemented method of claim 1 , wherein the spatially varying function is in reference coordinate system defined relative to the one or more input medical images.

5 . The computer-implemented method of claim 4 , wherein the reference coordinate system comprises a physical coordinate system of an image acquisition device that acquired the one or more input medical images.

6 . The computer-implemented method of claim 1 , wherein the one or more input medical images comprises a plurality of input medical images and the one or more pixelwise positional embedding images comprises a plurality of pixelwise positional embedding images.

7 . The computer-implemented method of claim 1 , wherein the one or more input medical images comprises a plurality of input medical images and the one or more pixelwise positional embedding images comprises a single pixelwise positional embedding image.

8 . The computer-implemented method of claim 1 , wherein:

encoding the patches extracted from the one or more input medical images with corresponding ones of the patches extracted from the one or more pixelwise positional embedding images into pixelwise positionally encoded features comprises encoding the patches extracted from the one or more input medical images with patch-wise positionally embedded features to generate patch-wise and pixelwise positionally encoded features; and

performing a medical imaging analysis task using a machine learning based network based on the pixelwise positionally encoded features comprises performing the medical imaging analysis task based on the patch-wise and pixelwise positionally encoded features.

9 . The computer-implemented method of claim 1 , wherein the machine learning based network is a vision transformer network.

10 . An apparatus comprising:

means for receiving one or more input medical images;

means for generating one or more pixelwise positional embedding images for the one or more input medical images using a spatially varying function;

means for extracting patches from the one or more input medical images and the one or more pixelwise positional embedding images;

means for encoding the patches extracted from the one or more input medical images with corresponding ones of the patches extracted from the one or more pixelwise positional embedding images into pixelwise positionally encoded features by:

combining each patch extracted from the one or more input medical images with its corresponding patch extracted from the one or more pixelwise positional embedding images, and

separately encoding the combined patches;

means for performing a medical imaging analysis task using a machine learning based network based on the pixelwise positionally encoded features; and

means for outputting results of the medical imaging analysis task.

11 . The apparatus of claim 10 , wherein the means for generating one or more pixelwise positional embedding images for the one or more input medical images using a spatially varying function comprises:

means for sampling a spatially varying function at a location of each pixel of the one or more input medical images.

12 . The apparatus of claim 10 , wherein the spatially varying function is a sinusoidal function.

13 . The apparatus of claim 10 , wherein the spatially varying function is in reference coordinate system defined relative to the one or more input medical images.

14 . The apparatus of claim 13 , wherein the reference coordinate system comprises a physical coordinate system of an image acquisition device that acquired the one or more input medical images.

15 . A non-transitory computer readable medium storing computer program instructions, the computer program instructions when executed by a processor cause the processor to perform operations comprising:

receiving one or more input medical images;

generating one or more pixelwise positional embedding images for the one or more input medical images using a spatially varying function;

extracting patches from the one or more input medical images and the one or more pixelwise positional embedding images;

encoding the patches extracted from the one or more input medical images with corresponding ones of the patches extracted from the one or more pixelwise positional embedding images into pixelwise positionally encoded features by:

combining each patch extracted from the one or more input medical images with its corresponding patch extracted from the one or more pixelwise positional embedding images, and

separately encoding the combined patches;

performing a medical imaging analysis task using a machine learning based network based on the pixelwise positionally encoded features; and

outputting results of the medical imaging analysis task.

16 . The non-transitory computer readable medium of claim 15 , wherein generating one or more pixelwise positional embedding images for the one or more input medical images using a spatially varying function comprises:

sampling a spatially varying function at a location of each pixel of the one or more input medical images.

17 . The non-transitory computer readable medium of claim 15 , wherein the one or more input medical images comprises a plurality of input medical images and the one or more pixelwise positional embedding images comprises a plurality of pixelwise positional embedding images.

18 . The non-transitory computer readable medium of claim 15 , wherein the one or more input medical images comprises a plurality of input medical images and the one or more pixelwise positional embedding images comprises a single pixelwise positional embedding image.

19 . The non-transitory computer readable medium of claim 15 , wherein:

encoding the patches extracted from the one or more input medical images with corresponding ones of the patches extracted from the one or more pixelwise positional embedding images into pixelwise positionally encoded features comprises encoding the patches extracted from the one or more input medical images with patch-wise positionally embedded features to generate patch-wise and pixelwise positionally encoded features; and

performing a medical imaging analysis task using a machine learning based network based on the pixelwise positionally encoded features comprises performing the medical imaging analysis task based on the patch-wise and pixelwise positionally encoded features.

20 . The non-transitory computer readable medium of claim 15 , wherein the machine learning based network is a vision transformer network.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: SIEMENS HEALTHCARE GMBH
To: SIEMENS HEALTHINEERS AG
Reel/Frame 066267/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2023
From: SIEMENS MEDICAL SOLUTIONS USA, INC.
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 065331/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2023
From: ZHAO, GENGYAN; GIBSON, ELI; COMANICIU, DORIN
To: SIEMENS MEDICAL SOLUTIONS USA, INC.
Reel/Frame 065292/0730 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: DAS, BADHAN KUMAR
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 065142/0616 →
Continuity (1)
Related Publication 20250104276A1 · Mar 27, 2025
References Cited (29)
US 20170164836A1 · Krishnaswamy · 2017 [cited by examiner]
US 20180232883A1 · Sethi · 2018 [cited by examiner]
US 20240203098A1 · Sultana · 2024 [cited by examiner]
KR 102376249B1 · 2022 [cited by examiner]
WO WO2010040396A1 · 2010 [cited by examiner]
Extended European Search Report (EESR) mailed Feb. 14, 2025 in corresponding European Patent Application No. 24202841.3. [cited by applicant]
Xiao Hanguang et al: “Transformers in medical image segmentation: A review”, Biomedical Signal Processing and Control, Elsevier, Amsterdam, NL, vol. 84, Mar. 7, 2023. [cited by applicant]
Gandhi Dhruvin et al: “A Vision Transformer Approach for Classification an A Small-Sized Medical Image Dataset”, 2022 5th International Conference on Advances in Science and Technology (ICAST), IEEE, Dec. 2, 2022 (Dec. … [cited by applicant]
Gai Lulu et al: “Using Vision Transformers in 3-D Medical Image Classifications”, 2022 IEEE International Conference on Image Processing (ICIP), IEEE, Oct. 16, 2022 (Oct. 16, 2022), pp. 696-700. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, arXiv:1706.03762, 2017, pp. 1-15. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale”, arXiv:2010.11929, 2021, pp. 1-22. [cited by applicant]
Bello et al., “Attention Augmented Convolutional Networks”, arXiv:1904.09925v5, 2020, pp. 1-13. [cited by applicant]
Yang et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, arXiv:1906.08237v2, 2020, pp. 1-18. [cited by applicant]
He et al., “DeBERTa: Decoding-enhanced BERT with Disentangled Attention”, arXiv:2006.03654v6, 2021, pp. 1-23. [cited by applicant]
Hatamizadeh et al., “UNETR: Transformers for 3D Medical Image Segmentation”, arXiv:2103.10504v1, 2021, pp. 1-11. [cited by applicant]
Hatamizadeh et al., “Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images”, arXiv:2201.01266v1, 2022, pp. 1-13. [cited by applicant]
Chen et al., “Transformers Improve Breast Cancer Diagnosis from Unregistered Multi-View Mammograms”, Diagnostics, 2022, pp. 1-14. [cited by applicant]
Mkindu et al., “Lung nodule detection in chest CT images based on vision transformer network with Bayesian optimization”, Biomedical Signal Processing and Control, 2023, pp. 1-8. [cited by applicant]
Chen et al., “TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation”, arXiv:2102.04306v1, 2021, pp. 1-13. [cited by applicant]
Valanarasu et al., “Medical Transformer: Gated Axial-Attention for Medical Image Segmentation”, arXiv:2102.10662v2, 2021, pp. 1-18. [cited by applicant]
Xie et al., “Universal Medical Self-Supervised Learning via Breaking Dimensionality Barrier”, arXiv:2112.09356v2, 2022, pp. 1-23. [cited by applicant]
Ma et al., “Transformer Network for Significant Stenosis Detection in CCTA of Coronary Arteries”, arXiv:2107.03035v3, 2021, pp. 1-10. [cited by applicant]
Simpson et al., “A large annotated medical image dataset for the development and evaluation of segmentation algorithms”, arXiv:1902.09063v1, 2019, pp. 1-15. [cited by applicant]
Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows”, arXiv:2103.14030v2, 2021, pp. 1-14. [cited by applicant]
Wilcoxon, “Individual Comparisons by Ranking Methods”, Breakthroughs in Statistics, Biometrics Bulletin, 1945, pp. 80-83. [cited by applicant]
Isensee et al., “Automated Design of Deep Learning Methods for Biomedical Image Segmentation”, arXiv:1904.08128v2, 2020, pp. 1-55. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, arXiv:1505.04597v1, 2015, pp. 1-8. [cited by applicant]
Futrega et al., “Optimized U-Net for Brain Tumor Segmentation”, arXiv:2110.03352v2, 2021, pp. 1-15. [cited by applicant]
Isensee et al., “nnU-Net: Self-adapting Framework for U-Net-Based Medical Image Segmentation”, arXiv:1809.10486v1, 2018, pp. 1-11. [cited by applicant]