IP Library › Granted Patent US 12,394,186
Granted Patent B2
US 12,394,186 · App. 18/126,318 · Granted Aug 19, 2025

Systems, methods, and apparatuses for implementing self-supervised domain-adaptive pre-training via a transformer for use with medical image classification

Inventors: DongAo Ma (Tempe, AZ); Jiaxuan Pang (Tempe, AZ); Nahid Ul Islam (Mesa, AZ); Mohammad Reza Hosseinzadeh Taher (Tempe, AZ); Fatemeh Haghighi (Tempe, AZ); Jianming Liang (Scottsdale, AZ)
Assignee: Arizona Board of Regents on behalf of Arizona State University
G06V10/774G06N3/0895G06T7/0012G06V10/764G06V10/776G06T2207/10116G06T2207/20081G06T2207/20092G06T2207/30004G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,186
App. No.
18/126,318
Granted
Aug 19, 2025
Kind
B2
Abstract

Described herein are systems, methods, and apparatuses for implementing self-supervised domain-adaptive pre-training via a transformer for use with medical image classification in the context of medical image analysis. An exemplary system includes means for receiving a first set of training data having non-medical photographic images; receiving a second set of training data with medical images; pre-training an AI model on the first set of training data with the non-medical photographic images; performing domain-adaptive pre-training of the AI model via self-supervised learning operations using the second set of training data having the medical images; generating a trained domain-adapted AI model by fine-tuning the AI model against the targeted medical diagnosis task using the second set of training data having the medical images; outputting the trained domain-adapted AI model; and executing the trained domain-adapted AI model to generate a predicted medical diagnosis from an input image not present within the training data.

Claims (80)

1. A system comprising:

a memory to store instructions;

a processor to execute the instructions stored in the memory;

wherein the system is specially configured to execute instructions via the processor for performing the following operations:

receiving at the system, a first set of training data which includes photographic images unrelated to a targeted medical diagnosis task;

receiving at the system, a second set of training data which includes a plurality of medical images derived from multiple distinct sources, wherein the plurality of medical images are configured with multiple inconsistent annotation and classification data;

executing instructions via the processor of the system for pre-training an AI model on the first set of training data which includes the photographic images by learning image classification from the photographic images within the first set of training data;

executing instructions via the processor of the system for performing domain-adaptive pre-training of the AI model via self-supervised learning operations, wherein the AI model previously trained on the first set of training data applies domain-adaptive learning to scale up data utilization of the second set of training data within the AI model's learned image classifications;

generating a trained domain-adapted AI model by fine-tuning the AI model against the targeted medical diagnosis task using the second set of training data which includes the plurality of medical images;

outputting the trained domain-adapted AI model; and

executing the trained domain-adapted AI model to generate a predicted medical diagnosis from an input image which forms no part of the first or second sets of training data.

2. The system of claim 1 :

wherein pre-training the AI model on the first set of training data generates a trained AI model; and

wherein performing domain-adaptive pre-training of the AI model via self-supervised learning operations generates a trained domain-adapted AI model by fine-tuning the trained AI model against the targeted medical diagnosis task using the second set of training data which includes the plurality of medical images.

3. The system of claim 1 , further comprising:

performing in-domain medical transfer learning for via the trained domain-adapted AI model by prioritizing all in-domain medical transfer learnings derived from the second set of training data which includes the plurality of medical images over any out-of-domain learnings derived from the pre-training of the AI model on the first set of training data generates a trained AI model.

4. The system of claim 3 , wherein prioritizing all in-domain medical transfer learnings is configured to improve performance of the trained domain-adapted AI model by reducing domain disparities between a source domain corresponding to the first set of training data and a target domain corresponding to the second set of training data which includes the plurality of medical images.

5. The system of claim 1 , wherein receiving the first set of training data which includes photographic images unrelated to a targeted medical diagnosis task comprises receiving photographic images corresponding to a first domain type which lacks medical imaging data.

6. The system of claim 1 , wherein receiving the second set of training data which includes a plurality of medical images comprises receiving medical imaging data corresponding to a second domain type having at least a sub-set of the plurality of medical images correlated to the targeted medical diagnosis task.

7. The system of claim 1 , wherein receiving the first set of training data which includes photographic images unrelated to a targeted medical diagnosis task comprises receiving a set of non-domain specific photographic images lacking any images correlated to the targeted medical diagnosis task.

8. The system of claim 1 , further comprising:

receiving, at the system, multiple medical imaging training data sets from multiple distinct sources;

aggregating the multiple medical imaging training data sets into a single aggregated medical imaging dataset; and

wherein, receiving at the system, the second set of training data which includes the plurality of medical images derived from multiple distinct sources comprises specifying the single aggregated medical imaging dataset as the second set of training data for performing the domain-adaptive pre-training of the AI model.

9. The system of claim 1 , wherein receiving the second set of training data which includes the plurality of medical images derived from multiple distinct sources includes training data incompatible with supervised learning for the AI model.

10. The system of claim 1 , wherein the plurality of medical images received with the second set of training data are configured with multiple inconsistent annotation and classification data including at least two or more of:

inconsistent annotations across a sub-set of the plurality of medical images for the targeted medical diagnosis task;

inconsistent annotations across a sub-set of the plurality of medical images for a common disease condition represented within the plurality of medical images;

inconsistent annotations across a sub-set of the plurality of medical images for a common human anatomical feature classified within the plurality of medical images;

inconsistent global level image annotations identifying disease conditions within an image and local level boxed-lesion labels identifying disease conditions within bounding boxes present within the plurality of medical images;

inconsistent use of expert annotations for the plurality of medical images with at least a first portion of the plurality of medical images including expert annotations and at least a second portion of the plurality of medical images lacking any expert annotations; and

inconsistent use of radiological reports associated with the plurality of medical images with at least a first portion of the plurality of medical images having radiological reports associated with them and at least a second portion of the plurality of medical images lacking any associated radiological reports.

11. The system of claim 1 , wherein performing the domain-adaptive pre-training of the AI model via the self-supervised learning operations,

bridges a domain gap between photographic images from the first set of training data representing a first domain and medical images from the second set of training data representing a second domain;

wherein the self-supervised learning operations operate without requiring expert annotation of the medical images from the second set of training data which exhibit inconsistent or missing labeling and inconsistent or missing expert annotations.

12. The system of claim 1 , wherein performing the domain-adaptive pre-training of the AI model via self-supervised learning operations, comprises:

applying continual pre-training on a large-scale domain-specific dataset represented within the second set of training data having the plurality of medical images derived from the multiple distinct sources;

wherein the continual pre-training is performed via self-supervised Masked Image Modeling (MIM) learning by the AI model; and

wherein the self-supervised Masked Image Modeling (MIM) learning by the AI model mitigates over-fitting in the targeted medical diagnosis task.

13. A computer-implemented method performed by a system having at least a processor and a memory therein for executing instructions, wherein the computer-implemented method comprises:

receiving at the system, a first set of training data which includes photographic images unrelated to a targeted medical diagnosis task;

receiving at the system, a second set of training data which includes a plurality of medical images derived from multiple distinct sources, wherein the plurality of medical images are configured with multiple inconsistent annotation and classification data;

executing instructions via the processor of the system for pre-training an AI model on the first set of training data which includes the photographic images by learning image classification from the photographic images within the first set of training data;

executing instructions via the processor of the system for performing domain-adaptive pre-training of the AI model via self-supervised learning operations, wherein the AI model previously trained on the first set of training data applies domain-adaptive learning to scale up data utilization of the second set of training data within the AI model's learned image classifications;

generating a trained domain-adapted AI model by fine-tuning the AI model against the targeted medical diagnosis task using the second set of training data which includes the plurality of medical images;

outputting the trained domain-adapted AI model; and

executing the trained domain-adapted AI model to generate a predicted medical diagnosis from an input image which forms no part of the first or second sets of training data.

14. The computer-implemented method of claim 13 :

wherein pre-training the AI model on the first set of training data generates a trained AI model; and

wherein performing domain-adaptive pre-training of the AI model via self-supervised learning operations generates a trained domain-adapted AI model by fine-tuning the trained AI model against the targeted medical diagnosis task using the second set of training data which includes the plurality of medical images.

15. The computer-implemented method of claim 13 , further comprising:

performing in-domain medical transfer learning for via the trained domain-adapted AI model by prioritizing all in-domain medical transfer learnings derived from the second set of training data which includes the plurality of medical images over any out-of-domain learnings derived from the pre-training of the AI model on the first set of training data generates a trained AI model; and

wherein prioritizing all in-domain medical transfer learnings is configured to improve performance of the trained domain-adapted AI model by reducing domain disparities between a source domain corresponding to the first set of training data and a target domain corresponding to the second set of training data which includes the plurality of medical images.

16. The computer-implemented method of claim 13 , wherein receiving the first set of training data which includes photographic images unrelated to a targeted medical diagnosis task comprises receiving photographic images corresponding to a first domain type which lacks medical imaging data; and

wherein receiving the second set of training data which includes a plurality of medical images comprises receiving medical imaging data corresponding to a second domain type having at least a sub-set of the plurality of medical images correlated to the targeted medical diagnosis task.

17. The computer-implemented method of claim 13 , wherein receiving the first set of training data which includes photographic images unrelated to a targeted medical diagnosis task comprises receiving a set of non-domain specific photographic images lacking any images correlated to the targeted medical diagnosis task.

18. The computer-implemented method of claim 13 , further comprising:

receiving, at the system, multiple medical imaging training data sets from multiple distinct sources;

aggregating the multiple medical imaging training data sets into a single aggregated medical imaging dataset; and

wherein, receiving at the system, the second set of training data which includes the plurality of medical images derived from multiple distinct sources comprises specifying the single aggregated medical imaging dataset as the second set of training data for performing the domain-adaptive pre-training of the AI model.

19. The computer-implemented method of claim 13 , wherein the plurality of medical images received with the second set of training data are configured with multiple inconsistent annotation and classification data including at least two or more of:

inconsistent annotations across a sub-set of the plurality of medical images for the targeted medical diagnosis task;

inconsistent annotations across a sub-set of the plurality of medical images for a common disease condition represented within the plurality of medical images;

inconsistent annotations across a sub-set of the plurality of medical images for a common human anatomical feature classified within the plurality of medical images;

inconsistent global level image annotations identifying disease conditions within an image and local level boxed-lesion labels identifying disease conditions within bounding boxes present within the plurality of medical images;

inconsistent use of expert annotations for the plurality of medical images with at least a first portion of the plurality of medical images including expert annotations and at least a second portion of the plurality of medical images lacking any expert annotations; and

inconsistent use of radiological reports associated with the plurality of medical images with at least a first portion of the plurality of medical images having radiological reports associated with them and at least a second portion of the plurality of medical images lacking any associated radiological reports.

20. A system comprising:

a memory to store instructions;

a processor to execute the instructions stored in the memory;

wherein the system is specially configured to execute instructions for systematically benchmarking vision transformers for use with chest x-ray classification, by performing operations including:

receiving first user input specifying multiple vision transformers;

receiving second user input specifying multiple training image datasets;

generating a list of all possible combinations of the multiple vision transformers specified and the multiple training image datasets specified according to the first and second user inputs;

retrieving a pre-trained base model for each of the multiple vision transformers specified and storing the pre-trained base models retrieved in the memory of the system for local execution;

retrieving the multiple training image datasets specified and storing the multiple training image datasets locally at the system;

initializing each of the pre-trained base models stored in memory using a standardized ImageNet dataset with randomized initialization weights to generate initialized vision transformer models corresponding to each of the pre-trained base models stored in memory;

executing Self-Supervised Learning (SSL) against each of the previously initialized vision transformer models using each of the multiple training image datasets corresponding to the list of all possible combinations previously generated to produce as output multiple SSL trained vision transformer models corresponding to the list of all possible combinations previously generated;

executing, via the processor of the system, each of the multiple SSL trained vision transformer models to generate image classification results as output from each of the multiple SSL trained vision transformer models; and

outputting from the system, a ranking of the image classification results generated as the output from each of the multiple SSL trained vision transformer models according to an area under the curve percentage calculation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2024
From: MA, DONGAO; PANG, JIAXUAN; ISLAM, NAHID UI; HOSSEINZADEH TAHER, MOHAMMAD REZA; HAGHIGHI, FATEMEH; LIANG, JIANMING
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 068668/0484 →
Continuity (3)
Provisional Application 63405262 · Sep 9, 2022
Provisional Application 63323988 · Mar 25, 2022
Related Publication 20230306723A1 · Sep 28, 2023
References Cited (85)
US 5577188A · Zhu · 1996 [cited by applicant]
US 10650286B2 · Ben-Ari · 2020 [cited by examiner]
US 11100647B2 · Nikolov · 2021 [cited by examiner]
US 11334992B2 · Schmidt-Richberg · 2022 [cited by examiner]
US 12051428B1 · Petrochuk · 2024 [cited by examiner]
US 20160093048A1 · Cheng · 2016 [cited by examiner]
US 20160224892A1 · Sawada · 2016 [cited by examiner]
US 20180225823A1 · Zhou · 2018 [cited by examiner]
US 20180240219A1 · Mentl · 2018 [cited by examiner]
US 20190073569A1 · Ben-Ari · 2019 [cited by examiner]
US 20190205766A1 · Krebs · 2019 [cited by examiner]
US 20200082534A1 · Nikolov · 2020 [cited by examiner]
US 20200130177A1 · Kolouri · 2020 [cited by examiner]
US 20210004646A1 · Guizilini · 2021 [cited by examiner]
US 20210056693A1 · Cheng · 2021 [cited by examiner]
US 20210118184A1 · Pillai · 2021 [cited by examiner]
US 20210150710A1 · Hosseinzadeh Taher et al. · 2021 [cited by applicant]
US 20210158580A1 · Barroyer · 2021 [cited by examiner]
US 20210248427A1 · Guo · 2021 [cited by examiner]
US 20220012891A1 · Nikolov · 2022 [cited by examiner]
US 20220114389A1 · Ghose · 2022 [cited by examiner]
US 20220148191A1 · Liu · 2022 [cited by examiner]
US 20220164610A1 · Otto · 2022 [cited by examiner]
US 20220198339A1 · Zhao · 2022 [cited by examiner]
US 20220254150A1 · Wolfe · 2022 [cited by examiner]
US 20220262105A1 · Zhou · 2022 [cited by examiner]
US 20230129779A1 · Sommer · 2023 [cited by examiner]
US 20230146292A1 · Zhou · 2023 [cited by examiner]
US 20230154164A1 · Ghesu · 2023 [cited by examiner]
US 20230237333A1 · Shao · 2023 [cited by examiner]
US 20230274420A1 · Seah · 2023 [cited by examiner]
US 20230281828A1 · Du · 2023 [cited by examiner]
US 20240029868A1 · Gulsun · 2024 [cited by examiner]
US 20240046453A1 · Jacob · 2024 [cited by examiner]
US 20240102979A1 · Tai · 2024 [cited by examiner]
US 20240203098A1 · Sultana · 2024 [cited by examiner]
US 20240211825A1 · Wu · 2024 [cited by examiner]
Azizi, S. et al., “Big self-supervised models advance medical image classification,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 3478-3488. [cited by applicant]
Bao, H. et al., “BEiT: BERT Pre-Training of Image Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision, arXiv preprint arXiv:2106.08254 (2021), 18 pages. [cited by applicant]
Bustos, A., et al., “Padchest: A large chest x-ray image dataset with multi-label annotated reports,” Medical Image Analysis vol. 66, 101797 (2020). [cited by applicant]
Candemir, S., et al., “Lung Segmentation in Chest Radiographs Using Anatomical Atlases with Nonrigid Registration,” IEEE Transactions on Medical Imaging, vol. 33(2), pp. 577-590 (2013). [cited by applicant]
Cao, H., et al., “Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation,” arXiv preprint arXiv:2105.05537 (2021). [cited by applicant]
Caron, M. et al., “Emerging Properties in Self-Supervised Vision Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9650-9660 (2021). [cited by applicant]
Chen, L., et al., “Identifying cardiomegaly in chest x-rays using dual attention network,” Applied Intelligence, vol. 52, pp. 11058-11067 (2022). [cited by applicant]
Chen, X. et al., “An Empirical Study of Training Self-supervised Vision Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 9640-9649. [cited by applicant]
Chen, Z., et al., “Masked Image Modeling Advances 3d Medical Image Analysis,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 1970-1980 (2023). [cited by applicant]
Chowdhury, M.E., et al., “Can AI Help in Screening Viral and COVID-19 Pneumonia?”, IEEE Access vol. 8, pp. 132665-132676 (2020). [cited by applicant]
Colak, E. et al., “The RSNA Pulmonary Embolism CT Dataset,” Radiology: Artificial Intelligence, vol. 3, No. 2, 2021, 7 pages. [cited by applicant]
Dosovitskiy, A. et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv preprint arXiv:2010.11929v2 [cs.CV] Jun. 3, 2021, 22 pages. [cited by applicant]
Haghighi, F. et al., “DiRA: Discriminative, Restorative, and Adversarial Learning for Self-Supervised Medical Image Analysis,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),… [cited by applicant]
Haghighi, F. et al., “Learning Semantics-enriched Representation via Self- discovery, Self-classification, and Self-restoration,” Medical Image Computing and Computer Assisted Intervention—MICCAI 2020, Lecture Notes in … [cited by applicant]
Haghighi, F. et al., “Transferable Visual Words: Exploiting the Semantics of Anatomical Patterns for Self-supervised Learning,” IEEE Transactions on Medical Imaging, vol. 40, No. 10, 2021, pp. 2857-2868. [cited by applicant]
Han, K., et al., “A Survey on Vision Transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, No. 1, pp. 87-110 (2022). [cited by applicant]
Hatamizadeh, A., et al., “Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images,” arXiv preprint arXiv:2201.01266v1 [eess.IV] (2022), 13 pages. [cited by applicant]
Haztamizadeh, A., et al., “UNETR: Transformers for 3D Medical Image Segmentation,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (2022), pp. 574-584. [cited by applicant]
He, K. et al., “Deep Residual Learning for Image Recognition”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 770-778. [cited by applicant]
He, K. et al., “Masked Autoencoders are Scalable Vision Learners,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2022), pp. 16000-16009. [cited by applicant]
Irvin, J. et al. “CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, No. 01, 2019, pp. 590-597. [cited by applicant]
Jaeger, S. et al., “Two public chest X-ray datasets for computer-aided screening of pulmonary diseases,” Quantitative Imaging in Medicine and Surgery, vol. 4, No. 6, 2014, pp. 475-477. [cited by applicant]
Johnson, A.E., et al., “MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports,” Scientific Data vol. 6, Article No. 317, (2019), pp. 1-8. [cited by applicant]
Khan, S., et al., “Transformers in Vision: A Survey,” ACM Computing Surveys (CSUR), vol. 54, Issue 10, Article 200, (2022), pp. 200:1-200:41. [cited by applicant]
Kim, E., et al., “XProtoNet: Diagnosis in Chest Radiography with Global and Local Explanations,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2021), pp. 15719-15728. [cited by applicant]
Li, Y., “Benchmarking Detection Transfer Learning with Vision Transformers,” arXiv preprint arXiv:2111.11429v1 [cs.CV] (2021), 9 pages. [cited by applicant]
Liu, Z. et al., “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 10012-10022. [cited by applicant]
Ma, D., et al., “Benchmarking Transformers for Medical Image Classification,” 2022, 14 pages. [cited by applicant]
Matsoukas, C., et al., “Is it Time to Replace CNNs with Transformers for Medical Images?”, arXiv preprint arXiv:2108.09038v1 [cs.CV] (2021), 6 pages. [cited by applicant]
Nguyen, H.Q., et al., “VinDR-CXR: An open dataset of chest X-rays with radiologist's annotations,” Scientific Data, vol. 9, Article 429 (2022), 7 pages. [cited by applicant]
Parvaiz, A., et al., “Vision Transformers in medical computer vision—A contemplative retrospection,” in Engineering Applications of Artificial Intelligence, vol. 122, 2023, 38 pages. [cited by applicant]
Paszke, A., et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” Advances in Neural Information Processing Systems vol. 32 (NeurIPS 2019), pp. 8024-8035. [cited by applicant]
Ridnik, T., et al., “Imagenet-21k Pretraining for the Masses,” arXiv preprint arXiv:210410972v4 [cs.CV] (2021), 20 pages. [cited by applicant]
Russakovsky, O., et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, vol. 115, 2015, 211-252. [cited by applicant]
Shamshad, F., et al., “Transformers in medical imaging: A survey,” in Medical Image Analysis, vol. 88, 2023, 102802, 40 pages. [cited by applicant]
Shiraishi, J., et al., “Development of a Digital Image Database for Chest Radiographs With and Without a Lung Nodule: Receiver Operating Characteristic Analysis of Radiologists' Detection of Pulmonary Nodules,” American… [cited by applicant]
Steiner, A., et al., “How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers,” arXiv preprint arXiv:2106:10270v2 [cs.CV] (2022), 16 pages. [cited by applicant]
Sun, C. et al., “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 843-852. [cited by applicant]
Taher, M.R.H., et al., “CAID: Context-Aware Instance Discrimination for Self-supervised Learning in Medical Imaging,” International Conference on Medical Imaging with Deep Learning (PMLR), vol. 172 (2022), pp. 535-551. [cited by applicant]
Taher, M.R.H., et al., “A Systematic Benchmarking Analysis of Transfer Learning for Medical Image Analysis,” Domain Adaptation and Representation Transfer, and Affordable Healthcare and AI for Resource Diverse Global He… [cited by applicant]
Touvron, H., “Training data-efficient image transformers & distillation through attention,” Proceedings of the 38th International Conference on Machine Learning, PMLR, vol. 139, 2021, pp. 10347-10357. [cited by applicant]
Wang, H., et al., “Triple attention learning for classification of 14 thoracic diseases using chest radiography,” Medical Image Analysis, vol. 67, 101846 (2021), 13 pages. [cited by applicant]
Wang, L., et al., “COVID-Net: a tailored deep convolutional neural network design for detection of COVID-19 cases from chest X-ray images,” Scientific Reports, vol. 10, Article No. 19549, (2020), 12 pages. [cited by applicant]
Wang, X., et al., “ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases,” Proceedings of the IEEE Conference on Computer Vision a… [cited by applicant]
Xie, Z. et al., “Self-Supervised Learning with Swin Transformers,” arXiv reprint arXiv:2015.04553v2 [cs.CV] (2021), 8 pages. [cited by applicant]
Xie, Z. et al., “SimMIM: A Simple Framework for Masked Image Modeling,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 9653-9663. [cited by applicant]
Zhai, X., et al., “Scaling Vision Transformers,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2021), pp. 12104-12113. [cited by applicant]
Zhou, Z., et al., “Models Genesis,” Medical Image Analysis, vol. 67, 101840 (2021). [cited by applicant]