IP Library Granted Patent US 12,562,260
Granted Patent B2
US 12,562,260 · App. 18/494,790 · Granted Feb 24, 2026

Multi-modal foundational models for medical images using modality-specific and cross-modality expert subnetworks

Inventors: Alexandru Constantin Serban (Constanta, RO); Florin-Cristian Ghesu (Baiersdorf, DE); Dominik Neumann (Erlangen, DE); Venkatesh Narasimha Murthy (Hillsborough, NJ); Bogdan Georgescu (Princeton, NJ)
Assignee: Siemens Healthineers AG
G16H30/40G06V10/774G06V10/95G06V20/50G06V10/82G06V2201/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,260
App. No.
18/494,790
Granted
Feb 24, 2026
Kind
B2
Abstract

Systems and methods for performing a medical imaging analysis task using a machine learning based model (e.g., a foundational model) are provided. One or more input medical images are received. One or more modality-specific expert subnetworks of the machine learning based model are selected for performing a first processing of the one or more input medical images. The first processing of the one or more input medical images is performed for performing a medical imaging analysis task using the one or more selected modality-specific expert subnetworks. Results of the first processing are merged into one or more sets of merged results. One or more cross-modality expert subnetworks of the machine learning based model are selected for performing a second processing of the one or more sets of merged results. The second processing of the one or more sets of merged results is performed for performing the medical imaging analysis task using the one or more selected cross-modality expert subnetworks. Results of the second processing are output.

Claims (45)

1 . A computer-implemented method comprising:

receiving one or more input medical images;

selecting one or more modality-specific expert subnetworks of a machine learning based model for performing a first processing of the one or more input medical images;

performing the first processing of the one or more input medical images for performing a medical imaging analysis task using the one or more selected modality-specific expert subnetworks;

merging results of the first processing into one or more sets of merged results;

selecting one or more cross-modality expert subnetworks of the machine learning based model for performing a second processing of the one or more sets of merged results;

performing the second processing of the one or more sets of merged results for performing the medical imaging analysis task using the one or more selected cross-modality expert subnetworks; and

outputting results of the second processing.

2 . The computer-implemented method of claim 1 , wherein the machine learning based model is trained via self-supervised learning.

3 . The computer-implemented method of claim 1 , wherein performing the first processing of the one or more input medical images for performing a medical imaging analysis task using the one or more selected modality-specific expert subnetworks comprises:

processing modality-specific data of the one or more input medical images using the one or more selected modality-specific expert subnetworks.

4 . The computer-implemented method of claim 1 , wherein performing the second processing of the one or more sets of merged results for performing the medical imaging analysis task using the one or more selected cross-modality expert subnetworks comprises:

integrating the one or more sets of merged results between each other.

5 . The computer-implemented method of claim 1 , wherein a number of modality-specific expert subnetworks in the machine learning based network is learned for each modality during training of the machine learning based network.

6 . The computer-implemented method of claim 1 , wherein a number of modality-specific expert subnetworks in the machine learning based network is predefined for each modality.

7 . The computer-implemented method of claim 1 , wherein a number of cross-modality expert subnetworks in the machine learning based network is learned during training of the machine learning based network.

8 . The computer-implemented method of claim 1 , wherein a number of layers and a number of parameters of each of the layer for each modality-specific expert subnetwork and each cross-modality expert subnetworks in the machine learning based network are defined based on resource availability.

9 . The computer-implemented method of claim 1 , wherein the machine learning based model is a foundational model.

10 . An apparatus comprising:

means for receiving one or more input medical images;

means for selecting one or more modality-specific expert subnetworks of a machine learning based model for performing a first processing of the one or more input medical images;

means for performing the first processing of the one or more input medical images for performing a medical imaging analysis task using the one or more selected modality-specific expert subnetworks;

means for merging results of the first processing into one or more sets of merged results;

means for selecting one or more cross-modality expert subnetworks of the machine learning based model for performing a second processing of the one or more sets of merged results;

means for performing the second processing of the one or more sets of merged results for performing the medical imaging analysis task using the one or more selected cross-modality expert subnetworks; and

means for outputting results of the second processing.

11 . The apparatus of claim 10 , wherein the machine learning based model is trained via self-supervised learning.

12 . The apparatus of claim 10 , wherein the means for performing the first processing of the one or more input medical images for performing a medical imaging analysis task using the one or more selected modality-specific expert subnetworks comprises:

means for processing modality-specific data of the one or more input medical images using the one or more selected modality-specific expert subnetworks.

13 . The apparatus of claim 10 , wherein the means for performing the second processing of the one or more sets of merged results for performing the medical imaging analysis task using the one or more selected cross-modality expert subnetworks comprises:

means for integrating the one or more sets of merged results between each other.

14 . The apparatus of claim 10 , wherein a number of modality-specific expert subnetworks in the machine learning based network is learned for each modality during training of the machine learning based network.

15 . A non-transitory computer readable medium storing computer program instructions, the computer program instructions when executed by a processor cause the processor to perform operations comprising:

receiving one or more input medical images;

selecting one or more modality-specific expert subnetworks of a machine learning based model for performing a first processing of the one or more input medical images;

performing the first processing of the one or more input medical images for performing a medical imaging analysis task using the one or more selected modality-specific expert subnetworks;

merging results of the first processing into one or more sets of merged results;

selecting one or more cross-modality expert subnetworks of the machine learning based model for performing a second processing of the one or more sets of merged results;

performing the second processing of the one or more sets of merged results for performing the medical imaging analysis task using the one or more selected cross-modality expert subnetworks; and

outputting results of the second processing.

16 . The non-transitory computer readable medium of claim 15 , wherein the machine learning based model is trained via self-supervised learning.

17 . The non-transitory computer readable medium of claim 15 , wherein a number of modality-specific expert subnetworks in the machine learning based network is predefined for each modality.

18 . The non-transitory computer readable medium of claim 15 , wherein a number of cross-modality expert subnetworks in the machine learning based network is learned during training of the machine learning based network.

19 . The non-transitory computer readable medium of claim 15 , wherein a number of layers and a number of parameters of each of the layer for each modality-specific expert subnetwork and each cross-modality expert subnetworks in the machine learning based network are defined based on resource availability.

20 . The non-transitory computer readable medium of claim 15 , wherein the machine learning based model is a foundational model.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: SIEMENS HEALTHCARE GMBH
To: SIEMENS HEALTHINEERS AG
Reel/Frame 066267/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2023
From: SIEMENS MEDICAL SOLUTIONS USA, INC.
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 065708/0200 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2023
From: SIEMENS S.R.L.
To: SIEMENS MEDICAL SOLUTIONS USA, INC.
Reel/Frame 065679/0401 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 27, 2023
From: SERBAN, ALEXANDRU CONSTANTIN
To: SIEMENS S.R.L.
Reel/Frame 065665/0792 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 27, 2023
From: GHESU, FLORIN-CRISTIAN; NEUMANN, DOMINIK
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 065665/0914 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2023
From: NARASIMHA MURTHY, VENKATESH; GEORGESCU, BOGDAN
To: SIEMENS MEDICAL SOLUTIONS USA, INC.
Reel/Frame 065604/0609 →
Continuity (1)
Related Publication 20250140382A1 · May 1, 2025
References Cited (9)
US 20190114773A1 · Song · 2019 [cited by examiner]
Extended European Search Report (EESR) mailed Mar. 5, 2025 in corresponding European Patent Application No. 24208292.3. [cited by applicant]
Guo Zhe et al: “Deep Learning-Based Image 1-15 INV. Segmentation on Multimodal Medical Imaging”, IEEE Transactions on Radiation and Plasma Medical Sciences, IEEE, vol. 3, No. 2, Mar. 1, 2019; pp. 162-169. [cited by applicant]
Radford et al., “Learning transferable visual models from natural language supervision”, arXiv:2103.00020v1, 2021, pp. 1-48. [cited by applicant]
Driess et al., “PaLM-E: An embodied multimodal language model”, arXiv:2303.03378v1, 2023, 18 pgs. [cited by applicant]
Ghesu et al., “Contrastive self-supervised learning from 100 million medical images”, arXiv:2201.01283v1, 2022, pp. 1-13. [cited by applicant]
Riquelme et al., “Scaling vision with sparse mixture of experts”, Advances in Neural Information Processing Systems, 2021, pp. 1-13. [cited by applicant]
Du et al., “GlaM: Efficient scaling of language models with mixture-of-experts”, Proceedings of the 39th International Conference on Machine Learning, 2022, 23 pgs. [cited by applicant]
Baevsky et al., “Efficient self-supervised learning with contextualized target representations for vision, speech and language”, Proceedings of the 40th International Conference on Machine Learning, 2023, pp. 1-14. [cited by applicant]