IP Library › Granted Patent US 12,651,332
Granted Patent B2
US 12,651,332 · App. 17/821,511 · Granted Jun 9, 2026

Self-supervised learning for modeling a 3D brain anatomical representation

Inventors: Youngjin Yoo (Princeton, NJ); Eli Gibson (Plainsboro, NJ); Gengyan Zhao (Plainsboro, NJ); Bogdan Georgescu (Princeton, NJ)
Assignee: Siemens Healthineers AG
G06T7/0012G06T7/62G06V10/765G06T2207/10028G06T2207/20221G06T2207/30016G06V2201/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,332
App. No.
17/821,511
Granted
Jun 9, 2026
Kind
B2
Abstract

Systems and methods for performing a medical imaging analysis task are provided. A plurality of 3D (three dimensional) patches extracted from a 3D input medical image is received. A set of local features is extracted from each of the plurality of 3D patches using a machine learning based local feature extractor network. Global features representing relationships between the sets of local features are determined. A medical imaging analysis task is performed on the 3D input medical image based on the global features. Results of the medical imaging analysis task are output.

Claims (45)

1 . A computer-implemented method comprising:

receiving a plurality of 3D (three dimensional) patches extracted from a 3D input medical image;

extracting a set of local features from each of the plurality of 3D patches using a machine learning based local feature extractor network;

reconstructing, using a machine learning based decoder network, the sets of local features output by a bottleneck layer of the machine learning based feature extractor network;

determining global features representing relationships between the sets of local features;

hierarchically combining the global features with the reconstructed sets of local features at a plurality of scales;

performing a medical imaging analysis task on the 3D input medical image based on the hierarchical combination of the global features and the decoded sets of local features; and

outputting results of the medical imaging analysis task.

2 . The computer-implemented method of claim 1 , wherein the plurality of 3D patches is of a fixed resolution and a fixed field-of-view.

3 . The computer-implemented method of claim 1 , wherein determining global features representing relationships between the sets of local features comprises:

encoding the sets of local features with positional information; and

determining the global features based on the positional information.

4 . The computer-implemented method of claim 3 , wherein the positional information comprises pairwise positional relationships between the sets of local features.

5 . The computer-implemented method of claim 1 , wherein performing a medical imaging analysis task on the 3 D input medical image based on the hierarchical combination of the global features and the decoded sets of local features comprises:

performing a segmentation task based on the hierarchical combination of the global features and the decoded sets of local features.

6 . The computer-implemented method of claim 1 , wherein the medical imaging analysis task comprises one of classification, detection, or segmentation.

7 . The computer-implemented method of claim 1 , wherein the local feature extractor network is trained with self-supervised learning.

8 . The computer-implemented method of claim 1 , wherein the 3D input medical image depicts a brain of a patient.

9 . An apparatus comprising:

means for receiving a plurality of 3D (three dimensional) patches extracted from a 3D input medical image;

means for extracting a set of local features from each of the plurality of 3D patches using a machine learning based local feature extractor network;

means for reconstructing, using a machine learning based decoder network, the sets of local features output by a bottleneck layer of the machine learning based feature extractor network;

means for determining global features representing relationships between the sets of local features;

means for hierarchically combining the global features with the reconstructed sets of local features at a plurality of scales;

means for performing a medical imaging analysis task on the 3D input medical image based on the hierarchical combination of the global features and the decoded sets of local features; and

means for outputting results of the medical imaging analysis task.

10 . The apparatus of claim 9 , wherein the plurality of 3D patches is of a fixed resolution and a fixed field-of-view.

11 . The apparatus of claim 9 , wherein the means for determining global features representing relationships between the sets of local features comprises:

means for encoding the sets of local features with positional information; and

means for determining the global features based on the positional information.

12 . The apparatus of claim 11 , wherein the positional information comprises pairwise positional relationships between the sets of local features.

13 . A non-transitory computer readable medium storing computer program instructions, the computer program instructions when executed by a processor cause the processor to perform operations comprising:

receiving a plurality of 3D (three dimensional) patches extracted from a 3D input medical image;

extracting a set of local features from each of the plurality of 3D patches using a machine learning based local feature extractor network;

reconstructing, using a machine learning based decoder network, the sets of local features output by a bottleneck layer of the machine learning based feature extractor network;

determining global features representing relationships between the sets of local features;

hierarchically combining the global features with the reconstructed sets of local features at a plurality of scales;

performing a medical imaging analysis task on the 3D input medical image based on the hierarchical combination of the global features and the decoded sets of local features; and

outputting results of the medical imaging analysis task.

14 . The non-transitory computer readable medium of claim 13 , wherein the plurality of 3D patches is of a fixed resolution and a fixed field-of-view.

15 . The non-transitory computer readable medium of claim 13 , wherein performing a medical imaging analysis task on the 3D input medical image based on the hierarchical combination of the global features and the decoded sets of local features comprises:

performing a segmentation task based on the hierarchical combination of the global features and the decoded sets of local features.

16 . The non-transitory computer readable medium of claim 13 , wherein the medical imaging analysis task comprises one of classification, detection, or segmentation.

17 . The non-transitory computer readable medium of claim 13 , wherein the local feature extractor network is trained with self-supervised learning.

18 . The non-transitory computer readable medium of claim 13 , wherein the 3D input medical image depicts a brain of a patient.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: SIEMENS HEALTHCARE GMBH
To: SIEMENS HEALTHINEERS AG
Reel/Frame 066267/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2022
From: SIEMENS MEDICAL SOLUTIONS USA, INC.
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 060925/0879 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2022
From: YOO, YOUNGJIN; GIBSON, ELI; ZHAO, GENGYAN; GEORGESCU, BOGDAN
To: SIEMENS MEDICAL SOLUTIONS USA, INC.
Reel/Frame 060907/0898 →
Continuity (1)
Related Publication 20240070853A1 · Feb 29, 2024
References Cited (22)
US 10665011B1 · Sunkavalli · 2020 [cited by examiner]
US 20140254936A1 · Sun · 2014 [cited by examiner]
US 20200311426A1 · Charlton · 2020 [cited by examiner]
US 20210150710A1 · Hosseinzadeh Taher · 2021 [cited by examiner]
US 20220262024A1 · Sun · 2022 [cited by examiner]
US 20220351863A1 · Wang · 2022 [cited by examiner]
CN 112966626A · 2021 [cited by examiner]
CN 114581462A · 2022 [cited by examiner]
WO WO2013073624A1 · 2013 [cited by examiner]
WO WO2019018063A1 · 2019 [cited by examiner]
Huang et al., “Learning hierarchical representations for face verification with convolutional deep belief networks”, IEEE Conference on Computer Vision and Pattern Recognition, 2012, 8 pgs. [cited by applicant]
Caron et al., “Unsupervised learning of visual features by contrasting cluster assignments”, Advances in Neural Information Processing Systems, 2020, 22 pgs. [cited by applicant]
Chen et al., “A simple framework for contrastive learning of visual representations”, International Conference on Machine Learning, 2020, 20 pgs. [cited by applicant]
Zhou et al., “Models Genesis”, Medical Image Analysis, 2022, pp. 1-26. [cited by applicant]
Jun et al., “Medical Transformer: Universal Brain Encoder for 3D MRI Analysis”, arXiv, 2021, 9 pgs. [cited by applicant]
Lambert et al., “Leveraging 3D information in unsupervised brain MRI segmentation”, IEEE 18th International Symposium on Biomedical Imaging, 2021, 4 pgs. [cited by applicant]
Roy et al., “Are 2.5D approaches superior to 3D deep networks in whole brain segmentation?”, Proceedings of Machine Learning Research, 2022, 17 pgs. [cited by applicant]
Taleb et al., “3D self-supervised methods for medical imaging”, Advances in Neural Information Processing Systems, 2020, 17 pgs. [cited by applicant]
Dosovitskiy et al., “Discriminative unsupervised feature learning with convolutional neural networks”, Advances in Neural Information Processing Systems, 2014, 13 pgs. [cited by applicant]
Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale”, arXiv, Computer Science, Computer Vision and Pattern Recognition, 2020, 22 pgs. [cited by applicant]
Ranftl et al., “Vision transformers for dense prediction”, Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, 10 pgs. [cited by applicant]
Wu et al., “Rethinking and improving relative position encoding for vision transformer”, Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, 13 pgs. [cited by applicant]