IP Library Granted Patent US 12,620,139
Granted Patent B1
US 12,620,139 · App. 17/690,531 · Granted May 5, 2026

Neural network-based image segmentation

Inventors: Ali Hatamizadeh (Los Angeles, CA); Vishwesh Nath (Nashville, TN); Yucheng Tang (Nashville, TN); Dong Yang (Pocatello, ID); Wenqi Li (London, GB); Holger Roth (Rockville, MD); Daguang Xu (Potomac, MD)
Assignee: NVIDIA Corporation
G06T9/002G06T3/60G06T7/11G06V10/25G06V10/40G06V10/7747G06V10/82G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,139
App. No.
17/690,531
Granted
May 5, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques are presented to perform segmentation on images. In at least one embodiment, one or more neural networks are used to segment an image based, at least in part, on one or more visual modifications of the image.

Claims (30)

1 . A processor, comprising:

one or more circuits to use one or more neural networks to segment an image, wherein the one or more neural networks comprise:

a transformer neural network that includes one or more encoders, wherein the one or more encoders are pre-trained to extract features from input data using self-attention between a sequence of different portions of the input data based, at least in part, on one or more proxy tasks and unlabeled data; and

a convolutional neural network that includes one or more decoders trained together with the one or more encoders as part of additional training of the one or more encoders to extract features from the image to segment and input to the one or more decoders based, at least in part, on labeled data.

2 . The processor of claim 1 , wherein the one or more proxy tasks include removing, from one or more sub-volumes of an input image volume of the input data, one or more mask regions and training the one or more encoders to predict image data that was removed from the one or more mask regions.

3 . The processor of claim 1 , wherein the one or more proxy tasks include predicting one or more sub-volumes of an input image volume of the input data, given one or more rotated versions of the one or more sub-volumes.

4 . The processor of claim 1 , wherein the one or more proxy tasks include a contrastive learning task to train the one or more encoders to differentiate between different regions of interest (ROIs) including different types of features in different views or portions of the input data.

5 . A system comprising:

one or more processors to use one or more neural networks to segment an image, wherein the one or more neural networks comprise:

a transformer network that includes one or more encoders pre-trained to extract features from input data using self-attention between a sequence of different portions of the input data based, at least in part, on one or more proxy tasks and unlabeled data; and

a convolutional neural network that includes one or more decoders trained together with the one or more encoders as part of additional training of the one or more encoders to extract features from the image to segment and input to the one or more decoders based, at least in part, on labeled data.

6 . The system of claim 5 , wherein the one or more proxy tasks include removing, from one or more sub-volumes of an input image volume of the input data, one or more mask regions and training the one or more encoders to predict image data that was removed from the one or more mask regions.

7 . The system of claim 5 , wherein the one or more proxy tasks include predicting one or more sub-volumes of an input image volume of the input data, given one or more rotated versions of the one or more sub-volumes.

8 . The system of claim 5 , wherein the one or more proxy tasks include a contrastive learning task to train the one or more encoders to differentiate between different regions of interest (ROIs) including different types of features in different views or portions of the input data.

9 . A method comprising:

using one or more neural networks to segment an image, wherein the one or more neural networks comprise:

transformer neural network that includes one or more encoders, wherein the one or more encoders are pre-trained to extract features from input data using self-attention between a sequence of different portions of the input data based, at least in part, on one or more proxy tasks and unlabeled data; and

a convolutional neural network that includes one or more decoders of trained together with the one or more encoders as part of additional training of the one or more encoders to extract features from the image to segment and input to the one or more decoders based, at least in part, on labeled data.

10 . The method of claim 9 , wherein the one or more proxy tasks include removing, from one or more sub-volumes of an input image volume of the input data, one or more mask regions and training the one or more encoders to predict image data that was removed from the one or more mask regions.

11 . The method of claim 9 , wherein the one or more proxy tasks include predicting one or more sub-volumes of an input image volume of the input data, given one or more rotated versions of the one or more sub-volumes.

12 . The method of claim 9 , wherein the one or more proxy tasks include a contrastive learning task to train the one or more encoders to differentiate between different regions of interest (ROIs) including different types of features in different views or portions of the input data.

13 . An image segmentation system, comprising:

one or more processors to use one or more neural networks to segment an image;

memory for storing network parameters for the one or more neural networks; and

wherein the one or more neural networks comprise:

a transformer neural network that includes one or more encoders, wherein the one or more encoders are pre-trained to extract features from input data using self-attention between a sequence of different portions of the input data based, at least in part, on one or more proxy tasks and unlabeled data; and

a convolutional neural network that includes one or more decoders trained together with the one or more encoders as part of additional training of the one or more encoders to extract features from the image to segment and input to the one or more decoders based, at least in part, on labeled data.

14 . The image segmentation system of claim 13 , wherein the one or more proxy tasks include removing, from one or more sub-volumes of an input image volume of the input data, one or more mask regions and training the one or more encoders to predict image data that was removed from the one or more mask regions.

15 . The image segmentation system of claim 13 , wherein the one or more proxy tasks include predicting one or more sub-volumes of an input image volume of the input data, given one or more rotated versions of the one or more sub-volumes.

16 . The image segmentation system of claim 13 , wherein the one or more proxy tasks include a contrastive learning task to train the one or more encoders to differentiate between different regions of interest (ROIs) including different types of features in different views or portions of the input data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2022
From: HATAMIZADEH, ALI; NATH, VISHWESH; TANG, YUCHENG; YANG, DONG; LI, WENQI; ROTH, HOLGER; XU, DAGUANG
To: NVIDIA CORPORATION
Reel/Frame 059225/0724 →
References Cited (69)
US 11748887B2 · Jampani · 2023 [cited by examiner]
US 11816185B1 · Roth · 2023 [cited by examiner]
US 11922628B2 · Zhou · 2024 [cited by examiner]
US 20180144244A1 · Masoud · 2018 [cited by examiner]
US 20190130229A1 · Lu · 2019 [cited by examiner]
US 20210374547A1 · Wang · 2021 [cited by examiner]
US 20220148162A1 · Lee · 2022 [cited by examiner]
US 20220156943A1 · Zhang · 2022 [cited by examiner]
US 20220261593A1 · Yu · 2022 [cited by examiner]
US 20230072400A1 · Bajpai · 2023 [cited by examiner]
US 20240005650A1 · Dippel · 2024 [cited by examiner]
Xu et al., “LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation,” Jul. 19, 2021, 10 Pages. [cited by applicant]
Yan et al., “Self-supervised Learning of Pixel-wise Anatomical Embeddings in Radiological Images,” Dec. 4, 2020, 16 Pages. [cited by applicant]
Zhai et al., “Scaling Vision Transformers,” Jun. 8, 2021, 31 Pages. [cited by applicant]
Zheng et al., “Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers,” 2020, 12 pages. [cited by applicant]
Zhou et al., “Models Genesis,” Medical Image Analysis, Dec. 16, 2020, 26 Pages. [cited by applicant]
Zhou et al., “nnFormer: Interleaved Transformer for Volumetric Segmentation,” Sep. 21, 2021, 18 Pages. [cited by applicant]
Zhou et al., “Prior-aware Neural Network for Partially-Supervised Multi-Organ Segmentation,” Aug. 21, 2019, 12 Pages. [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks,” In Proceedings of the IEEE International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Antonelli et al., “The Medical Segmentation Decathlon,” Jun. 10, 2021, 41 pages. [cited by applicant]
Armato et al., “The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A Completed Reference Database of Lung Nodules on CT Scans,” Medical Physics 38, 2011, 17 pages. [cited by applicant]
Atito et al., “SiT: Self-supervised Vision Transformer,” Nov. 14, 2021, 13 Pages. [cited by applicant]
Azizi et al., “Big Self-Supervised Models Advance Medical Image Classifications,” Apr. 1, 2021, 19 Pages. [cited by applicant]
Bakas, “Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge,” Nov. 5, 2018, 28 pages. [cited by applicant]
Bao et al., “BEIT: BERT Pre-Training of Image Transformers,” Jun. 15, 2021, 16 pages. [cited by applicant]
Cao et al., “Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation,” May 12, 2021, 14 Pages. [cited by applicant]
Caron et al., “Emerging Properties in Self-Supervised Vision Transformers,” May 24, 2021, 21 pages. [cited by applicant]
Chen et al., “An Empirical Study of Training Self-Supervised Vision Transformers,” Aug. 16, 2021, 10 pages. [cited by applicant]
Chen et al., “Deeplab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFS,” IEEE transactions on pattern analysis and machine intelligence, 40(4):, May 12, 2017, 14 pa… [cited by applicant]
Chen et al., “Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation,” Aug. 22, 2018, 18 pages. [cited by applicant]
Chen et al., “Self-supervised Learning for Medical Image Analysis Using Image Context Restoration,” Medical Image Analysis, Jul. 26, 2019, 12 Pages. [cited by applicant]
Chen et al., “TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation,” Feb. 8, 2021, 13 pages. [cited by applicant]
Cicek et al., “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation,” Jun. 21, 2016, 8 Pages. [cited by applicant]
Dai et al., “UP-DETR: Unsupervised Pre-training for Object Detection with Transformers,” Apr. 7, 2021, 11 Pages. [cited by applicant]
Desai et al., “Chest Imaging Representing a COVID-19 Positive Rural U.S. Population,” Scientific Data, 2020, 6 Pages. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” Oct. 22, 2020, 21 pages. [cited by applicant]
Gidaris et al., “Unsupervised Representation Learning by Predicting Image Rotations,” Mar. 21, 2018, 16 pages. [cited by applicant]
Grossberg et al., “Imaging and Clinical Data Archive for Head and Neck Squamous Cell Carcinoma Patients Treated with Radiotherapy,” Scientific Data, Jan. 4, 2018, 10 Pages. [cited by applicant]
Haghighi et al., “Transferable Visual Words: Exploiting the Semantics of Anatomical Patterns for Self-supervised earning,” Feb. 21, 2021, 15 Pages. [cited by applicant]
Hatamizadeh et al., “UNETR: Transformers for 3D Medical Image Segemntation,” IEEE/CVF Winter Conference on Applications of Computer Vision, Mar. 18, 2021, 11 pages. [cited by applicant]
He et al., “Momentum Contrast for Unsupervised Visual Representation Learning,” CVPR, 2020, 10 pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008. [cited by applicant]
Sensee et al., “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature Methods 18(2): 2021, 14 pages. [cited by applicant]
Jing et al., “Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey,” Feb. 16, 2019, 24 Pages. [cited by applicant]
Johnson et al., “Accuracy of CT Colonography for Detection of Large Adenomas and Cancers,” New England Journal of Medicine vol. 359, p. 1207, 2008, 13 Pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” Dec. 27, 2013, 14 pages. [cited by applicant]
Landman et al., Miccai Multi-Atlas Labeling Beyond the Cranial Vault-Workshop and Challenge, MICCAI Multi-Atlas Labeling Beyond Cranial Vault-Workshop Challenge, Apr. 15, 2015, 5 pages. [cited by applicant]
Liang et al., “SwinIR: Image Restoration Using Swin Transformer,” Aug. 23, 2021, 12 Pages. [cited by applicant]
Lin et al., “DS-TransUNet: Dual Swin Transformer U-Net for Medical Image Segmentation,” Jun. 12, 2021, 13 Pages. [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” CVPR, 2017, 9 pages. [cited by applicant]
Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” Mar. 25, 2021, 11 pages. [cited by applicant]
Luo et al., “Understanding the Effective Receptive Field in Deep Convolutional Neural Networks,” 29th Conference on Neutral Information Processing Systems (NIPS 2016), Barcelona, Spain, 2016, 9 Pages. [cited by applicant]
Noroozi et al., “Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles,” European Conference on Computer Vision, Jun. 26, 2016, 17 pages. [cited by applicant]
Oord et al., “Representation Learning with Contrastive Predictive Coding,” Jul. 10, 2018, 13 pages. [cited by applicant]
Pathak et al., “Context Encoders: Feature Learning by Inpainting,” IEEE Conference on Computer Vision and Pattern Recognition, Nov. 21, 2016, 12 pages. [cited by applicant]
Raghu et al., “Transfusion: Understanding Transfer Learning for Medical Imaging,” Oct. 29, 2019, 22 Pages. [cited by applicant]
Roth et al., “A Multi-Scale Pyramid of 3D Fully Convolutional Networks for Abdominal Multi-Organ Segmentation,” Jun. 6, 2018, 8 Pages. [cited by applicant]
Roth et al., “An Application of Cascaded 3D Fully Convolutional Networks for Medical Image Segmentation,” Mar. 20, 2018, 22 Pages. [cited by applicant]
Setio et al., “Validation, Comparison, and Combination of Algorithms for Automatic Detection of Pulmonary Nodules in Computed Tomography Images: The LUNA16 challenge,” Jul. 15, 2017, 17 Pages. [cited by applicant]
Sharir et al., “An Image is Worth 16x16 Words, What is a Video Worth?,” May 27, 2021, 11 Pages. [cited by applicant]
Singh et al., “SNIPER: Efficient Multi-Scale Training,” May 23, 2018, 10 Pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Taleb et al., “3D Self-Supervised Methods for Medical Imaging,” Nov. 2, 2020, 19 Pages. [cited by applicant]
Tang et al., “Body Part Regression With Self-Supervision,” IEEE Transactions On Medical Imaging, vol. 40, No. 5, May 2021, 9 Pages. [cited by applicant]
Tang et al., “High-resolution 3D Abdominal Segmentation with Random patch Network Fusion,” Medical Image Analysis, Dec. 16, 2020, 15 Pages. [cited by applicant]
Valanarasu et al., “Medical Transformer: Gated Axial-Attention for Medical Image Segmentation,” Jul. 6, 2021, 18 Pages. [cited by applicant]
Wang et al., “Self-supervised Image-text Pre-training With Mixed Data In Chest X-rays,” Mar. 30, 2021, 10 Pages. [cited by applicant]
Xie et al., “CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation,” Mar. 4, 2021, 13 Pages. [cited by applicant]