IP Library Granted Patent US 12,536,664
Granted Patent B2
US 12,536,664 · App. 16/913,085 · Granted Jan 27, 2026

Encoder regularization of a segmentation model

Inventor: Andriy Myronenko (San Mateo, CA)
Assignee: NVIDIA Corporation
G06T7/11G06F18/211G06F18/217G06V10/764G06V10/82G06F18/214G06T2207/10088G06T2207/20076G06T2207/20081G06T2207/20084G06T2207/30004G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,664
App. No.
16/913,085
Granted
Jan 27, 2026
Kind
B2
Abstract

A segmentation model is trained with an image reconstruction model that shares an encoding. During application of the segmentation model, the segmentation model may use the encoding and network layers trained for the segmentation without the image reconstruction model. The image reconstruction model may include a probabilistic representation of the image that represents the image based on a probability distribution. When training the model, the encoding layers of the model use a loss function including an error term from the segmentation model and from the autoencoder model. The image reconstruction model thus regularizes the encoding layers and improves modeling results and prevents overfitting, particularly for small training sizes.

Claims (50)

1 . A method comprising:

generating an encoding representation based, at least in part, on a first image;

generating a second image from the encoding representation;

generating a segmentation based, at least in part, on the encoding representation;

calculating an error based, at least in part, on the encoding representation, the segmentation, and the second image; and

updating one or more neural networks based, at least in part, on the error.

2 . The method of claim 1 , further comprising generating the second image using one or more autoencoders of one or more neural networks.

3 . The method of claim 1 , further comprising training one or more neural networks based, at least in part, on a first set of data computed based, at least in part, on the segmentation and a second set of data computed based, at least in part, on the second image.

4 . The method of claim 1 , further comprising generating the segmentation using one or more neural networks during training of the one or more neural networks, where the training is based, at least in part, on the second image and the segmentation.

5 . The method of claim 1 , wherein the second image is to be generated by one or more neural networks comprising at least an autoencoder, wherein the one or more neural networks are to generate the segmentation during training of the one or more neural networks.

6 . The method of claim 1 , further comprising one or more autoencoders to generate the second image, where one or more errors are to be computed based, at least in part, on data generated by the one or more autoencoders and the one or more errors are usable to generate the segmentation during training of the one or more neural networks.

7 . The method of claim 1 , wherein the first image comprises three-dimensional medical imaging data.

8 . One or more processors comprising:

circuitry to generate a second image and segmentation information based, at least in part, on an encoding representation of a first image, and to update a neural network based, at least in part, on the second image, the encoding representation, and the segmentation information.

9 . The one or more processors of claim 8 , wherein the circuitry is to train one or more neural networks based, at least in part, on the second image and the segmentation information.

10 . The one or more processors of claim 8 , wherein the circuitry is to calculate one or more first types of error values based, at least in part, on the segmentation information and one or more second types of error values based, at least in part, on the second image, and one or more neural networks are to be trained based, at least in part, on the one or more first types of error values and the one or more second types of error values.

11 . The one or more processors of claim 8 , wherein the second image is to be generated, at least in part, by one or more autoencoders.

12 . The one or more processors of claim 8 , wherein one or more circuits are to generate the segmentation based, at least in part, on the first image during training of one or more neural networks, where the one or more circuits are to train the one or more neural networks based, at least in part, on the segmentation and the second image.

13 . The one or more processors of claim 8 , wherein the circuitry is to train one or more neural networks to generate the segmentation information based, at least in part, on the first image, where the one or more neural networks are to generate the second image using one or more autoencoders.

14 . The one or more processors of claim 8 , wherein the segmentation information comprises one or more data values usable to identify one or more objects in the first image.

15 . The one or more processors of claim 8 , wherein the first image comprises medical imaging data and the second image comprises at least one or more portions of the first image.

16 . A system comprising:

one or more processors to:

generate an encoding representation based, at least in part, on a first image;

generate a second image from the encoding representation;

generate a segmentation based, at least in part, on the encoding representation;

calculate an error based, at least in part, on the encoding representation, the segmentation, and the second image; and

update one or more neural networks based, at least in part, on the error; and

one or more memory devices to at least partially store the one or more neural networks.

17 . The system of claim 16 , further comprising one or more autoencoders to generate, at least in part, the second image.

18 . The system of claim 16 , wherein the one or more processors are to generate the second image and the segmentation using one or more neural networks.

19 . The system of claim 16 , wherein the one or more processors are to generate the second image using one or more neural networks and the one or more neural networks comprise one or more autoencoders and the one or more processors are to train the one or more neural networks based, at least in part, on the second image and the segmentation.

20 . The system of claim 16 , further comprising a neural network usable to generate the segmentation and the second image, wherein one or more processors are to train the neural network based, at least in part, on one or more first types of error values calculated based, at least in part, on the segmentation and one or more second types of error values calculated based, at least in part, on the second image.

21 . The system of claim 16 , further comprising one or more neural networks to generate the segmentation based, at least in part, on the first image, where the one or more processors are to train the one or more neural networks based, at least in part, on the segmentation and the second image.

22 . The system of claim 16 , wherein the first image comprises medical imaging data and the segmentation is to indicate one or more objects in the first image.

23 . A non-transitory computer-readable medium having stored thereon one or more instructions, which if performed by one or more processors, cause the one or more processors to at least:

generate an encoding representation based, at least in part, on a first image;

generate a second image from the encoding representation;

generate a segmentation based, at least in part, on the encoding representation;

calculate an error based, at least in part, on the encoding representation, the segmentation, and the second image; and

update one or more neural networks based, at least in part, on the error; and

one or more memory devices to at least partially store the one or more neural networks.

24 . The non-transitory computer-readable medium of claim 23 , further comprising instructions which, if performed by the one or more processors, cause the one or more processors to generate the second image using one or more autoencoders of one or more neural networks, where the one or more neural networks are to be trained based, at least in part, on the second image and the segmentation.

25 . The non-transitory computer-readable medium of claim 23 , further comprising instructions which, if performed by the one or more processors, cause the one or more processors to generate the second image using one or more autoencoders, where one or more error values of the one or more autoencoders is to be used to train one or more neural networks with the segmentation.

26 . The non-transitory computer-readable medium of claim 23 , further comprising instructions which, if performed by the one or more processors, cause the one or more processors to compute a first set of error values using the segmentation and a second set of error values using the second image, the first set and the second set usable to train one or more neural networks comprising one or more autoencoders usable to generate the second image.

27 . The non-transitory computer-readable medium of claim 23 , further comprising one or more neural networks to generate the segmentation and the second image.

28 . The non-transitory computer-readable medium of claim 23 , further comprising one or more autoencoders to generate the second image.

29 . The non-transitory computer-readable medium of claim 23 , further comprising error information generated based, at least in part, on the second image and the segmentation, where the error information is to be usable to train one or more neural networks.

30 . The non-transitory computer-readable medium of claim 23 , wherein the segmentation is a set of data comprising data values to identify one or more portions of the first image.

31 . The non-transitory computer-readable medium of claim 23 , wherein the first image comprises magnetic resonance imaging data and the segmentation is usable to identify one or more objects in the first image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2020
From: MYRONENKO, ANDRIY
To: NVIDIA CORPORATION
Reel/Frame 053725/0756 →
Continuity (2)
Continuation 16223005 · Dec 17, 2018
Related Publication 20210012504A1 · Jan 14, 2021
References Cited (71)
US 4804831A · Baba et al. · 1989 [cited by applicant]
US 4907156A · Doi et al. · 1990 [cited by applicant]
US 5956435A · Buzug et al. · 1999 [cited by applicant]
US 6137531A · Kanzaki et al. · 2000 [cited by applicant]
US 6337926B2 · Takahashi et al. · 2002 [cited by applicant]
US 6586934B2 · Biglieri et al. · 2003 [cited by applicant]
US 6611629B2 · Bender et al. · 2003 [cited by applicant]
US 6683974B1 · Nagasawa et al. · 2004 [cited by applicant]
US 6819952B2 · Pfefferbaum et al. · 2004 [cited by applicant]
US 6888894B2 · Prakash et al. · 2005 [cited by applicant]
US 7050503B2 · Prakash et al. · 2006 [cited by applicant]
US 7602965B2 · Hong et al. · 2009 [cited by applicant]
US 7742650B2 · Xu et al. · 2010 [cited by applicant]
US 7809154B2 · Lienhart et al. · 2010 [cited by applicant]
US 7865866B2 · Kim et al. · 2011 [cited by applicant]
US 7958063B2 · Long et al. · 2011 [cited by applicant]
US 8094903B2 · Zhu et al. · 2012 [cited by applicant]
US 8116982B2 · Hunter et al. · 2012 [cited by applicant]
US 8345944B2 · Zhu et al. · 2013 [cited by applicant]
US 8620083B2 · King et al. · 2013 [cited by applicant]
US 8824758B2 · Yu et al. · 2014 [cited by applicant]
US 9383347B2 · Marugame · 2016 [cited by applicant]
US 9576219B2 · Piatrou et al. · 2017 [cited by applicant]
US 9584814B2 · Socek et al. · 2017 [cited by applicant]
US 9621781B2 · On et al. · 2017 [cited by applicant]
US 9883198B2 · Puri et al. · 2018 [cited by applicant]
US 10311302B2 · Kottenstette et al. · 2019 [cited by applicant]
US 10311334B1 · Florez Choque et al. · 2019 [cited by applicant]
US 10339685B2 · Fu et al. · 2019 [cited by applicant]
US 10346740B2 · Zhang et al. · 2019 [cited by applicant]
US 10373055B1 · Matthey-de-l'Endroit et al. · 2019 [cited by applicant]
US 10373056B1 · Andoni et al. · 2019 [cited by applicant]
US 10380484B2 · Goel et al. · 2019 [cited by applicant]
US 10426442B1 · Schnorr · 2019 [cited by applicant]
US 10499081B1 · Wang et al. · 2019 [cited by applicant]
US 10521902B2 · Avendi et al. · 2019 [cited by applicant]
US 10600184B2 · Golden et al. · 2020 [cited by applicant]
US 10607135B2 · Zhang et al. · 2020 [cited by applicant]
US 10740901B2 · Myronenko · 2020 [cited by examiner]
US 20060072799A1 · McLain · 2006 [cited by applicant]
US 20170076438A1 · Kottenstette et al. · 2017 [cited by applicant]
US 20170249739A1 · Kallenberg et al. · 2017 [cited by applicant]
US 20180239949A1 · Chander et al. · 2018 [cited by applicant]
US 20190073563A1 · Chapados et al. · 2019 [cited by applicant]
US 20190220691A1 · Valpola et al. · 2019 [cited by applicant]
US 20190223725A1 · Lu et al. · 2019 [cited by applicant]
Gelder et al., Jun. 27, 2018, “Autoencoders for Multi-Label Prostate MR Segmentation” (pp. 1-6). (Year: 2018). [cited by examiner]
Chen et al, “3D intracranial artery segmentation using a convolutional autoencoder” (pp. 714-717) (Year: 2017). [cited by examiner]
Oktay et al., “Anatomically Constrained Neural Networks (ACNN): Application to Cardiac Image Enhancement and Segmentation”. (pp. 1-13) (Year: 2017). [cited by examiner]
Abadi et al., “Tensorflow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,” Nov. 9, 2015, 19 pages. [cited by applicant]
Bakas et al., “Advancing the Cancer Genome Atlas Glioma MRI Collections with Expert Segmentation Labels and Radiomic Features,” Scientic Data, 2017, 13 pages. [cited by applicant]
Bakas et al., “Segmentation Labels and Radiomic Features for the Pre-Operative Scans of the TCGA-GBM Collection,” The Cancer Imaging Archive, 2017, 2 pages. [cited by applicant]
Bakas et al., “Segmentation Labels and Radiomic features for the Pre-operative Scans of the TCGA-LGG Collection,” The Cancer Imaging Archive, 2017, 2 pages. [cited by applicant]
Chen et al., “Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation,” Aug. 22, 2018, 18 pages. [cited by applicant]
Chollet, “Xception: Deep Learning with Depthwise Separable Convolutions.” IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2017, 8 pages. [cited by applicant]
Doersch, “Tutorial on Variational Autoencoders,” Stat, arXiv, Aug. 16, 2016, 23 pages. [cited by applicant]
He et al., “Identity Mappings in Deep Residual Networks,” European Conference on Computer Vision, Jul. 25, 2016, 15 pages. [cited by applicant]
He et al., “Mask R-CNN,” ICCV, 2017, 9 pages. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks,” In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, 9 pages. [cited by applicant]
Isensee et al., “No New-Net,” International Conference on Medical Image Computing and Computer Assisted Intervention, Sep. 27, 2018, 10 pages. [cited by applicant]
Kamnitsas et al., “Efficient Multi-scale 3D CNN with fully connected CRF for accurate brain lesion segmentation,” Medical Image Analysis, 2016, 18 pages. [cited by applicant]
Kamnitsas et al., “Ensembles of Multiple Models and Architectures for Robust Brain Tumour Segmentation,” 2017, 12 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” May 1, 2014, 14 pages. [cited by applicant]
Long, “Fully Convolutional Networks for Semantic Segmentation,” CVPR, 2015, 10 pages. [cited by applicant]
Menze et al., “The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS),” IEEE Transactions on Medical Imaging, vol. 34(10), Oct. 10, 2015, 32 pages. [cited by applicant]
Milletari et al., “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” Fourth International Conference on 3D Vision (3DV), Oct. 25, 2016, 11 pages. [cited by applicant]
Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, Oct. 5, 2015, 8 pages. [cited by applicant]
Wang et al., “Automatic Brain Tumor Segmentation using Cascaded Anisotropic Convolutional Neural Networks,” 2017, 13 pages. [cited by applicant]
Wu et al., “Group Normalization,” ECCV, 2018, 17 pages. [cited by applicant]
Zhou et al., “Learning Contextual and Attentive Information for Brain Tumor Segmentation,” 2019, 11 pages. [cited by applicant]
Mckinley et al., “Ensembles of Densely-Connected CNNs with Label-Uncertainty for Brain Tumor Segmentation,” International Conference on Medical Image Computing and Computer Assisted Intervention, 2018, 10 pages. [cited by applicant]