IP Library › Granted Patent US 12,573,028
Granted Patent B2
US 12,573,028 · App. 16/540,717 · Granted Mar 10, 2026

Neural network for image registration and image segmentation trained using a registration simulator

Inventors: Wentao Zhu (Mountain View, CA); Daguang Xu (Potomac, MD); Andriy Myronenko (San Mateo, CA); Ziyue Xu (Reston, VA)
Assignee: NVIDIA Corporation
G06T7/0012G06F30/20G06N3/08G06T7/11G06T7/344G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,028
App. No.
16/540,717
Granted
Mar 10, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to perform registration among images. In at least one embodiment, one or more neural networks are trained to indicate registration of features in common among at least two images by generating a first correspondence by simulating a registration process of registering an image and applying the at least two images and the first correspondence to a neural network to derive a second correspondence of the features in common among the at least two images.

Claims (82)

1 . A method, comprising:

using a processor comprising one or more circuits to;

compare a first image generated based, at least in part, on one or more pixel displacement values with a second image generated based, at least in part, on one or more other pixel displacement values; and

update a neural network based, at least in part, on a result of the comparison between the first image and the second image.

2 . The method of claim 1 , further comprising training the neural network to generate the first third image, wherein training the neural network comprises:

selecting a set of sampled transformation parameters from among a range of possible transformation parameters;

generating, from the set of sampled transformation parameters, a first displacement field including the one or more other pixel displacement values, the one or more other pixel displacement values indicating one or more displacements of pixels in a training image displaced according to the set of sampled transformation parameters;

transforming the training image into the second image according to the first displacement field;

generating, using the neural network, a second displacement field from the training image and the second image;

transforming the training image into the first image according to the second displacement field;

generating a loss function from a first difference between the first displacement field and the second displacement field; and

using the loss function to train the neural network.

3 . The method of claim 2 , wherein generating the loss function comprises generating the loss function from the first difference and from a second difference between the second image and the first image, wherein the loss function is a weighted sum of the first difference and the second difference.

4 . The method of claim 3 , wherein the first difference is a mean value over an image coordinate space of a norm of differences of elements between the first displacement field and the second displacement field, and wherein the second difference is a negative normalized cross-correlation between the second image and the first image.

5 . The method of claim 2 , wherein the range of possible transformation parameters to transform the training image into the second image comprises one or more of a range of rotation angles, a range of translations, a range of scale factors, and a range of elastic distortions, and wherein selecting the set of sampled transformation parameters comprises selecting a rotation angle within the range of rotation angles, selecting a translation within the range of translations, selecting a scale factor within the range of scale factors, and/or selecting an elastic distortion within the range of elastic distortions.

6 . The method of claim 2 , wherein the first training image is a two-dimensional or three-dimensional image of a biological environment.

7 . The method of claim 1 , further comprising training the neural network to generate the first image based, at least in part, on the one or more other pixel displacement values, wherein the training is performed in a supervised manner.

8 . A non-transitory computer-readable storage media storing executable instructions that, when executed by one or more processors of a computer system, cause the computer system to:

compare a first image generated based, at least in part, on one or more pixel displacement values with a second image generated based, at least in part, on one or more other pixel displacement values; and

update a neural network based, at least in part, on a result of the comparison between the first image and the second image.

9 . The non-transitory computer-readable storage media of claim 8 , wherein the executable instructions, when executed by the one or more processors of the computer system, further cause the computer system to train the neural network to generate the first image, wherein training the neural network comprises:

selecting a set of sampled transformation parameters from among a range of possible transformation parameters;

generating, from the set of sampled transformation parameters, a first displacement field including the other one or more pixel displacement values, the one or more other pixel displacement values indicating one or more displacements of pixels in a training image displaced according to the set of sampled transformation parameters;

transforming the training image into the second image according to the displacement field;

generating, using the neural network, a second displacement field from the training image and the second image;

transforming the training image into the first image according to the second displacement field;

generating a loss function from a first difference between the first displacement field and the second displacement field; and

using the loss function to train the neural network.

10 . The non-transitory computer-readable storage media of claim 9 , wherein generating the loss function comprises generating the loss function from the first difference and from a second difference between the second image and the first image, wherein the loss function is a weighted sum of the first difference and the second difference.

11 . The non-transitory computer-readable storage media of claim 10 , wherein the first difference is a mean value over an image coordinate space of a norm of differences of elements between the first displacement field and the second displacement field, and wherein the second difference is a negative normalized cross-correlation between the second image and the first image.

12 . The non-transitory computer-readable storage media of claim 9 , wherein the range of possible transformation parameters to transform the training image into the second image comprises one or more of a range of rotation angles, a range of translations, a range of scale factors, and a range of elastic distortions, and wherein selecting the set of sampled transformation parameters comprises selecting a rotation angle within the range of rotation angles, selecting a translation within the range of translations, selecting a scale factor within the range of scale factors, and/or selecting an elastic distortion within the range of elastic distortions.

13 . The non-transitory computer-readable storage media of claim 9 , wherein the training image is a two-dimensional or three-dimensional image of a biological environment.

14 . A processor, comprising:

one or more circuits to:

compare a first image generated based, at least in part, on one or more pixel displacement values with a second image generated based, at least in part, on one or more other pixel displacement values; and

update a neural network based, at least in part, on a result of the comparison between the first image and the second image.

15 . The processor of claim 14 , wherein the one or more circuits are further to train the neural network to generate the first image, wherein training the neural network comprises:

selecting a set of sampled transformation parameters from among a range of possible transformation parameters;

generating, from the set of sampled transformation parameters, a first displacement field including the one or more other pixel displacement values, the one or more other pixel displacement values indicating one or more displacements of pixels in a training image displaced according to the set of sampled transformation parameters;

transforming the training image into the second image according to the displacement field;

generating, using the neural network, a second displacement field from the training image and the second image;

transforming the training image into the first image according to the second displacement field;

generating a loss function from a first difference between the first displacement field and the second displacement field; and

using the loss function to train the neural network.

16 . The processor of claim 15 , wherein generating the loss function comprises generating the loss function from the first difference and from a second difference between the second image and the first image, wherein the loss function is a weighted sum of the first difference and the second difference.

17 . The processor of claim 16 , wherein the first difference is a mean value over an image coordinate space of a norm of differences of elements between the first displacement field and the second displacement field, and wherein the second difference is a negative normalized cross-correlation between the second image and the first image.

18 . The processor of claim 15 , wherein the range of possible transformation parameters to transform the training image into the second image comprises one or more of a range of rotation angles, a range of translations, a range of scale factors, and a range of elastic distortions, and wherein selecting the set of sampled transformation parameters comprises selecting a rotation angle within the range of rotation angles, selecting a translation within the range of translations, selecting a scale factor within the range of scale factors, and/or selecting an elastic distortion within the range of elastic distortions.

19 . The processor of claim 15 , wherein the training image is a two-dimensional or three-dimensional image of a biological environment.

20 . A computer system, comprising:

one or more processors and a computer-readable memory storing executable instructions that, when executed by the one or more processors, cause the computer system to:

compare a first image generated based, at least in part, on one or more pixel displacement values with a second image generated based, at least in part, on one or more other pixel displacement values; and

update one or more neural networks based, at least in part, on a result of the comparison between the first image and the second image.

21 . The computer system of claim 20 , wherein:

the executable instructions, when executed by the one or more processors, further cause the computer system to calculate parameters corresponding to the one or more neural networks, the calculation based, at least in part, on the comparison of the first image with the second image and providing a loss function for the one or more neural networks to optimize, the loss function including:

a first function of differences between a simulated displacement field and an output displacement field that the one or more neural networks outputs based on attempted registration of a training image and the second image; and

a second function of a first difference function of differences between a first displacement field including the one or more other pixel displacement values, the one or more other pixel displacement values indicating one or more displacements of pixels in the training image displaced according to a set of sampled transformation parameters, and a second displacement field, wherein the second displacement field is generated by the one or more neural networks from the output displacement field;

the computer system further comprises:

storage for the parameters;

storage for the set of sampled transformation parameters selected from among a range of possible transformation parameters; and

storage for the first displacement field; and

the parameters calculated based, at least in part, on a third function optimizing for the loss function.

22 . The computer system of claim 21 , wherein:

the loss function is a weighted sum of the first difference function and a second difference function of differences between the second image and the first image, the first image transformed from the training image according to the second displacement field; and

the first difference function is a mean value over an image coordinate space of a norm of differences of elements between the first displacement field and the second displacement field, and wherein the second difference function is a negative normalized cross-correlation between the second image and the first image.

23 . The computer system of claim 21 , wherein the range of possible transformation parameters to transform the training image into the second image comprises one or more of a range of rotation angles, a range of translations, a range of scale factors, and a range of elastic distortions, and wherein selecting the set of sampled transformation parameters comprises selecting an angle within the range of rotation angles, selecting a translation within the range of translations, selecting a scale factor within the range of scale factors, and/or selecting an elastic distortion within the range of elastic distortions.

24 . The computer system of claim 21 , wherein the training image is a two-dimensional or three-dimensional image of a biological environment.

25 . The computer system of claim 21 , further comprising:

a convolutional layer to modify a last neural network layer that provides a convolution of a last layer of the one or more neural networks and a predicted segmentation; and

a softmax activation layer to modify the predicted segmentation according to a softmax operation on the convolution.

26 . The computer system of claim 20 , wherein the executable instructions, when executed by the one or more processors, further cause the computer system to train the one or more neural networks to generate the first image, wherein training the one or more neural networks comprises:

selecting a set of sampled transformation parameters from among a range of possible transformation parameters;

generating, from the set of sampled transformation parameters, a first displacement field including the one or more other pixel displacement values, the one or more other pixel displacement values indicating one or more displacements of pixels in a training image displaced according to the set of sampled transformation parameters;

transforming the training image into the second image according to the first displacement field;

generating, using the one or more neural networks, a second displacement field from the training image and the second image;

transforming the training image into the first image according to the second displacement field;

generating a loss function from a first difference between the first displacement field and the second displacement field; and

using the loss function to train the one or more neural networks.

27 . The computer system of claim 26 , wherein:

generating the loss function comprises generating the loss function from the first difference and from a second difference between the second image and the first image, wherein the loss function is a weighted sum of the first difference and the second difference; and

the first difference is a mean value over an image coordinate space of a norm of differences of elements between the first displacement field and the second displacement field, and wherein the second difference is a negative normalized cross-correlation between the second image and the first image.

28 . The computer system of claim 26 , wherein the range of possible transformation parameters to transform the training image into the second image comprises one or more of a range of rotation angles, a range of translations, a range of scale factors, and a range of elastic distortions, and wherein selecting the set of sampled transformation parameters comprises selecting a rotation angle within the range of rotation angles, selecting a translation within the range of translations, selecting a scale factor within the range of scale factors, and/or selecting an elastic distortion within the range of elastic distortions.

29 . The computer system of claim 26 , wherein the training image is a two-dimensional or three-dimensional image of a biological environment.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded Jan 4, 2024
From: ZHU, WENTAO; XU, DAGUANG; MYRONENKO, ANDRIY; XU, ZIYUE
To: NVIDIA CORPORATION
Reel/Frame 066206/0417 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2019
From: ZHU, WENTAO; XU, DAGUANG; MYRONENKO, ANDRIY; XU, ZIYUE
To: NVIDIA CORPORATION
Reel/Frame 050438/0242 →
Continuity (1)
Related Publication 20210049757A1 · Feb 18, 2021
References Cited (63)
US 10318889B2 · Xu · 2019 [cited by applicant]
US 11550011B2 · Zhang et al. · 2023 [cited by applicant]
US 11557036B2 · Liao · 2023 [cited by examiner]
US 11593632B2 · Rippel · 2023 [cited by examiner]
US 20170178340A1 · Schadewaldt · 2017 [cited by examiner]
US 20170287109A1 · Tasfi · 2017 [cited by examiner]
US 20170336713A1 · Middlebrooks · 2017 [cited by examiner]
US 20170337682A1 · Liao · 2017 [cited by examiner]
US 20170345138A1 · Middlebrooks · 2017 [cited by examiner]
US 20180235577A1 · Buerger · 2018 [cited by examiner]
US 20180300850A1 · Johnson · 2018 [cited by examiner]
US 20180314716A1 · Kim · 2018 [cited by examiner]
US 20180373999A1 · Xu · 2018 [cited by examiner]
US 20190050999A1 · Piat · 2019 [cited by examiner]
US 20190164290A1 · Wang · 2019 [cited by examiner]
US 20190197662A1 · Sloan et al. · 2019 [cited by applicant]
US 20190205606A1 · Zhou · 2019 [cited by examiner]
US 20190205766A1 · Krebs et al. · 2019 [cited by applicant]
US 20190251401A1 · Shechtman · 2019 [cited by examiner]
US 20190385282A1 · Sasaki · 2019 [cited by examiner]
US 20200134366A1 · Xu · 2020 [cited by examiner]
US 20200151309A1 · Thuillier · 2020 [cited by examiner]
US 20200184660A1 · Shi · 2020 [cited by examiner]
US 20200202515A1 · Prasad · 2020 [cited by examiner]
US 20200226731A1 · Hu · 2020 [cited by examiner]
US 20200242736A1 · Liu · 2020 [cited by examiner]
US 20200250462A1 · Yang · 2020 [cited by examiner]
US 20200285888A1 · Borar · 2020 [cited by examiner]
US 20200285911A1 · Guo · 2020 [cited by examiner]
US 20200286229A1 · Ogino · 2020 [cited by examiner]
US 20200311923A1 · Walton · 2020 [cited by examiner]
US 20200327363A1 · Chen · 2020 [cited by examiner]
US 20200327639A1 · Walvoord · 2020 [cited by examiner]
US 20200342276A1 · Lee · 2020 [cited by examiner]
US 20200356827A1 · Dinerstein · 2020 [cited by examiner]
US 20200411164A1 · Donner · 2020 [cited by examiner]
US 20210049733A1 · Kang · 2021 [cited by examiner]
US 20210056343A1 · Toizumi · 2021 [cited by examiner]
US 20210174543A1 · Claessen · 2021 [cited by examiner]
US 20210192758A1 · Song · 2021 [cited by examiner]
US 20210201527A1 · Cai · 2021 [cited by examiner]
US 20210209775A1 · Song · 2021 [cited by examiner]
US 20210216878A1 · Norman · 2021 [cited by examiner]
US 20210224590A1 · Matsumoto · 2021 [cited by examiner]
US 20210248763A1 · Gao · 2021 [cited by examiner]
US 20210327041A1 · Asendorf · 2021 [cited by examiner]
US 20210397889A1 · Gong · 2021 [cited by examiner]
US 20220012846A1 · Dorta · 2022 [cited by examiner]
US 20220254071A1 · Ojha · 2022 [cited by examiner]
US 20230260141A1 · Schaefferkoetter · 2023 [cited by applicant]
CN 109272443A · 2019 [cited by applicant]
“An Unsupervised Learning Model for Deformable Medical Image Registration”; Balakrishnan, Guha, arXiv:1802.02604v3 [cs.CV] Apr. 20, 2018 (Year: 2018). [cited by examiner]
Balakrishnan et al., “VoxelMorph: A Learning Framework for Deformable Medical Image Registration,” IEEE Transactions on Medical Imaging 38(8):1788-1800, Feb. 4, 2019. [cited by applicant]
IEEE Computer Society, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” The Institute of Electrical and Electronics Engineers, Inc., Aug. 29, 2008, 70 pages. [cited by applicant]
International Electrotechnical Commission, “Functional safety of electrical/electronic/programmableelectronic safety-related systems,” IEC Standard 61508-1, Apr. 2014, 23 pages. [cited by applicant]
International Organization for Standardization, “Road vehicles—Functional safety,” ISO Standard 26262, https://www.iso.org/obp/ui/#/iso:std:iso:26262:-1:ed-1:v1:en, Nov. 11, 2011, 35 pages. [cited by applicant]
International Search Report and Written Opinion, mailed Oct. 23, 2020, in International Patent Application No. PCT/US2020/046252, 20 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Sokooti et al., “Nonrigid Image Registration Using Multi-scale 3D Convolutional Neural Networks,” Medical Image Computing and Computer Assisted Intervention (MICCAI '17), Sep. 11, 2017, 8 pages. [cited by applicant]
Zhu et al., “NeurReg: Neural Registration and Its Application to Image Segmentation,” Oct. 4, 2019, retrieved Jan. 4, 2021 from https://arxiv.org/pdf/1910.01763.pdf, 10 pages. [cited by applicant]
Office Action mailed May 23, 2025 in Chinese Patent Application No. 202080070936.5, NVIDIA Corporation, 47 pages (including translation). [cited by applicant]
Zhang Jinya, “Research on non-rigid medical image registration technology,” A Dissertation for doctor's degree, Univ Science & Technology of China, Mar. 2015. [cited by applicant]