Multi-modality and multi-scale feature aggregation for synthesizing SPECT image from fast SPECT scan and CT image
A computer-implemented method is provided for improving image quality. The method comprises: acquiring, using single-photon emission computed tomography (SPECT), a medical image of a subject, wherein the medical image is acquired with shortened acquisition time; and applying a deep learning network model to the medical image to generate an enhanced medical image.
1 . A computer-implemented method for improving image quality comprising:
(a) acquiring, using single-photon emission computed tomography (SPECT), a first medical image of a subject, wherein the first medical image is acquired with shortened acquisition time;
(b) combining the medical image acquired using SPECT with a second medical image acquired using computed tomography (CT) to generate a multi-modal input image; and
(c) applying a deep learning network model to the multi-modal input image and outputting an enhanced medical image, wherein the enhanced medical image has an image quality same as a SPECT image acquired with an acquisition time longer than the shortened acquisition time combined with a corresponding CT image or has a quantification accuracy improved over the first medical image, and wherein the deep learning network model comprises a U 2 -Net architecture including a plurality of residual blocks with different sizes to provide contextual information at different scales.
2 . The computer-implemented method of claim 1 , wherein an output generated by a decoder stage of the U 2 -Net architecture is fused with the multi-modal input image to generate the enhanced medical image.
3 . The computer-implemented method of claim 1 , wherein the deep learning network model is trained using training data comprising a SPECT image acquired using shortened acquisition time, a corresponding CT image and a SPECT image acquired using a standard acquisition time.
4 . The computer-implemented method of claim 1 , wherein the quantification accuracy is improved by training deep learning network model using a lesion attention mask.
5 . The computer-implemented method of claim 4 , wherein the lesion attention mask is included in a loss function of the deep learning network model.
6 . The computer-implemented method of claim 4 , wherein the lesion attention mask is generated from a SPECT image acquired using shortened acquisition time in the training data.
7 . The computer-implemented method of claim 6 , wherein the lesion attention mask is generated by filtering the SPECT image acquired using shortened acquisition time with a standardized uptake value (SUV) threshold.
8 . The computer-implemented method of claim 1 , wherein the first medical image and the second medical image are acquired using a SPECT/CT scanner.
9 . The computer-implemented method of claim 1 , wherein the enhanced medical image has an improved signal-noise ratio.
10 . A non-transitory computer-readable storage medium including instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:
(a) acquiring, using single-photon emission computed tomography (SPECT), a first medical image of a subject, wherein the first medical image is acquired with shortened acquisition time;
(b) combining the medical image acquired using SPECT with a second medical image acquired using computed tomography (CT) to generate a multi-modal input image; and
(c) applying a deep learning network model to the multi-modal input image and outputting an enhanced medical image, wherein the enhanced medical image has an image quality same as a SPECT image acquired with an acquisition time longer than the shortened acquisition time combined with a corresponding CT image or has a quantification accuracy improved over the first medical image, and wherein the deep learning network model comprises a U 2 -Net architecture including a plurality of residual blocks with different sizes to provide contextual information at different scales.
11 . The non-transitory computer-readable storage medium of claim 10 , wherein an output generated by a decoder stage of the U 2 -Net architecture is fused with the multi-modal input image to generate the enhanced medical image.
12 . The non-transitory computer-readable storage medium of claim 10 , wherein the deep learning network model is trained using training data comprising a SPECT image acquired using shortened acquisition time, a corresponding CT image and a SPECT image acquired using a standard acquisition time.
13 . The non-transitory computer-readable storage medium of claim 10 , wherein the quantification accuracy is improved by training deep learning network model using a lesion attention mask.
14 . The non-transitory computer-readable storage medium of claim 13 , wherein the lesion attention mask is included in a loss function of the deep learning network model.
15 . The non-transitory computer-readable storage medium of claim 13 , wherein the lesion attention mask is generated from a SPECT image acquired using shortened acquisition time in the training data.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the lesion attention mask is generated by filtering the SPECT image acquired using shortened acquisition time with a standardized uptake value (SUV) threshold.
17 . The non-transitory computer-readable storage medium of claim 10 , wherein the first medical image and the second medical image are acquired using a SPECT/CT scanner.
18 . The non-transitory computer-readable storage medium of claim 10 , wherein the enhanced medical image has an improved signal-noise ratio.