IP Library › Granted Patent US 12,417,566
Granted Patent B2
US 12,417,566 · App. 18/050,564 · Granted Sep 16, 2025

Convolution and transformer based compressive sensing

Inventors: Sam Tak Wu Kwong (Kowloon, HK); Dongjie Ye (Kowloon, HK); Zhangkai Ni (Kowloon, HK)
Assignee: City University of Hong Kong
G06T11/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,566
App. No.
18/050,564
Granted
Sep 16, 2025
Kind
B2
Abstract

A method for adaptive reconstruction of a compressively sensed data. The method contains the steps of receiving sensed data; conducting an initial reconstruction to the sensed data to obtain a plurality of first reconstruction patches; by a reconstruction module, conducting a progressive reconstruction to the sensed data to obtain a plurality of second reconstruction patches; summing the plurality of second reconstruction patches with the a plurality of first reconstruction patches to obtain final patches; and merging the final patches to obtain a reconstructed data. The progressive reconstruction further contains concatenating transformer features and convolution features to obtain the second reconstruction patches. The invention provides a hybrid network for adaptive sampling and reconstruction of CS, which integrates the advantages of leveraging both detailed spatial information from CNN and the global context provided by transformer for enhanced representation learning.

Claims (37)

1. A method for adaptive reconstruction of a compressively sensed data, comprising steps of:

a) receiving sensed data;

b) conducting an initial reconstruction to the sensed data to obtain a plurality of first reconstruction patches;

c) by a reconstruction module, conducting a progressive reconstruction to the sensed data to obtain a plurality of second reconstruction patches;

d) summing the plurality of second reconstruction patches with the a plurality of first reconstruction patches to obtain final patches; and

e) merging the final patches to obtain a reconstructed data,

wherein the progressive reconstruction further comprises concatenating transformer features and convolution features to obtain the second reconstruction patches.

2. The method according to claim 1 , wherein the reconstruction module comprises a Convolutional Neural Network (CNN) stem for producing the convolution features, and a transformer stem for producing the transformer features.

3. The method according to claim 2 , wherein the transformer stem comprises a first transformer block and a second transformer block; the CNN stem comprising a first convolution block corresponding to the first transformer block, and a second convolution block corresponding to the second transformer block; and wherein Step c) further comprising steps of:

f) generating a first transformer feature of the transformer features at the first transformer block based on the sensed data and an output of the first convolution block;

g) generating a second transformer feature of the transformer features at the second transformer block based on the first transformer feature and an output of the second convolution block.

4. The method of claim 3 , wherein at least one of the first and second convolution blocks comprises a plurality of convolution layers, followed by a leaky rectified linear unit (ReLU) and a batch norm layer.

5. The method of claim 3 , wherein at least one of the first and second transformer blocks is a window-based transformer.

6. The method of claim 5 , wherein at least one of the first and second transformer blocks comprises a multi-head self-attention (MSA) module, followed by a multi-layer perceptron (MLP) module.

7. The method according to claim 2 , wherein the reconstruction module further comprises an input projection module before the CNN stem and the transformer stem; and wherein Step c) further comprising a step of increasing a dimension of the sensed data inputted to the reconstruction module by the input projection module.

8. The method according to claim 7 , wherein the input projection module comprises a plurality of 1×1 convolution layers and a sub-pixel convolution layer.

9. The method according to claim 2 , wherein the reconstruction module further comprises an output projection module after the transformer stem; and wherein Step c) further comprising a step of projecting the transformer features into a single channel to obtain the plurality of second reconstruction patches.

10. The method according to claim 9 , wherein the output projection module comprises a plurality of convolution layers followed by a tanh action function.

11. The method according to claim 1 , wherein the step of conducting an initial reconstruction is performed on a linear initialization module.

12. The method according to claim 11 , wherein the linear initialization module comprises a 1×1 convolution layer and a sub-pixel convolution layer.

13. The method according to claim 1 , wherein the sensed data comprises a plurality of input convolutional patches.

14. An apparatus for adaptive reconstruction of a compressively sensed data, comprising:

a) one or more processors; and

b) a memory storing computer-executable instructions that, when executed, cause the one or more processors to

i) receive a sensed data;

ii) conduct an initial reconstruction to the sensed data to obtain a plurality of first reconstruction patches;

iii) conduct a progressive reconstruction to the sensed data to obtain a plurality of second reconstruction patches;

iv) sum the plurality of second reconstruction patches with the a plurality of first reconstruction patches to obtain final patches; and

v) merge the final patches to obtain a reconstructed data,

wherein the progressive reconstruction further comprises concatenating transformer features and convolution features to obtain the second reconstruction patches.

15. A non-transitory computer readable medium, comprising executable instructions that, when executed by at least one processor, direct the at least one processor to perform a method, the method comprising:

a) receiving a sensed data;

b) conducting an initial reconstruction to the sensed data to obtain a plurality of first reconstruction patches;

c) by a reconstruction module, conducting a progressive reconstruction to the sensed data to obtain a plurality of second reconstruction patches;

d) summing the plurality of second reconstruction patches with the a plurality of first reconstruction patches to obtain final patches; and

e) merging the final patches to obtain a reconstructed data,

wherein the progressive reconstruction further comprises concatenating transformer features and convolution features to obtain the second reconstruction patches.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: KWONG, SAM TAK WU; YE, DONGJIE; NI, ZHANGKAI
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 061576/0550 →
Continuity (1)
Related Publication 20240153161A1 · May 9, 2024
References Cited (84)
US 8666180B2 · Sen et al. · 2014 [cited by applicant]
US 9373159B2 · Amroabadi et al. · 2016 [cited by applicant]
US 9892527B2 · Ye · 2018 [cited by examiner]
US 11270477B2 · Nett · 2022 [cited by examiner]
US 11816765B2 · Andrew · 2023 [cited by examiner]
US 11875813B2 · Mesgarani · 2024 [cited by examiner]
US 12010302B2 · Kim · 2024 [cited by examiner]
US 12272375B2 · Ozturk · 2025 [cited by examiner]
US 20140205170A1 · Cierniak · 2014 [cited by examiner]
US 20170169586A1 · Ding · 2017 [cited by examiner]
US 20200234406A1 · Ren · 2020 [cited by examiner]
US 20210193112A1 · Cui · 2021 [cited by examiner]
US 20220277452A1 · Vaidyanathan · 2022 [cited by examiner]
US 20230104757A1 · Pramod · 2023 [cited by examiner]
US 20230109260A1 · Bald · 2023 [cited by examiner]
US 20240070853A1 · Yoo · 2024 [cited by examiner]
US 20240081705A1 · Zhao · 2024 [cited by examiner]
US 20250087041A1 · Linari · 2025 [cited by examiner]
CN 103595414 · 2014 [cited by applicant]
Dongjie Ye, Graduate Student Member, IEEE, Zhangkai Ni, Graduate Student Member, IEEE, Hanli Wang, Senior Member, IEEE, Jian Zhang, Member, IEEE, Shiqi Wang, Senior Member, IEEE, and Sam Kwong, Fellow, IEEE, CSformer: B… [cited by examiner]
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollar, and R. B. Girshick, “Early convolutions help transformers see better,” In Proceedings of the Advances in Neural Information Processing Systems, 2021. [cited by applicant]
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy, “Do vision transformers see like convolutional neural networks?” in Proceedings of the Advances in Neural Information Processing Systems, 2021. [cited by applicant]
D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in Proceedings of the IEEE Inter… [cited by applicant]
M. Bevilacqua, A. Roumy, C. Guillemot, and M.-L. A. Morel, “Low-complexity single-image super-resolution based on honnegative neighbor embedding,” in Proceedings of the British Machine Vision Conference, 2012, pp. 1-10. [cited by applicant]
R. Zeyde, M. Elad, and M. Protter, “On single image scale-up using sparse-representations,” in Proceedings of the International Conference on Curves and Surfaces, 2010, pp. 711-730. [cited by applicant]
J.-B. Huang, A. Singh, and N. Ahuja, “Single image super-resolution from transformed self-exemplars,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 5197-5206. [cited by applicant]
Z. Ni, H. Zeng, L. Ma, J. Hou, J. Chen, and K.-K. Ma, “A gabor feature-based quality assessment model for the screen content images,” IEEE Transactions on Image Processing, vol. 27, No. 9, pp. 4516-4528, 2018. [cited by applicant]
P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik, “Contour detection and hierarchical image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, No. 5, pp. 898-916, 2011. [cited by applicant]
K. G. Derpanis and R. Wildes, “Spacetime texture representation and recognition based on a spatiotemporal orientation analysis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, No. 6, pp. 1193-1… [cited by applicant]
P. Saisan, G. Doretto, Y. N. Wu, and S. Soatto, “Dynamic texture recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, vol. 2, 2001, pp. II-II. [cited by applicant]
R. Peteri, S. Fazekas, and M. J. Huiskes, “Dyntex: A comprehensive database of dynamic textures,” Pattern Recognition Letters, vol. 31, No. 12, pp. 1627-1632, 2010. [cited by applicant]
I. Hadji and R. P. Wildes, “A new large scale dynamic texture dataset with application to convnet understanding,” in Proceedings of the European Conference on Computer Vision, 2018, pp. 320-335. [cited by applicant]
Y. Fang, H. Zhu, K. Ma, Z. Wang, and S. Li, “Perceptual evaluation for multi-exposure image fusion of dynamic scenes,” IEEE Transactions on Image Processing, vol. 29, No. 1, pp. 1127-1138, 2020. [cited by applicant]
S. Kornblith, M. Norouzi, H. Lee, and G. Hinton, “Similarity of neural network representations revisited,” in Proceedings of the International Conference on Machine Learning, 2019, pp. 3519-3529. [cited by applicant]
D. L. Donoho, “Compressed sensing,” IEEE Transactions on Information Theory, vol. 52, No. 4, pp. 1289-1306, 2006. [cited by applicant]
J. Zhang, D. Zhao, C. Zhao, R. Xiong, S. Ma, and W. Gao, “Image compressive sensing recovery via collaborative sparsity,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 2, No. 3, pp. 380-391,… [cited by applicant]
J. Zhang, C. Zhao, D. Zhao, and W. Gao, “Image compressive sensing recovery using adaptively learned sparsifying basis via I0 minimization,” Signal Processing, vol. 103, pp. 114-126, 2014. [cited by applicant]
M. E. Ahsen and M. Vidyasagar, “Error bounds for compressed sensing algorithms with group sparsity: A unified approach,” Applied and Computational Harmonic Analysis, vol. 43, No. 2, pp. 212-232, 2017. [cited by applicant]
K. Kulkarni, S. Lohit, p. Turaga, R. Kerviche, and A. Ashok, “Reconnet: Non-iterative reconstruction of images from compressively sensed measurements,” in Proceedings of the IEEE Conference on Computer Vision and Patter… [cited by applicant]
C. A. Metzler, A. Mousavi, and R. G. Baraniuk, “Learned d-amp: Principled! neural network based compressive image recovery,” in Proceedings of the Advances in Neural Information Processing Systems, 2017, pp. 1773-1784. [cited by applicant]
W. Shi, F. Jiang, S. Zhang, and D. Zhao, “Deep networks for compressed image sensing,” in Proceedings of the IEEE International Conference on Multimedia and Expo, 2017, pp. 877-882. [cited by applicant]
J. Zhang and B. Ghanem, “Ista-net: Interpretable optimization-inspired deep network for image compressive sensing,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 1828-1837. [cited by applicant]
Y. Yang, J. Sun, H. Li, and Z. Xu, “Admm-csnet: A deep learning approach for image compressive sensing,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, No. 3, pp. 521-538, 2018. [cited by applicant]
M. Kabkab, P. Samangouei, and R. Chellappa, “Task-aware compressed sensing with generative adversarial networks,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, No. 1, 2018. [cited by applicant]
W. Shi, F. Jiang, S. Liu, and D. Zhao, “Image compressed sensing using convolutional neural network,” IEEE Transactions on Image Processing, vol. 29, pp. 375-388, 2020. [cited by applicant]
J. Zhang, C. Zhao, and W. Gao, “Optimization-inspired compact deep compressive sensing,” IEEE Journal of Selected Topics in Signal Processing, vol. 14, No. 4, pp. 765-774, 2020. [cited by applicant]
Z. Zhang, Y. Liu, J. Liu, F. Wen, and C. Zhu, “Amp-net: Denoising-based deep unfolding for compressive image sensing,” IEEE Transactions on Image Processing, vol. 30, pp. 1487-1500, 2020. [cited by applicant]
D. You, J. Zhang, J. Xie, B. Chen, and S. Ma, “Coast: Controllable arbitrary-sampling network for compressive sensing,” IEEE Transactions on Image Processing, vol. 30, pp. 6066-6080, 2021. [cited by applicant]
D. You, J. Xie, and J. Zhang, “Ista-net++: Flexible deep unfolding network for compressive sensing,” in Proceedings of the IEEE International Conference on Multimedia and Expo, 2021, pp. 1-6. [cited by applicant]
Y. Sun, J. Chen, Q. Liu, and G. Liu, “Learning image compressed sensing with sub-pixel convolutional generative adversarial network,” Pattern Recognition, vol. 98, p. 107051, 2020. [cited by applicant]
J. Chen, Y. Sun, Q. Liu, and R. Huang, “Learning memory augmented cascading network for compressed sensing of mages,” in Proceedings of the European Conference on Computer Vision, 2020, pp. 513-529. [cited by applicant]
Y. Sun, Y. Yang, Q. Liu, J. Chen, X.-T. Yuan, and G. Guo, “Learning non-locally regularized compressed sensing network with half-quadratic splitting,” IEEE Transactions on Multimedia, vol. 22, No. 12, pp. 3236-3248, 202… [cited by applicant]
A. Bora, A. Jalal, E. Price, and A. G. Dimakis, “Compressed sensing using generative models,” in Proceedings of the International Conference on Machine Learning, 2017, pp. 537-546. [cited by applicant]
A. Mousavi, G. Dasarathy, and R. G. Baraniuk, “Deepcodec: Adaptive sensing and recovery via deep convolutional neural networks,” arXiv preprint arXiv:1707.03386, 2017. [cited by applicant]
K. Xu, Z. Zhang, and F. Ren, “Lapran: A scalable laplacian pyramid reconstructive adversarial network for flexible compressive sensing reconstruction,” in Proceedings of the European Conference on Computer Vision, 2018,… [cited by applicant]
Y. Wu, M. Rosca, and T. Lillicrap, “Deep compressed sensing,” in Proceedings of the International Conference on Machine Learning. PMLR, 2019, pp. 6850-6860. [cited by applicant]
Y. Sun, J. Chen, Q. Liu, B. Liu, and G. Guo, “Dual-path attention network for compressed sensing image reconstruction,” IEEE Transactions on Image Processing, vol. 29, pp. 9482-9495, 2020. [cited by applicant]
H. Yao, F. Dai, S. Zhang, Y. Zhang, Q. Tian, and C. Xu, “Dr2-net: Deep residual reconstruction network for image compressive sensing,” Neurocomputing, vol. 359, pp. 483-493, 2019. [cited by applicant]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the Advances in Neural Information Processing Systems, 2017, pp. 5998-… [cited by applicant]
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image reco… [cited by applicant]
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. … [cited by applicant]
Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li, “Uformer: A general u-shaped transformer for image restoration,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 17683-176… [cited by applicant]
Y. Jiang, S. Chang, and Z. Wang, “Transgan: Two pure transformers can make one strong gan, and that can scale up,” vol. 34, 2021. [cited by applicant]
J. You, T. Ebrahimi, and A. Perkis, “Attention driven foveated video quality assessment,” IEEE Transactions on Image Processing, vol. 23, No. 1, pp. 200-213, 2014. [cited by applicant]
T. Chen, H. Liu, Z. Ma, Q. Shen, X. Cao, and Y. Wang, “End-to-end learnt image compression via non-local attention optimization and improved context modeling,” IEEE Transactions on Image Processing, vol. 30, pp. 3179-31… [cited by applicant]
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao, “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1-1… [cited by applicant]
S. S. Chen, D. L. Donoho, and M. A. Saunders, “Atomic decomposition by basis pursuit,” SIAM review, vol. 43, No. 1, pp. 129-159, 2001. [cited by applicant]
R. Tibshirani, “Regression shrinkage and selection via the lasso,” Journal of the Royal Statistical Society: Series B (Methodological), vol. 58, No. 1, pp. 267-288, 1996. [cited by applicant]
A. Beck and M. Teboulle, “A fast iterative shrinkage-thresholding algorithm for linear inverse problems,” SIAM Journal on Imaging Sciences, vol. 2, No. 1, pp. 183-202, 2009. [cited by applicant]
M. V. Afonso, J. M. Bioucas-Dias, and M. A. Figueiredo, “An augmented lagrangian approach to the constrained optimization formulation of imaging inverse problems,” IEEE Transactions on Image Processing, vol. 20, No. 3, … [cited by applicant]
C. Li, W. Yin, H. Jiang, and Y. Zhang, “An efficient augmented lagrangian method with applications to total variation minimization,” Computational Optimization and Applications, vol. 56, No. 3, pp. 507-530, 2013. [cited by applicant]
C. A. Metzler, A. Maleki, and R. G. Baraniuk, “From denoising to compressed sensing,” IEEE Transactions on Information Theory, vol. 62, No. 9, pp. 5117-5144, 2016. [cited by applicant]
D. L. Donoho, A. Maleki, and A. Montanari, “Message-passing algorithms for compressed sensing,” Proceedings of the National Academy of Sciences, vol. 106, No. 45, pp. 18914-18919, 2009. [cited by applicant]
M. Mardani, E. Gong, J. Y. Cheng, S. Vasanawala, G. Zaharchuk, M. Alley, N. Thakur, S. Han, W. Dally, J. M. Pauly et al., “Deep generative adversarial networks for compressed sensing automates mri,” arXiv preprint arXiv… [cited by applicant]
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” in Proceedings of the International Conference on Learning Representations, 2016. [cited by applicant]
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 7794-7803. [cited by applicant]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE International Conference on Computer Vision, 202… [cited by applicant]
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proceedings of the Advances in Neural Information Processing Systems, 2014, pp.… [cited by applicant]
Z. Ni, W. Yang, S. Wang, L. Ma, and S. Kwong, “Unpaired image enhancement with quality-attention generative adversarial network,” in Proceedings of the ACM International Conference on Multimedia, 2020, pp. 1697-1705. [cited by applicant]
Z. Ni, W. Yang, S. Wang, L. Ma, and S. Kwong, “Towards unsupervised deep image enhancement with generative adversarial network,” IEEE Transactions on Image Processing, vol. 29, pp. 9140-9151, 2020. [cited by applicant]
Y. Xie, J. Zhang, C. Shen, and Y. Xia, “Cotr: Efficiently bridging cnn and transformer for 3d medical image segmentation,” in Proceedings of the International conference on medical image computing and computer-assisted … [cited by applicant]
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” in Proceedings of the Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 3965-3977. [cited by applicant]
Z. Peng, W. Huang, S. Gu, L. Xie, Y. Wang, J. Jiao, and Q. Ye, “Conformer: Local features coupling global representations for visual recognition,” in Proceedings of the IEEE International Conference on Computer Vision, … [cited by applicant]
S. d'Ascoli, H. Touvron, M. L. Leavitt, A. S. Morcos, G. Biroli, and L. Sagun, “Convit: Improving vision transformers with soft convolutional inductive biases,” in Proceedings of the International Conference on Machine … [cited by applicant]