IP Library Granted Patent US 12,585,949
Granted Patent B2
US 12,585,949 · App. 17/826,606 · Granted Mar 24, 2026

System and method for designing efficient super resolution deep convolutional neural networks by cascade network training, cascade network trimming, and dilated convolutions

Inventors: Haoyu Ren (San Diego, CA); Mostafa El-Khamy (San Diego, CA); Jungwon Lee (San Diego, CA)
Assignee: Samsung Electronics Co., Ltd
G06N3/082G06N3/045G06N3/088G06T3/4053G06T5/70G06N3/047G06N3/048G06T7/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,949
App. No.
17/826,606
Granted
Mar 24, 2026
Kind
B2
Abstract

Apparatuses and methods of manufacturing same, systems, and methods are described. In one aspect, a method includes generating a convolutional neural network (CNN) by training a CNN having a plurality of convolutional layers, and performing cascade training on the trained CNN. The cascade training includes an iterative process of a plurality of stages, in which each stage includes inserting a residual block (ResBlock) and training the CNN with the inserted ResBlock.

Claims (34)

1 . A method, comprising:

generating a convolutional neural network (CNN), wherein generating the CNN comprises:

training a CNN; and

performing cascade training on the trained CNN,

wherein cascade training comprises an iterative process of a plurality of stages, in which each stage comprises:

inserting a residual block (ResBlock); and

training the CNN with the inserted ResBlock.

2 . The method of claim 1 , wherein the inserted ResBlock further includes a rectified linear unit layer between at least two convolutional layers.

3 . The method of claim 1 , wherein each of the stages further comprises replacing a convolutional layer with a depthwise separable convolutional layer.

4 . The method of claim 3 , wherein each of the stages further comprises initializing the depthwise separable convolutional layer with a random weight.

5 . The method of claim 4 , wherein each of the stages further comprises training the CNN with the replaced depthwise separable convolutional layer.

6 . The method of claim 1 , wherein a weight of the inserted ResBlock is randomly initialized.

7 . The method of claim 1 , wherein the CNN is trained on multiple color channels.

8 . The method of claim 1 , wherein the cascade training further comprises replacing the ResBlock with a depthwise separable residual block (DS-ResBlock).

9 . The method of claim 1 , further comprising denoising an image with the generated CNN.

10 . The method of claim 9 , wherein denoising the image comprises applying an edge-aware loss function to the image.

11 . An apparatus, comprising:

one or more non-transitory computer-readable media; and

at least one processor which, when executing instructions stored on the one or more non-transitory computer-readable media, performs the steps of:

generating a convolutional neural network (CNN) by:

training a CNN; and

performing cascade training on the trained CNN,

wherein cascade training comprises an iterative process of a plurality of stages, in which each of the stages comprises:

inserting a residual block (ResBlock); and

training the CNN with the inserted ResBlock.

12 . The apparatus of claim 11 , wherein the inserted ResBlock further includes a rectified linear unit layer between at least two convolutional layers.

13 . The apparatus of claim 11 , wherein each of the stages further comprises replacing one of a convolutional layer with a depthwise separable convolutional layer.

14 . The apparatus of claim 13 , wherein each of the stages further comprises initializing the depthwise separable convolutional layer with a random weight.

15 . The apparatus of claim 14 , wherein each of the stages stage further comprises training the CNN with the replaced depthwise separable convolutional layer.

16 . The apparatus of claim 11 , wherein a weight of the inserted ResBlock is randomly initialized.

17 . The apparatus of claim 11 , wherein the CNN is trained on multiple color channels.

18 . The apparatus of claim 11 , wherein the cascade training further comprises replacing the ResBlock with a depthwise separable residual block (DS-ResBlock).

19 . The apparatus of claim 11 , wherein the at least one processor, when executing the instructions, denoises an image with the generated CNN.

20 . The apparatus of claim 19 , wherein denoising the image comprises applying an edge-aware loss function to the image.

Continuity (6)
Continuation 16138279 · Sep 21, 2018
Continuation In Part 15655557 · Jul 20, 2017
Provisional Application 62692032 · Jun 29, 2018
Provisional Application 62674941 · May 22, 2018
Provisional Application 62471816 · Mar 15, 2017
Related Publication 20220300819A1 · Sep 22, 2022
References Cited (88)
US 7499588B2 · Jacobs et al. · 2009 [cited by applicant]
US 8566264B2 · Schafer et al. · 2013 [cited by applicant]
US 20150134583A1 · Tamatsu et al. · 2015 [cited by applicant]
US 20180189642A1 · Boesch · 2018 [cited by applicant]
US 20180260975A1 · Sunkavalli · 2018 [cited by applicant]
CN 102722712 · 2012 [cited by applicant]
CN 105960657 · 2014 [cited by applicant]
CN 103279933 · 2016 [cited by applicant]
CN 106204499 · 2016 [cited by applicant]
JP 2014049118 · 2014 [cited by applicant]
Foi, A. et al.: Practical poissonian-gaussian noise modeling and fitting for single-image raw-data. IEEE Transactions on Image Processing 17 (2008) 1737-1754. [cited by applicant]
Zhang, K. et al.: Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising. IEEE Transactions on Image Processing 26 (2017) 3142-3155. [cited by applicant]
Remez, T. et al.: Deep class-aware image denoising. In: 2017 International Conference on Sampling Theory and Applications, IEEE (2017) 138-142. [cited by applicant]
Remez, T. et al.: Deep convolutional denoising of low-light images. arXiv preprint arXiv:1701.01687 (2017). [cited by applicant]
Gu, S. et al.: Weighted nuclear norm minimization with application to image denoising. In: IEEE Conference on Computer Vision and Pattern Recognition. (2014) 2862-2869. [cited by applicant]
Dong, W. et al.: Nonlocally centralized sparse representation for image restoration. IEEE Transactions on Image Processing 22 (2013) 1620-1630. [cited by applicant]
Dabov, K. et al.: Image denoising by sparse 3-D transform-domain collaborative filtering. IEEE Transactions on image processing 16 (2007) 2080-2095. [cited by applicant]
Azzari, L. et al.: Variance stabilization for noisy+ estimate combination in iterative Poisson denoising. IEEE signal processing letters 23 (2016) 1086-1090. [cited by applicant]
Makitalo, M. et al.: Optimal inversion of the generalized Anscombe transformation for Poisson-Gaussian noise. IEEE Transactions on image processing 22 (2013) 91-103. [cited by applicant]
Hinton, G.E. et al., “A fast learning algorithm for deep belief nets”, Neural computation 18(7), 1527 (2006). [cited by applicant]
Chen, Y. et al.: Trainable nonlinear reaction diffusion: A flexible framework for fast and effective image restoration, IEEE Transactions on pattern analysis and machine intelligence 39 (2017) 1256-1272. [cited by applicant]
Schmidt, U. et al.: Cascades of regression tree fields for image restoration. IEEE Transactions on pattern analysis and machine intelligence 38 (2016) 677-689. [cited by applicant]
Burger, H.C. et al.: Image denoising: Can plain neural networks compete with BM3D? In: IEEE Conference on Computer Vision and Pattern Recognition, IEEE (2012) 2392-2399. [cited by applicant]
Zhang, K. et al.: FFDnet: Toward a fast and flexible solution for cnn based image denoising. arXiv preprint arXiv:1710.04026 (2017). [cited by applicant]
Toderici, G. et al.: Full resolution image compression with recurrent neural networks. In: ComputerVision and Pattern Recognition (CVPR) 2017 IEEE Conference on, IEEE (2017) 5435-5443. [cited by applicant]
Johnston, N. et al.: Improved lossy image compression with priming and spatially adaptive bit rates for recurrent networks. arXiv preprint arXiv:1703.10114 (2017). [cited by applicant]
Theis, L. et al.: Lossy image compression with compressive autoencoders. arXiv preprint arXiv:1703.00395 (2017). [cited by applicant]
He, K. et al.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. (2016) 770-778. [cited by applicant]
Lim, B. et al.: Enhanced deep residual networks for single image super-resolution. In: IEEE Conference on Computer Vision and Pattern Recognition Workshops. vol. 1. (2017) 3. [cited by applicant]
Ledig, C. et al.: Photo-realistic single image superresolution using a generative adversarial network. (2017). [cited by applicant]
Ren, H. et al.: CT-SRCNN: Cascade trained and trimmed deep convolutional neural networks for image super resolution. (2018). [cited by applicant]
Howard, A.G. et al.: Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017). [cited by applicant]
Sandler, M. et al.: Mobilenetv2: Inverted residuals and linear bottlenecks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (2018) 4510-4520. [cited by applicant]
Everingham, M., et al.: The pascal visual object classes (VOC) challenge. International journal of computer vision 88 (2010) 303-338. [cited by applicant]
Martin, D. et al.: A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics, In: IEEE International Conference on Computer Vision. vol. 2… [cited by applicant]
http://www.compression.cc/: Workshop and challenge on learned image compression (clic). (2018). [cited by applicant]
Remez, Tal et al., “Deep Class-Aware Image Denoising”, 2017 International Conference on Sampling Theory and Applications (SAMPTA) IEEE (2017), 23 pgs. [cited by applicant]
Chua, Kah Keong et al., Enhanced Image Super-Resolution Technique Using Convolutional Neural Network, IVIC 2013, LNCS 8237, pp. 157-164. [cited by applicant]
Girshick, Ross, Fast R-CNN, IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1440-1448. [cited by applicant]
Song, Youyi et al., A Deep Learning Based Framework for Accurate Segmentation of Cervical Cytoplasm and Nuclei, Engineering in Medicine and Biology Society, 2014 36th Annual International Conference of the IEEE, Aug. 26… [cited by applicant]
Bevilacqua, Marco et al., Neighbor Embedding Based Single-Image Super-Resolution Using Semi-Nonnegative Matrix Factorization, ICASSP, Mar. 2012, Kyoto, Japan, 4 pages. [cited by applicant]
Cai, Zhaowei et al., Learning Complexity-Aware Cascades for Deep Pedestrian Detection, 2015 IEEE International Conference on Computer Vision, pp. 3361-3369. [cited by applicant]
Chang, Hong et al., Super-Resolution Through Neighbor Embedding, Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 8 pages. [cited by applicant]
Dahl, Ryan et al., Pixel Recursive Super Resolution, Mar. 22, 2017, 21 pages. [cited by applicant]
Dong, Chao et al., Accelerating the Super-Resolution Convolutional Neural Network, Aug. 1, 2016, 17 pages. [cited by applicant]
Dong, Chao et al., Image Super-Resolution Using Deep Convolutional Networks IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, No. 2, Feb. 2016, pp. 295-307. [cited by applicant]
Dong, Chao et al., Learning a Deep Convolutional Network for Image Super-Resolution, European Conference Computer Vision 2014, 16 pages. [cited by applicant]
Dong, Chao et al., Learning a Deep Convolutional Network for image Super-Resolution, Supplemental Material, European Conference Computer Vision 2014, 13 pages. [cited by applicant]
Elsayed, Ahmed et al., Effect of Super Resolution on High Dimensional Features for Unsupervised Face Recognition in the Wild, May 13, 2017, 5 pages. [cited by applicant]
Gao, Xing et al., A Hybrid Wavelet Convolution Network With Sparse-Coding for Image Super-Resolution, 2016 IEEE ICIP, pp. 1439-1443. [cited by applicant]
Johnson, Justin et al., Perceptual Losses for Real-Time Style Transfer and Super-Resolution, Mar. 27, 2016, 18 pages. [cited by applicant]
Johnson, Justin et al., Perceptual Losses for Real-Time Style Transfer and Super-Resolution, Supplementary Material, Mar. 27, 2016, 5 pages. [cited by applicant]
Kim, Jiwon et al., Deeply-Recursive Convolutional Network for Image Super-Resolution, 2016 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1637-1645. [cited by applicant]
Kim, Jiwon et al., Accurate Image Super-Resolution Using Very Deep Convolutional Networks, 2016 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1646-1654. [cited by applicant]
Kim, Jaeyoung et al., Residual LSTM: Design of a Deep Recurrent Architecture for Distant Speech Recognition, Mar. 15, 2017, 5 pages. [cited by applicant]
Kim, Kwang In et al., Single-Image Super-resolution Using Sparse Regression and Natural Image Prior, IEEE TAMPI 2010, pp. 1127-1133. [cited by applicant]
Ledig, Christian et al., Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, Apr. 13, 2017, 19 pages. [cited by applicant]
Li, Yue et al., Convolutional Neural Network-Based Block Up-sampling for Intra Frame Coding, Submitted to IEEE Transactions on Circuits and Systems for Video Technology, Feb. 22, 2017, 13 pages. [cited by applicant]
Li, Xin et al., New Edge Directed Interpolation, IEEE TX on Img Proc, Oct. 2001, pp. 311-314. [cited by applicant]
Li, Xiaoxiao et al., Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade, CVPR 2017, 10 pages. [cited by applicant]
Li, Hao et al., Pruning Filters for Efficient Convnets, Published as a conference paper at ICLR 2017, Mar. 10, 2017, 13 pages. [cited by applicant]
Liu, Ding et al., Learning a Mixture of Deep Networks for Single Image Super-Resolution, Jan. 3, 2017, 12 pages. [cited by applicant]
Pang, Junbiao et al., Accelerate Convolutional Neural Networks for Binary Classification via Cascading Cost-Sensitive Feature, 2016 ICIP, pp. 1037-1041. [cited by applicant]
Ren, Jimmy et al., Accurate Single Stage Detector Using Recurrent Rolling Convolution, Apr. 19, 2017, 9 pages. [cited by applicant]
Ren, Mengye et al., Normalizing the Normalizers: Comparing and Extending Network Normalization Schemes, Published as a conference paper at ICLR 2017, Mar. 6, 2017, 15 pages. [cited by applicant]
Romano, Yaniv et al., RAISR: Rapid and Accurate Image Super Resolution, Oct. 4, 2016, 31 pages. [cited by applicant]
Romano, Yaniv et al., Supplementary Material for RAISR: Rapid and Accurate Image Super Resolution, 19 pages. [cited by applicant]
Shi, Wenzhe et al., Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network, 2016 IEEE Conference, Computer Vision and Pattern Recognition, pp. 1874-1883. [cited by applicant]
Sonderby, Casper Kaae et al., Amortised Map Inference for Image Super-Resolution, Published as a conference paper at ICLR 2017, Feb. 21, 2017, 17 pages,. [cited by applicant]
Timofte, Radu et al., A+: Adjusted Anchored Neighborhood Regression for Fast Super-Resolution, ACCV 2014, pp. 111-126. [cited by applicant]
Timofte, Radu et al., Anchored Neighborhood Regression for Fast Example-Based Super-Resolution, 2013 IEEE International Conference on Computer Vision, pp. 1920-1927. [cited by applicant]
Timofte, Radu et al., Seven ways to improve example-based single image super resolution, CVPR 2016, 9 pages. [cited by applicant]
Wang, Zhaowen et al., Deep Networks for Image Super-Resolution with Sparse Prior, 2015 IEEE International Conference on Computer Vision, pp. 370-378. [cited by applicant]
Yang, Chih-Yuan et al., Fast Direct Super-Resolution by Simple Functions, 2013 IEEE International Conference on Computer Vision, pp. 561-568. [cited by applicant]
Yang, Jianchao et al., Image Super-Resolution via Sparse Representation, IEEE Transactions on Image Processing, vol. 19, No. 11, Nov. 2010, pp. 2861-2873. [cited by applicant]
Yu, Fisher et al., Multi-Scale Context Aggregation by Dilated Convolutions, Published as a conference paper at ICLR 2016, Apr. 30, 2016, 13 pages. [cited by applicant]
Zeyde, Roman et al., On Single Image Scale-Up Using Sparse-Representations, 2011 International Conf. on Curves and Surfaces, pp. 711-730. [cited by applicant]
Tai, Yu-Wing et al., Super Resolution using Edge Prior and Single Image Detail Synthesis, 2010 IEEE, pp. 2400-2407. [cited by applicant]
Deshpande, Adit, A Beginner's Guide to Understanding Convolutional Neural Networks, https://adeshpande3.github.io/adeshpande3.github.io/A-Beginner's-Guide-To-Understanding-Conolutional-Neural-Networks, Jul. 20, 2016, 17… [cited by applicant]
Deshpande, Adit, A Beginner's Guide to Understanding Convolutional Neural Networks Part 2, https://adeshpande3.github.io/A-Beginner%27s-Guide-To-Understanding-Convolutional-Neural-Networks-Part-2/, Jul. 29, 2016, 8 page… [cited by applicant]
Deshpande, Adit, The 9 Deep Learning Papers You Need to Know About (Understanding CNNs Part 3), https://adeshpande3.github.io/adeshpande3.github.io/The-9-Deep-Learning-Papers-You-Need-To-Know-About.html, Aug. 24, 2016, … [cited by applicant]
Wikipedia entry, Convolutional neural network, 13 pages. [cited by applicant]
Wikipedia entry, Peak signal-to-noise ratio, 2 pages. [cited by applicant]
Wikipedia entry, Structural similarity, 5 pages. [cited by applicant]
Holschneider, M et al., A Real-Time Algorithm for Signal Analysis with the Help of the Wavelet Transform, Wavelets, Springer-Veriag, pp. 286-297. [cited by applicant]
Dong et al., Adaptive Cascade Deep Convolutional Neural Networks for Face Alignment, 2015, Computer Standards & Interfaces, pp. 105-112. [cited by applicant]
Yalta et al., “Sound Source Localization Using Deep Learning Models”, Journal of Robotics & Mechatronics, 2017, pp. 37-48. [cited by applicant]
Turchenkop et al. Course-Grain Parallelization of Neural Network-Based Face Detection Method, IEEE International Workshop on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Appliations, Sep. … [cited by applicant]