IP Library › Granted Patent US 12,347,068
Granted Patent B2
US 12,347,068 · App. 18/127,199 · Granted Jul 1, 2025

Method, device, and computer program product for image processing

Inventors: Zhisong Liu (Shenzhen, CN); Zijia Wang (Weifang, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T3/4061G06N3/0455G06T3/4046G06T3/4053
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,068
App. No.
18/127,199
Granted
Jul 1, 2025
Kind
B2
Abstract

Embodiments of the present disclosure relate to a method, a device, and a computer program product for image processing. The method includes: obtaining an encoding feature of a reference image and an encoding feature of an input image of a first resolution, wherein the reference image has a resolution greater than the first resolution. The method further includes: obtaining high-frequency information and low-frequency information on the input image by interpolating the input image; obtaining a first output feature based on the encoding feature of the reference image and the high-frequency information; and obtaining a second output feature based on the encoding feature of the input image and the low-frequency information. The method further includes: generating an output image of a second resolution based on the first output feature and the second output feature, wherein the second resolution is greater than the first resolution.

Claims (76)

1. A method for image processing, comprising:

obtaining an encoding feature of a reference image and an encoding feature of an input image of a first resolution, wherein the reference image has a resolution greater than the first resolution;

obtaining high-frequency information and low-frequency information on the input image by interpolating the input image;

obtaining a first output feature based on the encoding feature of the reference image and the high-frequency information;

obtaining a second output feature based on the encoding feature of the input image and the low-frequency information; and

generating an output image of a second resolution based on the first output feature and the second output feature, wherein the second resolution is greater than the first resolution.

2. The method according to claim 1 , wherein obtaining the encoding feature of the reference image and the encoding feature of the input image comprises:

processing the reference image using a first encoder to obtain the encoding feature of the reference image; and

processing the input image using a second encoder to obtain the encoding feature of the input image.

3. The method according to claim 2 , wherein processing the reference image using the first encoder comprises:

processing, in the first encoder, features of a hidden layer in the first encoder by means of reparameterization; and

inputting the processed features of the hidden layer to a fully connected layer in the first encoder.

4. The method according to claim 1 , wherein interpolating the input image comprises:

receiving a scale set by a user as a continuous scale, wherein the scale set by the user comprises an integer scale or a non-integer scale; and

performing interpolation on the input image in the continuous scale.

5. The method according to claim 4 , further comprising:

determining continuous scales of the input image in width and in height that are set differently from each other.

6. The method according to claim 1 , wherein generating the output image of the second resolution comprises:

combining the first output feature and the second output feature to obtain a combined output feature; and

processing the combined output feature using a decoder to obtain the output image.

7. The method according to claim 1 , further comprising:

obtaining a combined encoding feature based on the encoding feature of the reference image and an encoding feature of a first training image;

obtaining an encoding feature of a second training image;

obtaining high-frequency information and low-frequency information on the second training image by interpolating the second training image, wherein the first training image corresponds to the second training image, and the first training image has a resolution greater than that of the second training image; and

obtaining a third training image based on the combined encoding feature, the encoding feature of the second training image, the high-frequency information on the second training image, and the low-frequency information on the second training image.

8. The method according to claim 7 , wherein obtaining the third training image comprises:

obtaining a third output feature based on the combined encoding feature and the high-frequency information on the second training image;

obtaining a fourth output feature based on the encoding feature of the second training image and the low-frequency information on the second training image; and

obtaining the third training image based on the third output feature and the fourth output feature.

9. The method according to claim 8 , further comprising:

obtaining the encoding feature of the reference image and the encoding feature of the first training image using a first encoder;

obtaining the encoding feature of the second training image using a second encoder; and

processing the third output feature and the fourth output feature using a decoder to obtain the third training image.

10. The method according to claim 9 , further comprising:

determining relative entropy of the third training image to the reference image; and

training the first encoder, the second encoder, and the decoder based on the relative entropy.

11. The method according to claim 10 , wherein training the first encoder, the second encoder, and the decoder further comprises:

in a first training stage, fixing parameters of the decoder and training the first encoder and the second encoder; and

in a second training stage, fixing parameters of the first encoder and the second encoder and training the parameters of the decoder.

12. An electronic device, comprising:

a processor; and

a memory coupled to the processor, wherein the memory has instructions stored therein which, when executed by the processor, cause the electronic device to execute actions comprising:

obtaining an encoding feature of a reference image and an encoding feature of an input image of a first resolution, wherein the reference image has a resolution greater than the first resolution;

obtaining high-frequency information and low-frequency information on the input image by interpolating the input image;

obtaining a first output feature based on the encoding feature of the reference image and the high-frequency information;

obtaining a second output feature based on the encoding feature of the input image and the low-frequency information; and

generating an output image of a second resolution based on the first output feature and the second output feature, wherein the second resolution is greater than the first resolution.

13. The electronic device according to claim 12 , wherein obtaining the encoding feature of the reference image and the encoding feature of the input image comprises:

processing the reference image using a first encoder to obtain the encoding feature of the reference image; and

processing the input image using a second encoder to obtain the encoding feature of the input image.

14. The electronic device according to claim 13 , wherein processing the reference image using the first encoder comprises:

processing, in the first encoder, features of a hidden layer in the first encoder by means of reparameterization; and

inputting the processed features of the hidden layer to a fully connected layer in the first encoder.

15. The electronic device according to claim 12 , wherein interpolating the input image comprises:

receiving a scale set by a user as a continuous scale, wherein the scale set by the user comprises an integer scale or a non-integer scale; and

performing interpolation on the input image in the continuous scale.

16. The electronic device according to claim 15 , wherein the actions further comprise:

determining continuous scales of the input image in width and in height that are set differently from each other.

17. The electronic device according to claim 12 , wherein generating the output image of the second resolution comprises:

combining the first output feature and the second output feature to obtain a combined output feature; and

processing the combined output feature using a decoder to obtain the output image.

18. The electronic device according to claim 12 , wherein the actions further comprise:

obtaining a combined encoding feature based on the encoding feature of the reference image and an encoding feature of a first training image;

obtaining an encoding feature of a second training image;

obtaining high-frequency information and low-frequency information on the second training image by interpolating the second training image, wherein the first training image corresponds to the second training image, and the first training image has a resolution greater than that of the second training image; and

obtaining a third training image based on the combined encoding feature, the encoding feature of the second training image, the high-frequency information on the second training image, and the low-frequency information on the second training image.

19. The electronic device according to claim 18 , wherein obtaining the third training image comprises:

obtaining a third output feature based on the combined encoding feature and the high-frequency information on the second training image;

obtaining a fourth output feature based on the encoding feature of the second training image and the low-frequency information on the second training image; and

obtaining the third training image based on the third output feature and the fourth output feature.

20. A computer program product that is tangibly stored on a non-transitory computer-readable medium and comprises machine-executable instructions, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform the following actions:

obtaining an encoding feature of a reference image and an encoding feature of an input image of a first resolution, wherein the reference image has a resolution greater than the first resolution;

obtaining high-frequency information and low-frequency information on the input image by interpolating the input image;

obtaining a first output feature based on the encoding feature of the reference image and the high-frequency information;

obtaining a second output feature based on the encoding feature of the input image and the low-frequency information; and

generating an output image of a second resolution based on the first output feature and the second output feature, wherein the second resolution is greater than the first resolution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: LIU, ZHISONG; WANG, ZIJIA; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 063131/0208 →
Priority Claims (1)
CN 202310181589.9 · Feb 20, 2023 · national
Continuity (1)
Related Publication 20240281926A1 · Aug 22, 2024
References Cited (29)
US 20210342974A1 · Zhang · 2021 [cited by examiner]
US 20220092350A1 · Liou · 2022 [cited by examiner]
US 20230098437A1 · Li · 2023 [cited by examiner]
Liu, Zhi-Song, Wan-Chi Siu, and Li-Wen Wang. “Variational Auto Encoder for Reference based Image Super-Resolution.” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). IEEE, 2021. (Yea… [cited by examiner]
Sun, Wanjie, and Zhenzhong Chen. “Learning discrete representations from reference images for large scale factor image super-resolution.” IEEE Transactions on Image Processing 31 (2022): 1490-1503. (Year: 2022). [cited by examiner]
Daniel, Tal, and Aviv Tamar. “Soft-IntroVAE: Analyzing and Improving the Introspective Variational Autoencoder.” arXiv preprint arXiv:2012.13253v2 (2020). (Year: 2021). [cited by examiner]
Chen, Yinbo, Sifei Liu, and Xiaolong Wang. “Learning Continuous Image Representation with Local Implicit Image Function.” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2021. (Year: 20… [cited by examiner]
Fritsche, Manuel, Shuhang Gu, and Radu Timofte. “Frequency separation for real-world super-resolution.” 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). IEEE, 2019. (Year: 2019). [cited by examiner]
Li, Shanshan, et al. “Frequency separation network for image super-resolution.” IEEE Access 8 (2020): 33768-33777. (Year: 2020). [cited by examiner]
Prost, Jean, et al. “Diverse super-resolution with pretrained deep hiererarchical variational autoencoders.” arXiv preprint arXiv: 2205.10347v2 (2022). (Year: 2022). [cited by examiner]
C. Dong et al., “Image Super-Resolution Using Deep Convolutional Networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, arXiv:1501.00092v3, Jul. 31, 2015, 14 pages. [cited by applicant]
J. Kim et al., “Accurate Image Super-Resolution Using Very Deep Convolutional Networks,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), arXiv:1511.04587v2, Nov. 11, 2016, 9 pages. [cited by applicant]
W.-S. Lai et al., “Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), arXiv:1704.03915v2, Oct. 9, 2017, 9 pages. [cited by applicant]
C. Ledig et al., “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), arXiv:1609.04802v5, May 25, 2017, 19 pages. [cited by applicant]
B. Lim et al., “Enhanced Deep Residual Networks for Single Image Super-Resolution,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), arXiv:1707.02921v1, Jul. 10, 2017, 9 pages. [cited by applicant]
Y. Zhang et al., “Image Super-Resolution Using Very Deep Residual Channel Attention Networks,” European Conference on Computer Vision, arXiv:1807.02758v2, Jul. 12, 2018, 16 pages. [cited by applicant]
Y. Zhang et al., “Residual Dense Network for Image Super-Resolution,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:1802.08797v2, Mar. 27, 2018, 10 pages. [cited by applicant]
M. Haris et al., “Deep Back-Projection Networks for Single Image Super-resolution,” IEEE Transactions on Pattern Analysis and Machine Intelligence, arXiv:1904.05677v2, Jun. 13, 2020, 14 pages. [cited by applicant]
Z.-S. Liu et al., “Hierarchical Back Projection Network for Image Super-Resolution,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:1906.06874v2, Jun. 20, 2019, 10 pages. [cited by applicant]
Z.-S. Liu et al., “Image Super-Resolution via Attention based Back Projection Networks,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:1910.04476v, Oct. 10, 2019, 9 pages. [cited by applicant]
T. Dai et al., “Second-order Attention Network for Single Image Super-Resolution,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 1, 2019, 10 pages. [cited by applicant]
B. Niu et al., “Single Image Super-Resolution via a Holistic Attention Network,” European Conference on Computer Vision, arXiv:2008.08767v1, Aug. 20, 2020, 16 pages. [cited by applicant]
I. J. Goodfellow et al., “Generative Adversarial Nets,” Advances in Neural Information Processing Systems, arXiv:1406.2661v1, Jun. 10, 2014, 9 pages. [cited by applicant]
Z. Liu et al., “Photo-Realistic Image Super-Resolution via Variational Autoencoders,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 4, Apr. 2021, 15 pages. [cited by applicant]
Z.-S. Liu et al., “Reference Based Face Super-Resolution,” IEEE Access, vol. 7, Sep. 23, 2019, pp. 129112-129126. [cited by applicant]
J. Engel et al., “Latent Constraints: Learning to Generate Conditionally from Unconditional Generative Models,” arXiv:1711.05772v2, Dec. 21, 2017, 22 pages. [cited by applicant]
Z.-S. Liu et al., “Unsupervised Real Image Super-Resolution via Generative Variational AutoEncoder,” IEEE International Conference on Computer Vision and Pattern Recognition Workshop, arXiv:2004.12811v1, Apr. 27, 2020, … [cited by applicant]
X. Wang et al., “ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition, Sep. 2018, 16 pages. [cited by applicant]
U.S. Appl. No. 17/993,328, filed in the name of Zhisong Liu et al. on Nov. 23, 2022, and entitled “Method, Electronic Device, and Computer Program Product for Image Processing.” [cited by applicant]
Cited By (1)
US 12,444,055