IP Library › Granted Patent US 12,511,713
Granted Patent B2
US 12,511,713 · App. 17/927,863 · Granted Dec 30, 2025

Image super-resolution reconstructing

Inventors: Shuxin Zheng (Redmond, WA); Chang Liu (Redmond, WA); Di He (Redmond, WA); Guolin Ke (Redmond, WA); Jiang Bian (Redmond, WA); Tie-Yan Liu (Bejing, CN)
Assignee: Microsoft Technology Licensing, LLC
G06T3/4053G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,713
App. No.
17/927,863
Granted
Dec 30, 2025
Kind
B2
Abstract

According to implementations of the subject matter described herein, a solution is proposed for super-resolution image reconstructing. According to the solution, an input image with first resolution is obtained. An invertible neural network is trained using the input image, wherein the invertible neural network is configured to generate an intermediate image with second resolution and first high-frequency information based on the input image, the second resolution being lower than the first resolution. Subsequently, an output image with third resolution is generated based on the input image and second high-frequency information by using an inverse network of the trained invertible neural network, the second high-frequency information conforming to a predetermined distribution, and the third resolution being higher than the first resolution. The solution can effectively process a low-resolution image obtained by an unknown downsampling method, thereby obtaining a high-quality and high-resolution image.

Claims (56)

1 . A computer-implemented method, comprising:

obtaining an input image of a first resolution;

training an invertible neural network with the input image, the invertible neural network being configured to generate an intermediate image of a second resolution and first high-frequency information based on the input image, the second resolution being lower than the first resolution;

wherein training an invertible neural network with the input image comprises:

determining a plurality of target functions based on the input image and the intermediate image by determining a first target function based on a difference between a pixel distribution in the intermediate image and a pixel distribution in an image block of the input image, the image block being of the second resolution; and

generating, using an inverse network of the trained invertible neural network, an output image of a third resolution based on second high-frequency information conforming to a predetermined distribution and the input image, the third resolution being greater than the first resolution.

2 . The method of claim 1 , wherein training an invertible neural network with the input image further comprises:

determining a total target function for training the invertible neural network by combining at least some of the plurality of target functions; and

determining network parameters for the invertible neural network by minimizing the total target function.

3 . The method of claim 2 , wherein determining the first target function comprises:

distinguishing, by using a discriminator, whether a pixel of the intermediate image belongs to the intermediate image or the image block; and

determining the first target function based on the distinguishing.

4 . The method of claim 2 , wherein determining a plurality of target functions comprises:

determining a second target function based on a difference between a distribution of the first high-frequency information and the predetermined distribution.

5 . The method of claim 2 , wherein determining a plurality of target functions comprises:

generating, by using an inverse network of the invertible neural network, a reconstructed image of the first resolution based on third high-frequency information conforming to the predetermined distribution and the intermediate image; and

determining a third target function based on a difference between the input image and the reconstructed image.

6 . The method of claim 2 , wherein determining the plurality of target functions comprises:

obtaining a reference image corresponding to semantics of the input image, the reference image being of the second resolution; and

determining a fourth target function based on a difference between the intermediate image and the reference image.

7 . The method of claim 1 , wherein the invertible neural network comprises a transforming module and at least one invertible network unit, and wherein generating the output image comprises:

generating, based on the input image and the second high-frequency information and by using the at least one invertible network unit, a low-frequency component and a high-frequency component to be merged, the low-frequency component representing semantics of the input image and the high-frequency component being related to the semantics; and

merging, by using the transforming module, the low-frequency component and the high-frequency component into the output image.

8 . The method of claim 7 , wherein the transforming module comprises at least one of:

an invertible convolution block; and

a wavelet transforming module.

9 . A device, comprising:

a processing unit; and

a memory coupled to the processing unit and comprising instructions stored thereon which, when executed by the processing unit, cause the device to perform acts comprising:

obtaining an input image of a first resolution;

training an invertible neural network with the input image, the invertible neural network being configured to generate an intermediate image of a second resolution and first high-frequency information based on the input image, the second resolution being lower than the first resolution;

wherein training an invertible neural network with the input image comprises:

determining a plurality of target functions based on the input image and the intermediate image by determining a first target function based on a difference between a pixel distribution in the intermediate image and a pixel distribution in an image block of the input image, the image block being of the second resolution; and

generating, using an inverse network of the trained invertible neural network, an output image of a third resolution based on second high-frequency information conforming to a predetermined distribution and the input image, the third resolution being greater than the first resolution.

10 . The device of claim 9 , wherein training an invertible neural network with the input image further comprises:

determining a total target function for training the invertible neural network by combining at least some of the plurality of target functions; and

determining network parameters for the invertible neural network by minimizing the total target function.

11 . The device of claim 10 , wherein determining the first target function comprises:

distinguishing, by using a discriminator, whether a pixel of the intermediate image belongs to the intermediate image or the image block; and

determining the first target function based on the distinguishing.

12 . The device of claim 10 , wherein determining a plurality of target functions comprises:

determining a second target function based on a difference between a distribution of the first high-frequency information and the predetermined distribution.

13 . A computer program product being tangibly stored in a non-transitory computer readable storage medium and comprising machine-executable instructions which, when executed by a device, cause the device to perform acts comprising:

obtaining an input image of a first resolution;

training an invertible neural network with the input image, the invertible neural network being configured to generate an intermediate image of a second resolution and first high-frequency information based on the input image, the second resolution being lower than the first resolution;

wherein training an invertible neural network with the input image comprises:

determining a plurality of target functions based on the input image and the intermediate image by determining a first target function based on a difference between a pixel distribution in the intermediate image and a pixel distribution in an image block of the input image, the image block being of the second resolution; and

generating, using an inverse network of the trained invertible neural network, an output image of a third resolution based on second high-frequency information conforming to a predetermined distribution and the input image, the third resolution being greater than the first resolution.

14 . The computer program product of claim 13 , wherein training an invertible neural network with the input image further comprises:

determining a total target function for training the invertible neural network by combining at least some of the plurality of target functions; and

determining network parameters for the invertible neural network by minimizing the total target function.

15 . The computer program product of claim 14 , wherein determining the first target function comprises:

distinguishing, by using a discriminator, whether a pixel of the intermediate image belongs to the intermediate image or the image block; and

determining the first target function based on the distinguishing.

16 . The computer program product of claim 14 , wherein determining a plurality of target functions comprises:

determining a second target function based on a difference between a distribution of the first high-frequency information and the predetermined distribution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2022
From: ZHENG, SHUXIN; LIU, CHANG; HE, DI; KE, GUOLIN; BIAN, JIANG; LIU, TIE-YAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061962/0001 →
Priority Claims (1)
CN 202010621955.4 · Jun 30, 2020 · national
Continuity (1)
Related Publication 20230206396A1 · Jun 29, 2023
References Cited (19)
US 8594464B2 · Liu · 2013 [cited by applicant]
US 9692939B2 · Irani et al. · 2017 [cited by applicant]
US 10489887B2 · El-Khamy · 2019 [cited by examiner]
US 20130016920A1 · Matsuda et al. · 2013 [cited by applicant]
US 20180075581A1 · Shi et al. · 2018 [cited by applicant]
CN 110136063A · 2019 [cited by applicant]
CN 110310227A · 2019 [cited by applicant]
WO 2019102476A2 · 2019 [cited by applicant]
Xiao, Mingging, Shuxin Zheng, Chang Liu, Yaolong Wang, Di He, Guolin Ke, Jiang Bian, Zhouchen Lin, and Tie-Yan Liu. ‘Invertible Image Rescaling’. arXiv [Eess.|V], 2020. arXiv. htto://arxiv.org/abs/2005.05650 (Year: 2020… [cited by examiner]
Bell-Kligler, et al., “Blind Super-Resolution Kernel Estimation using an Internal-GAN”, In Journal of Computing Research Repository, Sep. 2019, pp. 1-10. [cited by applicant]
He, et al., “A soft MAP framework for blind super-resolution image reconstruction”, In Journal of Image and Vision Computing, vol. 27, Issue 4, Mar. 1, 2009, 10 Pages. [cited by applicant]
Li, et al., “Multi-Scale Invertible Network for Image Super-Resolution”, In Proceedings of the ACM Multimedia Asia, Dec. 15, 2019, 6 Pages. [cited by applicant]
Michaeli, et al., “Nonparametric blind super-resolution”, In Proceedings of the IEEE International Conference on Computer Vision, Dec. 1, 2013, pp. 945-952. [cited by applicant]
“International Search Report and Written Opinion issued in PCT Application No. PCT/US21/031467”, Mailed Date: Aug. 20, 2021, 9 Pages. [cited by applicant]
Qin, et al., “Blind Single-Image Super Resolution Reconstruction with Gaussian Blur and Pepper & Salt Noise”, In Journal of Computers, vol. 9, Issue 4, Apr. 2014, pp. 896-902. [cited by applicant]
Xiao, et al., “Invertible Image Rescaling”, In Repository of arXiv:2005.05650v1, May 12, 2020, 27 Pages. [cited by applicant]
First Office Action Received for Chinese Application No. 202010621955.4, mailed on Feb. 26, 2025, 18 pages. (English Translation Provided). [cited by applicant]
Rejection Decision Received for Chinese Application No. 202010621955.4, mailed on Jul. 25, 2025, 18 pages (English Translation Provided). [cited by applicant]
Communication pursuant to Article 94(3) received in European Application No. 21730006.0, mailed on Sep. 3, 2025, 4 Pages. [cited by applicant]