IP Library › Granted Patent US 11,908,103
Granted Patent B2
US 11,908,103 · App. 17/363,280 · Granted Feb 20, 2024

Multi-scale-factor image super resolution with micro-structured masks

Inventors: Wei Jiang (Sunnyvale, CA); Wei Wang (San Jose, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
G06T3/4053G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,908,103
App. No.
17/363,280
Granted
Feb 20, 2024
Kind
B2
Abstract

There is included a method and apparatus comprising computer code configured to cause a processor or processors to perform obtaining an input low resolution (LR) image comprising a height, a width, and a number of channels, implementing a feature learning deep neural network (DNN) configured to compute a feature tensor based on the input LR image, generating, by an upscaling DNN, a high resolution (HR) image, having a higher resolution than the input LR image, based on the feature tensor computed by the feature learning DNN, wherein a networking structure of the upscaling DNN differs depending on different scale factors, and wherein a networking structure of the feature learning DNN is a same structure for each of the different scale factors.

Claims (78)

1. A method for image processing, the method performed by at least one processor and comprising:

obtaining an input low resolution (LR) image comprising a height, a width, and a number of channels;

implementing a feature learning deep neural network (DNN) configured to compute a feature tensor based on the input LR image;

generating, by an upscaling DNN, a high resolution (HR) image, having a higher resolution than the input LR image, based on the feature tensor computed by the feature learning DNN,

wherein a networking structure of the upscaling DNN differs depending on different scale factors, and

wherein a networking structure of the feature learning DNN is a same structure for each of the different scale factors.

2. The method according to claim 1 , further, at a test stage, comprising:

generating masked weight coefficients for the feature learning DNN based on a target scale factor; and

selecting a subnetwork of the upscaling DNN for the target scale factor based on selected weight coefficients and computing the feature tensor through inference computation.

3. The method according to claim 2 , further comprising:

generating the HR image based on passing the feature tensor through an upscaling module using the selected weight coefficients.

4. The method according to claim 1 ,

wherein weight coefficients of at least one of the feature learning DNN and the upscaling DNN comprise a 5-dimensional (5D) tensor with a size of c 1 , k 1 , k 2 , k 3 , c 2 ,

wherein an input of a layer of the at least one of the feature learning DNN and the upscaling DNN comprises a 4-dimensional (4D) tensor A with a size of h 1 , w 1 , d 1 , c 1 ,

wherein an output of the layer is a 4D tensor B with a size of h 2 , w 2 , d 2 , c 2 , and

wherein each of c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , c 1 , h 2 , w 2 , d 2 , and c 2 are integer numbers greater than or equal to 1,

wherein h 1 , w 1 , d 1 are a height, a weight, and a depth of the tensor A,

wherein h 2 , w 2 , d 2 are a height, a weight, and a depth of the tensor B,

wherein c 1 and c 2 are numbers of input and output channels respectively, and

wherein k 1 , k 2 , k 3 are sizes of a convolution kernel and correspond to height, weight, and depth axes respectively.

5. The method according to claim 4 , further comprising:

reshaping the 5D tensor to a 3-dimensional (3D) tensor; and

reshaping the 5D tensor to a 2-dimensional (2D) matrix.

6. The method according to claim 5 ,

wherein the 3D tensor is of a size c′ 1 , c′ 2 , k, where c′ 1 ×c′ 2 ×k=c 1 ×c 2 ×k 1 ×k 2 ×k 3 , and

wherein the 2D matrix is of a size c′ 1 , c′ 2 , where c′ 1 ×c′ 2 =c 1 ×c 2 ×k 1 ×k 2 ×k 3 .

7. The method according to claim 6 , further, at a training stage, comprising:

fixing weight coefficients that are masked;

obtaining updated weight coefficients based on a learning process on weights of the feature learning DNN and weights of the upscaling DNN through a weight filling module; and

conducting, based on the updated weight coefficients, a micro-structured pruning process to obtain a model instance and masks.

8. The method according to claim 7 ,

wherein the learning process comprises reinitializing the weight coefficients that have zero values by setting those weight coefficients to any of random initial values and corresponding weights of a previously learned model.

9. The method according to claim 7 , wherein the micro-structured pruning process comprises:

computing losses for each of a plurality of micro-structured blocks of at least one of the 3D tensor and the 2D matrix; and

ranking the micro-structured blocks based on the computed losses.

10. The method according to claim 9 , wherein the micro-structured pruning process further comprises:

determining to stop the micro-structured pruning process based on whether a distortion loss reaches a threshold.

11. An apparatus for image processing, the apparatus comprising:

at least one memory configured to store computer program code;

at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code including:

obtaining code configured to cause the at least one processor to obtain an input low resolution (LR) image comprising a height, a width, and a number of channels;

implementing code configured to cause the at least one processor to implement a feature learning deep neural network (DNN) configured to compute a feature tensor based on the input LR image;

generating code configured to cause the at least one processor to generate, by an upscaling DNN, a high resolution (HR) image, having a higher resolution than the input LR image, based on the feature tensor computed by the feature learning DNN,

wherein a networking structure of the upscaling DNN differs depending on different scale factors, and

wherein a networking structure of the feature learning DNN is a same structure for each of the different scale factors.

12. The apparatus according to claim 11 , wherein the generating code is further, at a test stage, code configured to cause the at least one processor to:

generate masked weight coefficients for the feature learning DNN based on a target scale factor; and

select a subnetwork of the upscaling DNN for the target scale factor based on selected weight coefficients and computing the feature tensor through inference computation.

13. The apparatus according to claim 12 , wherein generating the HR image is based on passing the feature tensor through an upscaling module using the selected weight coefficients.

14. The apparatus according to claim 11 ,

wherein weight coefficients of at least one of the feature learning DNN and the upscaling DNN comprise a 5-dimensional (5D) tensor with a size of c 1 , k 1 , k 2 , k 3 , c 2 ,

wherein an input of a layer of the at least one of the feature learning DNN and the upscaling DNN comprises a 4-dimensional (4D) tensor A with a size of h 1 , w 1 , d 1 , c 1 ,

wherein an output of the layer is a 4D tensor B with a size of h 2 , w 2 , d 2 , c 2 , and

wherein each of c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , c 1 , h 2 , w 2 , d 2 , and c 2 are integer numbers greater than or equal to 1,

wherein h 1 , w 1 , d 1 are a height, a weight, and a depth of the tensor A,

wherein h 2 , w 2 , d 2 are a height, a weight, and a depth of the tensor B,

wherein c 1 and c 2 are numbers of input and output channels respectively, and

wherein k 1 , k 2 , k 3 are sizes of a convolution kernel and correspond to height, weight, and depth axes respectively.

15. The apparatus according to claim 14 , further comprising:

reshaping code configured to cause the at least one processor to reshape the 5D tensor to a 3-dimensional (3D) tensor and to reshape the 5D tensor to a 2-dimensional (2D) matrix.

16. The apparatus according to claim 15 ,

wherein the 3D tensor is of a size c′ 1 , c′ 2 , k, where c′ 1 ×c′ 2 ×k=c 1 ×c 2 ×k 1 ×k 2 ×k 3 , and

wherein the 2D matrix is of a size c′ 1 , c′ 2 , where c′ 1 ×c′ 2 =c 1 ×c 2 ×k 1 ×k 2 ×k 3 .

17. The apparatus according to claim 16 , further, at a training stage, comprising:

fixing code configured to cause the at least one processor to fix weight coefficients that are masked;

obtaining code configured to cause the at least one processor to obtain updated weight coefficients based on a learning process on weights of the feature learning DNN and weights of the upscaling DNN through a weight filling module; and

conducting code configured to cause the at least one processor to conduct, based on the updated weight coefficients, a micro-structured pruning process to obtain a model instance and masks.

18. The apparatus according to claim 17 ,

wherein the learning process comprises reinitializing the weight coefficients that have zero values by setting those weight coefficients to any of random initial values and corresponding weights of a previously learned model.

19. The apparatus according to claim 17 , wherein the micro-structured pruning process comprises:

computing losses for each of a plurality of micro-structured blocks of at least one of the 3D tensor and the 2D matrix; and

ranking the micro-structured blocks based on the computed losses.

20. A non-transitory computer readable medium storing a program causing a computer to execute a process, the process comprising:

obtaining an input low resolution (LR) image comprising a height, a width, and a number of channels;

implementing a feature learning deep neural network (DNN) configured to compute a feature tensor based on the input LR image;

generating, by an upscaling DNN, a high resolution (HR) image, having a higher resolution than the input LR image, based on the feature tensor computed by the feature learning DNN,

wherein a networking structure of the upscaling DNN differs depending on different scale factors, and

wherein a networking structure of the feature learning DNN is a same structure for each of the different scale factors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2021
From: JIANG, WEI; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 056779/0576 →
Continuity (2)
Provisional Application 63065608 · Aug 14, 2020
Related Publication 20220051367A1 · Feb 17, 2022