IP Library › Granted Patent US 12,236,556
Granted Patent B2
US 12,236,556 · App. 17/762,199 · Granted Feb 25, 2025

Video resolution enhancement method, storage medium, and electronic device

Inventors: Lijie Zhang (Beijing, CN); Dan Zhu (Beijing, CN)
Assignee: BOE TECHNOLOGY GROUP CO., LTD.
G06T3/4053G06T5/20G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,556
App. No.
17/762,199
Granted
Feb 25, 2025
Kind
B2
Abstract

Provided are a method for enhancing a video resolution enhancement, a computer readable storage medium, and an electronic device. The method includes: obtaining multiple image frames as input data, and obtaining initial data by performing feature extraction on the input data using a first three-dimensional convolutional layer; obtaining first feature data by performing down-sampling on the initial data at a preset multiple; obtaining first reference data by performing a convolution operation on the first feature data using a second three-dimensional convolutional layer to merge the first feature data as one frame; and obtaining first output data by performing up-sampling on the first reference data at the preset multiple.

Claims (91)

1. A method for enhancing a video resolution, comprising:

obtaining multiple frames of images as input data, and obtaining initial data by performing feature extraction on the input data using a first three-dimensional convolutional layer;

obtaining first feature data by performing down-sampling on the initial data at a preset multiple;

obtaining first reference data by performing a convolution operation on the first feature data using a second three-dimensional convolutional layer to merge the first feature data into one frame; and

obtaining first output data by performing up-sampling on the first reference data at the preset multiple;

wherein the method further comprises:

performing an Nth super-resolution operation on the first feature data, the super-resolution operation comprising a down-sampling operation, a first feature extraction operation, a merging operation, a second feature extraction operation, and an up-sampling operation, wherein

the down-sampling operation comprises performing down-sampling on the first feature data at the preset multiple;

the first feature extraction operation comprises performing the first feature extraction operation on the down-sampled first feature data by using the first three-dimensional convolution layer to obtain third feature data;

the merging operation comprises performing a convolution operation on the third feature data by using the second three-dimensional convolutional layer to merge the third feature data into one frame to obtain second reference data;

the second feature extraction operation comprises performing the second feature extraction operation on stacked data of the second reference data and (N+1)th output result by using the first three-dimensional convolution layer to obtain fourth feature data; and

the up-sampling operation comprises performing up-sampling on the fourth feature data at the preset multiple to obtain third output data; and

updating the first reference data with the third output data;

wherein an input of Nth down-sampling operation is an output of the first feature extraction operation of (N−1)th super-resolution operation, and N is a positive integer starting from 1.

2. The method according to claim 1 , wherein the obtaining the first feature data by performing down-sampling on the initial data at the preset multiple comprises:

performing down-sampling on the initial data at the preset multiple, and

obtaining the first feature data by performing feature extraction on the down-sampled initial data using the first three-dimensional convolutional layer.

3. The method according to claim 1 , further comprising:

obtaining second output data by performing a convolution operation on the initial data using the second three-dimensional convolution layer to merge the initial data into one frame;

obtaining target data by stacking the second output data and the first output data and performing a convolution operation using the first three-dimensional convolution layer; and

obtaining a super-resolution image by performing multiple up-sampling on the target data.

4. The method according to claim 1 , wherein the obtaining the first output data by performing up-sampling on the first reference data at the preset multiple comprises:

obtaining second feature data by performing feature extraction on the first reference data using the first three-dimensional convolution layer; and

obtaining the first output data by performing up-sampling on the second feature data at the preset multiple.

5. The method according to claim 1 , wherein the updating the first reference data with the third output data comprises:

updating the third output data output by (N−1)th up-sampling operation using an output of Nth up-sampling operation; and

obtaining a updated first reference data by stacking the third output data and the first reference data.

6. The method according to claim 1 , wherein at least one characteristic parameter of the input data comprises a quantity of channels, a batch size, and a height, width and time of each frame of the image; and

wherein the batch size is a quantity of the input data simultaneously input.

7. The method according to claim 1 , wherein the preset multiple is an even number.

8. A non-transitory computer-readable storage medium on which a computer program is stored, wherein the program, when executed by a processor, implements a method for enhancing a video resolution comprising:

obtaining multiple frames of images as input data, and obtaining initial data by performing feature extraction on the input data using a first three-dimensional convolutional layer;

obtaining first feature data by performing down-sampling on the initial data at a preset multiple;

obtaining first reference data by performing a convolution operation on the first feature data using a second three-dimensional convolutional layer to merge the first feature data into one frame; and

obtaining first output data by performing up-sampling on the first reference data at the preset multiple;

wherein the method further comprises:

performing an Nth super-resolution operation on the first feature data, the super-resolution operation comprising a down-sampling operation, a first feature extraction operation, a merging operation, a second feature extraction operation, and an up-sampling operation; wherein

the down-sampling operation comprises performing down-sampling on the first feature data at the preset multiple;

the first feature extraction operation comprises performing the first feature extraction operation on the down-sampled first feature data by using the first three-dimensional convolution layer to obtain third feature data;

the merging operation comprises performing a convolution operation on the third feature data by using the second three-dimensional convolutional layer to merge the third feature data into one frame to obtain second reference data;

the second feature extraction operation comprises performing the second feature extraction operation on stacked data of the second reference data and (N+1)th output result by using the first three-dimensional convolution layer to obtain fourth feature data; and

the up-sampling operation comprises performing up-sampling on the fourth feature data at the preset multiple to obtain third output data; and

updating the first reference data with the third output data;

wherein an input of Nth down-sampling operation is an output of the first feature extraction operation of (N−1)th super-resolution operation, and N is a positive integer starting from 1.

9. The non-transitory computer-readable storage medium according to claim 8 , wherein the obtaining the first feature data by performing down-sampling on the initial data at the preset multiple comprises:

performing down-sampling on the initial data at the preset multiple, and

obtaining the first feature data by performing feature extraction on the down-sampled initial data using the first three-dimensional convolutional layer.

10. The non-transitory computer-readable storage medium according to claim 8 , wherein the obtaining the first output data by performing up-sampling on the first reference data at the preset multiple comprises:

obtaining second feature data by performing feature extraction on the first reference data using the first three-dimensional convolution layer; and

obtaining the first output data by performing up-sampling on the second feature data at the preset multiple.

11. The non-transitory computer-readable storage medium according to claim 8 , wherein the updating the first reference data with the third output data comprises:

updating the third output data output by (N−1)th up-sampling operation using an output of Nth up-sampling operation; and

obtaining a updated first reference data by stacking the third output data and the first reference data.

12. The non-transitory computer-readable storage medium according to claim 8 , wherein at least one characteristic parameter of the input data comprises a quantity of channels, a batch size, and a height, width and time of each frame of the image; and

wherein the batch size is a quantity of the input data simultaneously input.

13. An electronic device, comprising:

one or more processors; and

a memory configured to store one or more programs which, when executed by the one or more processors, cause the one or more processors to:

obtain multiple frames of images as input data, and obtain initial data by performing feature extraction on the input data using a first three-dimensional convolutional layer;

obtain first feature data by performing down-sampling on the initial data at a preset multiple;

obtain first reference data by performing a convolution operation on the first feature data using a second three-dimensional convolutional layer to merge the first feature data into one frame; and

obtain first output data by performing up-sampling on the first reference data at the preset multiple;

wherein the one or more processors are further configured to:

perform an Nth super-resolution operation on the first feature data, the super-resolution operation comprising a down-sampling operation, a first feature extraction operation, a merging operation, a second feature extraction operation, and an up-sampling operation; wherein

the down-sampling operation comprises performing down-sampling on the first feature data at the preset multiple;

the first feature extraction operation comprises performing the first feature extraction operation on the down-sampled first feature data by using the first three-dimensional convolution layer to obtain third feature data;

the merging operation comprises performing a convolution operation on the third feature data by using the second three-dimensional convolutional layer to merge the third feature data into one frame to obtain second reference data;

the second feature extraction operation comprises performing the second feature extraction operation on stacked data of the second reference data and (N+1)th output result by using the first three-dimensional convolution layer to obtain fourth feature data; and

the up-sampling operation comprises performing up-sampling on the fourth feature data at the preset multiple to obtain third output data; and

update the first reference data with the third output data;

wherein an input of Nth down-sampling operation is an output of the first feature extraction operation of (N−1)th super-resolution operation, and N is a positive integer starting from 1.

14. The electronic device according to claim 13 , wherein the one or more processors are further configured to:

perform down-sampling on the initial data at the preset multiple, and

obtain the first feature data by performing feature extraction on the down-sampled initial data using the first three-dimensional convolutional layer.

15. The electronic device according to claim 13 , wherein the one or more processors are further configured to:

obtain second output data by performing a convolution operation on the initial data using the second three-dimensional convolution layer to merge the initial data into one frame;

obtain target data by stacking the second output data and the first output data and performing a convolution operation using the first three-dimensional convolution layer; and

obtain a super-resolution image by performing multiple up-sampling on the target data.

16. The electronic device according to claim 13 , wherein the one or more processors are further configured to:

obtain second feature data by performing feature extraction on the first reference data using the first three-dimensional convolution layer; and

obtain the first output data by performing up-sampling on the second feature data at the preset multiple.

17. The electronic device according to claim 13 , wherein the one or more processors are further configured to:

update the third output data output by (N−1)th up-sampling operation using an output of Nth up-sampling operation; and

obtain a updated first reference data by stacking the third output data and the first reference data.

18. The electronic device according to claim 13 , wherein at least one characteristic parameter of the input data comprises a quantity of channels, a batch size, and a height, width and time of each frame of the image; and

wherein the batch size is a quantity of the input data simultaneously input.

19. The electronic device according to claim 13 , wherein the preset multiple is an even number.

20. The non-transitory computer-readable storage medium according to claim 17 , wherein method further comprises:

obtaining second output data by performing a convolution operation on the initial data using the second three-dimensional convolution layer to merge the initial data into one frame;

obtaining target data by stacking the second output data and the first output data and performing a convolution operation using the first three-dimensional convolution layer; and

obtaining a super-resolution image by performing multiple up-sampling on the target data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2024
From: ZHANG, LIJIE; ZHU, DAN
To: BOE TECHNOLOGY GROUP CO., LTD.
Reel/Frame 069295/0312 →
Priority Claims (1)
CN 202010326998.X · Apr 23, 2020 · national
Continuity (1)
Related Publication 20220292638A1 · Sep 15, 2022
References Cited (22)
US 20180075581A1 · Shi · 2018 [cited by examiner]
US 20210383169A1 · Wang · 2021 [cited by examiner]
CN 107527044A · 2017 [cited by applicant]
CN 108830790A · 2018 [cited by applicant]
CN 109003229A · 2018 [cited by applicant]
CN 109118430A · 2019 [cited by applicant]
CN 109272450A · 2019 [cited by applicant]
CN 109360151A · 2019 [cited by applicant]
CN 109829855A · 2019 [cited by applicant]
CN 109862370A · 2019 [cited by applicant]
CN 110136066A · 2019 [cited by applicant]
CN 110276721A · 2019 [cited by applicant]
CN 110322400A · 2019 [cited by applicant]
CN 110717851A · 2020 [cited by applicant]
CN 111028150A · 2020 [cited by applicant]
EP 3617947A1 · 2020 [cited by applicant]
KR 20190130478A · 2019 [cited by applicant]
WO 2018053340A1 · 2018 [cited by applicant]
Chadha, Aaron, Alhabib Abbas, and Yiannis Andreopoulos. “Video classification with CNNs: Using the codec as a spatio-temporal activity sensor.” IEEE Transactions on Circuits and Systems for Video Technology 29.2 (2017):… [cited by examiner]
International Search Report and Written Opinion mailed on Jul. 19, 2021, corresponding PCT/CN2021/088187, 8 pages. [cited by applicant]
Office Action issued on Feb. 25, 2022, in corresponding Chinese patent Application No. 202010326998.X, 21 pages. [cited by applicant]
Lin Qi et al., “Video Super-Resolution Method Based on Multi-Beale Characteristics Residual Learning Convolutional Neural Network”, Journal of Signal Processing, vol. 36 No. 1, Jan. 2020, total 8 pages, DOI: 10.16798/j.… [cited by applicant]
Cited By (2)
US 12,718,322 US 12,739,343