IP Library › Granted Patent US 12,400,287
Granted Patent B2
US 12,400,287 · App. 17/962,581 · Granted Aug 26, 2025

Method, electronic device, and computer program product for video processing utilizing template and associated image conversion model

Inventors: Qiang Chen (Shanghai, CN); Pedro Fernandez (Shanghai, CN)
Assignee: Dell Products L.P.
G06T3/40G06T5/50G06T2207/10016G06T2207/20016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,287
App. No.
17/962,581
Granted
Aug 26, 2025
Kind
B2
Abstract

Embodiments of the present disclosure relate to a method, an electronic device, and a computer program product for video processing. The method includes: converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames that are more similar to the template. The method also includes: converting the second group of image frames to a third group of image frames having a higher resolution to generate a higher-resolution video. In this manner, low-resolution image frames can be first converted to image frames that are more suitable for resolution conversion, so that a high-resolution video of a higher quality can be obtained when reconstructing the high-resolution video.

Claims (56)

1. A method for video processing, comprising:

converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames, wherein the similarity between the second group of image frames and the template is higher than that between the first group of image frames and the template;

converting the second group of image frames having the first resolution to a third group of image frames having a second resolution, wherein the second resolution is higher than the first resolution; and

generating a second video having the second resolution based on the third group of image frames;

wherein converting the second group of image frames having the first resolution to a third group of image frames having a second resolution comprises:

converting the second group of image frames having the first resolution to the third group of image frames having the second resolution using an image conversion model associated with the template.

2. The method according to claim 1 , wherein converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames comprises:

determining a mapping model based on the sample frame and the template; and

converting the first group of image frames to the second group of image frames using the mapping model.

3. The method according to claim 2 , further comprising:

converting the third group of image frames having the second resolution to a fourth group of image frames having the second resolution using a reverse mapping model corresponding to the mapping model.

4. The method according to claim 2 , further comprising:

updating, if determining that the similarity between at least one image frame in the second group of image frames and the template is below a threshold, the mapping model based on another sample frame corresponding to the first group of image frames and the template for use in updating the at least one image frame in the second group of image frames.

5. The method according to claim 2 , wherein determining a mapping model based on the sample frame and the template comprises:

determining the mapping model based on a comparison between characteristics of the sample frame and characteristics of the template.

6. The method according to claim 5 , wherein the characteristics comprise at least one of the following:

brightness, color distribution, a scene in an image frame, and a posture of an object in the image frame.

7. The method according to claim 1 , further comprising:

selecting the template from a plurality of templates based on the sample frame, each of the plurality of templates being associated with a corresponding one of a plurality of image conversion models.

8. The method according to claim 1 , wherein the image conversion model associated with the template is trained based on a training data set, the training data set comprising training frames of the first resolution and training frames of the second resolution that have similar characteristics to the template.

9. The method according to claim 1 , further comprising:

identifying the first group of image frames among a plurality of image frames of the first video based on characteristics of the image frames.

10. An electronic device, comprising:

at least one processor; and

memory coupled to the at least one processor, wherein the memory has instructions stored therein which, when executed by the at least one processor, cause the electronic device to perform actions comprising:

converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames, wherein the similarity between the second group of image frames and the template is higher than that between the first group of image frames and the template;

converting the second group of image frames having the first resolution to a third group of image frames having a second resolution, wherein the second resolution is higher than the first resolution; and

generating a second video having the second resolution based on the third group of image frames;

wherein converting the second group of image frames having the first resolution to a third group of image frames having a second resolution comprises:

converting the second group of image frames having the first resolution to the third group of image frames having the second resolution using an image conversion model associated with the template.

11. The electronic device according to claim 10 , wherein converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames comprises:

determining a mapping model based on the sample frame and the template; and

converting the first group of image frames to the second group of image frames using the mapping model.

12. The electronic device according to claim 11 , wherein the actions further comprise:

converting the third group of image frames having the second resolution to a fourth group of image frames having the second resolution using a reverse mapping model corresponding to the mapping model.

13. The electronic device according to claim 11 , wherein the actions further comprise:

updating, if determining that the similarity between at least one image frame in the second group of image frames and the template is below a threshold, the mapping model based on another sample frame corresponding to the first group of image frames and the template for use in updating the at least one image frame in the second group of image frames.

14. The electronic device according to claim 11 , wherein determining a mapping model based on the sample frame and the template comprises:

determining the mapping model based on a comparison between characteristics of the sample frame and characteristics of the template.

15. The electronic device according to claim 14 , wherein the characteristics comprise at least one of the following:

brightness, color distribution, a scene in an image frame, and a posture of an object in the image frame.

16. The electronic device according to claim 10 , wherein the actions further comprise:

selecting the template from a plurality of templates based on the sample frame, each of the plurality of templates being associated with a corresponding one of a plurality of image conversion models.

17. The electronic device according to claim 10 , wherein the image conversion model associated with the template is trained based on a training data set, the training data set comprising training frames of the first resolution and training frames of the second resolution that have similar characteristics to the template.

18. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising:

converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames, wherein the similarity between the second group of image frames and the template is higher than that between the first group of image frames and the template;

converting the second group of image frames having the first resolution to a third group of image frames having a second resolution, wherein the second resolution is higher than the first resolution; and

generating a second video having the second resolution based on the third group of image frames;

wherein converting the second group of image frames having the first resolution to a third group of image frames having a second resolution comprises:

converting the second group of image frames having the first resolution to the third group of image frames having the second resolution using an image conversion model associated with the template.

19. The computer program product according to claim 18 , wherein converting, based on a sample frame in a first video having a first resolution as well as a template, a first group of image frames in the first video that correspond to the sample frame to a second group of image frames comprises:

determining a mapping model based on the sample frame and the template; and

converting the first group of image frames to the second group of image frames using the mapping model;

wherein the third group of image frames having the second resolution is converted to a fourth group of image frames having the second resolution using a reverse mapping model corresponding to the mapping model.

20. The computer program product according to claim 18 , wherein the actions further comprise:

selecting the template from a plurality of templates based on the sample frame, each of the plurality of templates being associated with a corresponding one of a plurality of image conversion models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2022
From: CHEN, QIANG; FERNANDEZ, PEDRO
To: DELL PRODUCTS L.P.
Reel/Frame 061358/0580 →
Priority Claims (1)
CN 202211132092.X · Sep 16, 2022 · national
Continuity (1)
Related Publication 20240095878A1 · Mar 21, 2024
References Cited (28)
US 10701394B1 · Caballero et al. · 2020 [cited by applicant]
US 12033301B2 · Liu · 2024 [cited by examiner]
US 20130169863A1 · Smith et al. · 2013 [cited by applicant]
US 20180139458A1 · Wang et al. · 2018 [cited by applicant]
US 20190130530A1 · Schroers et al. · 2019 [cited by applicant]
US 20200162789A1 · Ma et al. · 2020 [cited by applicant]
US 20210006782A1 · Ramchandran et al. · 2021 [cited by applicant]
US 20230099034A1 · Le · 2023 [cited by examiner]
US 20240005518A1 · Bandwar · 2024 [cited by examiner]
US 20240020790A1 · Gu · 2024 [cited by examiner]
Iizuka, Satoshi, and Edgar Simo-Serra. “Deepremaster: temporal source-reference attention networks for comprehensive video enhancement.” ACM Transactions on Graphics (TOG) 38.6 (2019): 1-13. [cited by examiner]
Lee, Royson, Stylianos I. Venieris, and Nicholas D. Lane. “Deep neural network-based enhancement for image and video streaming systems: A survey and future directions.” ACM Computing Surveys (CSUR) 54.8 (2021): 1-30. [cited by examiner]
Kauff, Peter, and Oliver Schreer. “An immersive 3D video-conferencing system using shared virtual team user environments.” Proceedings of the 4th international conference on Collaborative virtual environments. 2002. [cited by examiner]
He, Zhenyi, et al. “Gazechat: Enhancing virtual conferences with gaze-aware 3d photos.” The 34th Annual ACM Symposium on User Interface Software and Technology. 2021. [cited by examiner]
Wikipedia, “Google Stadia,” https://en.wikipedia.org/wiki/Google_Stadia, Aug. 11, 2021, 15 pages. [cited by applicant]
Wikipedia, “Video Super Resolution,” https://en.wikipedia.org/wiki/Video_Super_Resolution, Jun. 27, 2021, 18 pages. [cited by applicant]
Amazon Web Services, “AI Video Super Resolution,” https://www.amazonaws.cn/en/solutions/ai-super-resolution-on-aws/, Feb. 2020, 6 pages. [cited by applicant]
Wikipedia, “GeForce Now,” https://en.wikipedia.org/wiki/GeForce_Now, Jun. 6, 2021, 5 pages. [cited by applicant]
Wikipedia, “Xbox Cloud Gaming,” https://en.wikipedia.org/wiki/Xbox_Cloud_Gaming, Aug. 9, 2021, 7 pages. [cited by applicant]
C. Faulkner, “Microsoft's xCloud game streaming is now widely available on iOS and PC,” https://www.theverge.com/2021/6/28/22554267/microsoft-xcloud-game-streaming-xbox-pass-ios-iphone-ipad-pc, Jun. 28, 2021, 4 pages. [cited by applicant]
Wikipedia, “Nvidia Shield TV,” https://en.wikipedia.org/wiki/Nvidia_Shield_TV, Jun. 24, 2021, 3 pages. [cited by applicant]
U.S. Appl. No. 17/400,350 filed in the name of Qiang Chen et al. on Aug. 12, 2021, and entitled “Method, Electronic Device, and Computer Program Product for Video Processing.” [cited by applicant]
U.S. Appl. No. 17/400,382 filed in the name of Pedro Fernandez Orellana et al. on Aug. 12, 2021, and entitled “Method, Electronic Device, and Computer Program Product for Image Processing.” [cited by applicant]
U.S. Appl. No. 17/520,908 filed in the name of Qiang Chen et al. on Nov. 8, 2021, and entitled “Method, System, and Computer Program Product for Streaming.” [cited by applicant]
U.S. Appl. No. 17/572,203 filed in the name of Pedro Fernandez Orellana et al. on Jan. 10, 2022, and entitled “Method, Device, and Computer Program Product for Video Processing.” [cited by applicant]
U.S. Appl. No. 17/665,649 filed in the name of Pedro Fernandez Orellana et al. on Feb. 7, 2022, and entitled “Computer-Implemented Method, Device, and Computer Program Product.” [cited by applicant]
U.S. Appl. No. 17/672,369 filed in the name of Qiang Chen et al. on Feb. 15, 2022, and entitled “Method for Generating Metadata, Image Processing Method, Electronic Device, and Program Product.” [cited by applicant]
U.S. Appl. No. 17/857,739 filed in the name of Qiang Chen et al. on Jul. 5, 2022, and entitled “Method, Electronic Device, and Computer Program Product for Video Processing.” [cited by applicant]