IP Library › Granted Patent US 10,733,714
Granted Patent B2
US 10,733,714 · App. 15/946,531 · Granted Aug 4, 2020

Method and apparatus for video super resolution using convolutional neural network with two-stage motion compensation

Inventors: Mostafa El-Khamy (San Diego, CA); Haoyu Ren (San Diego, CA); Jungwon Lee (San Diego, CA)
Assignee: Samsung Electronics Co., Ltd
G06T5/50G06K9/6256G06K9/6857G06T3/0093G06T3/4046G06T3/4053G06T2207/10016G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,733,714
App. No.
15/946,531
Granted
Aug 4, 2020
Kind
B2
Abstract

A method and an apparatus are provided. The method includes receiving a video with a first plurality of frames having a first resolution; generating a plurality of warped frames from the first plurality of frames based on a first type of motion compensation; generating a second plurality of frames having a second resolution, wherein the second resolution is of higher resolution than the first resolution, wherein each of the second plurality of frames having the second resolution is derived from a subset of the plurality of warped frames using a convolutional network; and generating a third plurality of frames having the second resolution based on a second type of motion compensation, wherein each of the third plurality of frames having the second resolution is derived from a fusing a subset of the second plurality of frames.

Claims (48)

1. A method, comprising:

receiving a video with a first plurality of frames having a first resolution;

generating a plurality of warped frames from the first plurality of frames based on a first type of motion compensation;

generating a second plurality of frames having a second resolution,

wherein the second resolution is of higher resolution than the first resolution,

wherein each of the second plurality of frames having the second resolution is derived from a subset of the plurality of warped frames using a convolutional network; and

generating a third plurality of frames having the second resolution based on a second type of motion compensation,

wherein each of the third plurality of frames having the second resolution is derived from fusing a subset of the second plurality of frames.

2. The method of claim 1 , wherein generating the plurality of warped frames comprises:

learning a low resolution (LR) optical flow of neighboring frames around a reference frame in an LR video with respect to the reference frame in the LR video using a convolutional neural network (CNN); and

generating warped LR frames of the neighboring frames around the reference frame in the LR video based on the LR optical flow.

3. The method of claim 2 , wherein learning the LR optical flow is comprised of training an ensemble of deep fully convolutional neural networks (CNNs) on ground truth optical flow data and applying the ensemble of deep fully CNNs directly to frames of the LR video.

4. The method of claim 1 , wherein generating the second plurality of frames having a second resolution comprises using a deep convolutional neural network wherein the layers deploy a three dimensional (3D) convolutional operator.

5. The method of claim 1 , wherein generating the third plurality of frames comprises:

learning a high resolution (HR) optical flow from HR intermediate frames using a convolutional neural network (CNN);

generating warped HR frames from consecutive frames of the HR intermediate frames based on an HR optical flow with respect to a reference intermediate frame; and

applying a weighted fusion to the warped HR frames to generate refined frames of an HR video.

6. The method of claim 5 , wherein learning the HR optical flow comprises training an ensemble of deep fully convolutional neural networks (CNNs) on ground truth optical flow data and applying the ensemble of deep fully CNNs directly to frames of the HR intermediate frames.

7. The method of claim 5 , wherein applying the weighted fusion comprises identifying pixels in neighboring frames that correspond to those in the reference intermediate frame based on the learned HR optical flow for each neighboring frame, and applying the weighted fusion based on Gaussian weights and motion penalization based on a magnitude of the optical flow.

8. The method of claim 1 , wherein generating the second plurality of frames comprises:

learning a high resolution (HR) optical flow by performing super resolution (SR) on a low resolution (LR) optical flow;

generating warped HR frames from consecutive frames of intermediate HR frames based on the HR optical flow; and

applying a weighted fusion to the warped HR frames to generate refined frames of a HR video.

9. The method of claim 8 , wherein performing SR on the LR optical flow comprises training a deep full convolutional neural network to perform SR on the LR optical flow where ground truth HR optical flow is estimated by applying HR optical flow on ground truth HR frames.

10. A apparatus, comprising:

a receiver configured to receive a video with a first plurality of frames having a first resolution;

a first resolution compensation device configured to generate a plurality of warped frames from the first plurality of frames based on a first type of motion compensation;

a multi-image spatial super resolution (SR) device configured to generate a second plurality of frames having a second resolution,

wherein the second resolution is of higher resolution than the first resolution,

wherein each of the second plurality of frames having the second resolution is derived from a subset of the plurality of warped frames using a convolutional network; and

a second resolution motion compensation device configured to generate a third plurality of frames having the second resolution based on a second type of motion compensation,

wherein each of the third plurality of frames having the second resolution is derived from fusing a subset of the second plurality of frames.

11. The apparatus of claim 10 , wherein the first resolution motion compensation device is further configured to:

learn a low resolution (LR) optical flow of neighboring frames around a reference frame in an LR video with respect to the reference frame in the LR video using a convolutional neural network (CNN); and

generate warped LR frames of the neighboring frames around the reference frame in the LR video based on the LR optical flow.

12. The apparatus of claim 11 , wherein the first resolution motion compensation device is further configured to train an ensemble of deep fully convolutional neural networks (CNNs) on ground truth optical flow data and applying the ensemble of deep fully CNNs directly to frames of the LR video.

13. The apparatus of claim 10 , wherein the multi-image spatial SR device is further configured to use a deep convolutional neural network wherein the layers deploy a three dimensional (3D) convolutional operator.

14. The apparatus of claim 10 , wherein the second resolution motion compensation device is further configured to:

learn a high resolution (HR) optical flow from HR intermediate frames using a convolutional neural network (CNN);

generate warped HR frames from consecutive frames of the HR intermediate frames based on an HR optical flow with respect to a reference intermediate frame; and

apply a weighted fusion to the warped HR frames to generate refined frames of an HR video.

15. The apparatus of claim 14 , wherein the second resolution motion compensation device is further configured to train an ensemble of deep fully convolutional neural networks (CNNs) directly to frames of the HR intermediate frames.

16. The apparatus of claim 14 , wherein the second resolution motion compensation device is further configured to identify pixels in neighboring frames that correspond to those in the reference intermediate frame based on the learned HR optical flow for each neighboring frame, and apply the weighted fusion based on Gaussian weights and motion penalization based on a magnitude of the optical flow.

17. The apparatus of claim 10 , wherein the second resolution motion compensation device is further configured to:

learn a high resolution (HR) optical flow by performing SR on a low resolution (LR) optical flow;

generate warped HR frames from consecutive frames of intermediate HR frames based on the HR optical flow; and

apply a weighted fusion to the warped HR frames to generate refined frames of an HR video.

18. The apparatus of claim 17 , wherein the second resolution motion compensation device is further configured to perform SR on the LR optical flow by training a deep full convolutional neural network to perform SR on the LR optical flow where ground truth HR optical flow is estimated by applying HR optical flow on ground truth HR frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2018
From: EL-KHAMY, MOSTAFA; REN, HAOYU; LEE, JUNGWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 045732/0402 →
Continuity (2)
Provisional Application 62583633 · Nov 9, 2017
Related Publication 20190139205A1 · May 9, 2019
Cited By (1)
US 12,307,756