IP Library › Granted Patent US 11,601,661
Granted Patent B2
US 11,601,661 · App. 17/394,504 · Granted Mar 7, 2023

Deep loop filter by temporal deformable convolution

Inventors: Wei Jiang (San Jose, CA); Wei Wang (Palo Alto, CA); Zeqiang Li (Los Gatos, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
H04N19/42G06K9/6232G06V20/46H04N19/136H04N19/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,601,661
App. No.
17/394,504
Granted
Mar 7, 2023
Kind
B2
Abstract

A method, apparatus and storage medium for performing video coding are provided. The method includes obtaining a plurality of image frames in a video sequence; determining a feature map for each of the plurality of image frames and determining an offset map based on the feature map; determining an aligned feature map by performing a temporal deformable convolution (TDC) on the feature map and the offset map; and generating a plurality of aligned frames based on the aligned feature map.

Claims (55)

1. A method of performing video coding using one or more neural networks with a loop filter, the method comprising:

obtaining a plurality of image frames in a video sequence;

determining a feature map for each of the plurality of image frames;

selecting a reference frame among the plurality of image frames, the reference frame being a frame to which other frames in the plurality of image frames are to be aligned;

concatenating a reference feature map of the reference frame and a feature map of each of the other frames in the plurality of image frames and passing a concatenated feature map through an offset generation Deep Neural Network (DNN), to generate an offset map;

determining an aligned feature map by performing a temporal deformable convolution (TDC) on the feature map and the offset map; and

generating a plurality of aligned frames based on the aligned feature map.

2. The method of claim 1 , further comprising:

synthesizing the plurality of aligned frames to output a plurality of high-quality frames corresponding to the plurality of image frames.

3. The method of claim 1 , further comprising:

determining an alignment loss indicating an error of misalignment between the feature map and the aligned feature map,

wherein the one or more neural networks are trained by the alignment loss.

4. The method of claim 1 , wherein the obtaining the plurality of image frames comprises stacking the plurality of image frames to obtain a 4-dimensional (4D) input tensor.

5. The method of claim 1 , wherein the plurality of image frames are further processed using at least one of a Deblocking Filter (DF), a Sample-Adaptive Offset (SAO), an Adaptive Loop Filter (ALF) or a Cross-Component Adaptive Filter (CCALF).

6. The method of claim 2 , wherein the plurality of high-quality image frames are assessed to determine reconstruction quality of the plurality of image frames,

wherein the reconstruction quality of the plurality of image frames are back-propagated in the one or more neural networks, and

wherein the one or more neural networks are trained by the reconstruction quality of the plurality of image frames.

7. The method of claim 1 , further comprising determining a discrimination loss indicating an error in a classification of whether each of the plurality of image frames is an original image frame or a high-quality frame, and

wherein one or more neural networks implemented in the apparatus are trained by the discrimination loss.

8. The method of claim 1 , wherein the determining the aligned feature map comprises using a temporal deformable convolution deep neural network (TDC DNN),

wherein the TDC DNN comprises a plurality of TDC layers in a stack, and

wherein each of the plurality of TDC layers is followed by a non-linear activation layer including a Rectified Linear Unit (ReLU).

9. An apparatus comprising:

at least one memory storing computer program code; and

at least one processor configured to access the at least one memory and operate as instructed by the computer program code, the computer program code comprising:

obtaining code configured to cause the at least one processor to obtain a plurality of image frames in a video sequence;

determining code configured to cause the at least one processor to:

determine a feature map for each of the plurality of image frames;

select a reference frame among the plurality of image frames, the reference frame being a frame to which other frames in the plurality of image frames are to be aligned;

concatenate a reference feature map of the reference frame and a feature map of each of the other frames in the plurality of image frames and passing a concatenated feature map through an offset generation Deep Neural Network (DNN), to generate an offset map; and

determine an aligned feature map by performing a temporal deformable convolution (TDC) on the feature map and the offset map; and

generating code configured to cause the at least one processor to generate a plurality of aligned frames based on the aligned feature map.

10. The apparatus of claim 9 , wherein the generating code is further configured to cause the at least one processor to synthesize the plurality of aligned frames to output a plurality of high-quality frames corresponding to the plurality of image frames.

11. The apparatus of claim 9 , wherein the determining code is further configured to cause the at least one processor to determine an alignment loss indicating an error of misalignment between the feature map and the aligned feature map, and

wherein one or more neural networks implemented in the apparatus are trained by the alignment loss.

12. The apparatus of claim 9 , wherein the obtaining code is further configured to cause the at least one processor to arrange the plurality of image frames in a stack to obtain a 4-dimensional (4D) input tensor.

13. The apparatus of claim 9 , further comprising:

processing code configured to cause the at least one processor to process the plurality of image frames using at least one of a Deblocking Filter (DF), a Sample-Adaptive Offset (SAO), an Adaptive Loop Filter (ALF) or a Cross-Component Adaptive Filter (CCALF).

14. The apparatus of claim 10 , wherein the plurality of high-quality image frames are assessed to determine reconstruction quality of the plurality of image frames,

wherein the reconstruction quality of the plurality of image frames are back-propagated to one or more neural networks, and

wherein the one or more neural networks are trained by the reconstruction quality of the plurality of image frames.

15. The apparatus of claim 9 , wherein the determining code is further configured to cause the at least one processor to determine a discrimination loss indicating an error in a classification of whether each of the plurality of image frames is an original image frame or a high-quality frame, and

wherein one or more neural networks implemented in the apparatus are trained by the discrimination loss.

16. The apparatus of claim 9 , wherein the determining code is further configured to cause the at least one processor to determine the aligned feature map using a temporal deformable convolution deep neural network (TDC DNN),

wherein the TDC DNN comprises a plurality of TDC layers in a stack, and

wherein each of the plurality of TDC layers is followed by a non-linear activation layer including a Rectified Linear Unit (ReLU).

17. A non-transitory computer-readable storage medium storing computer program code, the computer program code, when executed by at least one processor, the at least one processor is configured to:

obtain a plurality of image frames in a video sequence;

determine a feature map for each of the plurality of image frames;

selecting a reference frame among the plurality of image frames, the reference frame being a frame to which other frames in the plurality of image frames are to be aligned;

concatenating a reference feature map of the reference frame and a feature map of each of the other frames in the plurality of image frames and passing a concatenated feature map through an offset generation Deep Neural Network (DNN), to generate an offset map;

determine an aligned feature map by performing a temporal deformable convolution (TDC) on the feature map and the offset map; and

generate a plurality of aligned frames based on the aligned feature map.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the at least one processor is further configured to:

synthesize the plurality of aligned frames to output a plurality of high-quality frames corresponding to the plurality of image frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2021
From: JIANG, WEI; WANG, WEI; LI, ZEQIANG; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 057090/0229 →
Continuity (2)
Provisional Application 63090126 · Oct 9, 2020
Related Publication 20220116633A1 · Apr 14, 2022
Cited By (1)
US 12,664,770