Method and apparatus for video coding using an in-loop filter based on a transformer
A method and an apparatus are disclosed for video using an in-loop filter based on Transformer. The video coding method and the apparatus apply a current video block to an attention module of a Transformer, which is a deep learning model. The video coding method and the apparatus utilize the resultant Transformer-based in-loop filter.
1 . A method performed by a video decoding device for enhancing a picture quality of a reconstructed frame, the method comprising:
obtaining an input region of a preset size from the reconstructed frame which is a reconstruction of an original frame and has been reconstructed in advance by the video decoding device; and
generating an enhanced video region that approximates the original frame by feeding the input region into an in-loop filter that is deep learning-based,
wherein the in-loop filter comprises K consecutive Transformer blocks, and K is a natural number,
wherein generating the enhanced video region includes converting, by using the K consecutive Transformer blocks, an input image into a final output feature based on an attention operation, and
wherein at least one of the K consecutive Transformer blocks performs the attention operation based on a query vector, a key vector, and a value vector.
2 . The method of claim 1 , wherein:
the in-loop filter further comprises a first convolutional neural network (CNN); and
generating the enhanced video region further includes
generating an input feature by feeding the input region into the first CNN, and
converting, by using the K consecutive Transformer blocks, the input feature into the final output feature based on the attention operation.
3 . The method of claim 2 , wherein the in-loop filter further comprises a second CNN, and wherein generating the enhanced video region includes feeding the final output feature into the second CNN to generate the enhanced video region.
4 . The method of claim 2 , wherein converting the input feature includes:
partitioning an input feature of each of transformer blocks into patches, and applying an attenuation operation to each of the patches to generate an output feature for the input feature of each of the transformer block.
5 . The method of claim 4 , wherein converting the input feature includes:
setting each of the patches as a query;
calculating, based on a similarity between two patches used in the attention operation, a self-attention score for each of the patches, and attention scores between each of the patches and other patches; and
weighted summing the attention scores to generate an attention value for each of the patches.
6 . The method of claim 4 , wherein converting the input feature includes:
when two patches used in the attention operation are not equal in size, applying a padding to equalize the two patches in size.
7 . The method of claim 5 , wherein converting the input feature includes:
not calculating the attention scores when the two patches used in the attention operation are not present together in the input feature.
8 . The method of claim 5 , wherein converting the input feature includes:
not calculating the attention scores when the two patches used in the attention operation reside in different data processing units, which are coding tree units (CTUs) or virtual pipeline data units (VPDUs).
9 . The method of claim 4 , wherein the input feature includes:
not applying the attenuation operation to each of the patches when each of the patches falls outside a boundary of a preset data processing unit; and
not applying the attenuation operation to a partial region of each of the patches based on a mask that indicates the partial region outside the boundary when the partial region falls outside the boundary.
10 . The method of claim 4 , wherein converting the input feature includes:
not applying the attention operation to each of the patches when each of the patches is present in a later order of decoding than a current block; and
not applying the attenuation operation to a partial region of each of the patches based on a mask that indicates the partial region in the later order when the partial region is present in the later order.
11 . The method of claim 4 , wherein each of the Transformer blocks includes:
consecutive Transformer layers; and
one optional convolution layer,
wherein each of the consecutive Transformer layers includes one encoder layer and one decoder layer.
12 . The method of claim 11 , wherein each of the Transformer blocks is formed as a residual block of the consecutive Transformer layers by using a skip connection.
13 . A method performed by a video encoding device for enhancing a picture quality of a reconstructed frame, the method comprising:
obtaining an input region of a preset size from the reconstructed frame, which is a reconstruction of an original frame and has been reconstructed in advance by the video encoding device; and
generating an enhanced video region that approximates the original frame by feeding the input region into an in-loop filter that is deep learning-based,
wherein the in-loop filter comprises K consecutive Transformer blocks, and K is a natural number,
wherein generating the enhanced video region includes converting, by using the K consecutive Transformer blocks, an input image into a final output feature based on an attention operation, and
wherein at least one of the K consecutive Transformer blocks performs the attention operation based on a query vector, a key vector, and a value vector.
14 . The method of claim 13 , wherein:
the in-loop filter further comprises a first convolutional neural network (CNN); and
generating the enhanced video region further includes
generating an input feature by feeding the input region into the first CNN, and
converting, by using the K consecutive Transformer blocks, the input feature into the final output feature based on the attention operation.
15 . The method of claim 14 , wherein:
the in-loop filter further comprises a second CNN; and
generating the enhanced video region includes feeding the final output feature into the second CNN to generate the enhanced video region.
16 . A non-transitory computer-readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising:
obtaining an input region of a preset size from a reconstructed frame, which is a reconstruction of an original frame and has been reconstructed in advance by a video encoding device; and
generating an enhanced video region that approximates the original frame by feeding the input region into an in-loop filter that is deep learning-based,
wherein the in-loop filter comprises K consecutive Transformer blocks, and K is a natural number,
wherein generating the enhanced video region includes converting, by using the K consecutive Transformer blocks, an input image into a final output feature based on an attention operation, and
wherein at least one of the K consecutive Transformer blocks performs the attention operation based on a query vector, a key vector, and a value vector.