Method for decoder-side motion vector derivation using spatial correlation
View Patent ↗A method for decoder-side motion vector derivation utilizes a spatial correlation. A video coding method and apparatus minimize discontinuities at block boundaries, in order to overcome disadvantages of motion prediction that performs motion compensation on a per block basis. The video coding method and the apparatus derive, during decoder-side motion vector derivation, motion vectors by taking into account spatial correlation of the current block with surrounding blocks rather than deriving motion vectors by considering only the cost of a current and prediction block.
1 . A method performed by a video decoding device for refining a motion vector of a current block, the method comprising:
obtaining an initial motion vector according to an inter-prediction mode of the current block;
calculating a template matching cost by applying a template matching method to the current block and a plurality of reference blocks present in a search range of a reference picture with the initial motion vector as a reference;
calculating a discontinuity measure at a boundary of each of the reference blocks;
calculating a combined cost by a weighted summation with weights on the template matching cost and the discontinuity measure;
selecting from the search range a reference block having a minimum of the combined cost; and
generating a final motion vector by refining the initial motion vector based on motion information between the selected reference block and the current block,
wherein a sum of the weights is equal to 1, and the weights each have a value in a range from 0 to 1.
2 . The method of claim 1 , wherein the inter-prediction mode includes:
an advanced motion vector prediction (AMVP) mode, a merge mode, a combined intra/inter prediction mode, or a geometric partitioning mode.
3 . The method of claim 1 , wherein obtaining the initial motion vector includes:
decoding information on the inter-prediction mode from a bitstream, and then using the information to generate the initial motion vector.
4 . The method of claim 1 , wherein the template matching method includes:
calculating the template matching cost between the current block and each of the reference blocks by using a similarity between a neighboring template of the current block and a neighboring template of each of the reference blocks.
5 . The method of claim 1 , wherein the discontinuity measure is calculated by applying a method of utilizing spatial correlation.
6 . The method of claim 5 , wherein the method of utilizing spatial correlation includes:
calculating the discontinuity measure at the boundary within the reference block between samples and neighboring samples.
7 . The method of claim 1 , wherein the search range has a sample range of a predetermined size in horizontal and vertical directions with the initial motion vector as a reference.
8 . A method performed by a video encoding device for refining a motion vector of a current block, the method comprising:
determining an inter-prediction mode of the current block and an initial motion vector according to the inter-prediction mode;
calculating a template matching cost by applying a template matching method to the current block and a plurality of reference blocks present in a search range of a reference picture with the initial motion vector as a reference;
calculating a discontinuity measure at a boundary of each of the reference blocks;
calculating a combined cost by a weighted summation with weights on the template matching cost and the discontinuity measure; and
selecting from the search range a reference block having a minimum of the combined cost; and
generating a final motion vector by refining the initial motion vector based on motion information between the selected reference block and the current block,
wherein a sum of the weights is equal to 1, and the weights each have a value in a range from 0 to 1.
9 . The method of claim 8 , further comprising:
encoding information on the inter-prediction mode, information on the initial motion vector, and information on the reference picture.
10 . The method of claim 8 , wherein the inter-prediction mode includes:
an advanced motion vector prediction (AMVP) mode, a merge mode, a combined intra/inter prediction mode, or a geometric partitioning mode.
11 . The method of claim 8 , wherein the template matching method includes:
calculating the template matching cost between the current block and each of the reference blocks by using a similarity between a neighboring template of the current block and a neighboring template of each of the reference blocks.
12 . The method of claim 8 , wherein the discontinuity measure is calculated by applying a method of utilizing spatial correlation.
13 . The method of claim 12 , wherein the method of utilizing spatial correlation includes:
calculating the discontinuity measure at the boundary within the reference block between samples and neighboring samples.
14 . A non-transitory computer-recording medium storing instructions, when executed by a processor, to perform an encoding method for generating a bitstream comprising:
determining an inter-prediction mode of a current block and an initial motion vector according to the inter-prediction mode;
calculating a template matching cost by applying a template matching method to the current block and a plurality of reference blocks present in a search range of a reference picture with the initial motion vector as a reference;
calculating a discontinuity measure at a boundary of each of the reference blocks;
calculating a combined cost by a weighted summation with weights on the template matching cost and the discontinuity measure;
selecting from the search range a reference block having a minimum of the combined cost;
and generating a final motion vector by refining the initial motion vector based on motion information between the selected reference block and the current block,
wherein a sum of the weights is equal to 1, and the weights each have a value in a range from 0 to 1.
15 . The non-transitory computer-readable recording medium of claim 14 , wherein the inter-prediction mode includes:
an advanced motion vector prediction (AMVP) mode, a merge mode, a combined intra/inter prediction mode, or a geometric partitioning mode.
16 . The non-transitory computer-readable recording medium of claim 14 , wherein obtaining the initial motion vector includes:
decoding information on the inter-prediction mode from a bitstream, and then using the information to generate the initial motion vector.
17 . The non-transitory computer-readable recording medium of claim 14 , wherein the template matching method includes:
calculating the template matching cost between the current block and each of the reference blocks by using a similarity between a neighboring template of the current block and a neighboring template of each of the reference blocks.
18 . The non-transitory computer-readable recording medium of claim 14 , wherein the discontinuity measure is calculated by applying a method of utilizing spatial correlation.
19 . The non-transitory computer-readable recording medium of claim 18 , wherein the method of utilizing spatial correlation includes:
calculating the discontinuity measure at the boundary within the reference block between samples and neighboring samples.